Toolso.AI
Toolso.AI
All ToolsCategoriesTrendingLatest ToolsPricingBlog
Toolso.AI
Toolso.AI
Toolso.AI
Toolso.AI

Discover the best AI tools to boost your productivity

GitHubGitHubTwitterX (Twitter)YouTubeYouTubeTikTokEmail

Popular Categories

  • AI Writing
  • AI Image
  • AI Video
  • AI Coding
  • More Categories

Explore

  • Latest Tools
  • Popular Tools
  • More Tools
  • Submit Tool
  • Pricing

About

  • About Us
  • Contact
  • Blog
  • Changelog

Legal

  • Cookie Policy
  • Privacy Policy
  • Terms of Service
  • Refund Policy
© 2026 Toolso.AI All Rights Reserved
Limited timeLimited-time offerFeatured Listing24h priority review · No backlink · 30 days featured$29.90then $59.90Price rises to $59.90 after Oct 31Ends in--:--:--Submit now
  1. Home
  2. All Tools
  3. Developer Tools
  4. Together AI
Together AI interface preview
Together AI logo

Together AI

Together AI runs open-weight models through a serverless, OpenAI-compatible API and adds dedicated endpoints, fine-tuning, GPU clusters and sandboxes, all paid from prepaid credits.

Developer ToolsAI DevelopmentAI Training Platform#Batch Processing#Sdk#LLM
Try for Free
Saves
Visits
Views
Pricing
Freemium
Published
Oct 6, 2026
Domain
together.ai
Community rating

Used this tool? Rate it

Rate this tool

Together AI Product Information

Try for Free
Tool Information
Saves
Visits
Views
Pricing
Freemium
Published
Oct 6, 2026
Domain
together.ai
Community rating

Used this tool? Rate it

Rate this tool

Featured Tools

Related Tools

Try for Free

What is Together AI?

Together AI is a cloud platform for running, customizing and training open-source and open-weight AI models through APIs and rentable GPU infrastructure. The company markets it as the AI Native Cloud, a full-stack AI platform powered by cutting-edge research. For most developers the entry point is an open-source LLM API.

The service is operated by Together Computer, Inc., a Delaware corporation whose terms cover APIs and web interfaces to host, use, fine-tune and train large AI models. Independent tech press has described Together AI as an AI neocloud that rents out Nvidia GPU clusters and reported an $800 million Series C at an $8.3 billion valuation. The same report says the company claims thousands of paying customers, naming Cursor, Cognition and Decagon among them.

Three facts shape most decisions: there is no free trial and access starts with a $5 credit purchase, prompts are stored by default unless zero data retention is enabled, and shared serverless capacity is best-effort.

Core features

Together AI groups its products into inference, compute and model shaping, all billed from the same prepaid credit balance.

Inference options

  • Serverless inference: Developers call open-weight models through a single API, with nothing to deploy or manage, and pay only for the tokens they use. The catalog covers text, image, video, code and voice models.
  • Batch inference: Large asynchronous jobs can scale to 30 billion tokens per model, using any serverless model or a private deployment.
  • Provisioned Throughput: Committed inference capacity comes with token-based pricing, reserved throughput and a 99% uptime SLA.
  • Dedicated endpoints and containers: Dedicated Model Inference runs a model on reserved, isolated hardware, while Dedicated Container Inference targets generative media workloads such as video, audio and image models.

The company says it achieved nearly 2x faster serverless inference for gpt-oss-20B than the next fastest provider. That is a vendor benchmark, and the model was later scheduled for removal from serverless.

Fine-tuning and model shaping

  • Bring a model: Together AI says any open-source model from Hugging Face Hub can be fine-tuned without format conversions.
  • Tool-calling data: Tool-calling training accepts existing agent logs as-is, with function definitions and tool calls placed directly in the dataset.
  • Vision fine-tuning: Vision models train directly on raw PNG, JPEG or WEBP images without special preprocessing, using LoRA or full fine-tuning.
  • Portable results: A finished model can be hosted on a Together dedicated endpoint or downloaded as a standalone checkpoint.

LLM fine-tuning output therefore runs on dedicated hardware or leaves the platform as weights, rather than joining the per-token serverless catalog.

Compute, sandboxes and storage

  • GPU Clusters: Customers can reserve from 8 to more than 4,000 GPUs for training or serving.
  • Code Sandbox: Sandbox VMs range from 2 to 64 vCPUs and 1 to 128 GB of RAM for AI development environments.
  • Managed Storage: Managed Storage offers object storage and parallel filesystems with zero egress fees.

Guide

A first API call needs a funded account, an API key and either the official SDK or an existing OpenAI client:

  1. Register in the Together AI console and add at least $5 in credits from the billing settings.
  2. Create a key on the project's API keys page; new API keys are shown only once, so store the value right away.
  3. Install the official Python or TypeScript SDK, or keep an existing OpenAI client and change its base URL, then set the key as an environment variable.
  4. Run the quickstart example, which sends a chat completion request to MiniMax M3 and prints the response.
  5. For offline workloads, the batch API runs many independent requests asynchronously from a single uploaded JSONL file.

Fine-tuning cost can be checked before any money is spent, because the CLI prints the estimated price and asks for confirmation before the job is submitted.

Together AI use cases and examples

  • Offline batch work: Batch jobs suit latency-tolerant tasks such as classifying a large dataset, running evaluations, generating synthetic data and offline summarization.
  • Real-time voice agents: A Together AI customer story on Decagon's sub-second voice AI lists a 6x cost reduction per turn versus gpt-5 mini.
  • Lower-latency production APIs: A customer testimonial from Salesforce AI Research cites a 2x latency reduction and costs cut by roughly a third.
  • Research and training runs: GPU Clusters are pitched for spinning up a cluster in minutes, testing a hypothesis and shutting it down when the experiment ends.
  • Coding agents on open models: Together Link runs six coding agents on Together-hosted models, including Claude Code, Codex, OpenCode and Pi Code in the terminal plus Claude Desktop and ChatGPT Desktop.

The customer figures above are published by Together AI on its own pages and were not independently verified for this listing.

Who is it for

Together AI fits developers and ML teams who work comfortably with APIs, SDKs and command-line tools and who want open-weight models without running their own inference stack.

  • Prototyping and variable traffic: Serverless inference is pitched as best for variable or unpredictable traffic, rapid prototyping, and cost-sensitive or early-stage production workloads.
  • Steady, latency-sensitive production: Dedicated Model Inference is aimed at predictable traffic, latency-sensitive applications and high-throughput production workloads.
  • Funded startups: A selection-based accelerator offers up to $15K in platform credits for startups that raised up to $5M, up to $30K for $5M–$10M, and up to $50K above $10M.

Choosing between the two inference modes comes down to utilization. Dedicated model inference is usually cheaper when a replica would stay busy most of the day. Serverless is usually cheaper when traffic is low or bursty enough that a dedicated replica would sit idle most of the time. The pricing page also notes that most teams start with serverless inference and move to dedicated endpoints at scale.

It is a weaker fit for anyone who wants to try models before paying, since there is no free trial, and for apps built on OpenAI's Assistants API, which the compatibility layer does not implement.

Platforms

Together AI is used through a web console, SDKs, a CLI and a REST API rather than a consumer app.

  • SDKs and APIs: Together AI publishes official SDKs for Python and TypeScript, and the OpenAI SDK or plain REST calls from any language also work.
  • OpenAI compatibility: The API is compatible with the OpenAI REST API and SDKs across chat, completions, vision, image generation, text-to-speech and embeddings.
  • Batch tooling: Batch jobs can be started from the console, through the API or SDK, or with the CLI.
  • Cluster management: GPU cluster access is managed with project-level RBAC through the CLI, SDK, API, Terraform or the web console.
  • Regions: Serverless endpoints do not offer region selection. Private networking and VPC-based deployments, including EU regions, are available through dedicated setups.

Pricing

Together AI pricing is usage-based and fully prepaid. Together AI does not currently offer a free trial, and platform access requires a minimum $5 credit purchase. A positive credit balance is required to use the platform. Prepaid balance credits currently have no expiration date. Auto-recharge can top up the balance automatically, but only when the default payment method is a credit or debit card. Under the Terms of Service, fees paid are non-refundable unless the terms or an order form specify otherwise.

Serverless token prices

Serverless prices are quoted per 1M tokens and change often. The representative rows below were captured on October 5, 2026.

ModelInputCached inputOutput
MiniMax M3$0.30$0.06$1.20
Kimi K3$3.00$0.30$15.00
gpt-oss-120B$0.15Not listed$0.60
Llama 3.3 70B$1.04Not listed$1.04

The Batch API advertises up to 50% off serverless rates and a separate rate limit pool. The batch documentation's discounted-model table, however, lists only a Llama 3.3 70B Turbo variant and Whisper Large v3 at 50% off, and models not listed run at standard rates.

GPU clusters

All GPU cluster prices are per GPU per hour.

GPUPreemptibleOn-demandReserved 7–30 days31–90 days91–180 days181+ days
NVIDIA HGX H100$1.99$3.99$3.69$3.45$3.19Contact us
NVIDIA HGX H200$2.99$5.99$4.99$4.15$3.99Contact us
NVIDIA HGX B200$4.09$8.19$7.99$7.79$6.79Contact us

Reservations are charged for the full reserved duration once the cluster is provisioned, and usage beyond the reserved capacity is billed at on-demand rates. Startup accelerator credits do not apply to Reserved GPU Clusters.

Dedicated inference

A dedicated replica bills only while it is ready and able to serve traffic, so provisioning and cold-start time are not charged. The H100 rate differs between official pages: the pricing page's dedicated inference table lists $5.49 per GPU-hour on demand. A docs changelog entry, by contrast, says H100 80GB dedicated endpoint hardware is now $3.99 per hour, down from $5.49.

Fine-tuning

Fine-tuning is billed on tokens processed, meaning training dataset size times epochs plus any evaluation tokens, and each job carries a per-model minimum charge. The first fine-tuning table on the pricing page, under tabs labeled LoRA and full fine-tuning, lists Qwen3.5 0.8B at $0.34 per 1M tokens for supervised fine-tuning, $0.84 for DPO and a $4.00 minimum charge. Further down the same table, GLM-5.2 lists $40.00 for supervised fine-tuning, $100.00 for DPO and a $60.00 minimum. Cancelled or early-stopped jobs are charged for completed steps only. Hosting a fine-tuned model on a dedicated endpoint is billed separately by the minute.

Together AI alternatives

Comparable open-model inference providers differ mainly in how they bill fine-tuned models and whether new accounts start with free credits.

  • Fireworks AI: Fireworks AI states that fine-tuned models are served for the same price as base models. Its serverless inference is billed per token on a pre-paid, usage-based basis.
  • Baseten: Baseten says new accounts come with credits to experiment with deployments for free. Its dedicated deployments charge only for the compute used, down to the minute.

Against those options, Together AI asks for a $5 first purchase with no free trial and bills hosting of fine-tuned models per minute on dedicated hardware. In exchange, one account covers serverless models, batch jobs, dedicated endpoints, GPU clusters, sandboxes and storage.

Limitations

Service limits

  • Best-effort serverless: Serverless performance is best-effort, and committed throughput requires Provisioned Throughput.
  • Throttling: When demand is high, Together may limit requests, and clients can receive 429 Too Many Requests or 503 Service Unavailable responses.
  • Model churn: One changelog entry scheduled openai/gpt-oss-20b and other models for removal from serverless on September 14, 2026. All but one of the models in that entry stayed available through on-demand dedicated endpoints, which bill for hardware time instead of tokens.
  • Partial OpenAI compatibility: Assistants, Threads and Runs from the OpenAI API are not implemented. OpenAI model names such as gpt-4o or text-embedding-3-large return a 404 on Together.
  • Batch caps: A single batch can hold up to 50,000 requests. Batch input, output and error files are kept for 7 days, so results must be downloaded within that window.

Billing and support boundaries

  • Zero balance: If the balance reaches zero, API access is suspended until credits are added.
  • On-demand clusters: For fully on-demand GPU clusters, running out of credits means the cluster is first paused and then decommissioned if credits are not restored.
  • Support tiers: The support page still describes support by tier, with community Discord help for Build tier users. A changelog entry, however, says the Build Tiers 1–5, Scale and Enterprise tier labels have been retired.

Privacy, data use and ownership

  • Default storage: By default, Together stores prompts and model responses and may use them for product improvements.
  • Training: Use of customer data for training models is a separate opt-in that is off by default. The privacy policy states that collected data will not be used to train the company's models without explicit opt-in consent.
  • Zero data retention: Zero data retention is not enabled by default. With ZDR on, prompts and outputs are not stored by Together or used for any secondary purpose. ZDR applies only from the moment it is enabled and does not affect data processed earlier. Files uploaded for fine-tuning or the Batch API remain stored until they are deleted through the API, CLI or web console.
  • Passthrough models: Passthrough models, which forward requests to a third-party provider, are allowed by default and fall under that provider's own data policy.
  • Hosted third-party models: Models from third-party authors such as DeepSeek, Qwen or Mistral run on Together's own infrastructure and do not call out to the model author.
  • Ownership and compliance: Customers exclusively own their content and output under the Terms of Service. Together AI lists SOC 2 Type II, HIPAA-aligned options, and encryption in transit and at rest.

The independent media and review sources checked for this page showed no product-level outage or security report, and review samples were under 20 reviews, too few for a quality conclusion.

FAQ

Q1. Does Together AI have a free tier or free trial?

No. Access requires buying at least $5 in credits, and no free trial is offered. Startups accepted into the accelerator program can receive platform credits instead, though those credits exclude Reserved GPU Clusters.

Q2. Can I use the OpenAI SDK with Together AI?

Yes, for chat, completions, vision, image generation, text-to-speech and embeddings: point the OpenAI client at Together's base URL and switch to Together's namespaced model IDs. Assistants, Threads and Runs are not available, and batch jobs use Together's own Batch API.

Q3. Does Together AI train on my prompts?

Not by default. Training use is an opt-in setting that stays off unless enabled. Prompts and responses are still stored by default unless zero data retention is turned on.

Q4. What happens when my credits run out?

API access is suspended until you add credits. On-demand GPU clusters are paused and later decommissioned if credits are not restored, while existing GPU cluster reservations remain active until their scheduled end date.

Q5. Can I take a fine-tuned model elsewhere?

Yes. After a job completes, the model can be served on a Together dedicated endpoint or downloaded as a checkpoint for local inference or another host.

Know a Similar Tool?
If you know other great AI tools, feel free to submit them to us