Together AI is a cloud platform for running, customizing and training open-source and open-weight AI models through APIs and rentable GPU infrastructure. The company markets it as the AI Native Cloud, a full-stack AI platform powered by cutting-edge research. For most developers the entry point is an open-source LLM API.
The service is operated by Together Computer, Inc., a Delaware corporation whose terms cover APIs and web interfaces to host, use, fine-tune and train large AI models. Independent tech press has described Together AI as an AI neocloud that rents out Nvidia GPU clusters and reported an $800 million Series C at an $8.3 billion valuation. The same report says the company claims thousands of paying customers, naming Cursor, Cognition and Decagon among them.
Three facts shape most decisions: there is no free trial and access starts with a $5 credit purchase, prompts are stored by default unless zero data retention is enabled, and shared serverless capacity is best-effort.
Together AI groups its products into inference, compute and model shaping, all billed from the same prepaid credit balance.
The company says it achieved nearly 2x faster serverless inference for gpt-oss-20B than the next fastest provider. That is a vendor benchmark, and the model was later scheduled for removal from serverless.
LLM fine-tuning output therefore runs on dedicated hardware or leaves the platform as weights, rather than joining the per-token serverless catalog.
A first API call needs a funded account, an API key and either the official SDK or an existing OpenAI client:
Fine-tuning cost can be checked before any money is spent, because the CLI prints the estimated price and asks for confirmation before the job is submitted.
The customer figures above are published by Together AI on its own pages and were not independently verified for this listing.
Together AI fits developers and ML teams who work comfortably with APIs, SDKs and command-line tools and who want open-weight models without running their own inference stack.
Choosing between the two inference modes comes down to utilization. Dedicated model inference is usually cheaper when a replica would stay busy most of the day. Serverless is usually cheaper when traffic is low or bursty enough that a dedicated replica would sit idle most of the time. The pricing page also notes that most teams start with serverless inference and move to dedicated endpoints at scale.
It is a weaker fit for anyone who wants to try models before paying, since there is no free trial, and for apps built on OpenAI's Assistants API, which the compatibility layer does not implement.
Together AI is used through a web console, SDKs, a CLI and a REST API rather than a consumer app.
Together AI pricing is usage-based and fully prepaid. Together AI does not currently offer a free trial, and platform access requires a minimum $5 credit purchase. A positive credit balance is required to use the platform. Prepaid balance credits currently have no expiration date. Auto-recharge can top up the balance automatically, but only when the default payment method is a credit or debit card. Under the Terms of Service, fees paid are non-refundable unless the terms or an order form specify otherwise.
Serverless prices are quoted per 1M tokens and change often. The representative rows below were captured on October 5, 2026.
| Model | Input | Cached input | Output |
|---|---|---|---|
| MiniMax M3 | $0.30 | $0.06 | $1.20 |
| Kimi K3 | $3.00 | $0.30 | $15.00 |
| gpt-oss-120B | $0.15 | Not listed | $0.60 |
| Llama 3.3 70B | $1.04 | Not listed | $1.04 |
The Batch API advertises up to 50% off serverless rates and a separate rate limit pool. The batch documentation's discounted-model table, however, lists only a Llama 3.3 70B Turbo variant and Whisper Large v3 at 50% off, and models not listed run at standard rates.
All GPU cluster prices are per GPU per hour.
| GPU | Preemptible | On-demand | Reserved 7–30 days | 31–90 days | 91–180 days | 181+ days |
|---|---|---|---|---|---|---|
| NVIDIA HGX H100 | $1.99 | $3.99 | $3.69 | $3.45 | $3.19 | Contact us |
| NVIDIA HGX H200 | $2.99 | $5.99 | $4.99 | $4.15 | $3.99 | Contact us |
| NVIDIA HGX B200 | $4.09 | $8.19 | $7.99 | $7.79 | $6.79 | Contact us |
Reservations are charged for the full reserved duration once the cluster is provisioned, and usage beyond the reserved capacity is billed at on-demand rates. Startup accelerator credits do not apply to Reserved GPU Clusters.
A dedicated replica bills only while it is ready and able to serve traffic, so provisioning and cold-start time are not charged. The H100 rate differs between official pages: the pricing page's dedicated inference table lists $5.49 per GPU-hour on demand. A docs changelog entry, by contrast, says H100 80GB dedicated endpoint hardware is now $3.99 per hour, down from $5.49.
Fine-tuning is billed on tokens processed, meaning training dataset size times epochs plus any evaluation tokens, and each job carries a per-model minimum charge. The first fine-tuning table on the pricing page, under tabs labeled LoRA and full fine-tuning, lists Qwen3.5 0.8B at $0.34 per 1M tokens for supervised fine-tuning, $0.84 for DPO and a $4.00 minimum charge. Further down the same table, GLM-5.2 lists $40.00 for supervised fine-tuning, $100.00 for DPO and a $60.00 minimum. Cancelled or early-stopped jobs are charged for completed steps only. Hosting a fine-tuned model on a dedicated endpoint is billed separately by the minute.
Comparable open-model inference providers differ mainly in how they bill fine-tuned models and whether new accounts start with free credits.
Against those options, Together AI asks for a $5 first purchase with no free trial and bills hosting of fine-tuned models per minute on dedicated hardware. In exchange, one account covers serverless models, batch jobs, dedicated endpoints, GPU clusters, sandboxes and storage.
The independent media and review sources checked for this page showed no product-level outage or security report, and review samples were under 20 reviews, too few for a quality conclusion.
No. Access requires buying at least $5 in credits, and no free trial is offered. Startups accepted into the accelerator program can receive platform credits instead, though those credits exclude Reserved GPU Clusters.
Yes, for chat, completions, vision, image generation, text-to-speech and embeddings: point the OpenAI client at Together's base URL and switch to Together's namespaced model IDs. Assistants, Threads and Runs are not available, and batch jobs use Together's own Batch API.
Not by default. Training use is an opt-in setting that stays off unless enabled. Prompts and responses are still stored by default unless zero data retention is turned on.
API access is suspended until you add credits. On-demand GPU clusters are paused and later decommissioned if credits are not restored, while existing GPU cluster reservations remain active until their scheduled end date.
Yes. After a job completes, the model can be served on a Together dedicated endpoint or downloaded as a checkpoint for local inference or another host.