Fireworks (styled Fireworks AI in its documentation) is a cloud platform for running and training open AI models. Developers call hosted models through an LLM inference API, rent dedicated GPUs for their own deployments, or fine-tune open models on their own data and serve the result on the same infrastructure.
The vendor pitches the platform as "Foundational Infrastructure for Specialized Intelligence," an owned learning loop that turns leading open source models and a customer's data into an edge that sharpens with every iteration. Fireworks does not build foundation models from scratch: co-founder and CEO Lin Qiao has described the business as helping companies fine-tune existing models for their particular needs. Qiao previously led the AI platform development team at Meta.
For scale, independent technology press reported in September 2026 that Fireworks announced in July that its annualized revenue had reached $1 billion, a fivefold increase from the year before.
These options make fine-tuning open models a pipeline into serving, not a separate product: the homepage states that every checkpoint deploys to production in seconds.
A typical first project moves from serverless testing to cost controls:
firectl quota list to check the account's current rate limits, GPU quotas, spend limits and usage before planning a launch.Customer quotes on the Fireworks homepage are vendor-selected testimonials, not independent case studies, but they show the workloads the platform targets:
It is a weaker fit in two cases. Teams that need a fixed model version get less from serverless, because serverless models are managed by the Fireworks team and may be updated or deprecated as new models are released. Use by minors is prohibited by the terms unless a parent or legal guardian supervises it.
Billing is prepaid: Auto Reload can top up the balance automatically and a monthly spend limit caps usage. Contracted customers may have the option to move to post-paid billing.
Prices are in USD per 1 million tokens, shown as input / cached input / output, as captured on October 5, 2026.
| Model | Standard | Priority |
|---|---|---|
| Kimi K3 | $3.00 / $0.30 / $15.00 | $3.75 / $0.375 / $18.75 |
| DeepSeek V4.1 Flash | $0.30 / $0.006 / $1.20 | $0.375 / $0.0075 / $1.50 |
| GLM 5.3 | $1.40 / $0.26 / $4.40 | $1.75 / $0.325 / $5.50 |
| OpenAI GPT OSS 120B | $0.15 / $0.015 / $0.60 | $0.18 / $0.018 / $0.72 |
Text and vision models not listed individually are priced by size: less than 4B parameters costs $0.10, 4B to 16B costs $0.20, and more than 16B costs $0.90 per 1M tokens, with no separate cached rate. Batch inference is billed at 50% of serverless pricing on both input and output. Beginning September 1, 2026, launched US-only models are priced at 1.5x the base model serverless prices.
On-demand deployments are billed per GPU second, with no extra charges for start-up times.
| GPU | Per minute | Per hour |
|---|---|---|
| H100 80 GB | $0.134 | $8.00 |
| H200 141 GB | $0.134 | $8.00 |
| B200 180 GB | $0.217 | $13.00 |
| B300 288 GB | $0.250 | $15.00 |
| GB300 288 GB | $0.334 | $20.00 |
Region-restricted deployments are priced at a 1.5x premium and go through sales.
Managed supervised and preference fine-tuning is priced per 1M training tokens:
| Base model size | LoRA SFT | LoRA DPO | Full-param SFT | Full-param DPO |
|---|---|---|---|---|
| Up to 16B | $0.50 | $1.00 | $1.00 | $2.00 |
| 16.1B to 80B | $3.00 | $6.00 | $6.00 | $12.00 |
| 80B to 300B | $6.00 | $12.00 | $12.00 | $24.00 |
| Over 300B | $10.00 | $20.00 | $20.00 | $40.00 |
Training tokens are estimated as the number of tokens in the training dataset multiplied by the number of epochs. Dedicated Training API jobs are priced per GPU hour at the on-demand rates.
Serving costs read differently depending on the page. The pricing page says you can serve fine-tuned models for the same price as base models, while the billing FAQ says trained LoRA models require a dedicated deployment to serve, billed per GPU second.
Fireworks bills dedicated GPUs per second, while Baseten meters to the minute, a difference that matters most for short, bursty deployments.
By default, Fireworks does not log or store prompt or generation data for any open models, without explicit user opt-in; it does log request metadata such as token counts. With prompt caching active, some prompt data can stay in volatile memory for several minutes.
The Response API is an exception: it stores conversations by default, and stored conversation data is automatically deleted after 30 days. The terms also limit the Zero Data Retention commitment to inference; it does not cover features that keep data by design, such as training, fine-tuning or agent features.
On training use, the two documents are worded differently. The Terms of Service say Fireworks will not use your Content to train its own models or to improve the Service. The Privacy Notice says Fireworks does not use prompts, training data or API inputs to train or improve its models without your explicit opt-in. As between the parties, you own all right, title and interest in your Content, which the terms define as your inputs and outputs.
Enterprise admins can switch on an account-wide Zero Data Retention policy, which rejects anything that would persist customer content, including batch inference jobs, fine-tuning and RL jobs, dataset uploads and FireRouter virtual models. When FireRouter sends a turn to Claude or GPT, the closed model runs on your own Anthropic or OpenAI account.
These are vendor statements, not independent audits:
No free plan appears on the pricing pages checked. Beyond the $1 credit mentioned in the billing FAQ, continued use requires a payment method and prepaid credits.
No. LoRA models need an on-demand deployment billed by GPU time; one deployment can host up to 100 LoRA adapters, which spreads that cost.
Serverless inference pricing bills each token, so idle time costs nothing. On-demand deployments are described as cheaper under high utilization; they are billed by GPU-second rather than per token.