Cerebras is an AI inference platform that runs open-weight language models on the company's own wafer-scale processors. Developers reach it through Cerebras Inference, a cloud API with a browser Playground; larger organizations can reserve Dedicated capacity or install Cerebras systems in their own data centers, and the same hardware also backs training and fine-tuning services.
The public cloud tier serves a regularly updated lineup of models on a shared endpoint that you call with an API key.
Speed is the central pitch. Cerebras presents its newest rack-scale system, the CS-4, as delivering up to 30x faster inference than GPUs, a vendor figure rather than an independent measurement.
The company behind the product, Cerebras Systems, raised $5.5 billion in a May 2026 IPO, and independent business reporting lists OpenAI, G42, Mohamed bin Zayed University of Artificial Intelligence and Amazon Web Services among its customers.
Model fidelity is documented in unusual detail. Cerebras states that every Shared Inference model is the original, unpruned version. Weights are stored with selective weight-only quantization in partial 16-, 8- or 4-bit formats, while sensitive layers stay at full precision and activations, attention and the KV cache remain unquantized. For anyone comparing fast AI inference providers, this spells out what is actually being served behind the speed figures.
Prompt caching works automatically on all supported API requests, with no cache breakpoints or code changes. Cerebras guarantees a cache time-to-live of 5 minutes, and entries may persist up to 1 hour depending on system load. The benefit is capacity rather than price: cached tokens are excluded from the uncached rate limit but are not discounted.
Dedicated Inference reserves capacity for a single organization and is the option Cerebras points to for production work that needs predictable performance. It also lets teams deploy fine-tuned weights alongside the standard model versions. Each deployment can be tuned for capacity, draft models, model configuration and quantization, and it adds fine-tuning, weight management and service tier controls on top of every Shared Inference capability. The supported families on the dedicated side include Qwen 3.8, Qwen3 and Qwen3-Coder, Llama 3 and Llama 4, GLM 4.X and 5.X, Kimi K2.X and DeepSeek V3.X.
The Batch API processes groups of requests asynchronously, and batch requests are guaranteed to complete within 24 hours.
Cerebras Code is a coding subscription meant to be used from your own editor. Its product page says it runs GLM 4.7 for code generation at more than 1,000 tokens per second, although the API deprecations log tells a different story.
Cerebras Training Cloud is described as a way to train and fine-tune models from 1 billion to tens of trillions of parameters with the same simple code.
The CS-4 system runs on the WSE-3 Turbo processor, which Cerebras describes as having four trillion transistors and 900,000 AI cores delivering 250 PFLOPS of compute and 43.2 petabytes per second of memory bandwidth. The CS-4 page says first shipments begin this quarter but does not name a date.
Cerebras frames its use cases around workloads where waiting for tokens hurts the product:
The most visible deployment is external. In February 2026, independent technology press reported that OpenAI's GPT-5.3-Codex-Spark, a smaller Codex model designed for faster inference and real-time collaboration, runs on Cerebras' Wafer Scale Engine 3.
Cerebras suits teams for whom response speed changes the product, but the right entry point depends on the stage:
Support also differs by tier: Developer accounts rely on a community Discord, while Enterprise customers get a dedicated account team.
It is a weaker fit for anyone expecting a permanently free tier, very long context windows on the self-serve endpoint, or a wide self-serve model choice; the specifics sit in the pricing, limitation and feature sections.
Cerebras API pricing on the self-serve Developer tier is pay-as-you-go and starts with a free $5 credit, while Enterprise pricing is quoted on request.
| Model (Developer tier) | Input | Output |
|---|---|---|
| GPT OSS 120B | $0.35/M tokens | $0.75/M tokens |
| Qwen 3.8 27B | $0.99/M tokens | $1.49/M tokens |
Cerebras Code Pro is listed at $50 for up to 24 million tokens per day, aimed at indie developers and weekend projects. Cerebras Code Max is listed at $200 for up to 120 million tokens per day, aimed at full-time development and multi-agent systems. Both paid plans were marked sold out as captured on October 4, 2026, and the plan cards checked for this page do not state a billing period.
Training Cloud is sold either per hour, with Cerebras sizing the time a submitted workload needs, or per model, with Cerebras experts designing and fine-tuning a model on your dataset; the Training Cloud page lists these options without rates.
Other vendors also build custom silicon for fast inference instead of renting standard GPUs:
Compared with these, Cerebras stands out for its wafer-scale processor, an OpenAI-compatible API with a $5 self-serve trial, and the option to install CS-4 systems on-premises, while its self-serve catalog is limited to two models.
User reports point the same way. In a public developer discussion of the Qwen 3.8 27B launch on Cerebras, several developers said the 128k context Cerebras allows is too small for long agentic or coding tasks. Others in the same thread said the 150K TPM limit on the public endpoint makes the model hard to use for many coding tasks. An IT-media opinion column from September 2025, written while Cerebras Code ran Qwen3 Coder, reported that generation flew for 10 to 20 seconds before the per-minute token cap triggered 429 errors until the minute reset.
Cerebras says its performance comparisons are based on third-party benchmarking or internal testing, and that observed speed gains over GPU systems may vary by workload, configuration, date and model.
API and Playground access stop, but API keys, projects and settings stay intact, and a pay-as-you-go purchase reactivates access on the Developer tier, which has no hourly or daily token caps.
Cerebras says it serves the original weights for existing model IDs without modification and would release any future pruned variants under separate model IDs.
Yes. The OpenAI-compatible model field takes an exact model ID on Shared Inference and your stable endpoint ID on Dedicated Inference, which routes each request to its active deployment.