Cartesia is a voice AI company that sells real-time speech models and a hosted voice-agent platform to developers, and it calls itself an AI lab building real-time voice models and agents. Cartesia pitches one API for speech generation, transcription and production voice agents, fast enough for live conversation.
The product line has three layers. Cartesia Sonic is the text-to-speech model, Ink is the speech-to-text model, and Managed Agents combines both with a large language model to run phone and in-app conversations. Voice cloning sits on top of Sonic, and a free browser voice changer is offered next to the paid API.
Its founding team met as PhDs at Stanford AI Lab, where they say they invented State Space Models (SSMs), the architecture Cartesia uses in place of standard transformers. The company credits SSMs for its speed and cost, which is its own explanation rather than a measured result.
Cartesia offers two AI voice cloning tiers. An Instant Voice Clone starts from 10 seconds of audio, Sonic 3.6 and newer can use up to 60 seconds to keep an accent, and uploaded files are capped at 16 MB. A Pro Voice Clone fine-tunes the model on at least 30 minutes of one speaker, and training takes up to 3 hours. Each Pro clone is trained on one language at a time, so a multilingual brand voice needs a separately recorded dataset for every language.
Managed Agents join Ink-2, an LLM the customer picks and Sonic-3.6 into one realtime voice agent that handles turn-taking, interruptions, audio streaming and tool calls. Agents can use three kinds of tools: built-in system tools such as ending a call, webhook tools that send requests to the customer's server, and client tools that run in the connected browser, mobile app or server. Cartesia runs the LLM integration itself, so no separate model-provider account or API key is needed. A built-in evaluation framework adds live testing, system metrics and custom call analytics.
Through the API, a Pro Voice Clone takes four steps: create a dataset, upload audio files to it, create a fine-tune from that dataset, then list the voices the fine-tune produced.
Cartesia's examples center on phone agents and cloned voices for media.
For scale, Cartesia cites Bolna, which it says handles 500,000+ daily voice calls across Indian languages on its low-latency infrastructure; this is a vendor-published customer story rather than an independent audit.
Cartesia fits teams that put voice inside their own software, not people looking for a finished consumer app.
It is a weaker fit for teams that need streaming transcription outside Ink's five languages, for anyone who expects voice cloning on the Free plan, and for buyers who need refundable subscriptions.
Cartesia bills voice models in credits and agent calls in dollars, and every plan includes unlimited workspace seats. Monthly plans as captured on October 6, 2026:
| Plan | Price per month | Credits per month | Prepaid agent dollars | What the plan adds |
|---|---|---|---|---|
| Free | $0 | 20K | $1 | Text to speech and speech to text access |
| Pro | $5 | 100K | $5 | Commercial use license and instant voice cloning |
| Startup | $49 | 1.25M | $49 | Everything in Pro plus Pro voice cloning |
| Scale | $299 | 8M | $299 | Priority support and high concurrency limits |
| Enterprise | Custom | Custom | Custom | Volume pricing, custom concurrency, DPAs and BAAs, SSO |
On Free, Pro, Startup and Scale, Sonic-3.6 allowances work out to about 27, 133, 1,667 and 10,667 minutes a month, with 2, 3, 5 and 15 concurrent text-to-speech requests. Ink-2 allowances come to roughly 1h 51m, 9h 16m, 115h 44m and 740h 44m of transcription, with 8, 12, 20 and 60 concurrent requests.
Cartesia's site says its models are ranked #1 on a third-party speech arena leaderboard and speech-to-text leaderboard. That independent arena's provider-voice text-to-speech leaderboard, based on blind listener votes and captured on October 6, 2026, placed Eleven v4 Turbo (Elo 1334), Eleven v4 (1321) and Qwen-Audio-3.1-TTS-Plus (1292) ahead of Sonic 3.6 (1278), with Gemini 3.8 Flash TTS (1275) just behind. The two claims did not match on that date, and this comparison covers only the provider-voice view of the arena.
For transcription, Cartesia claims Ink-2 outperforms Deepgram Flux, Soniox RT-V4, AssemblyAI RT Pro and ElevenLabs Scribe-2-realtime across line recordings, accented speech, noisy audio and earnings calls. The benchmark is the vendor's own, but it names the streaming models Cartesia treats as direct competitors.
Yes, in part. The AI voice changer is free to try with no sign-up; free conversions have a daily limit, and uploads must be WAV or MP3 files under 15 MB.
No. Only successful requests consume credits, and errors are not charged.
Yes, with an extra step. Each instant clone is built from one language, so the clone is made in its original language first and other languages are added through the Add Voice Accents API. Each added voice accent costs 225 credits.
The Bytes endpoint can return raw, WAV or MP3 audio, while the SSE and WebSocket endpoints return raw audio only. Supported sample rates run from 8000 to 48000 Hz.
Cartesia states that its compliance posture includes SOC 2 Type II, HIPAA eligibility with BAAs available for healthcare customers, and GDPR compliance, with zero data retention available to enterprise customers on eligible services.