Toolso.AI
Toolso.AI
All ToolsCategoriesTrendingLatest ToolsPricingBlog
Toolso.AI
Toolso.AI
Toolso.AI
Toolso.AI

Discover the best AI tools to boost your productivity

GitHubGitHubTwitterX (Twitter)YouTubeYouTubeTikTokEmail

Popular Categories

  • AI Writing
  • AI Image
  • AI Video
  • AI Coding
  • More Categories

Explore

  • Latest Tools
  • Popular Tools
  • More Tools
  • Submit Tool
  • Pricing

About

  • About Us
  • Contact
  • Blog
  • Changelog

Legal

  • Cookie Policy
  • Privacy Policy
  • Terms of Service
  • Refund Policy
© 2026 Toolso.AI All Rights Reserved
Limited timeLimited-time offerFeatured Listing24h priority review · No backlink · 30 days featured$29.90then $59.90Price rises to $59.90 after Oct 31Ends in--:--:--Submit now
  1. Home
  2. All Tools
  3. Voice Speech
  4. Cartesia
Cartesia interface preview
Cartesia logo

Cartesia

Cartesia provides Sonic text-to-speech in 44 languages, Ink streaming speech-to-text, voice cloning and hosted voice agents through one API, built for developers who need speech fast enough for live calls.

Voice SpeechDeveloper ToolsVoice Generation#Text To Speech#Voice Cloning#Real Time
View Pricing
Saves
Visits
Views
Pricing
Paid
Published
Oct 6, 2026
Domain
cartesia.ai
Community rating

Used this tool? Rate it

Rate this tool

Cartesia Product Information

View Pricing
Tool Information
Saves
Visits
Views
Pricing
Paid
Published
Oct 6, 2026
Domain
cartesia.ai
Community rating

Used this tool? Rate it

Rate this tool

Featured Tools

Related Tools

View Pricing

What is Cartesia?

Cartesia is a voice AI company that sells real-time speech models and a hosted voice-agent platform to developers, and it calls itself an AI lab building real-time voice models and agents. Cartesia pitches one API for speech generation, transcription and production voice agents, fast enough for live conversation.

The product line has three layers. Cartesia Sonic is the text-to-speech model, Ink is the speech-to-text model, and Managed Agents combines both with a large language model to run phone and in-app conversations. Voice cloning sits on top of Sonic, and a free browser voice changer is offered next to the paid API.

Its founding team met as PhDs at Stanford AI Lab, where they say they invented State Space Models (SSMs), the architecture Cartesia uses in place of standard transformers. The company credits SSMs for its speed and cost, which is its own explanation rather than a measured result.

Core features

Sonic text-to-speech

  • 44 languages: Sonic 3.6 is described as Cartesia's fastest and most natural text-to-speech model, with native support for 44 languages.
  • Latency: Cartesia says Sonic delivers sub-90ms model latency, and streaming lets playback begin before the full response has been generated.
  • Expressive delivery: By default Sonic reads the emotional subtext of a transcript and adjusts its delivery, and non-verbal sounds such as laughter can be written straight into the transcript.
  • Pronunciation control: Custom pronunciation dictionaries set how proper nouns and domain terms are spoken.
  • Version pinning: The sonic-3.6 model ID always points to the latest stable snapshot, a dated snapshot never changes once released, and sonic-preview is a beta model that is not meant for production.

Ink speech-to-text

  • Turn detection built in: Ink 2 needs no separate voice-activity detector, because it emits turn events (start, update, eager end, resume and end) that tell an agent when to listen and when to reply.
  • Structured data: Cartesia says Ink has the lowest word error rate of any streaming speech-to-text model and natively handles phone numbers, dates, emails, currencies and UUIDs.

Voice cloning

Cartesia offers two AI voice cloning tiers. An Instant Voice Clone starts from 10 seconds of audio, Sonic 3.6 and newer can use up to 60 seconds to keep an accent, and uploaded files are capped at 16 MB. A Pro Voice Clone fine-tunes the model on at least 30 minutes of one speaker, and training takes up to 3 hours. Each Pro clone is trained on one language at a time, so a multilingual brand voice needs a separately recorded dataset for every language.

Managed Agents

Managed Agents join Ink-2, an LLM the customer picks and Sonic-3.6 into one realtime voice agent that handles turn-taking, interruptions, audio streaming and tool calls. Agents can use three kinds of tools: built-in system tools such as ending a call, webhook tools that send requests to the customer's server, and client tools that run in the connected browser, mobile app or server. Cartesia runs the LLM integration itself, so no separate model-provider account or API key is needed. A built-in evaluation framework adds live testing, system metrics and custom call analytics.

Guide

Create and use an instant voice clone

  1. Pick a speaker whose voice you own or have explicit permission to clone.
  2. Record one speaker in a quiet room with no music, echo or background noise, in the language and style the clone should use; the mood of the clip carries into an instant clone's voice.
  3. Open Instant Voice Clone in the Playground, record or upload the clip and choose the language spoken, or call the cloning API when cloning is part of an application.
  4. Use the clone's voice ID in text-to-speech requests, whether the voice was made in the Playground or through the API.

Train a Pro Voice Clone through the API

Through the API, a Pro Voice Clone takes four steps: create a dataset, upload audio files to it, create a fine-tune from that dataset, then list the voices the fine-tune produced.

Cartesia use cases and examples

Cartesia's examples center on phone agents and cloned voices for media.

  • Customer support: An agent authenticates callers, pulls live account data and resolves billing questions, order status and account issues without hold times or transfers.
  • Recruiting: An agent calls applicants right away, screens them and pushes a qualification summary into the applicant tracking system before the call ends.
  • Fraud detection: Agents place real-time outbound verification calls on suspicious transactions and handle step-up authentication.
  • Games: Studios can build character dialogue from a voice actor's clone and read new lines from text as the story changes.
  • Dubbing: A custom voice reads translated scripts in supported languages, with pronunciation and timing reviewed before the audio goes into the video.

For scale, Cartesia cites Bolna, which it says handles 500,000+ daily voice calls across Indian languages on its low-latency infrastructure; this is a vendor-published customer story rather than an independent audit.

Who is it for

Cartesia fits teams that put voice inside their own software, not people looking for a finished consumer app.

  • Early-stage companies: A grant program accepts startups at any stage and considers students and one-off projects case by case. Accepted startups get 12 months on Scale with 8 million shared Sonic and Ink credits each month plus $299 a month in Managed Agents credits. In return, participants share their logo and let Cartesia use their name and logo in marketing.
  • Regulated, high-volume and public-sector buyers: Cartesia sends teams to sales when they run more than 50M credits, need on-prem, VPC or OEM deployment, need a BAA or zero data retention, or buy for government.

It is a weaker fit for teams that need streaming transcription outside Ink's five languages, for anyone who expects voice cloning on the Free plan, and for buyers who need refundable subscriptions.

Platforms

  • Browser Playground: Every model can be tried in the browser, which is also where API keys are created.
  • API endpoints: The Sonic text-to-speech API is served over three endpoints, /tts/bytes, /tts/sse and /tts/websocket.
  • SDKs: Official SDKs cover the TTS, STT and voice APIs in JavaScript/TypeScript and Python.
  • Coding agents: Cartesia MCP works in Cursor, Claude Code and other MCP clients to list voices, run TTS and STT and manage pronunciation dictionaries without custom scripts.
  • Agent frameworks: A LiveKit integration offers realtime rooms and agents through the Cartesia plugin or LiveKit Inference. Pipecat ships official Cartesia TTS and STT services for voice and multi-modal agents.
  • Telephony: Agents connect to numbers provisioned by Cartesia, numbers imported from Twilio, or numbers routed through a SIP carrier.
  • Deployment: Cloud traffic runs through regional API endpoints for in-region processing, and the models can also run in a customer's data center or on custom hardware. Under enterprise contracts, Sonic-3.6 can be deployed on-prem (including air-gapped sites), in a customer VPC on AWS, GCP or Azure, or embedded through OEM licensing.
  • Marketplace: Cartesia also sells one subscription covering Sonic, Ink and voice agents through AWS Marketplace.

Pricing

Cartesia bills voice models in credits and agent calls in dollars, and every plan includes unlimited workspace seats. Monthly plans as captured on October 6, 2026:

PlanPrice per monthCredits per monthPrepaid agent dollarsWhat the plan adds
Free$020K$1Text to speech and speech to text access
Pro$5100K$5Commercial use license and instant voice cloning
Startup$491.25M$49Everything in Pro plus Pro voice cloning
Scale$2998M$299Priority support and high concurrency limits
EnterpriseCustomCustomCustomVolume pricing, custom concurrency, DPAs and BAAs, SSO

On Free, Pro, Startup and Scale, Sonic-3.6 allowances work out to about 27, 133, 1,667 and 10,667 minutes a month, with 2, 3, 5 and 15 concurrent text-to-speech requests. Ink-2 allowances come to roughly 1h 51m, 9h 16m, 115h 44m and 740h 44m of transcription, with 8, 12, 20 and 60 concurrent requests.

  • Text-to-speech metering: One minute of Sonic audio uses 750-800 credits, because 1 credit equals 1 character.
  • Speech-to-text metering: Ink-2 costs 3 credits per second of audio. Rates depend on the model and on batch versus realtime endpoints, and silence is billed even when no transcript comes back.
  • Voice agents: Calls cost $0.06 per minute, plus $0.014 per minute on a Cartesia-provided phone number. The prepaid agent dollars cover about 17 minutes on Free, 1 hour 23 minutes on Pro, 13 hours 37 minutes on Startup and 83 hours 3 minutes on Scale, before managed phone number costs.
  • Overages: Paid users can switch on overages at $65 per 1M credits on Pro, $45 on Startup and $38 on Scale; with overages off, new requests are blocked once credits run out.
  • Rollover: Unused credits roll over, up to 2x the monthly plan rate.
  • Pro Voice Clone slots: Training a Pro clone is free and its speech costs 1 credit per character; Startup includes 2 slots, Scale 4, and Free and Pro include none.
  • Renewal and cancellation: Subscriptions renew monthly from the purchase date, and after a mid-period cancel or downgrade the account keeps its tier until the period ends.
  • Refunds: Except where the Terms say otherwise, subscription payments are nonrefundable and partially used periods earn no credit.

Cartesia alternatives for speech synthesis and transcription

Cartesia's site says its models are ranked #1 on a third-party speech arena leaderboard and speech-to-text leaderboard. That independent arena's provider-voice text-to-speech leaderboard, based on blind listener votes and captured on October 6, 2026, placed Eleven v4 Turbo (Elo 1334), Eleven v4 (1321) and Qwen-Audio-3.1-TTS-Plus (1292) ahead of Sonic 3.6 (1278), with Gemini 3.8 Flash TTS (1275) just behind. The two claims did not match on that date, and this comparison covers only the provider-voice view of the arena.

For transcription, Cartesia claims Ink-2 outperforms Deepgram Flux, Soniox RT-V4, AssemblyAI RT Pro and ElevenLabs Scribe-2-realtime across line recordings, accented speech, noisy audio and earnings calls. The benchmark is the vendor's own, but it names the streaming models Cartesia treats as direct competitors.

Limitations

  • Conflicting cloning access: The voice-cloning page says cloning requires a paid plan, with Instant Voice Cloning from Pro and Pro Voice Cloning from Startup. The Sonic FAQ instead describes the Free tier as for evaluation and personal use, with Instant Voice Cloning included. The same FAQ puts Professional Voice Cloning at 15–30 minutes of reference audio, while the cloning documentation sets 30 minutes as the minimum.
  • Fixed speed and volume on Pro clones: A Pro Voice Clone learns pacing and loudness from its dataset, so request-time speed and volume controls have no effect on it.
  • Five transcription languages: Ink-2 supports English, Spanish, French, Hindi and Japanese, far fewer than Sonic's 44.
  • Unclear LLM charges for agents: The agents LLM documentation states free LLM usage until October 1, 2026, a date before this capture. The pricing page still lists LLM usage during calls as free for UI-created agents for a limited time. Per-model token prices, providers and average latency are returned by the agents models endpoint.
  • US-only provisioned numbers: Cartesia provisions US phone numbers only; numbers in other countries come through Twilio or a SIP carrier.
  • Calling compliance stays with the customer: Customers are solely responsible for making every call, text or prerecorded message comply with applicable law, including the Telephone Consumer Protection Act and Do-Not-Call rules.
  • Cloning safeguards, then and now: A December 2024 independent IT media report found few apparent safeguards on the Sonic voice cloner, which only required a checkbox agreeing to the terms of service; Cartesia said it had automated and manual review and was working on voice verification and watermarking. Cartesia now states that all voice clones require verified consent from the speaker and that cloning public figures or others without consent is prohibited. None of the cloning pages captured on October 6, 2026 describes how that consent is verified.
  • Thin review base: A cloud marketplace listing showed 4.3 out of 5 from 41 ratings as captured on October 6, 2026, all imported from an outside business-software review site, with no reviews left on the marketplace itself.

Privacy, data use and voice rights

  • Training on customer content: Unless otherwise agreed, inputs, outputs and user interactions may be used by Cartesia to train and improve its models. An online form lets users exclude certain categories of their content, but the opt-out only applies going forward and does not undo earlier use.
  • Zero data retention: Zero Data Retention is available only to Enterprise customers for the TTS and STT APIs, and it excludes voice cloning, Pro Voice Cloning and voice creation.
  • Disclosure duties for apps: An app built on Cartesia must tell end users that they are talking to AI rather than a human and that their conversations are recorded and may be shared with Cartesia and its third-party LLM providers.
  • Commercial use: Outputs may not be used commercially unless the subscription tier expressly allows it; in the plan table, the commercial use license starts at Pro.
  • Voice and impersonation rules: The Acceptable Use Policy bans impersonating any person, celebrities included, or any business, and allows only your own voice or other people's voices with explicit consent. The Terms also bar posting the voice of a deceased person or a political candidate without express permission.
  • Minors and politics: Outputs may not relate to or represent anyone under 18. The policy also rules out legitimate political advertising, campaigning, lobbying and fundraising.

FAQ

Q1. Can I try Cartesia without creating an account?

Yes, in part. The AI voice changer is free to try with no sign-up; free conversions have a daily limit, and uploads must be WAV or MP3 files under 15 MB.

Q2. Do failed API requests use credits?

No. Only successful requests consume credits, and errors are not charged.

Q3. Can a cloned voice speak other languages?

Yes, with an extra step. Each instant clone is built from one language, so the clone is made in its original language first and other languages are added through the Add Voice Accents API. Each added voice accent costs 225 credits.

Q4. Which audio formats does Sonic return?

The Bytes endpoint can return raw, WAV or MP3 audio, while the SSE and WebSocket endpoints return raw audio only. Supported sample rates run from 8000 to 48000 Hz.

Q5. Is Cartesia SOC 2 or HIPAA compliant?

Cartesia states that its compliance posture includes SOC 2 Type II, HIPAA eligibility with BAAs available for healthcare customers, and GDPR compliance, with zero data retention available to enterprise customers on eligible services.

Know a Similar Tool?
If you know other great AI tools, feel free to submit them to us