Deepgram is a voice AI platform delivered as developer APIs: speech-to-text for live and recorded audio, text-to-speech, a Voice Agent API and Audio Intelligence analysis. Deepgram combines speech-to-text, text-to-speech and LLM orchestration in a single Voice Agent API, which the company says reduces complexity, latency and cost.
In January 2026, Deepgram said it had raised $130 million in a Series C round led by AVP at a $1.3 billion valuation. The company also reported that more than 1,300 organizations use its voice AI products and models, including the meeting notetaker Granola, the voice agent startup Vapi and Twilio.
Deepgram's speech-to-text API accepts streaming audio for real-time transcription or recorded files for batch processing.
Accuracy and speed figures are vendor-reported. Deepgram says Nova-3 has a 54.2% lower word error rate for streaming and 47.4% lower for batch processing than competitors. It also says it delivers transcripts in under 300 milliseconds. Keyterm prompting is described as improving recognition of critical words with up to 90% higher keyword recall.
The Voice Agent API includes barge-in detection, turn-taking prediction, function calling and mid-session control. Teams can bring their own LLM or TTS provider while keeping Deepgram's orchestration and streaming pipeline.
Deepgram provides managed LLMs for the OpenAI, Anthropic, Google and NVIDIA provider types, so no endpoint is needed for those. Groq and Amazon Bedrock models need your own endpoint because Deepgram does not manage those LLMs. Managed models carry a pricing tier: claude-sonnet-4-5 is listed as Advanced and claude-haiku-4-5 as Standard, and the tier sets the per-minute rate.
Audio Intelligence adds summarization, topic detection, intent recognition and sentiment analysis on top of transcripts. Sentiment analysis labels positive, neutral or negative sentiment at word, sentence and transcript level. These features run on lightweight, task-specific models fine-tuned on conversational data rather than on general large language models.
Deepgram's speech-to-text API is aimed at transcription in customer support, healthcare, media and conversational AI.
Deepgram is sold as APIs and SDKs, so the buyer is usually a team writing code.
It is a weaker fit for non-English narration on Flux TTS or for projects that need a very large voice library.
Deepgram pricing is usage-based and split by product and model rather than by seat.
Pay As You Go starts with a free $200 credit and then bills per use; Growth starts at $4K+ per year. Pay As You Go has no minimums, no expiration and needs no credit card to start. Growth saves up to 20% through pre-paid annual credits that are redeemed against actual usage. Enterprise terms are quoted by sales.
| Service | Pay As You Go | Growth |
|---|---|---|
| Speech-to-text (REST / WebSocket / Whisper Cloud) | up to 50 / 150 / 5 | up to 50 / 225 / 5 |
| Text-to-speech (REST + WebSocket) | up to 45 | up to 60 |
| Voice Agent (WebSocket) | up to 45 | up to 60 |
| Audio Intelligence (REST) | up to 10 | up to 10 |
Streaming speech-to-text rates are currently labelled as limited-time promotional rates, so regular prices are shown in brackets.
| Model | Mode | Pay As You Go | Growth |
|---|---|---|---|
| Flux English | Streaming | $0.0065/min (regular $0.0077) | $0.0057/min (regular $0.0065) |
| Flux Multilingual | Streaming | $0.0078/min | $0.0068/min |
| Nova-3 Monolingual | Streaming | $0.0048/min (regular $0.0077) | $0.0042/min (regular $0.0065) |
| Nova-3 Multilingual | Streaming | $0.0058/min (regular $0.0092) | $0.0050/min (regular $0.0078) |
| Nova-3 Monolingual | Pre-recorded | $0.0043/min | $0.0036/min |
| Nova-3 Multilingual | Pre-recorded | $0.0052/min | $0.0043/min |
| Whisper Large | Pre-recorded | $0.0048/min | $0.0048/min |
Add-ons are charged on top of the model rate:
Billing is per second, so a 14-second file is charged as exactly 14 seconds. Multichannel audio is billed on total processed duration, so a 10-minute stereo file counts as 20 minutes.
| Model | Pay As You Go | Growth |
|---|---|---|
| Flux TTS | $0.0450 per 1k characters | $0.0405 per 1k characters |
| Aura-2 | $0.030 per 1k characters | $0.027 per 1k characters |
| Aura-1 | $0.0150 per 1k characters | $0.0135 per 1k characters |
Until December 31, 2026, Flux TTS spending is matched with credits up to $500 on Pay As You Go and Growth. A free Flux TTS build period ran through September 12, 2026, and standard pricing has applied since September 13, 2026. The text-to-speech product page FAQ says Deepgram TTS starts at $0.15 per 1,000 characters, which does not match the Aura-1 rate in the table above.
| Voice Agent tier | Pay As You Go | Growth |
|---|---|---|
| Standard | $0.075/min | $0.068/min |
| Standard with your own TTS | $0.065/min | $0.051/min |
| Custom with your own LLM and TTS | $0.050/min | $0.041/min |
| Advanced | $0.163/min | $0.146/min |
| Advanced with your own TTS | $0.122/min | $0.110/min |
Voice Agent minutes are calculated on WebSocket connection time. Summarization is priced at $0.0003 per 1k input tokens and $0.0006 per 1k output tokens on Pay As You Go.
Most published comparisons in this category come from the vendors themselves.
API requests take part in the Model Improvement Program by default, and Deepgram retains them. The Terms allow Deepgram to use customer input and output to improve its services, including training and testing its models.
Yes. New Pay As You Go accounts start with a $200 credit and no credit card requirement; after that, usage is billed per minute, per 1,000 characters or per minute of agent connection time.
By default, yes: requests join the Model Improvement Program and are retained. Adding mip_opt_out=true to a request excludes it and stops retention of its content.
Deepgram offers a dedicated EU endpoint whose requests are not routed outside the EU. Full in-region processing also requires opting out of model improvement, and Whisper models are not available there.