Toolso.AI
Toolso.AI
All ToolsCategoriesTrendingLatest ToolsPricingBlog
Toolso.AI
Toolso.AI
Toolso.AI
Toolso.AI

Discover the best AI tools to boost your productivity

GitHubGitHubTwitterX (Twitter)YouTubeYouTubeTikTokEmail

Popular Categories

  • AI Writing
  • AI Image
  • AI Video
  • AI Coding
  • More Categories

Explore

  • Latest Tools
  • Popular Tools
  • More Tools
  • Submit Tool
  • Pricing

About

  • About Us
  • Contact
  • Blog
  • Changelog

Legal

  • Cookie Policy
  • Privacy Policy
  • Terms of Service
  • Refund Policy
© 2026 Toolso.AI All Rights Reserved
Limited timeLimited-time offerFeatured Listing24h priority review · No backlink · 30 days featured$29.90then $59.90Price rises to $59.90 after Oct 31Ends in--:--:--Submit now
  1. Home
  2. All Tools
  3. Voice Speech
  4. Deepgram
Deepgram interface preview
Deepgram logo

Deepgram

Deepgram sells speech-to-text, text-to-speech and Voice Agent APIs billed per minute or per 1,000 characters, with a $200 starter credit, regional endpoints and self-hosting for qualifying enterprise customers.

Voice SpeechDeveloper ToolsAI Development#Text To Speech#Api#Transcription
Try for Free
Saves
Visits
Views
Pricing
Freemium
Published
Oct 6, 2026
Domain
deepgram.com
Community rating

Used this tool? Rate it

Rate this tool

Deepgram Product Information

Try for Free
Tool Information
Saves
Visits
Views
Pricing
Freemium
Published
Oct 6, 2026
Domain
deepgram.com
Community rating

Used this tool? Rate it

Rate this tool

Featured Tools

Related Tools

Try for Free

What is Deepgram?

Deepgram is a voice AI platform delivered as developer APIs: speech-to-text for live and recorded audio, text-to-speech, a Voice Agent API and Audio Intelligence analysis. Deepgram combines speech-to-text, text-to-speech and LLM orchestration in a single Voice Agent API, which the company says reduces complexity, latency and cost.

In January 2026, Deepgram said it had raised $130 million in a Series C round led by AVP at a $1.3 billion valuation. The company also reported that more than 1,300 organizations use its voice AI products and models, including the meeting notetaker Granola, the voice agent startup Vapi and Twilio.

Core features

Deepgram's speech-to-text API accepts streaming audio for real-time transcription or recorded files for batch processing.

Speech-to-text models

  • Flux STT: Flux STT is a conversational speech recognition model for real-time voice agents, with built-in turn detection and interruption handling in 10 languages. Its multilingual option covers English, Spanish, French, German, Hindi, Russian, Portuguese, Japanese, Italian and Dutch.
  • Nova-3: Nova-3 is positioned for production transcription with noise robustness and multilingual support in 50+ languages. The pricing page gives a lower figure, saying Nova models support 45+ languages with speaker diarization, smart formatting, keyterm prompting and automatic language detection.
  • Nova-2: Nova-2 remains recommended for languages not yet supported by Nova-3 and for filler-word identification.
  • Industry-tuned models: Industry-tuned models target vocabulary and structure in domains such as healthcare, legal and finance.
  • Custom models: Custom models can be trained on proprietary or novel datasets for edge-case audio.

Accuracy and speed figures are vendor-reported. Deepgram says Nova-3 has a 54.2% lower word error rate for streaming and 47.4% lower for batch processing than competitors. It also says it delivers transcripts in under 300 milliseconds. Keyterm prompting is described as improving recognition of critical words with up to 90% higher keyword recall.

Text-to-speech

  • Flux TTS: Flux TTS keeps tone, pacing and emotional register across a whole session, with no SSML, style tags or prompt engineering. Deepgram says Flux TTS output starts in as low as 80ms, even under production load. When a caller barges in, the API reports what the caller actually heard.
  • Aura-2: Aura-2 is the widest-language model in the lineup, with voices across seven languages.
  • Aura: Aura is the first-generation Deepgram TTS model and offers English voices only.

Voice Agent API

The Voice Agent API includes barge-in detection, turn-taking prediction, function calling and mid-session control. Teams can bring their own LLM or TTS provider while keeping Deepgram's orchestration and streaming pipeline.

Deepgram provides managed LLMs for the OpenAI, Anthropic, Google and NVIDIA provider types, so no endpoint is needed for those. Groq and Amazon Bedrock models need your own endpoint because Deepgram does not manage those LLMs. Managed models carry a pricing tier: claude-sonnet-4-5 is listed as Advanced and claude-haiku-4-5 as Standard, and the tier sets the per-minute rate.

Audio Intelligence

Audio Intelligence adds summarization, topic detection, intent recognition and sentiment analysis on top of transcripts. Sentiment analysis labels positive, neutral or negative sentiment at word, sentence and transcript level. These features run on lightweight, task-specific models fine-tuned on conversational data rather than on general large language models.

Guide

  1. Create an account: Sign up for a free Deepgram account to get an API key.
  2. Set the model and language: Speech-to-text models default to English unless the language parameter says otherwise. Text-to-speech needs a model on every request, with Flux TTS served on /v2/speak and Aura-2 and Aura on /v1/speak.
  3. Opt out of model training per request: Add mip_opt_out=true as a query parameter for speech-to-text and text-to-speech, or set "mip_opt_out": true in the Voice Agent Settings message. Opted-out requests show mip_opt_out as true under Usage > Logs in the Deepgram Console.
  4. Pick the agent's LLM: The Voice Agent's LLM model is set in that same Settings message.

Deepgram use cases and examples

Deepgram's speech-to-text API is aimed at transcription in customer support, healthcare, media and conversational AI.

  • Contact centers and speech analytics: Speech analytics converts audio into text to analyze conversations and detect intent.
  • Medical transcription: Healthcare transcription targets medical terms and specialized keywords.
  • Media and podcasts: Media transcription covers podcasts, videos and broadcasts with captions and summaries.
  • Support and outbound voice agents: The Deepgram text-to-speech API is positioned for outbound voice agents, with interruption handling and cross-turn context.
  • Restaurant ordering: Deepgram acquired OfOne, a startup that built voice AI ordering for quick-service restaurants.

Who is it for

Deepgram is sold as APIs and SDKs, so the buyer is usually a team writing code.

  • Developers and product teams: The self-serve route is aimed at developers and product teams that want flexible APIs.
  • Platforms and partners: A partner route targets platforms embedding Deepgram voice AI.
  • Regulated enterprises: Custom models and contracts target enterprises with unique workflows and compliance needs.
  • Teams planning to self-host: Self-hosting is built for teams with existing DevOps resources.

It is a weaker fit for non-English narration on Flux TTS or for projects that need a very large voice library.

Platforms

  • Cloud API: The hosted option is a multi-tenant cloud service running on Deepgram infrastructure. API rate limits are published separately for North America, Europe, Australia and India endpoints.
  • Dedicated and VPC: The Voice Agent API can be deployed fully managed, dedicated single-tenant, in a VPC or self-hosted.
  • Self-hosted: Self-hosting is available for Premium customers with unique business requirements. Enterprise self-hosted containers run in a customer VPC or on on-premise hardware and require NVIDIA GPUs. Deepgram can be self-hosted on cloud infrastructure, bare metal or through Amazon SageMaker marketplace listings.
  • SDKs: Deepgram SDKs exist for Java, JavaScript, Python, Go and .NET.
  • Voice agent frameworks: Deepgram TTS integrates with LiveKit, Pipecat, Vapi and other voice agent orchestration frameworks through streaming APIs.
  • Telephony: The Voice Agent API supports 8kHz mu-law telephony audio over bidirectional WebSocket streaming. Deepgram documents Twilio integrations plus integrations for Genesys Cloud CX and Amazon Connect.
  • Playground: Models can be tested in the Playground, and starter apps are published on GitHub.

Pricing

Deepgram pricing is usage-based and split by product and model rather than by seat.

Plans and concurrency

Pay As You Go starts with a free $200 credit and then bills per use; Growth starts at $4K+ per year. Pay As You Go has no minimums, no expiration and needs no credit card to start. Growth saves up to 20% through pre-paid annual credits that are redeemed against actual usage. Enterprise terms are quoted by sales.

ServicePay As You GoGrowth
Speech-to-text (REST / WebSocket / Whisper Cloud)up to 50 / 150 / 5up to 50 / 225 / 5
Text-to-speech (REST + WebSocket)up to 45up to 60
Voice Agent (WebSocket)up to 45up to 60
Audio Intelligence (REST)up to 10up to 10

Speech-to-text rates

Streaming speech-to-text rates are currently labelled as limited-time promotional rates, so regular prices are shown in brackets.

ModelModePay As You GoGrowth
Flux EnglishStreaming$0.0065/min (regular $0.0077)$0.0057/min (regular $0.0065)
Flux MultilingualStreaming$0.0078/min$0.0068/min
Nova-3 MonolingualStreaming$0.0048/min (regular $0.0077)$0.0042/min (regular $0.0065)
Nova-3 MultilingualStreaming$0.0058/min (regular $0.0092)$0.0050/min (regular $0.0078)
Nova-3 MonolingualPre-recorded$0.0043/min$0.0036/min
Nova-3 MultilingualPre-recorded$0.0052/min$0.0043/min
Whisper LargePre-recorded$0.0048/min$0.0048/min

Add-ons are charged on top of the model rate:

  • Redaction costs $0.0020/min on Pay As You Go and $0.0017/min on Growth.
  • Keyterm Prompting costs $0.0013/min on Pay As You Go and $0.0012/min on Growth.
  • Smart Formatting is included on both plans.

Billing is per second, so a 14-second file is charged as exactly 14 seconds. Multichannel audio is billed on total processed duration, so a 10-minute stereo file counts as 20 minutes.

Text-to-speech rates

ModelPay As You GoGrowth
Flux TTS$0.0450 per 1k characters$0.0405 per 1k characters
Aura-2$0.030 per 1k characters$0.027 per 1k characters
Aura-1$0.0150 per 1k characters$0.0135 per 1k characters

Until December 31, 2026, Flux TTS spending is matched with credits up to $500 on Pay As You Go and Growth. A free Flux TTS build period ran through September 12, 2026, and standard pricing has applied since September 13, 2026. The text-to-speech product page FAQ says Deepgram TTS starts at $0.15 per 1,000 characters, which does not match the Aura-1 rate in the table above.

Voice Agent and Audio Intelligence rates

Voice Agent tierPay As You GoGrowth
Standard$0.075/min$0.068/min
Standard with your own TTS$0.065/min$0.051/min
Custom with your own LLM and TTS$0.050/min$0.041/min
Advanced$0.163/min$0.146/min
Advanced with your own TTS$0.122/min$0.110/min

Voice Agent minutes are calculated on WebSocket connection time. Summarization is priced at $0.0003 per 1k input tokens and $0.0006 per 1k output tokens on Pay As You Go.

Credits, overages and refunds

  • Auto-load: Auto-load is on by default and reloads $100 when the remaining credit reaches $10.
  • Refund window: Refunds can be requested for unused credits purchased within the last 30 days. Cancelling otherwise does not entitle you to a refund of unused credits except as stated on the Pricing List at the time.
  • Expiry: Credits expire when the account is closed or the subscription is cancelled or terminated.
  • Growth overages: Growth overages are billed at the Growth rate plus a premium instead of switching to Pay As You Go pricing. The pricing FAQ calls this the 10% overage fee. Overages are charged weekly in arrears to a card on file.
  • Transfers: The pricing FAQ says purchased credits may be transferred between accounts in the same organization on request. The Terms of Service, by contrast, state that credits are not transferable and have no cash value.
  • Startups: Accepted startups can receive up to $100,000 in credits over 12 months through the startup program.

Deepgram alternatives

Most published comparisons in this category come from the vendors themselves.

  • AssemblyAI: Deepgram's own AssemblyAI comparison says both vendors document self-hosting and a BAA pathway and treats a like-for-like test on your own audio as the deciding factor.
  • OpenAI Whisper: Whisper is also available inside Deepgram as Whisper Large for pre-recorded audio. A voice-agent developer published a test of 13 speech-to-text providers on 100 real customer calls. In that test Deepgram Flux had the lowest word error rate at 15.86% and OpenAI Whisper the highest at 39.78%. Postcode recognition exceeded 50% word error rate for every provider, Deepgram included. It is a single team's test on its own call data, not a general ranking.
  • ElevenLabs: Deepgram positions its TTS for real-time voice agents and concedes that ElevenLabs offers a larger voice library and broader language coverage, suited to content creation.
  • Cartesia: Deepgram describes Cartesia as known for fast synthesis, expressive control through text tags and a broad multilingual voice library.

Limitations

Product and usage limits

  • Voice languages: Flux TTS is English-only; Spanish, German, French, Dutch, Italian and Japanese voices require Aura-2.
  • Whisper in the EU: Whisper models are not supported on the EU endpoint.
  • Whisper concurrency: The API rate-limit documentation lists Whisper Cloud at up to 3 concurrent pre-recorded requests in North America and not available in other regions, while the pricing page shows up to 5.
  • Projects: Rate limits apply per project, and secondary projects on self-serve accounts are limited to one concurrent stream.
  • Over the limit: On Pay As You Go, requests over the concurrency limit may be queued or rejected.
  • Faulty uploads: Customers pay for every processed file, including faulty uploads caused by their own systems.

Usage rules in the Terms

  • The Terms prohibit using the services or output for competitive purposes, including model training, benchmarking and competitive analysis.
  • Users must not present audio output as human-made and must give legally required notices that it is AI-generated.
  • Output may not be unique, and Deepgram makes no representation that users own similar output created for other customers.
  • Protected health information may not be sent unless a business associate agreement has been signed.

Data use and privacy

API requests take part in the Model Improvement Program by default, and Deepgram retains them. The Terms allow Deepgram to use customer input and output to improve its services, including training and testing its models.

  • Opt-out: Setting mip_opt_out=true on a request excludes it from retention, and on a regional endpoint keeps it in-region apart from third-party Voice Agent providers. Opted-out requests get zero data retention for audio, text, transcripts and synthesized audio once the response is returned. Account-wide or deployment-wide opt-out enforcement is available by contacting Deepgram.
  • Logs: Request metadata and usage logs remain retrievable for 90 days, with summarized usage kept longer.
  • Residency: A full in-region guarantee requires both a regional endpoint and mip_opt_out=true. EU endpoint requests are never routed outside the EU and fail rather than fall back if the region is unavailable.
  • Ownership and sharing: Deepgram does not claim ownership of customer input and output. Deepgram says it does not sell customer data, use it for advertising profiles or redistribute it without permission.
  • Policy scope: The website privacy notice does not apply to the processing of customer data.
  • Self-hosted data flow: In a typical self-hosted deployment, no audio, transcripts or request-identifying content is sent to Deepgram.
  • Certifications: Deepgram states it is SOC 2 Type 1 and Type 2 certified. Deepgram signs business associate agreements for Enterprise customers handling ePHI.

FAQ

Q1. Is there a free way to try Deepgram?

Yes. New Pay As You Go accounts start with a $200 credit and no credit card requirement; after that, usage is billed per minute, per 1,000 characters or per minute of agent connection time.

Q2. Does Deepgram use my audio to train its models?

By default, yes: requests join the Model Improvement Program and are retained. Adding mip_opt_out=true to a request excludes it and stops retention of its content.

Q3. Can processing stay inside the EU?

Deepgram offers a dedicated EU endpoint whose requests are not routed outside the EU. Full in-region processing also requires opting out of model improvement, and Whisper models are not available there.

Know a Similar Tool?
If you know other great AI tools, feel free to submit them to us