Home » AI Tools » Generative AI » Deepgram

Deepgram

Deepgram is a speech generation tool with an attached text intelligence feature for sentiment analysis.

UpdatedAugust 24, 2026

Starting Price

Pay Per Credit

TRIAL

$200 free credits

TL;DR

  • Nova-3 delivers transcripts in under 300ms with 99.85% accuracy on my West African accent in my test.
  • Failed to detect sarcasm in my customer service email test, misinterpreting a passive-aggressive complaint as positive sentiment.
  • Offers 7+ accent options, including Irish, Singaporean, and Indian voices, each with distinct personality traits.

Best For

  • Call center teams analyzing 1,000+ hours/month of recorded support calls
  • Developers building real-time voice agents with sub-300ms latency requirements
  • Data engineers processing high-volume audio archives with batch transcription pipelines
  • Medical transcriptionists documenting patient consultations in noisy clinic environments

Alternatives

  • AssemblyAI
  • Rev.ai
  • Gladia

    Pricing

    Deepgram pricing
    • Pay-As-You-Go: New accounts get a $200 free credit before moving to standard per-minute rates. Deepgram’s flagship model, Nova-3, costs roughly $0.0077 per minute on Pay-As-You-Go, which works out to about $0.46 per hour. Deepgram bills by the second rather than rounding up to the nearest minute, which saves money on short audio clips processed at scale.
    • Growth: It requires a minimum annual prepayment of $4,000 or more, in exchange for discounted usage rates. At that tier, Nova-3 drops to around $0.0065 per minute, a savings of roughly 16% compared to Pay-As-You-Go. This tier pays off once your product moves past the testing phase and starts generating consistent monthly usage.
    • Enterprise pricing: This isn’t published. Instead, it’s negotiated directly with Deepgram’s sales team, and it typically includes custom model training, dedicated support, and on-premises deployment options.

    One limitation I noted is that beyond the base transcription rate, Deepgram charges separately for add-ons. Summarization, for example, is billed per 1,000 input and output tokens, at roughly $0.0003 per 1,000 input tokens on Pay-As-You-Go. Topic detection and sentiment analysis also carry their own charges. 

    So, if your product uses several of these Audio Intelligence features together, your real cost per audio hour can climb well above the headline transcription rate. Based on my testing with a typical podcast workflow using summarization and sentiment analysis, the real cost climbed to $1.20 per hour once add-ons were included, more than twice the headline rate.

    For comparison, AssemblyAI’s starting rate is much lower; its Universal-2 model starts at $0.15 per hour, versus Deepgram Nova-3’s $0.46 per hour on Pay-As-You-Go. That’s roughly a three-times difference at base rates. Therefore, lower-budget teams should run the numbers for their specific use case before choosing a provider.

Overview

Deepgram launched in 2015. Since then, it has grown into one of the more recognizable names in speech AI. The company builds its own deep learning models rather than relying on third-party engines, which gives it tighter control over performance. 

As a result, Deepgram can tune its models for specific use cases, such as noisy call center audio or medical terminology. Today, Deepgram counts NASA, Spotify, Citi, and Twilio among its customers, and the company says it has processed over 50,000 years of audio and transcribed more than 1 trillion words. 

Notably, NASA uses Deepgram to transcribe communications between the ISS and Mission Control, plus low-quality audio from underwater astronaut training exercises.

Deepgram’s flagship model is Nova-3. It’s built specifically for speed and cost efficiency, and it supports real-time multilingual transcription. In fact, independent testing gives Deepgram high marks here.

On Product Hunt and G2, Deepgram is rated 4.6 and 4.9 out of 5. In my testing, Deepgram proved to be a strong tool for fast, real-time transcription, though not without some sharp edges I’ll cover below.

Also read: The Future of Voice Synthesis with AI

Key Features

1. Speech-to-text

Deepgram is a standard transcription tool that records voices and returns text transcripts in real-time. I hit the record button, and the live transcript appeared almost instantly. In my 15-minute test recording, Deepgram missed 3 words out of approximately 2,000, giving it roughly 99.85% accuracy on my West African accent. But to be fair, I did stumble over some of my words.

Deepgram transcription

2. Text-to-speech

Deepgram has a list of voices to choose from: American, Irish, British, Indian, Filipino, Australian, and Singaporean. That’s a wider accent range than most competitors offer, which matters if your product serves a global audience. The diversity of accents means that Deepgram can be used for global audiences to boost content appeal. There are other languages as well, mostly European languages and then Japanese.

In my testing, I noticed that each voice has a distinct personality that carries over to the speech. I typed in a script of a manager questioning an employee and found voice ‘Colin’ more suitable than voice ‘Kit.’ Colin had that authoritarian edge, while Kit sounded like a friend making a casual mention.

Using Deepgram's text to speech feature

3. Voice Agent

Deepgram’s voice agent is essentially a front-end interface with options for use cases to choose from. I went with the customer support representative and gave a hypothetical complaint. In my test, replies landed in well under 300 milliseconds, so it was easy to have a natural-flowing conversation. One limitation I noticed was the unmistakable robotic tone. Anyone who would rather speak to a human agent would be turned off by this.

Voice agent

4. Text Intelligence

This is a natural language feature that analyzes text to get meaning, sentiment, intent, and insights. Although Deepgram is primarily a voice tool, the text intelligence exists so developers can perform high-level content analysis.

In my testing, I pasted in an email from a customer upset at their purchase. The email had a passive-aggressive tone to it, laden with sarcasm, and I was hoping that Deepgram would detect it for what it was. It didn’t. 

The email contained sentences like “I particularly enjoyed the thrilling suspense of pressing the power button four times and receiving nothing in return except silence and a cold, empty cup.” This was taken literally and misunderstood as a happy purchase.

It seems that Deepgram is unable to make out emotional subtexts and subtleties, especially when they mean the opposite. Aside from its inability to detect sarcasm as the medium used by the customer to express their displeasure, Deepgram was able to give an accurate summary of the email.

Deepgram text intelligence

Don’t rely on Deepgram for decoding sentiment or tone in customer-facing workflows, it missed sarcasm entirely in my test.

The Bottom Line

I’d pick Deepgram for real-time voice agents where latency matters more than sentiment accuracy, but I wouldn’t trust it for customer service email analysis based on what I saw. However, anyone who prioritizes text intelligence, with delicate operations such as customer service, should explore another tool.

    FAQ

    1) Does Deepgram Work With Medical Terminology?
    Yes. Deepgram offers domain-tuned models, including options built for medical audio, that improve recognition of clinical terms, drug names, and jargon that general-purpose models often miss. For the highest accuracy, pair a medical-tuned model with custom vocabulary so it recognizes practice-specific terms too.

    2) Is Deepgram Better Than Whisper?
    It depends on your priority. Deepgram is better for faster transcription. Batch transcription runs 30 to 90 times faster than Whisper, and Deepgram's streaming latency stays under 300 milliseconds. Whisper, meanwhile, has better raw accuracy, especially for messy or accented audio, and it remains fully open source.

    3) Is Deepgram AI Free?
    Not exactly, but new users get a meaningful head start. Deepgram provides a $200 free credit for new accounts, which lets you test the API before paying anything. After that credit runs out, though, you'll move to standard Pay-As-You-Go pricing.

    4) Who Uses Deepgram?
    A wide range of companies do. Deepgram's customer list includes Citi, Vapi, Groq, Twilio, and Spotify, along with NASA, which uses the platform to transcribe space-to-ground communications. In total, the company serves over 500 customers, ranging from early-stage startups to large enterprises.

    5) What Is Deepgram Used For?
    Most commonly, teams use Deepgram to power voice agents, transcribe meetings, analyze call center conversations, and caption media at scale. It also supports niche use cases, such as NASA's audio search through historical mission recordings.

    6) Which Is Better, AssemblyAI or Deepgram?
    It depends on what you need. AssemblyAI has better pricing and analytics; it has faster latency and a lower word error rate than Deepgram across 4 million production calls. Deepgram, however, is still the better pick for real-time transcription and deployment flexibility.