Voice AI infrastructure, including speech recognition, speech generation, and voice-agent APIs
TL;DR. Deepgram is a San Francisco-based voice AI infrastructure company founded in 2015. It sells speech-to-text, text-to-speech, audio intelligence, and voice-agent APIs. Its Nova-3 model supports 50+ languages, and Flux targets real-time agents. Developers, contact centers, and healthcare and financial-services teams use it. It holds SOC 2 Type I and Type II certifications.
Deepgram’s 2026 Series C valued the company at $1.3 billion, reflecting investor interest in real-time voice AI infrastructure. It builds speech models and APIs that let developers and businesses add voice capabilities to products.
Founded in 2015, Deepgram grew from machine-learning research on waveform data from a dark-matter detector in China. Co-founders Scott Stephenson, Adam Sypniewski, and Noah Shutty were University of Michigan physicists. The company first targeted speech-to-text, then expanded into text-to-speech, audio intelligence, and voice-agent APIs. Its positioning centers on real-time voice AI infrastructure for developers and enterprise applications.
On January 13, 2026, Deepgram announced a $130 million Series C at a $1.3 billion valuation. AVP led the round. Participating investors included Alumni Ventures, Citi Ventures, Princeville Capital, Twilio, SAP, Madrona, Tiger, Wing, and Y Combinator. TechCrunch reported more than $215 million raised to date. The YC company profile lists 115 employees.
TechCrunch identifies San Francisco as its base. Deepgram describes itself as remote-first. Its staff work across 20-plus U.S. states and five-plus countries.
Key links: Website, Documentation, GitHub, and company blog. The employee figure comes from YC’s profile and may not reflect a current count.
Primary sources: funding announcement, company history, YC profile, and 2025 adoption.
Additional sources: leadership, pricing, and data policy.
More than 1,300 organizations used Deepgram’s voice AI products when it raised $130 million in January 2026. The round valued the company at $1.3 billion, signaling strong investor confidence in voice infrastructure.
Deepgram has a credible foothold in production voice workflows, especially contact centers and real-time voice agents. Its Five9 case study reports that a major healthcare provider doubled user authentication rates after switching to Deepgram transcription. Deepgram also reports 2-4 times higher accuracy than other Five9 speech-to-text options for alphanumeric inputs. These are vendor-reported results, not independent benchmarks.
The company competes directly with AssemblyAI for developer speech APIs. OpenAI Whisper and hosted Whisper services offer another path for teams prioritizing model choice or self-hosting. Cloud providers also compete for buyers already invested in their platforms. Deepgram’s distinction is a specialized speech stack spanning speech-to-text, text-to-speech, and voice-agent APIs. Its likely moat is model execution plus integrations, rather than exclusive control of the category.
Competition is moving quickly, and headline accuracy does not settle production performance. Independent benchmark analysis emphasizes latency, multilingual coverage, and workload-specific testing. Deepgram’s results should therefore be tested against each buyer’s audio, languages, and response-time needs.
Deepgram fits Chiri Atlas users who build call automation, customer support, or voice-agent products. These products need streaming speech APIs. We recommend shortlisting it alongside AssemblyAI and an open-source Whisper option. Then run a workload-specific test. The January 2026 funding and reported customer base support a positive outlook.
Intense competition remains a key risk. Continued execution across its expanded product stack is another key risk.
Primary Category: AI Infrastructure Categories: AI Infrastructure, AI Customer Service Tags: speech-to-text, text-to-speech, voice agents, conversational ai, real-time audio, transcription, speech recognition, developer api, multilingual, developer sdk
Deepgram fits AI Infrastructure because developers access speech models and voice-agent orchestration through APIs and SDKs. Its Nova-3 speech-to-text model supports more than 50 languages. The product also targets customer-service use cases through conversational voice agents and voicebots.
SOC 2 Type I and Type II reports anchor Deepgram’s security posture. Configurable data retention, regional processing, and self-hosted deployment support it. Customers must configure retention and routing controls to meet zero-retention or data-residency requirements.
Deepgram says it is GDPR-ready, CCPA-compliant, and PCI-compliant. HIPAA customers may request a Business Associate Agreement if they qualify. Deepgram’s compliance documentation is available on request, so these statements do not mean every certification or report is publicly downloadable.
By default, API requests join Deepgram’s Model Improvement Program. Deepgram then retains request content for model improvement. Setting mip_opt_out=true excludes request audio, text, transcripts, and generated audio from retention after processing. Request metadata and usage logs remain retrievable for 90 days. They do not contain that content.
Regional processing requires a regional endpoint. Full in-region handling also requires opting out. Deepgram lists endpoints for the EU, Australia, and India. Deepgram says requests fail if a region is unavailable. They do not route elsewhere.
Deepgram states that it encrypts data in transit and at rest, including TLS 1.3 and AES-256. Its security policy describes role-based access controls, two-factor authentication, and vulnerability and patch management. It also describes pre-release code scans, daily backups, and incident-response procedures. Eligible customers can also self-host. Deepgram says self-hosted request content stays within the customer deployment.
Deepgram also says it does not retain that content after each request. Its published policy describes customer notification if an incident affects customer data. However, the sources reviewed do not establish a confirmed breach history.
Flux adds turn awareness and interruption handling to Deepgram’s speech stack, while Nova-3 targets general transcription. Developers can combine speech recognition, speech generation, and conversation orchestration through APIs, or use Audio Intelligence to analyze conversations.
Speech-to-text supports recorded and streaming audio. Nova-3 offers multilingual support. It also offers features such as keyterm prompting, speaker diarization, redaction, and formatting. Flux targets real-time voice agents and includes built-in turn detection. Audio Intelligence adds summarization, sentiment analysis, intent recognition, and topic detection.
Aura-2 supplies text-to-speech. The Voice Agent API combines speech-to-text, an LLM, and speech generation. Users can bring another LLM or TTS provider.
Deepgram lists a free $200 credit and pay-as-you-go access. It also lists a Growth plan starting at $4,000 per year. Listed streaming rates include promotional Nova-3 monolingual transcription at $0.0048 per minute. Flux English costs $0.0065 per minute. The Voice Agent product page lists $4.50 per hour.
Prices and promotions can change. Check the live pricing page before budgeting.
The API is developer-focused, with REST and WebSocket options and official SDKs for Python, JavaScript, Go, .NET, and Java. The Voice Agent API supports function calling and deployment as managed, single-tenant, in-VPC, or self-hosted. Deepgram introduced Nova-3 in February 2025; Aura-2 and Flux expand its speech generation and conversational agent offerings. The reviewed product material describes APIs and developer tools, not consumer desktop or mobile apps.
Deepgram’s clearest fit is teams building voice products or modernizing customer interactions, from early-stage startups to large enterprises. Developers and product teams embed its speech-to-text, text-to-speech, and voice-agent APIs. Contact-center platforms, customer-service leaders, and enterprise IT teams are also explicit audiences.
Four common applications stand out. The first is live transcription and agent guidance. The second is contact-center analytics, including sentiment, call intent, and quality monitoring. The third is automated voice agents for routine customer requests and self-service. The fourth is clinical transcription for medical terminology and patient encounters.
Financial-services providers also use conversational agents. There, on-premises deployment can keep sensitive customer audio within controlled infrastructure.
Deepgram’s materials show three clear verticals: contact centers and conversational AI, healthcare, and financial services. Five9’s case study describes contact-center self-service. It reports that one major healthcare provider doubled user authentication rates after switching to Deepgram transcription. These are vendor-published examples, not independent adoption estimates.
There is no single stated company-size sweet spot. Deepgram says its customers range from startups to NASA. It describes hundreds of enterprise customers. Its startup program offers selected AI builders up to $100,000 in API credits over 12 months. Enterprise plans support high-throughput use and dedicated or self-hosted deployment.
The company’s materials do not identify a primary geographic market. They describe multilingual and varied-accent use cases. They also describe an early-access EU-hosted speech-to-text endpoint. This suggests deployment flexibility rather than a single-region focus.
Deepgram’s adoption reached 200,000+ developers and 400+ enterprise customers by the start of 2025, according to the company. The research team could not verify a current G2 rating or review count.
Review feedback most often praises fast transcription, low latency, accuracy, and API integration. PeerSpot reviewers also value configurable terminology, multilingual recognition, and support. Complaints point to weaker speaker identification and accuracy with multilingual accents. Users also mention live-transcription stability, concurrency limits, language gaps, and setup friction. These are reviewer reports, not controlled benchmarks.
Deepgram reports 3.3x annual usage growth over four years. It also reports more than 50,000 years of audio processed. In its Five9 case study, Deepgram says a healthcare provider doubled user authentication rates after switching. Deepgram cites better recognition of numbers and alphanumeric data. This is a vendor-reported case result, not an independently audited metric.
Adoption signals are strong, but the review evidence has limits. G2 blocks direct access to its review pages. PeerSpot’s narrative does not provide a clear review count. Prospective teams should test speaker labels, accents, concurrency, and streaming stability on their own audio before rollout.
$200 in free credit, followed by per-minute billing, anchors Deepgram’s self-serve offer. New accounts need no credit card or minimum spend for that credit. The credit has no expiration date. Growth starts at $4,000 per year in prepaid credits. It advertises savings up to 20%.
The pricing page lists promotional streaming Nova-3 monolingual rates. These are $0.0048 per minute for Pay As You Go and $0.0042 for Growth. Listed regular rates are $0.0077 and $0.0065. For pre-recorded Nova-3 monolingual audio, the rates are $0.0043 and $0.0036 per minute. These streaming promotions may change.
Deepgram does not publish enterprise rates on this page. Custom speech-to-text models require contacting sales.
Revenue figures are estimates, not company-reported results. Sacra estimates $100 million in ARR in August 2026. Sacra also reports 115% year-over-year ARR growth as Deepgram entered 2025. Deepgram reported positive cash flow at the end of 2024. It also reported more than 400 enterprise customers.
It reported 3.3x usage growth across the prior four years. In January 2026, it raised $130 million in Series C funding at a $1.3 billion valuation. Grand View Research provides broader market context. It forecasts the voice and speech recognition market will reach $53.67 billion by 2030. That industry estimate is not Deepgram’s company-specific TAM.
Deepgram’s leadership grew from a University of Michigan physics research team: three former physicists founded the company in 2015.
| Name | Title | Background |
|---|---|---|
| Scott Stephenson | Co-founder and CEO | Earned a University of Michigan PhD in particle physics. His research involved building an underground dark-matter detector and analyzing waveforms. |
| Adam Sypniewski | Co-founder and CTO | Earned a University of Michigan PhD in experimental astrophysics. He applied machine learning to research on dark energy. |
| Noah Shutty | Co-founder | YC identifies him as one of the three founders and a former University of Michigan physicist. The sources reviewed do not specify a current executive title. |
Deepgram’s leadership team combines scientific research experience with company-building. Stephenson’s team used machine learning for waveform analysis before applying deep learning to audio. Madrona’s 2024 interview describes Stephenson as CEO and co-founder, and recounts the shift from dark-matter physics to speech recognition.
Deepgram describes its culture as self-motivated, positive, passionate, and customer-focused. Its stated values include curiosity, personal authenticity, shared growth, and being human. The company says it is remote-first, with employees across 20+ U.S. states and 5+ countries. That footprint indicates geographic distribution, but the official page does not provide a current headcount. The reviewed official sources do not identify recent executive hires or named board members and advisors.
Chiri Score: 81/100
| Dimension | Score | Rationale |
|---|---|---|
| Enterprise readiness | 82/100 | Single-tenant, in-VPC, and self-hosted deployment plus 1,300+ organizations support enterprise use, though enterprise pricing and terms stay behind sales. |
| Security posture | 78/100 | SOC 2 Type I/II, TLS 1.3, AES-256, and BAAs are strong. However, default model-improvement data retention requires opt-out configuration. |
| Product depth | 85/100 | The stack spans STT, TTS, audio intelligence, and a voice-agent API with BYO LLM. Reviewers flag diarization and accent weaknesses. |
| Momentum | 86/100 | Deepgram raised a $130M Series C at a $1.3B valuation in January 2026. This and reported positive cash flow signal strong momentum in a fast-moving category. |
| Pricing transparency | 74/100 | Per-minute rates and free credit are public, but promotional rates change and enterprise and custom-model pricing are unpublished. |
Best for:
Developers building real-time voice agents that need streaming STT with turn detection
Contact-center platforms needing transcription, agent assist, and call analytics
Healthcare teams needing medical transcription under a BAA
Financial-services firms requiring on-premises or in-VPC deployment for sensitive audio
Startups wanting low-cost per-minute pricing and $200 free credit
Not for:
Teams needing strong speaker diarization accuracy without workload testing
Buyers requiring fully public enterprise pricing before contacting sales
Organizations wanting open-source model weights to modify and self-host freely
Users seeking consumer desktop or mobile transcription apps
Teams unwilling to configure opt-out settings for zero-retention requirements
| Competitor | Chiri verdict | Edge |
|---|---|---|
| AssemblyAI | Both target developer speech APIs; Deepgram has the edge on a full voice-agent stack with TTS and self-hosted deployment. | This tool |
| OpenAI Whisper | Whisper wins on open-source model control and free self-hosting; Deepgram wins on managed streaming, turn detection, and compliance. | Tie |
| Google Cloud Speech-to-Text | Deepgram offers a specialized real-time voice stack; Google suits teams consolidating on its cloud with broader platform integration. | This tool |
| Amazon Transcribe | Deepgram leads on voice-agent orchestration and low-latency focus; AWS fits buyers prioritizing existing AWS procurement and tooling. | This tool |
Yes. Deepgram holds SOC 2 Type I and Type II certifications. It also describes itself as GDPR-ready, CCPA-compliant, and PCI-compliant. Compliance documentation is available on request, not by public download.
Deepgram offers a Business Associate Agreement (BAA) to eligible HIPAA customers on request. A BAA is a contract that governs how a vendor handles protected health information. Deepgram also markets medical transcription for clinical workflows.
New accounts get $200 in free credit with no card, minimum, or expiration. Streaming Nova-3 monolingual costs $0.0048/minute pay-as-you-go at promotional rates, against a $0.0077 regular rate. Pre-recorded Nova-3 costs $0.0043/minute. The Growth plan starts at $4,000 per year. Enterprise rates require contacting sales.
By default, API requests join Deepgram's Model Improvement Program. Customers opt out by setting mip_opt_out=true. Opted-out content is retained only as long as processing requires. Metadata and usage logs stay retrievable for 90 days and contain no audio or transcripts.
Deepgram competes directly with AssemblyAI for developer speech APIs. OpenAI Whisper and hosted Whisper services compete on model choice and self-hosting. Cloud providers such as Google Cloud, AWS, and Microsoft Azure compete for buyers already committed to their platforms.
Deepgram suits enterprise voice workloads. It offers SOC 2 reports, EU, Australia, and India regional endpoints, and single-tenant, in-VPC, or self-hosted deployment. More than 1,300 organizations use its products. Buyers must configure retention and routing to meet data-residency needs.
Nova-3 supports more than 50 languages. Deepgram provides official SDKs for Python, JavaScript, Go, .NET, and Java. It exposes REST and WebSocket APIs for recorded and streaming audio.
Flux is Deepgram's speech-to-text model for real-time voice agents. It includes built-in turn detection, which identifies when a speaker finishes talking. It also handles interruptions. Flux English streaming lists at $0.0065 per minute.
Reviewed by Chiri Atlas Research Desk (AI Tooling Analyst) on 2026-10-02.