Skip to main content
Updated July 12, 2026
Real-time transcription models turn speech into text as you talk, trading off accuracy, latency, price, and how fast they detect when a speaker finishes. Vendor latency claims rarely mean the same thing, so we ranked 15 streaming models on one independent benchmark. Our primary score is final-transcript accuracy from the Artificial Analysis streaming speech-to-text benchmark, normalized to a 0-100 scale where higher is better. It covers English-language audio only, so treat it as a strong baseline, not the last word: telephony, accents, and other languages can shift the order.

Best Real-Time Transcription Models


Highest-accuracy multilingual streaming
The most accurate streaming model in the benchmark, with clean live partials and wide language support - the default pick when transcript quality matters most.
Score 100Price License ProprietaryTime to final 0.141s
  • Top-tier final accuracy paired with unusually clean, stable partial transcripts, so words hold their place as you speak instead of rewriting themselves.
  • Language coverage is broad and detection is automatic, which makes it the safest choice when accuracy across many languages is the priority.
  • It is proprietary and API-only, with no self-host route, and sits at the pricier end of the field.
  • Diarization is comparatively weak, so for clean multi-speaker separation you may prefer AssemblyAI or a dedicated diarization step.
Accuracy-first voice agents
Ties for the top accuracy spot and adds genuine semantic endpointing, making it one of the strongest picks built specifically for voice agents.
Score 100Price License ProprietaryTime to final 0.211s
  • Final-transcript accuracy is at the very top of the field, and its semantic endpointing judges when you have actually finished a thought rather than just paused.
  • That combination makes it one of the most convincing streaming models for building responsive voice agents.
  • It is English-only, which rules it out for multilingual products, and it is proprietary and API-only.
  • If you need broad language coverage, ElevenLabs Scribe v2 or a multilingual model like Qwen3 or Nemotron will serve you better.
Budget multilingual accuracy
One of the cheapest ways to get near-top final accuracy, as long as you can live with rough, unstable live partials.
Score 98Price License ProprietaryTime to final 0.476s
  • You get final accuracy close to the best models here at one of the lowest prices in the field, plus strong multilingual and dialect coverage.
  • For high-volume, cost-sensitive transcription where the finished transcript matters more than the live feed, it is hard to beat on value.
  • Its live partials are rough and unstable - fine if you only consume the final transcript, but a poor fit for interfaces where users watch words appear as they talk.
  • For steady live captions, Soniox v5, Cartesia Ink 2, or ElevenLabs are better.
Low-cost streaming with turn detection
A cheap newcomer with accurate final transcripts and built-in turn detection, though its live partials lag well behind the accuracy leaders.
Score 96Price License ProprietaryTime to final 0.373s
  • Accurate final transcripts and built-in turn detection at a low price, from a newcomer clearly aiming at the voice-agent market.
  • If you want inexpensive streaming and mostly care about the committed transcript, it is a credible option worth testing.
  • Like Qwen3, its live partials trail the accuracy leaders, so it is weaker for interfaces that display text as you talk.
  • It is also very new, so integrations and tooling are thinner than Deepgram’s or AssemblyAI’s. Proprietary and API-only.
Tunable accuracy for voice agents
A top-accuracy incumbent with a rare accuracy-versus-latency switch, though diarization is a paid, slower add-on rather than a core strength.
Score 93Price License ProprietaryTime to final 0.445s
  • Among the most accurate streaming models, with an explicit switch between maximum accuracy and minimum latency that few rivals offer, so you can tune the same model to the job.
  • Stable, immutable transcripts make it dependable for live captioning.
  • Speaker diarization is a paid add-on and has historically been slower than the core transcription, so multi-speaker work costs more and lags.
  • If diarization is central, test it carefully; if raw value matters more, Soniox v5 undercuts it heavily.
Best price-to-performance streaming
The value outlier here - near-top accuracy and among the fastest finals at the lowest price, and still oddly absent from most roundups.
Score 89Price License ProprietaryTime to final 0.054s
  • It combines near-top accuracy, among the fastest finals in the field, and the lowest price of any highlighted model, which is a genuinely rare mix.
  • There is also a first-party mobile app, so it is one of the few here you can try without writing code.
  • The main hesitation is maturity: it is a smaller, newer vendor than the incumbents, which matters for risk-averse enterprise buyers.
  • Feature depth like advanced diarization is still catching up to AssemblyAI and the larger clouds. Proprietary and API-first.
Broad multilingual coverage
Broad language coverage from a major cloud ASR, but slow finalization and steep list pricing make it hard to recommend for latency-sensitive work.
Score 86Price License ProprietaryTime to final 1.276s
  • Very broad language coverage plus the enterprise controls - data residency, regional endpoints, compliance - that regulated products often require.
  • As a pure recognition model it is accurate across a wide multilingual range, which is its real reason to exist.
  • Finalization is the slowest of any model here, which disqualifies it for latency-sensitive voice agents, and its list pricing is steep until you reach very high volume.
  • For real-time work, Soniox v5 or Deepgram are far better fits.
Realtime voice-app transcription
Solid, general-purpose realtime transcription that you pay a heavy premium for - most teams should downshift to a cheaper OpenAI transcribe model.
Score 85Price License ProprietaryTime to final 0.688s
  • General-purpose accuracy that holds up well across everyday speech, delivered through a mature, well-documented realtime interface built for conversational voice apps.
  • For products that want transcription which just works without configuration, it is a dependable default.
  • It is by far the most expensive option here, and it does not lead on accuracy or latency to justify that premium.
  • Most teams should drop to a cheaper transcribe model such as GPT-4o Transcribe, or move to Soniox v5 for value.
Cheapest low-latency streaming
The cheapest option on this list pairs low latency with built-in turn detection, a strong budget pick for real-time voice work.
Score 82Price License ProprietaryTime to final 0.082s
  • The lowest price on the list combined with very low latency is a strong pairing, and built-in turn detection means you get endpointing without bolting on a separate voice-activity step.
  • For cost-conscious real-time voice work, it punches above its price.
  • Accuracy is mid-pack rather than class-leading, so for demanding transcription you will want ElevenLabs, Cartesia, or AssemblyAI.
  • It is proprietary and API-only, and as a newer entrant its ecosystem is thinner than the incumbents’.
Self-hostable streaming accuracy
The standout open-weight pick - it reaches hosted-grade accuracy at low latency and can run on your own GPU, minus diarization and turn detection.
Score 80Price License Open weightTime to final 0.682s
  • The strongest open-weight option here: it reaches hosted-grade accuracy at low, configurable latency and is genuinely multilingual.
  • You can call the hosted API or run the weights yourself, which gives you privacy and data control while shifting cost to your own hardware.
  • There is no diarization or turn detection in the realtime model, so voice agents need extra components around it. Self-hosting needs a capable GPU, not a laptop, so “open” does not mean effortless.
  • Cartesia or Deepgram Flux handle turn-taking natively.
Enterprise multilingual transcription
A mature enterprise ASR with deep customization and compliance features, but it trails the newer wave on latency and real-time value.
Score 80Price License ProprietaryTime to final 0.625s
  • Deep customization - custom vocabulary, pronunciation, and tuned models - plus broad language support and the compliance and data-handling controls enterprises need.
  • For regulated, large-scale deployments that value configurability over raw speed, it remains a serious option.
  • It trails the newer wave on latency and, at standard real-time rates, on price, so it is a weak value pick for greenfield projects.
  • For faster or cheaper streaming, Soniox v5, Deepgram, or Inworld are stronger. Proprietary and API-based.
The most flexible self-hosted pick - open weights, tunable latency, and multilingual coverage in a compact 0.6B model.
Score 79Price License Open weightTime to final 0.418s
  • Open weights with runtime-selectable latency let you dial the accuracy-speed trade without swapping models, and it is multilingual and light enough to self-host at high concurrency.
  • Running it yourself keeps audio on your infrastructure and makes cost depend on your hardware rather than an API meter.
  • You own the deployment: serving, scaling, and updates are on you, which is real work versus a managed API. Peak accuracy trails the top hosted models.
  • If you want open weights without the ops, Voxtral Mini offers a hosted API too.
Managed enterprise streaming
A dependable managed streaming service that now trails newer models on accuracy, latency, and price, with little to pull you toward it.
Score 74Price License ProprietaryTime to final 0.620s
  • A mature, heavily operated managed service with predictable behavior, custom vocabulary, and the scale and reliability large deployments count on.
  • It handles high-volume streaming transcription dependably across a wide set of languages.
  • It now trails newer models on accuracy and latency while costing more than most, so there is little reason to start here on the merits.
  • For better accuracy, latency, or price, Soniox v5, Deepgram, or ElevenLabs all lead it.
Fast voice-agent default
The long-standing voice-agent default: very fast and reliable on clean audio, though newer models have caught and passed it on accuracy.
Score 64Price License ProprietaryTime to final 0.066s
  • Fast, reliable streaming that made it the long-time default for voice agents, with strong tooling, mature SDKs, and consistent low-latency behavior on clean audio.
  • For straightforward English voice pipelines, it is still a safe, well-supported workhorse.
  • On noisy, accented, or telephony audio its accuracy slips more than the newer leaders, and several models now beat it on final quality.
  • If accuracy is the priority, Soniox v5, ElevenLabs, or Cartesia are stronger; for turn-taking, look at Flux.
Fastest endpointing for agents
Built for turn-taking rather than raw accuracy - it delivers the fastest finals and fused end-of-turn detection, a deliberate voice-agent trade.
Score 55Price License ProprietaryTime to final 0.021s
  • Purpose-built for conversational turn-taking: it fuses transcription with end-of-turn detection and delivers the fastest finals in the field, so agents can respond the moment you actually stop talking.
  • For latency-critical voice agents, that focus is the whole point.
  • It is English-only in this model, and on final-transcript accuracy it sits at the back of this list, so transcription-quality work is better served elsewhere.
  • For higher accuracy, look at Soniox v5 or ElevenLabs; for multilingual, choose a different model entirely.

How to Choose

When choosing between these models, consider:
  • Access: Decide first whether you need a managed API, a first-party app, or a self-hosted model, because that choice drives cost, privacy, latency, and setup work. Most models here are API-only. Only Voxtral Mini and Nemotron 3.5 offer a real self-host route, and Soniox and Azure add a first-party app or portal for trying the model without code.
  • Quality: We use final-transcript accuracy from the Artificial Analysis streaming benchmark, scored 0-100 where higher is better. Watch the partial-versus-final split: a few models (Qwen3, Grok) produce excellent final transcripts but rough live partials, which is invisible in a single accuracy number and matters if users watch text appear as they speak.
  • Price: We compare on price per hour of streaming audio. Streaming costs more than batch, committed and volume tiers swing prices widely, and features like diarization are often billed on top - so confirm the tier and add-ons before you budget.
  • Time to final transcript: This is seconds from the end of speech to the final transcript; lower matters most for voice agents, where the practical target is a sub-500ms end-to-end response. Raw latency is only half of it - how well a model detects that a speaker has finished (its endpointing) shapes the felt responsiveness just as much.
For most teams building voice agents, start with Soniox v5 for value, Cartesia Ink 2 or Deepgram Flux when turn-taking is the hard part, and ElevenLabs Scribe v2 when transcript quality outweighs everything. For private or offline deployments, Voxtral Mini and Nemotron 3.5 are the two open-weight picks worth real testing.

Other Models We Considered

OpenAI GPT-4o Transcribe (OpenAI) — Cheaper OpenAI transcription than the realtime model, capable but not latency-first.Speechmatics Realtime Enhanced (Speechmatics) — Strong multilingual and enterprise specialist, but final accuracy trails the leaders.Smallest Pulse (Smallest.ai) — Very fast finalization, but accuracy and pricing trail the top value picks.OpenAI Whisper Large v3 (OpenAI) — A great open local baseline, but not truly streaming without wrappers.NVIDIA Parakeet Unified EN 0.6B (NVIDIA) — Local favorite for accuracy and speed, but English-only and GPU-bound.Kyutai STT (Kyutai) — Open streaming that runs on a typical machine, but no managed API.Moonshine v2 Streaming (Moonshine AI) — On-device streaming for CPU and phones, but off-benchmark with no managed API.Gladia Solaria 1 Realtime (Gladia) — A recognizable API option, but the slowest finalization in the benchmark.Rev AI Streaming (Rev AI) — An established, affordable baseline, but the weakest accuracy among current models.

Frequently Asked Questions

For pure transcript quality, ElevenLabs Scribe v2 Realtime and Cartesia Ink 2 lead on accuracy. But the model most teams should try first is Soniox v5, which pairs near-top accuracy with among the fastest finals at the lowest price on this list.
Soniox v5 Real-Time. It is the rare model that is accurate, fast, and cheap at the same time, and there is a first-party app if you want to try it before writing any code. Move to ElevenLabs or Cartesia only when transcript quality has to be the best available.
Not really. OpenAI’s original Whisper is batch-native - it processes fixed audio chunks, so “streaming” wrappers repeatedly recompute overlapping windows, which is slow and jittery. If you want real streaming, use a streaming-native model such as Voxtral Mini Transcribe Realtime, Nemotron 3.5 ASR Streaming, or Kyutai STT.
Voxtral Mini Transcribe Realtime is the strongest open-weight pick, though it needs a capable GPU. Nemotron 3.5 ASR Streaming is smaller but still needs the supported high-end NVIDIA stack. For ordinary-machine or phone deployments, Kyutai STT and Moonshine v2 are the more practical options.
It depends on the hard part. If turn-taking and endpointing are what break your agent, Deepgram Flux fuses transcription with end-of-turn detection and delivers the fastest finals. If you want accuracy plus native endpointing, Cartesia Ink 2 is excellent. For the best overall value, Soniox v5.
They are built for different jobs. Flux is tuned for conversational turn-taking and the fastest possible finals, which is what voice agents need. Nova-3 is the general-purpose streaming workhorse and scores higher on final-transcript accuracy in our benchmark. Choose Flux for responsiveness, Nova-3 for broader transcription.
Treat them as a strong starting point, not a guarantee. The benchmark covers English-language audio and does not directly measure 8kHz telephony, multilingual quality, diarization, or full voice-agent latency, so your own workload can reorder the results. Always test the shortlist on your own audio before committing.