Skip to main content
Updated July 12, 2026
Transcription models turn recorded audio into text. For prerecorded files, the real trade-off is accuracy against speed, price, and features like speaker labels - and today’s leaders score so close that the wrong pick is easy to make. We ranked 15 on a shared accuracy benchmark.

Best Transcription Models


Scribe v2

Best overall batch transcription
This is the strongest all-around transcription candidate when accuracy and rich output both matter.
Score 98Price License ProprietarySpeed 31.9x
  • Near-leading accuracy paired with the things transcripts actually need: speaker diarization across many voices, word-level timestamps, audio event tags, and broad language coverage.
  • There’s also a real upload interface, so you can run files without writing code. For most mixed-content jobs, it’s the safe default.
  • The base rate looks cheap until you switch on extras like entity detection or keyterm prompting, which carry surcharges.
  • If you only need fast, plain English transcripts, Pulse Pro or Parakeet TDT 0.6B V3 do that for less and quicker.
Fast accuracy-first batch jobs
A top-accuracy model that also runs unusually fast on long audio, so it’s the pick when you need both and can accept preview status.
Score 98Price License ProprietarySpeed 261.2x
  • It sits with the most accurate models here while clearing hours of audio in a fraction of the time most rivals take, which is rare - accuracy and throughput usually pull against each other.
  • Language auto-detection and phrase biasing help on messy, multi-speaker recordings.
  • It’s a public-preview endpoint with no production SLA yet, and it has no speaker diarization - a real gap for interviews and meetings.
  • If you need speaker labels, Scribe v2 or Universal-3.5 Pro are the safer calls.

Pulse Pro

Fast low-cost English transcription
A standout if your audio is English and you want top accuracy, high speed, and a low price without paying for extras you won’t use.
Score 98Price License ProprietarySpeed 292.3x
  • It lands accuracy, speed, and cost in the same place, which is unusual - most models make you give up one to get another.
  • For high-volume English transcription where you just need clean text back quickly, it’s one of the strongest options here.
  • Pulse Pro is English-only and file-based, so it’s out for multilingual work or streaming.
  • The ecosystem is smaller and less established with no end-user app, so you’re committing to an API from a less proven vendor - weigh it against Soniox v5 Async if you need languages.
Private multilingual deployment
The best-scoring open-weight option here, and the one to pick when you need to keep audio in-house and can bring serious hardware.
Score 96Price License Open weightSpeed 54.9x
  • Open weights under a permissive license mean you can run it on your own machines, keep sensitive audio private, and pay no per-minute fee.
  • It’s genuinely multilingual and doubles as an audio-understanding model, so it can summarize or answer questions about a clip, not just transcribe it.
  • The 24B weights are heavy - realistically a high-end GPU or aggressive quantization, not a casual local install. Clip length is capped, and there’s no built-in diarization.
  • For open weights that run on a laptop, Parakeet TDT 0.6B V3 or Whisper Large v3 Turbo fit better.
Multimodal audio analysis
Reach for this when you want to reason about audio - summaries, Q&A, structured notes - rather than get a faithful word-for-word transcript.
Score 96Price License ProprietarySpeed 7.1x
  • It understands audio, not just transcribes it: ask for a summary, action items, or speaker-attributed notes in one call, and it handles very long files thanks to a huge context window.
  • For turning a recording into structured output, it’s more flexible than any dedicated ASR model here.
  • It’s the slowest model here and priced well above dedicated transcribers, it’s still preview, and it tends to condense rather than transcribe verbatim - with timestamps that drift on long files.
  • For accurate, timestamped transcripts, Scribe v2 or Universal-3.5 Pro are better.
Feature-rich production transcription
A strong, well-rounded choice for production pipelines that need broad language support, diarization, and more control over difficult terminology.
Score 95Price License ProprietarySpeed 99.3x
  • AssemblyAI’s current async flagship supports 18 languages, native code switching, contextual prompting, and its latest diarization.
  • The surrounding audio-intelligence tools - sentiment, topics, entities, and redaction - can turn a transcript into something directly usable in a product.
  • Its performance numbers here are inherited from the predecessor, so treat its exact rank as provisional. Add-ons also stack on the base rate.
  • For directly benchmarked multilingual choices, compare Scribe v2, Soniox v5 Async, or Speechmatics Enhanced.
Noisy European business audio
A specialist tuned for messy, real-world business audio in a handful of European languages, not a broad general-purpose transcriber.
Score 95Price License ProprietarySpeed 60.2x
  • Purpose-built for the hard stuff: contact-center calls, meetings, and accented, multi-speaker recordings in its core European languages, where it holds accuracy that general models lose.
  • Diarization and language detection come bundled. If your audio is noisy business speech in those languages, it’s a sharp fit.
  • Coverage is narrow - a few European languages - and on clean, formal, or read-aloud audio it actually trails Gladia’s older Solaria-1, which spans far more languages.
  • It also costs more than most models here. For broad multilingual work, look at Solaria-1 or Soniox v5 Async.
Transcription plus audio reasoning
A multimodal model that transcribes well inside a broader audio-and-video reasoning workflow, but it isn’t a dedicated transcription tool.
Score 94Price License ProprietarySpeed 97.9x
  • Strong accuracy and throughput inside a model that also reasons over audio and video, so you can transcribe and then summarize, translate, or answer questions in the same workflow.
  • Language breadth is wide. It’s a fit when transcription is one step in a larger multimodal task.
  • It’s not dedicated ASR, so you don’t get turnkey word-level timestamps or diarization, and token-based pricing means you estimate cost per hour rather than pay a flat rate.
  • For plain transcription, Voxtral Mini Transcribe 2 or Deepgram Nova-3 are simpler and more predictable.
Low-cost dedicated transcription
A no-frills dedicated transcription endpoint with a simple flat price - a clean pick when you just want accurate transcripts back cheaply.
Score 94Price License ProprietarySpeed 80.7x
  • Solid accuracy at a low, flat per-minute price, with built-in diarization, word-level timestamps, and custom-term biasing. It handles long files in a single request.
  • For straightforward batch transcription without platform complexity, it’s one of the better value picks here.
  • It’s proprietary despite the Voxtral family’s open-weight reputation, so there’s no self-hosting here. Language coverage is limited, and overlapping speech tends to collapse to one speaker.
  • If you need many languages or audio-intelligence features, Universal-3.5 Pro or Soniox v5 Async go further.
Low-cost multilingual files
One of the cheapest ways to get accurate, multilingual transcripts with diarization and translation bundled in - if you can live with modest speed.
Score 93Price License ProprietarySpeed 19.6x
  • Broad language coverage with native code-switching, plus diarization, timestamps, and translation all included in one low rate - no per-feature surcharges.
  • It’s strong on hard audio: noisy, telephony, accented, multi-speaker. For cost-sensitive multilingual batch work, the all-in pricing is hard to beat.
  • Measured throughput is on the slow side, so it’s not ideal for huge, time-sensitive batches. The first-party app hides model selection, so exact async-v5 control lives in the API.
  • If you need speed, Parakeet TDT 0.6B V3 or Deepgram Nova-3 clear files far faster.
Low-friction general transcription
A simple, capable transcription endpoint that’s easy to reach for, but it doesn’t lead specialists on accuracy, price, or speed.
Score 92Price License ProprietarySpeed 31.5x
  • A clean, well-documented endpoint that handles accents and background noise well and accepts a prompt to steer names and terminology.
  • Broad language coverage and dead-simple integration make it a low-effort default when you want decent transcripts without evaluating a specialist provider.
  • The base model returns no word or segment timestamps, ruling it out for captioning and alignment work, and users report occasional dropped words on tough audio.
  • On accuracy, price, and speed, Scribe v2, Pulse Pro, and Voxtral Mini Transcribe 2 all beat it.
Accent-rich enterprise transcription
The pick when accents and dialects are the problem, with enterprise deployment options most hosted-only rivals don’t offer.
Score 92Price License ProprietarySpeed 61.6x
  • Its single global model per language holds up across accents and dialects that trip up others, and it covers a broad language set. Container and private-cloud deployment make it viable for regulated, data-sensitive work, and diarization and translation are built in.
  • A dependable choice for varied, accented audio.
  • It costs more than commodity transcription APIs, and its per-model pricing is opaque, so confirm your rate before committing. Brand mindshare is lower than Deepgram or Whisper.
  • If you don’t need accent robustness or on-prem, Universal-3.5 Pro or Soniox v5 Async cost less.
Laptop-friendly local speed
The standout when you want to run transcription yourself: tiny, extremely fast, and genuinely runnable on a laptop.
Score 92Price License Open weightSpeed 958.6x
  • At just 0.6B parameters it’s extremely fast and light enough to run on a typical laptop, including Apple Silicon, with 25-language support, word- and segment-level timestamps, and punctuation.
  • There’s also an exact hosted route if you’d rather not self-host. For local or high-volume transcription, it’s a standout.
  • Accuracy is good but not best-in-class, and it slips on non-English, accented, or noisy audio. There’s no built-in diarization, and the license requires attribution.
  • For the highest accuracy, Scribe v2 or MAI-Transcribe-1.5 win; for easier setup, Whisper Large v3 Turbo is friendlier.
Easiest local Whisper option
The most practical way into the Whisper ecosystem: nearly as accurate as full Large v3, far lighter, and easy to run locally.
Score 90Price License Open weightSpeed 145.9x
  • It keeps most of full Large v3’s accuracy while running several times faster and lighter, so it runs on a typical laptop or CPU through a mature ecosystem of tools.
  • A permissive license, 99-language support, and near-free hosted access make it the easiest open Whisper to actually use.
  • It’s an older architecture that now trails newer models on accuracy and speed, and it can hallucinate text during silence or music. There’s no built-in diarization.
  • For higher local accuracy, full Whisper Large v3 helps; for raw speed, Parakeet TDT 0.6B V3 is far quicker.
Highest-throughput hosted API
The fastest proprietary API we measured, with mature prerecorded features - a throughput play, not an accuracy leader.
Score 88Price License ProprietarySpeed 562.7x
  • Very high measured throughput and a mature, well-documented prerecorded stack with diarization, formatting, and keyword features.
  • If you’re processing large volumes of audio and need results back fast and reliably from a hosted API, few models keep up with its speed.
  • Accuracy trails the leaders, so it’s the wrong pick when transcript quality is paramount. Its clean rate is the prerecorded pay-as-you-go price, not the cheaper streaming tier.
  • For more accuracy at similar or lower cost, Scribe v2, Universal-3.5 Pro, or Soniox v5 Async are stronger.

How to Choose

When choosing between these models, weigh four things:
  • Access: Decide first whether you’ll use a hosted API, a first-party app, or run the model yourself, because that choice drives cost, privacy, latency, and setup work more than any single benchmark. Only Voxtral Small, Parakeet TDT 0.6B V3, and Whisper Large v3 Turbo are realistic self-host options; the rest are hosted.
  • Quality: The score is a 0-100 index built from Artificial Analysis’s AA-WER v2 benchmark, which blends conversational, parliamentary, and earnings-call English audio and rewards lower word error. Treat it as an English-accuracy proxy - it doesn’t fully capture multilingual breadth, diarization, timestamps, noisy telephony, or long-file reliability.
  • Price: We use current US dollars per hour of prerecorded audio for the scored route. Token-billed models like Gemini 3.1 Pro and Qwen3.5-Omni-Plus are converted to a comparable hourly figure, and add-ons like diarization or entity detection can push real cost above the base rate.
  • Speed Factor: How many seconds of audio each model transcribes per second of processing. If you’re clearing large batches, this matters as much as price - Parakeet TDT 0.6B V3 and Deepgram Nova-3 are in a different league from Gemini 3.1 Pro.

Other Models We Considered

Whisper Large v3 (OpenAI) — A bit more accurate than Turbo, but heavier and slower to run.Solaria-1 (Gladia) — Broader 100+ language coverage than Solaria-3, and better on clean audio.GPT-4o Mini Transcribe (OpenAI) — Cheaper and faster than the full model, but noticeably less accurate.Amazon Transcribe (Amazon) — Familiar cloud baseline, but the specialist models here are more accurate and faster.Chirp 3 (Google) — Google Cloud’s broad-language transcription, solid but behind the top picks.Canary-Qwen-2.5B (NVIDIA) — Strong English local accuracy, but it needs a capable NVIDIA GPU.Qwen3-ASR-1.7B (Alibaba) — Broad multilingual open model for self-hosting, but its accuracy is unproven here.Granite Speech 4.1 2B (IBM) — Compact, openly licensed local model, but hard to compare on the same benchmark.Gemini 3 Flash (Google) — A faster, cheaper Gemini for audio, but an older preview now superseded.Fun-ASR Realtime (Alibaba) — Tops the benchmark on paper, but it’s realtime-only and outside batch scope.

Frequently Asked Questions

For most mixed-content work, Scribe v2 is our top overall pick - near-leading accuracy with the diarization, timestamps, and language coverage real transcripts need. MAI-Transcribe-1.5 and Pulse Pro match it on raw accuracy and are much faster, so consider them when throughput matters - just note MAI’s preview status and Pulse Pro’s English-only limit.
If you want one safe default, Scribe v2. If your audio is English and you care about cost and speed, Pulse Pro or a hosted Whisper Large v3 Turbo will do the job for less. Match the model to your audio rather than chasing the top score.
For a typical laptop, Parakeet TDT 0.6B V3 is the fastest and lightest, and Whisper Large v3 Turbo is the easiest with the biggest ecosystem. Voxtral Small scores higher and is genuinely multilingual, but its 24B weights need a high-end GPU or heavy quantization.
Hosted, Whisper Large v3 Turbo and Parakeet TDT 0.6B V3 are the cheapest per hour, and Soniox v5 Async bundles diarization and translation into a very low rate. Self-hosting Parakeet or Whisper drops the cost to just your own compute.
Soniox v5 Async and Speechmatics Enhanced cover broad language sets with diarization built in, and Scribe v2 spans many languages with rich output. For noisy European business calls specifically, Solaria-3 is tuned for that; for the widest coverage, Gladia’s older Solaria-1 still leads.
You need diarization. Scribe v2, Universal-3.5 Pro, and Voxtral Mini Transcribe 2 all handle it well. Avoid MAI-Transcribe-1.5 here - it’s fast and accurate but has no speaker diarization.
Partly. Our score is English-only and rewards low word error on conversational, parliamentary, and earnings audio. It won’t tell you how a model handles your languages, accents, background noise, overlapping speakers, or long files - test a shortlist on your own audio before committing.