Skip to main content
Updated July 12, 2026
Speech-to-speech models take audio in and talk back in real time - the engines behind voice agents and live assistants. The catch: the best-sounding, smartest models often aren’t the fastest to respond, and price swings widely. We compared 15 on quality, speed, price, and access.

Best Speech-to-Speech Models


Complex production voice agents
The most capable speech-to-speech model in this comparison, and the one to beat for demanding, tool-driven voice agents that can’t afford to drift.
Score 100Price License ProprietaryTime to first audio 1.14s
  • It leads on speech reasoning and conversational dynamics at once, so it follows multi-step instructions, handles interruptions cleanly, and stays coherent through long, messy calls.
  • When the task is hard and the agent has to think, act, and talk without losing the thread, this is the pick.
  • It’s priced well above the cheaper realtime tiers, so high-volume, simple flows burn budget fast - GPT-Realtime mini or Qwen3.5 Omni Flash Realtime fit those better.
  • And confirm you want this exact model, since the newer GPT-Realtime-2.1 is worth testing beside it.
Tool-heavy phone agents
xAI’s strongest voice model, and the one we’d reach for when an agent has to call tools and take real actions mid-conversation.
Score 97Price License ProprietaryTime to first audio 1.25s
  • It posts the strongest agentic results in the suite, so it stays reliable when a call turns into actual work - looking things up, triggering functions, and pushing a task forward while still sounding natural.
  • For phone agents that do more than chat, it’s near the very top.
  • It edges just behind the top scorer overall, so for the hardest reasoning you might still prefer GPT-Realtime-2.
  • Keep it distinct from Grok Voice Agent, which starts faster but is noticeably weaker at both reasoning and tool use.
High-quality model to watch
A near-top quality result with no public deployment route, making this a model to monitor rather than one you can choose today.
Score 96Price License ProprietaryTime to first audio 1.39s
  • In the benchmark data it’s a genuine front-runner, matching the best on speech reasoning and natural back-and-forth.
  • If Alibaba ships a documented public deployment, it could move straight into the top tier of models you’d actually build on.
  • Right now there’s no verified public API, price, or app for the exact scored model, so you can’t ship it today.
  • Treat it as a watch-list entry, and don’t confuse it with the separate Fun-Audio-Chat project, which isn’t the same model.
Fast, high-end realtime voice
Fast and strong, but priced at the top of this list - hard to justify for a new build when GPT-Realtime-2 is cheaper and better.
Score 89Price License ProprietaryTime to first audio 0.82s
  • It’s among the highest-accuracy models here, and it starts talking quickly, so exchanges feel responsive without sacrificing much reasoning.
  • As a low-latency, high-quality realtime voice model it holds up well on its own - the issue is what it costs, not what it does.
  • The price is the problem: it sits at the top of the range while GPT-Realtime-2 scores higher and costs far less.
  • For almost any new project, start with GPT-Realtime-2 instead; there’s little reason to reach for 1.5.
Low-cost reasoning voice agents
Strong reasoning and native audio-to-audio at a genuinely low hourly price, held back mainly by how slowly it starts talking.
Score 83Price License ProprietaryTime to first audio 2.98s
  • It pairs solid speech reasoning with native audio-to-audio, live tool use, and one of the lower price points among the capable models.
  • For reasoning-heavy voice agents where you care more about answer quality and cost than instant response, it’s a smart, affordable choice.
  • Its weak spot is the slow first response - the high-quality configuration is among the laggiest here, which hurts on quick back-and-forth.
  • If latency is your priority, Deepslate Opal or the fast OpenAI tiers feel far more immediate.
Proven baseline realtime voice
The familiar, sub-second realtime model many teams already know - still capable, but both pricier and weaker than the newer GPT-Realtime-2.
Score 82Price License ProprietaryTime to first audio 0.98s
  • It responds in under a second and handles natural conversation reliably, which is why it became a common default for voice agents.
  • Nothing about it is broken, and it stays a dependable, well-understood option for straightforward spoken interactions.
  • It’s been overtaken: GPT-Realtime-2 is more capable and much cheaper, and GPT-Realtime-1.5 answers faster.
  • There’s no strong reason to start a new build here - keep it only where it’s already wired in and working.
Hosted multilingual speech reasoning
The strongest speech-reasoner we’ve seen at a fraction of the top-tier price, as long as you can live with a slow first response.
Score 82Price License ProprietaryTime to first audio 2.64s
  • It’s excellent at reasoning out loud, with function calling, search, broad multilingual support, and clean interruption handling - and it’s one of the cheapest capable models to run.
  • For multilingual, reasoning-led voice work on a tight budget, it’s hard to beat.
  • It’s slow to start talking, so it’s a poor fit for snappy, interactive agents. The listed price also covers input audio only, not a full session, so real costs run higher.
  • For low latency, look at the fast OpenAI or Qwen Flash tiers.
Open-weight speech reasoning
The standout open-weight pick - strong speech reasoning under an Apache-2.0 license, with first-party app and API routes if you’d rather not self-host.
Score 81Price License Open weightTime to first audio 1.51s
  • It’s the rare open-weight model that competes with hosted leaders on reasoning, and the permissive license lets you deploy it however your compliance needs dictate.
  • First-party app and API routes mean you can use it immediately without standing up your own infrastructure.
  • ”Open weight” here doesn’t mean easy - the official self-hosting path is infrastructure-heavy and multi-GPU, not a local-machine setup.
  • If you want a voice you truly run yourself, PersonaPlex fits better; if you just want it hosted, the API route is the practical choice.
Enterprise voice agents
A solid, mid-tier voice model built for production agents - streaming, tools, retrieval, and interruptions - though it trails the top scorers on raw quality.
Score 74Price License ProprietaryTime to first audio 1.14s
  • It’s built for real voice-agent work: low-latency streaming, tool use, retrieval, interruption handling, and multilingual support, all geared for production from the start.
  • For teams that want a dependable, feature-complete agent model rather than the highest benchmark score, it delivers.
  • On pure quality it sits mid-pack, behind GPT-Realtime-2, Grok Voice Think Fast 1.0, and the Gemini and Qwen reasoning models.
  • The listed price covers input audio only, and long calls need a session-continuation pattern that adds engineering work.
Fast everyday voice agents
xAI’s quick, practical voice agent - sub-second responses and an easy build path, but clearly a step below its Think Fast sibling on quality.
Score 71Price License ProprietaryTime to first audio 0.78s
  • It answers fast, near the quickest here, and comes with a straightforward builder and API, so you can stand up a responsive voice agent without much fuss.
  • For everyday, latency-sensitive assistants that don’t need frontier reasoning, it’s a reasonable pick.
  • It’s meaningfully weaker than Grok Voice Think Fast 1.0 on both reasoning and tool use, so don’t mix the two up.
  • If your agent does real work mid-call, step up to Think Fast; if you only need speed, other fast tiers compete on price.
Lowest-latency voice responses
The fastest model here by a clear margin on first response, with EU hosting and flexible integration routes, but only mid-tier on quality.
Score 68Price License ProprietaryTime to first audio 0.44s
  • Nothing else starts talking as quickly, so conversations feel genuinely instant - the closest to human turn-taking in this group.
  • REST, WebSocket, and SIP routes plus EU-based hosting make it easy to slot into telephony and privacy-sensitive setups.
  • That speed comes with only middling reasoning and weaker agentic performance, so it’s not the one for complex, tool-heavy tasks.
  • The ecosystem is smaller and public pricing is thin. For more capability at similar latency, weigh the fast OpenAI tiers.
Fast, low-cost realtime chat
The budget-friendly, low-latency OpenAI realtime option - great at natural conversation, but a real step down in reasoning and tool use.
Score 58Price License ProprietaryTime to first audio 0.81s
  • It’s quick to respond and handles everyday back-and-forth smoothly at a lower cost than the flagship.
  • For high-volume, lightweight voice - simple Q&A, routing, casual assistants - it’s an efficient workhorse that keeps conversations feeling natural.
  • Push it toward multi-step reasoning or serious tool use and it falls well short of GPT-Realtime-2 and the reasoning-led models.
  • Note it’s the older mini - GPT-Realtime-2.1 mini is a newer, distinct option worth testing before you commit.
Cheap high-volume voice agents
The speed-and-value play from the Qwen line - very cheap, quick to respond, and multilingual, but noticeably weaker at reasoning than Omni Plus.
Score 53Price License ProprietaryTime to first audio 0.79s
  • It’s among the cheapest models here and starts talking fast, with broad language coverage.
  • For high-volume, cost-sensitive voice where you need many concurrent sessions more than deep reasoning, it stretches a budget further than almost anything else on this list.
  • Its reasoning is well behind Qwen3.5 Omni Plus Realtime, so it’s the wrong tool for complex, multi-step conversations.
  • Treat it as the speed-and-volume option; when answers have to be right, step up to Omni Plus or a top-tier model.
Full-duplex enterprise evaluation
A full-duplex enterprise model you can trial, but it’s early-access evaluation software - not something to build a product on yet.
Score 38Price License ProprietaryTime to first audio n/a
  • The full-duplex design - listening and speaking at once - is its most interesting trait, and a trial endpoint lets you evaluate it directly.
  • For teams exploring where always-on, interruptible voice could go, it’s worth a look.
  • It’s proprietary early-access under an evaluation license, not a normal release, and real deployment expects H100-class infrastructure.
  • Overall quality also lands near the bottom here. For something you can actually ship today, almost everything above it is a safer bet.
Controllable local voice personas
The one genuinely local, open-weight pick with real persona and voice control - if you’ve got high-end hardware and don’t need strong reasoning.
Score 33Price License Open weightTime to first audio n/a
  • It’s open weight, full-duplex, and unusually good at conversational dynamics, with persona and voice conditioning you can actually steer.
  • If you want a private, customizable voice you run yourself and you have the GPU for it, nothing else here offers this mix.
  • Reasoning is weak, so it’s not for agents that need to think problems through, and official guidance targets A100/H100-class hardware, so “local” means a high-end rig, not a laptop.
  • For capability, any hosted leader is far ahead.
  • Run locally — If you have a high-end machine, you can run it with PersonaPlex after downloading weights from Hugging Face.

How to Choose

When choosing between these models, consider:
  • Access: First decide whether you want an app, an API, or a model you run yourself, because that changes cost, privacy, latency, and setup work. Most main picks are hosted services. Step-Audio R1.1 supports infrastructure-heavy self-hosting, PersonaPlex is the main high-end local option, and the weaker Moshi also runs on a typical machine.
  • Quality: We use the Artificial Analysis Speech to Speech benchmark suite as the main score - an equal-weighted look at speech reasoning, conversational dynamics, and agentic voice performance. It measures how good the model is, not how fast it responds.
  • Price: We compare USD per hour of input audio. We use Artificial Analysis’s calculated hourly cost where available; otherwise the value is the listed input-audio rate, so output charges may still apply.
  • Time to First Audio: This is how quickly audio starts, averaged across benchmark runs - not total call latency. Lower feels more human. A strong model can still start slowly (Qwen3.5 Omni Plus Realtime and Gemini 3.1 Flash Live both do), which matters a lot for snappy, interactive agents.

Other Models We Considered

GPT-Realtime-2.1 (OpenAI) — Current full-size OpenAI API model; test it beside GPT-Realtime-2.GPT-Realtime-2.1 mini (OpenAI) — Newer low-cost realtime option for lighter voice workloads.GPT-Live-1 (OpenAI) — Powers ChatGPT Voice for paid users, with no API to build on yet.GPT-Live-1 mini (OpenAI) — The free ChatGPT Voice model, also without an API yet.Gemini 2.5 Flash Native Audio Dialog Thinking (Google) — Stronger reasoning than newer Flash, but much slower to respond.Gemini 2.5 Flash Native Audio Dialog (Google) — Fast starts, but weaker reasoning than current options.Qwen3 Omni Realtime (Alibaba Cloud) — Earlier Qwen realtime model, now behind the 3.5 versions.Qwen3 Omni Flash (Alibaba Cloud) — Older, slower-starting Qwen option with weaker overall value.GPT-4o Realtime (OpenAI) — Legacy realtime model for existing builds, not new ones.GPT-4o mini Realtime (OpenAI) — Legacy mini model; newer Realtime mini tiers are easier picks.Moshi (Kyutai) — Runs on a normal machine, but far behind the hosted leaders on quality.

Frequently Asked Questions

GPT-Realtime-2. It’s the top scorer and the most reliable at complex, tool-driven conversations, so it’s the default recommendation for demanding production voice agents. The main reason not to use it is cost on very high-volume, simple traffic.
For most new builds, GPT-Realtime-2 if you want peak quality, or Gemini 3.1 Flash Live if you want strong reasoning at a much lower hourly price and can accept a slower start. For high-volume lightweight voice, GPT-Realtime mini or Qwen3.5 Omni Flash Realtime keep costs down.
Step-Audio R1.1. It competes with hosted leaders on reasoning under an Apache-2.0 license, and you can use it through first-party app and API routes. Just know that self-hosting it is infrastructure-heavy, not a local-machine task.
PersonaPlex is our main local pick, and it needs a high-end GPU (A100/H100-class), not a laptop. Moshi runs on a typical machine but is far weaker. For serious quality, a hosted model is still the better route.
Deepslate Opal starts talking faster than anything else on this list, which makes conversations feel close to instant. It’s only mid-tier on reasoning, though, so it’s best where responsiveness matters more than deep capability.
For quality, mostly yes - the score tracks how well a model reasons and holds a conversation. But it says nothing about speed. Always check Time to First Audio too, since a high-scoring model like Qwen3.5 Omni Plus Realtime can still feel sluggish in a live call.
They’re close. GPT-Realtime-2 edges ahead on overall quality and the hardest reasoning, while Grok Voice Think Fast 1.0 posts the strongest agentic, tool-using results. If your agent mainly takes actions and calls tools, test Think Fast; for the toughest reasoning, GPT-Realtime-2.
GPT-Realtime-2 for most cases - it’s more capable than both and much cheaper than GPT-Realtime. If you need lower latency or lower cost within the same realtime family, look at GPT-Realtime-1.5 for speed or GPT-Realtime mini for budget.