> ## Documentation Index
> Fetch the complete documentation index at: https://usefulai.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Best Speech-to-Speech Models in 2026

> Compare the best speech-to-speech models in 2026 for real-time voice agents by quality, first-audio latency, price, and local deployment.

<div className="uai-updated-row">Updated July 12, 2026</div>

Speech-to-speech models take audio in and talk back in real time - the engines behind voice agents and live assistants. The catch: the best-sounding, smartest models often aren't the fastest to respond, and price swings widely. We compared 15 on quality, speed, price, and access.

## Best Speech-to-Speech Models

<div className="uai-overview-table uai-overview-table--ranked">
  |  # | Model                                                                                                                                                                                                                                                           | Best for                             | Score <Tooltip tip="UsefulAI's 0-100 score uses the Artificial Analysis Speech to Speech benchmark suite across reasoning, conversation, and agentic performance."><span className="uai-tip-icon"><Icon icon="circle-info" size={12} color="currentColor" /><span className="uai-sr-only">About score</span></span></Tooltip> | Price <Tooltip tip="Hourly input-audio price for the represented route. Output audio, text tokens, tools, and session overhead can increase total cost."><span className="uai-tip-icon"><Icon icon="circle-info" size={12} color="currentColor" /><span className="uai-sr-only">About price</span></span></Tooltip> | License <Tooltip tip="Proprietary means no public model weights. Open weight means weights are available, though exact licenses and commercial-use terms vary."><span className="uai-tip-icon"><Icon icon="circle-info" size={12} color="currentColor" /><span className="uai-sr-only">About license</span></span></Tooltip> |
  | -: | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------ | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------: | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
  |  1 | <a href="https://developers.openai.com/api/docs/models/gpt-realtime-2" target="_blank" rel="noreferrer"><img src={"/images/icons/48/openai.com.png"} alt="" noZoom />GPT-Realtime-2</a>                                                                         | Complex production voice agents      |                                                                                                                                                                                                                                                                                                                           100 |                                                                                                                                                                                                                                                                                                           \$4.14/hr | <span className="uai-badge uai-badge--zinc">Proprietary</span>                                                                                                                                                                                                                                                               |
  |  2 | <a href="https://x.ai/news/grok-voice-think-fast-1" target="_blank" rel="noreferrer"><img src={"/images/icons/48/x.ai.png"} alt="" noZoom />Grok Voice Think Fast 1.0</a>                                                                                       | Tool-heavy phone agents              |                                                                                                                                                                                                                                                                                                                            97 |                                                                                                                                                                                                                                                                                                           \$3.00/hr | <span className="uai-badge uai-badge--zinc">Proprietary</span>                                                                                                                                                                                                                                                               |
  |  3 | <a href="https://www.linkedin.com/posts/alibaba-tongyi-lab_we-are-honored-to-share-that-our-fun-series-activity-7465748752973336576-o_tL" target="_blank" rel="noreferrer"><img src={"/images/icons/48/qwen.ai.png"} alt="" noZoom />Fun-Realtime-Audiochat</a> | High-quality model to watch          |                                                                                                                                                                                                                                                                                                                            96 |                                                                                                                                                                                                                                                                                                                 n/a | <span className="uai-badge uai-badge--zinc">Proprietary</span>                                                                                                                                                                                                                                                               |
  |  4 | <a href="https://developers.openai.com/api/docs/models/gpt-realtime-1.5" target="_blank" rel="noreferrer"><img src={"/images/icons/48/openai.com.png"} alt="" noZoom />GPT-Realtime-1.5</a>                                                                     | Fast, high-end realtime voice        |                                                                                                                                                                                                                                                                                                                            89 |                                                                                                                                                                                                                                                                                                          \$11.44/hr | <span className="uai-badge uai-badge--zinc">Proprietary</span>                                                                                                                                                                                                                                                               |
  |  5 | <a href="https://ai.google.dev/gemini-api/docs/models/gemini-3.1-flash-live-preview" target="_blank" rel="noreferrer"><img src={"/images/icons/48/google.com.png"} alt="" noZoom />Gemini 3.1 Flash Live</a>                                                    | Low-cost reasoning voice agents      |                                                                                                                                                                                                                                                                                                                            83 |                                                                                                                                                                                                                                                                                                           \$1.75/hr | <span className="uai-badge uai-badge--zinc">Proprietary</span>                                                                                                                                                                                                                                                               |
  |  6 | <a href="https://developers.openai.com/api/docs/models/gpt-realtime" target="_blank" rel="noreferrer"><img src={"/images/icons/48/openai.com.png"} alt="" noZoom />GPT-Realtime</a>                                                                             | Proven baseline realtime voice       |                                                                                                                                                                                                                                                                                                                            82 |                                                                                                                                                                                                                                                                                                          \$11.08/hr | <span className="uai-badge uai-badge--zinc">Proprietary</span>                                                                                                                                                                                                                                                               |
  |  7 | <a href="https://www.alibabacloud.com/help/en/model-studio/realtime" target="_blank" rel="noreferrer"><img src={"/images/icons/48/qwen.ai.png"} alt="" noZoom />Qwen3.5 Omni Plus Realtime</a>                                                                  | Hosted multilingual speech reasoning |                                                                                                                                                                                                                                                                                                                            82 |                                                                                                                                                                                                                                                                                                           \$0.16/hr | <span className="uai-badge uai-badge--zinc">Proprietary</span>                                                                                                                                                                                                                                                               |
  |  8 | <a href="https://huggingface.co/stepfun-ai/Step-Audio-R1.1" target="_blank" rel="noreferrer"><img src={"/images/icons/48/stepfun.ai.png"} alt="" noZoom />Step-Audio R1.1</a>                                                                                   | Open-weight speech reasoning         |                                                                                                                                                                                                                                                                                                                            81 |                                                                                                                                                                                                                                                                                                           \$0.06/hr | <span className="uai-badge uai-badge--emerald">Open weight</span>                                                                                                                                                                                                                                                            |
  |  9 | <a href="https://docs.aws.amazon.com/bedrock/latest/userguide/model-card-amazon-nova-2-sonic.html" target="_blank" rel="noreferrer"><img src={"/images/icons/48/aws.amazon.com.png"} alt="" noZoom />Amazon Nova 2 Sonic</a>                                    | Enterprise voice agents              |                                                                                                                                                                                                                                                                                                                            74 |                                                                                                                                                                                                                                                                                                           \$0.27/hr | <span className="uai-badge uai-badge--zinc">Proprietary</span>                                                                                                                                                                                                                                                               |
  | 10 | <a href="https://docs.x.ai/developers/model-capabilities/audio/voice-agent" target="_blank" rel="noreferrer"><img src={"/images/icons/48/x.ai.png"} alt="" noZoom />Grok Voice Agent</a>                                                                        | Fast everyday voice agents           |                                                                                                                                                                                                                                                                                                                            71 |                                                                                                                                                                                                                                                                                                           \$3.00/hr | <span className="uai-badge uai-badge--zinc">Proprietary</span>                                                                                                                                                                                                                                                               |
  | 11 | <a href="https://docs.deepslate.eu/opal" target="_blank" rel="noreferrer"><img src={"/images/icons/48/deepslate.eu.png"} alt="" noZoom />Deepslate Opal</a>                                                                                                     | Lowest-latency voice responses       |                                                                                                                                                                                                                                                                                                                            68 |                                                                                                                                                                                                                                                                                                           \$6.48/hr | <span className="uai-badge uai-badge--zinc">Proprietary</span>                                                                                                                                                                                                                                                               |
  | 12 | <a href="https://developers.openai.com/api/docs/models/gpt-realtime-mini" target="_blank" rel="noreferrer"><img src={"/images/icons/48/openai.com.png"} alt="" noZoom />GPT-Realtime mini</a>                                                                   | Fast, low-cost realtime chat         |                                                                                                                                                                                                                                                                                                                            58 |                                                                                                                                                                                                                                                                                                           \$3.04/hr | <span className="uai-badge uai-badge--zinc">Proprietary</span>                                                                                                                                                                                                                                                               |
  | 13 | <a href="https://www.alibabacloud.com/help/en/model-studio/realtime" target="_blank" rel="noreferrer"><img src={"/images/icons/48/qwen.ai.png"} alt="" noZoom />Qwen3.5 Omni Flash Realtime</a>                                                                 | Cheap high-volume voice agents       |                                                                                                                                                                                                                                                                                                                            53 |                                                                                                                                                                                                                                                                                                           \$0.16/hr | <span className="uai-badge uai-badge--zinc">Proprietary</span>                                                                                                                                                                                                                                                               |
  | 14 | <a href="https://build.nvidia.com/nvidia/nemotron-voicechat/modelcard" target="_blank" rel="noreferrer"><img src={"/images/icons/48/nvidia.com.png"} alt="" noZoom />Nemotron Voicechat</a>                                                                     | Full-duplex enterprise evaluation    |                                                                                                                                                                                                                                                                                                                            38 |                                                                                                                                                                                                                                                                                                                 n/a | <span className="uai-badge uai-badge--zinc">Proprietary</span>                                                                                                                                                                                                                                                               |
  | 15 | <a href="https://huggingface.co/nvidia/personaplex-7b-v1" target="_blank" rel="noreferrer"><img src={"/images/icons/48/nvidia.com.png"} alt="" noZoom />PersonaPlex</a>                                                                                         | Controllable local voice personas    |                                                                                                                                                                                                                                                                                                                            33 |                                                                                                                                                                                                                                                                                                                 n/a | <span className="uai-badge uai-badge--emerald">Open weight</span>                                                                                                                                                                                                                                                            |
</div>

<label className="uai-overview-more">
  <input type="checkbox" className="uai-overview-toggle" />

  <span className="uai-overview-more-open"><span className="uai-overview-more-count">Show more</span><Icon icon="chevron-down" size={13} /></span>
  <span className="uai-overview-more-close"><span>Show less</span><Icon icon="chevron-up" size={13} /></span>
</label>

***

<div className="uai-itemcard" role="article">
  <div className="uai-itemcard-head">
    <span className="uai-itemcard-icon">
      <img src="https://mintcdn.com/usefulai/Te6KzZ86-OxPuEC2/images/icons/144/openai.com.png?fit=max&auto=format&n=Te6KzZ86-OxPuEC2&q=85&s=745b8837f7535bc53cd70fc2f7024d58" alt="" noZoom loading="lazy" width="144" height="144" data-path="images/icons/144/openai.com.png" />
    </span>

    <div className="uai-itemcard-identity">
      <div className="uai-itemcard-row uai-itemcard-row--title">
        ## [GPT-Realtime-2](https://developers.openai.com/api/docs/models/gpt-realtime-2)

        <span className="uai-itemcard-byline">OpenAI</span>
      </div>

      <div className="uai-itemcard-row">
        <span className="uai-itemcard-note uai-itemcard-note--blue">Complex production voice agents</span>
      </div>
    </div>

    <div className="uai-itemcard-end">
      <a href="https://developers.openai.com/api/docs/models/gpt-realtime-2" target="_blank" rel="noreferrer" aria-label="Visit GPT-Realtime-2" className="uai-itemcard-cta uai-itemcard-cta--blue no-underline">Visit OpenAI</a>
    </div>
  </div>

  <div className="uai-itemcard-body">
    <div className="uai-itemcard-summary">
      The most capable speech-to-speech model in this comparison, and the one to beat for demanding, tool-driven voice agents that can't afford to drift.
    </div>

    <div className="uai-itemcard-facts" aria-label="GPT-Realtime-2 facts">
      <span>Score <strong>100</strong></span>
      <span>Price <strong>{"$4.14/hr"}</strong></span>
      <span>License <span className="uai-badge uai-badge--zinc">Proprietary</span></span>
      <span>Time to first audio <strong>1.14s</strong></span>
    </div>

    <div className="uai-itemcard-details-group">
      <div className="uai-itemcard-details">
        <input type="checkbox" id="speech-to-speech-gpt-realtime-2-strengths" className="uai-itemcard-details-toggle" />

        <label htmlFor="speech-to-speech-gpt-realtime-2-strengths"><span>Strengths</span><span className="uai-itemcard-details-chevron" /></label>

        <div className="uai-itemcard-details-body">
          <ul>
            <li>It leads on speech reasoning and conversational dynamics at once, so it follows multi-step instructions, handles interruptions cleanly, and stays coherent through long, messy calls.</li>
            <li>When the task is hard and the agent has to think, act, and talk without losing the thread, this is the pick.</li>
          </ul>
        </div>
      </div>

      <div className="uai-itemcard-details">
        <input type="checkbox" id="speech-to-speech-gpt-realtime-2-tradeoffs" className="uai-itemcard-details-toggle" />

        <label htmlFor="speech-to-speech-gpt-realtime-2-tradeoffs"><span>Tradeoffs</span><span className="uai-itemcard-details-chevron" /></label>

        <div className="uai-itemcard-details-body">
          <ul>
            <li>It's priced well above the cheaper realtime tiers, so high-volume, simple flows burn budget fast - GPT-Realtime mini or Qwen3.5 Omni Flash Realtime fit those better.</li>
            <li>And confirm you want this exact model, since the newer GPT-Realtime-2.1 is worth testing beside it.</li>
          </ul>
        </div>
      </div>

      <div className="uai-itemcard-details">
        <input type="checkbox" id="speech-to-speech-gpt-realtime-2-how-to-access" className="uai-itemcard-details-toggle" />

        <label htmlFor="speech-to-speech-gpt-realtime-2-how-to-access"><span>How to access</span><span className="uai-itemcard-details-chevron" /></label>

        <div className="uai-itemcard-details-body">
          <ul>
            <li><strong>API</strong> — Accessible via the <a href="https://developers.openai.com/api/docs/guides/realtime" target="_blank" rel="noreferrer" className="underline underline-offset-2">OpenAI Realtime API</a>.</li>
          </ul>
        </div>
      </div>
    </div>
  </div>
</div>

<div className="uai-itemcard" role="article">
  <div className="uai-itemcard-head">
    <span className="uai-itemcard-icon">
      <img src="https://mintcdn.com/usefulai/sV7VJe4pqO2Le0pu/images/icons/144/x.ai.png?fit=max&auto=format&n=sV7VJe4pqO2Le0pu&q=85&s=421f0adc6c1e753bf2ad0cd1661abbc7" alt="" noZoom loading="lazy" width="144" height="144" data-path="images/icons/144/x.ai.png" />
    </span>

    <div className="uai-itemcard-identity">
      <div className="uai-itemcard-row uai-itemcard-row--title">
        ## [Grok Voice Think Fast 1.0](https://x.ai/news/grok-voice-think-fast-1)

        <span className="uai-itemcard-byline">xAI</span>
      </div>

      <div className="uai-itemcard-row">
        <span className="uai-itemcard-note uai-itemcard-note--blue">Tool-heavy phone agents</span>
      </div>
    </div>

    <div className="uai-itemcard-end">
      <a href="https://x.ai/news/grok-voice-think-fast-1" target="_blank" rel="noreferrer" aria-label="Visit Grok Voice Think Fast 1.0" className="uai-itemcard-cta uai-itemcard-cta--blue no-underline">Visit xAI</a>
    </div>
  </div>

  <div className="uai-itemcard-body">
    <div className="uai-itemcard-summary">
      xAI's strongest voice model, and the one we'd reach for when an agent has to call tools and take real actions mid-conversation.
    </div>

    <div className="uai-itemcard-facts" aria-label="Grok Voice Think Fast 1.0 facts">
      <span>Score <strong>97</strong></span>
      <span>Price <strong>{"$3.00/hr"}</strong></span>
      <span>License <span className="uai-badge uai-badge--zinc">Proprietary</span></span>
      <span>Time to first audio <strong>1.25s</strong></span>
    </div>

    <div className="uai-itemcard-details-group">
      <div className="uai-itemcard-details">
        <input type="checkbox" id="speech-to-speech-grok-voice-think-fast-1-0-strengths" className="uai-itemcard-details-toggle" />

        <label htmlFor="speech-to-speech-grok-voice-think-fast-1-0-strengths"><span>Strengths</span><span className="uai-itemcard-details-chevron" /></label>

        <div className="uai-itemcard-details-body">
          <ul>
            <li>It posts the strongest agentic results in the suite, so it stays reliable when a call turns into actual work - looking things up, triggering functions, and pushing a task forward while still sounding natural.</li>
            <li>For phone agents that do more than chat, it's near the very top.</li>
          </ul>
        </div>
      </div>

      <div className="uai-itemcard-details">
        <input type="checkbox" id="speech-to-speech-grok-voice-think-fast-1-0-tradeoffs" className="uai-itemcard-details-toggle" />

        <label htmlFor="speech-to-speech-grok-voice-think-fast-1-0-tradeoffs"><span>Tradeoffs</span><span className="uai-itemcard-details-chevron" /></label>

        <div className="uai-itemcard-details-body">
          <ul>
            <li>It edges just behind the top scorer overall, so for the hardest reasoning you might still prefer GPT-Realtime-2.</li>
            <li>Keep it distinct from Grok Voice Agent, which starts faster but is noticeably weaker at both reasoning and tool use.</li>
          </ul>
        </div>
      </div>

      <div className="uai-itemcard-details">
        <input type="checkbox" id="speech-to-speech-grok-voice-think-fast-1-0-how-to-access" className="uai-itemcard-details-toggle" />

        <label htmlFor="speech-to-speech-grok-voice-think-fast-1-0-how-to-access"><span>How to access</span><span className="uai-itemcard-details-chevron" /></label>

        <div className="uai-itemcard-details-body">
          <ul>
            <li><strong>App</strong> — Available in <a href="https://x.ai/voice" target="_blank" rel="noreferrer" className="underline underline-offset-2">xAI Voice Agent Builder</a>.</li>
            <li><strong>API</strong> — Accessible via the <a href="https://docs.x.ai/developers/model-capabilities/audio/voice-agent" target="_blank" rel="noreferrer" className="underline underline-offset-2">xAI Voice Agent API</a>.</li>
          </ul>
        </div>
      </div>
    </div>
  </div>
</div>

<div className="uai-itemcard" role="article">
  <div className="uai-itemcard-head">
    <span className="uai-itemcard-icon">
      <img src="https://mintcdn.com/usefulai/Te6KzZ86-OxPuEC2/images/icons/144/qwen.ai.png?fit=max&auto=format&n=Te6KzZ86-OxPuEC2&q=85&s=77ec239e207f3d7865895d11af129b39" alt="" noZoom loading="lazy" width="144" height="144" data-path="images/icons/144/qwen.ai.png" />
    </span>

    <div className="uai-itemcard-identity">
      <div className="uai-itemcard-row uai-itemcard-row--title">
        ## [Fun-Realtime-Audiochat](https://www.linkedin.com/posts/alibaba-tongyi-lab_we-are-honored-to-share-that-our-fun-series-activity-7465748752973336576-o_tL)

        <span className="uai-itemcard-byline">Alibaba Cloud</span>
      </div>

      <div className="uai-itemcard-row">
        <span className="uai-itemcard-note uai-itemcard-note--blue">High-quality model to watch</span>
      </div>
    </div>

    <div className="uai-itemcard-end">
      <a href="https://www.linkedin.com/posts/alibaba-tongyi-lab_we-are-honored-to-share-that-our-fun-series-activity-7465748752973336576-o_tL" target="_blank" rel="noreferrer" aria-label="View Fun-Realtime-Audiochat on LinkedIn" className="uai-itemcard-cta uai-itemcard-cta--blue no-underline">View on LinkedIn</a>
    </div>
  </div>

  <div className="uai-itemcard-body">
    <div className="uai-itemcard-summary">
      A near-top quality result with no public deployment route, making this a model to monitor rather than one you can choose today.
    </div>

    <div className="uai-itemcard-facts" aria-label="Fun-Realtime-Audiochat facts">
      <span>Score <strong>96</strong></span>
      <span>Price <strong>{"n/a"}</strong></span>
      <span>License <span className="uai-badge uai-badge--zinc">Proprietary</span></span>
      <span>Time to first audio <strong>1.39s</strong></span>
    </div>

    <div className="uai-itemcard-details-group">
      <div className="uai-itemcard-details">
        <input type="checkbox" id="speech-to-speech-fun-realtime-audiochat-strengths" className="uai-itemcard-details-toggle" />

        <label htmlFor="speech-to-speech-fun-realtime-audiochat-strengths"><span>Strengths</span><span className="uai-itemcard-details-chevron" /></label>

        <div className="uai-itemcard-details-body">
          <ul>
            <li>In the benchmark data it's a genuine front-runner, matching the best on speech reasoning and natural back-and-forth.</li>
            <li>If Alibaba ships a documented public deployment, it could move straight into the top tier of models you'd actually build on.</li>
          </ul>
        </div>
      </div>

      <div className="uai-itemcard-details">
        <input type="checkbox" id="speech-to-speech-fun-realtime-audiochat-tradeoffs" className="uai-itemcard-details-toggle" />

        <label htmlFor="speech-to-speech-fun-realtime-audiochat-tradeoffs"><span>Tradeoffs</span><span className="uai-itemcard-details-chevron" /></label>

        <div className="uai-itemcard-details-body">
          <ul>
            <li>Right now there's no verified public API, price, or app for the exact scored model, so you can't ship it today.</li>
            <li>Treat it as a watch-list entry, and don't confuse it with the separate Fun-Audio-Chat project, which isn't the same model.</li>
          </ul>
        </div>
      </div>
    </div>
  </div>
</div>

<div className="uai-itemcard" role="article">
  <div className="uai-itemcard-head">
    <span className="uai-itemcard-icon">
      <img src="https://mintcdn.com/usefulai/Te6KzZ86-OxPuEC2/images/icons/144/openai.com.png?fit=max&auto=format&n=Te6KzZ86-OxPuEC2&q=85&s=745b8837f7535bc53cd70fc2f7024d58" alt="" noZoom loading="lazy" width="144" height="144" data-path="images/icons/144/openai.com.png" />
    </span>

    <div className="uai-itemcard-identity">
      <div className="uai-itemcard-row uai-itemcard-row--title">
        ## [GPT-Realtime-1.5](https://developers.openai.com/api/docs/models/gpt-realtime-1.5)

        <span className="uai-itemcard-byline">OpenAI</span>
      </div>

      <div className="uai-itemcard-row">
        <span className="uai-itemcard-note uai-itemcard-note--blue">Fast, high-end realtime voice</span>
      </div>
    </div>

    <div className="uai-itemcard-end">
      <a href="https://developers.openai.com/api/docs/models/gpt-realtime-1.5" target="_blank" rel="noreferrer" aria-label="Visit GPT-Realtime-1.5" className="uai-itemcard-cta uai-itemcard-cta--blue no-underline">Visit OpenAI</a>
    </div>
  </div>

  <div className="uai-itemcard-body">
    <div className="uai-itemcard-summary">
      Fast and strong, but priced at the top of this list - hard to justify for a new build when GPT-Realtime-2 is cheaper and better.
    </div>

    <div className="uai-itemcard-facts" aria-label="GPT-Realtime-1.5 facts">
      <span>Score <strong>89</strong></span>
      <span>Price <strong>{"$11.44/hr"}</strong></span>
      <span>License <span className="uai-badge uai-badge--zinc">Proprietary</span></span>
      <span>Time to first audio <strong>0.82s</strong></span>
    </div>

    <div className="uai-itemcard-details-group">
      <div className="uai-itemcard-details">
        <input type="checkbox" id="speech-to-speech-gpt-realtime-1-5-strengths" className="uai-itemcard-details-toggle" />

        <label htmlFor="speech-to-speech-gpt-realtime-1-5-strengths"><span>Strengths</span><span className="uai-itemcard-details-chevron" /></label>

        <div className="uai-itemcard-details-body">
          <ul>
            <li>It's among the highest-accuracy models here, and it starts talking quickly, so exchanges feel responsive without sacrificing much reasoning.</li>
            <li>As a low-latency, high-quality realtime voice model it holds up well on its own - the issue is what it costs, not what it does.</li>
          </ul>
        </div>
      </div>

      <div className="uai-itemcard-details">
        <input type="checkbox" id="speech-to-speech-gpt-realtime-1-5-tradeoffs" className="uai-itemcard-details-toggle" />

        <label htmlFor="speech-to-speech-gpt-realtime-1-5-tradeoffs"><span>Tradeoffs</span><span className="uai-itemcard-details-chevron" /></label>

        <div className="uai-itemcard-details-body">
          <ul>
            <li>The price is the problem: it sits at the top of the range while GPT-Realtime-2 scores higher and costs far less.</li>
            <li>For almost any new project, start with GPT-Realtime-2 instead; there's little reason to reach for 1.5.</li>
          </ul>
        </div>
      </div>

      <div className="uai-itemcard-details">
        <input type="checkbox" id="speech-to-speech-gpt-realtime-1-5-how-to-access" className="uai-itemcard-details-toggle" />

        <label htmlFor="speech-to-speech-gpt-realtime-1-5-how-to-access"><span>How to access</span><span className="uai-itemcard-details-chevron" /></label>

        <div className="uai-itemcard-details-body">
          <ul>
            <li><strong>API</strong> — Accessible via the <a href="https://developers.openai.com/api/docs/guides/realtime" target="_blank" rel="noreferrer" className="underline underline-offset-2">OpenAI Realtime API</a>.</li>
          </ul>
        </div>
      </div>
    </div>
  </div>
</div>

<div className="uai-itemcard" role="article">
  <div className="uai-itemcard-head">
    <span className="uai-itemcard-icon">
      <img src="https://mintcdn.com/usefulai/Ez-pJDkPpztLE7Cr/images/icons/144/google.com.png?fit=max&auto=format&n=Ez-pJDkPpztLE7Cr&q=85&s=11e0c725f841c45443116a184b63beac" alt="" noZoom loading="lazy" width="144" height="144" data-path="images/icons/144/google.com.png" />
    </span>

    <div className="uai-itemcard-identity">
      <div className="uai-itemcard-row uai-itemcard-row--title">
        ## [Gemini 3.1 Flash Live](https://ai.google.dev/gemini-api/docs/models/gemini-3.1-flash-live-preview)

        <span className="uai-itemcard-byline">Google</span>
      </div>

      <div className="uai-itemcard-row">
        <span className="uai-itemcard-note uai-itemcard-note--blue">Low-cost reasoning voice agents</span>
      </div>
    </div>

    <div className="uai-itemcard-end">
      <a href="https://ai.google.dev/gemini-api/docs/models/gemini-3.1-flash-live-preview" target="_blank" rel="noreferrer" aria-label="Visit Gemini 3.1 Flash Live" className="uai-itemcard-cta uai-itemcard-cta--blue no-underline">Visit Google</a>
    </div>
  </div>

  <div className="uai-itemcard-body">
    <div className="uai-itemcard-summary">
      Strong reasoning and native audio-to-audio at a genuinely low hourly price, held back mainly by how slowly it starts talking.
    </div>

    <div className="uai-itemcard-facts" aria-label="Gemini 3.1 Flash Live facts">
      <span>Score <strong>83</strong></span>
      <span>Price <strong>{"$1.75/hr"}</strong></span>
      <span>License <span className="uai-badge uai-badge--zinc">Proprietary</span></span>
      <span>Time to first audio <strong>2.98s</strong></span>
    </div>

    <div className="uai-itemcard-details-group">
      <div className="uai-itemcard-details">
        <input type="checkbox" id="speech-to-speech-gemini-3-1-flash-live-strengths" className="uai-itemcard-details-toggle" />

        <label htmlFor="speech-to-speech-gemini-3-1-flash-live-strengths"><span>Strengths</span><span className="uai-itemcard-details-chevron" /></label>

        <div className="uai-itemcard-details-body">
          <ul>
            <li>It pairs solid speech reasoning with native audio-to-audio, live tool use, and one of the lower price points among the capable models.</li>
            <li>For reasoning-heavy voice agents where you care more about answer quality and cost than instant response, it's a smart, affordable choice.</li>
          </ul>
        </div>
      </div>

      <div className="uai-itemcard-details">
        <input type="checkbox" id="speech-to-speech-gemini-3-1-flash-live-tradeoffs" className="uai-itemcard-details-toggle" />

        <label htmlFor="speech-to-speech-gemini-3-1-flash-live-tradeoffs"><span>Tradeoffs</span><span className="uai-itemcard-details-chevron" /></label>

        <div className="uai-itemcard-details-body">
          <ul>
            <li>Its weak spot is the slow first response - the high-quality configuration is among the laggiest here, which hurts on quick back-and-forth.</li>
            <li>If latency is your priority, Deepslate Opal or the fast OpenAI tiers feel far more immediate.</li>
          </ul>
        </div>
      </div>

      <div className="uai-itemcard-details">
        <input type="checkbox" id="speech-to-speech-gemini-3-1-flash-live-how-to-access" className="uai-itemcard-details-toggle" />

        <label htmlFor="speech-to-speech-gemini-3-1-flash-live-how-to-access"><span>How to access</span><span className="uai-itemcard-details-chevron" /></label>

        <div className="uai-itemcard-details-body">
          <ul>
            <li><strong>API</strong> — Accessible via the <a href="https://ai.google.dev/gemini-api/docs/live" target="_blank" rel="noreferrer" className="underline underline-offset-2">Gemini Live API</a>.</li>
          </ul>
        </div>
      </div>
    </div>
  </div>
</div>

<div className="uai-itemcard" role="article">
  <div className="uai-itemcard-head">
    <span className="uai-itemcard-icon">
      <img src="https://mintcdn.com/usefulai/Te6KzZ86-OxPuEC2/images/icons/144/openai.com.png?fit=max&auto=format&n=Te6KzZ86-OxPuEC2&q=85&s=745b8837f7535bc53cd70fc2f7024d58" alt="" noZoom loading="lazy" width="144" height="144" data-path="images/icons/144/openai.com.png" />
    </span>

    <div className="uai-itemcard-identity">
      <div className="uai-itemcard-row uai-itemcard-row--title">
        ## [GPT-Realtime](https://developers.openai.com/api/docs/models/gpt-realtime)

        <span className="uai-itemcard-byline">OpenAI</span>
      </div>

      <div className="uai-itemcard-row">
        <span className="uai-itemcard-note uai-itemcard-note--blue">Proven baseline realtime voice</span>
      </div>
    </div>

    <div className="uai-itemcard-end">
      <a href="https://developers.openai.com/api/docs/models/gpt-realtime" target="_blank" rel="noreferrer" aria-label="Visit GPT-Realtime" className="uai-itemcard-cta uai-itemcard-cta--blue no-underline">Visit OpenAI</a>
    </div>
  </div>

  <div className="uai-itemcard-body">
    <div className="uai-itemcard-summary">
      The familiar, sub-second realtime model many teams already know - still capable, but both pricier and weaker than the newer GPT-Realtime-2.
    </div>

    <div className="uai-itemcard-facts" aria-label="GPT-Realtime facts">
      <span>Score <strong>82</strong></span>
      <span>Price <strong>{"$11.08/hr"}</strong></span>
      <span>License <span className="uai-badge uai-badge--zinc">Proprietary</span></span>
      <span>Time to first audio <strong>0.98s</strong></span>
    </div>

    <div className="uai-itemcard-details-group">
      <div className="uai-itemcard-details">
        <input type="checkbox" id="speech-to-speech-gpt-realtime-strengths" className="uai-itemcard-details-toggle" />

        <label htmlFor="speech-to-speech-gpt-realtime-strengths"><span>Strengths</span><span className="uai-itemcard-details-chevron" /></label>

        <div className="uai-itemcard-details-body">
          <ul>
            <li>It responds in under a second and handles natural conversation reliably, which is why it became a common default for voice agents.</li>
            <li>Nothing about it is broken, and it stays a dependable, well-understood option for straightforward spoken interactions.</li>
          </ul>
        </div>
      </div>

      <div className="uai-itemcard-details">
        <input type="checkbox" id="speech-to-speech-gpt-realtime-tradeoffs" className="uai-itemcard-details-toggle" />

        <label htmlFor="speech-to-speech-gpt-realtime-tradeoffs"><span>Tradeoffs</span><span className="uai-itemcard-details-chevron" /></label>

        <div className="uai-itemcard-details-body">
          <ul>
            <li>It's been overtaken: GPT-Realtime-2 is more capable and much cheaper, and GPT-Realtime-1.5 answers faster.</li>
            <li>There's no strong reason to start a new build here - keep it only where it's already wired in and working.</li>
          </ul>
        </div>
      </div>

      <div className="uai-itemcard-details">
        <input type="checkbox" id="speech-to-speech-gpt-realtime-how-to-access" className="uai-itemcard-details-toggle" />

        <label htmlFor="speech-to-speech-gpt-realtime-how-to-access"><span>How to access</span><span className="uai-itemcard-details-chevron" /></label>

        <div className="uai-itemcard-details-body">
          <ul>
            <li><strong>API</strong> — Accessible via the <a href="https://developers.openai.com/api/docs/guides/realtime" target="_blank" rel="noreferrer" className="underline underline-offset-2">OpenAI Realtime API</a>.</li>
          </ul>
        </div>
      </div>
    </div>
  </div>
</div>

<div className="uai-itemcard" role="article">
  <div className="uai-itemcard-head">
    <span className="uai-itemcard-icon">
      <img src="https://mintcdn.com/usefulai/Te6KzZ86-OxPuEC2/images/icons/144/qwen.ai.png?fit=max&auto=format&n=Te6KzZ86-OxPuEC2&q=85&s=77ec239e207f3d7865895d11af129b39" alt="" noZoom loading="lazy" width="144" height="144" data-path="images/icons/144/qwen.ai.png" />
    </span>

    <div className="uai-itemcard-identity">
      <div className="uai-itemcard-row uai-itemcard-row--title">
        ## [Qwen3.5 Omni Plus Realtime](https://www.alibabacloud.com/help/en/model-studio/realtime)

        <span className="uai-itemcard-byline">Alibaba Cloud</span>
      </div>

      <div className="uai-itemcard-row">
        <span className="uai-itemcard-note uai-itemcard-note--blue">Hosted multilingual speech reasoning</span>
      </div>
    </div>

    <div className="uai-itemcard-end">
      <a href="https://www.alibabacloud.com/help/en/model-studio/realtime" target="_blank" rel="noreferrer" aria-label="Visit Qwen3.5 Omni Plus Realtime" className="uai-itemcard-cta uai-itemcard-cta--blue no-underline">Visit Alibaba Cloud</a>
    </div>
  </div>

  <div className="uai-itemcard-body">
    <div className="uai-itemcard-summary">
      The strongest speech-reasoner we've seen at a fraction of the top-tier price, as long as you can live with a slow first response.
    </div>

    <div className="uai-itemcard-facts" aria-label="Qwen3.5 Omni Plus Realtime facts">
      <span>Score <strong>82</strong></span>
      <span>Price <strong>{"$0.16/hr"}</strong></span>
      <span>License <span className="uai-badge uai-badge--zinc">Proprietary</span></span>
      <span>Time to first audio <strong>2.64s</strong></span>
    </div>

    <div className="uai-itemcard-details-group">
      <div className="uai-itemcard-details">
        <input type="checkbox" id="speech-to-speech-qwen3-5-omni-plus-realtime-strengths" className="uai-itemcard-details-toggle" />

        <label htmlFor="speech-to-speech-qwen3-5-omni-plus-realtime-strengths"><span>Strengths</span><span className="uai-itemcard-details-chevron" /></label>

        <div className="uai-itemcard-details-body">
          <ul>
            <li>It's excellent at reasoning out loud, with function calling, search, broad multilingual support, and clean interruption handling - and it's one of the cheapest capable models to run.</li>
            <li>For multilingual, reasoning-led voice work on a tight budget, it's hard to beat.</li>
          </ul>
        </div>
      </div>

      <div className="uai-itemcard-details">
        <input type="checkbox" id="speech-to-speech-qwen3-5-omni-plus-realtime-tradeoffs" className="uai-itemcard-details-toggle" />

        <label htmlFor="speech-to-speech-qwen3-5-omni-plus-realtime-tradeoffs"><span>Tradeoffs</span><span className="uai-itemcard-details-chevron" /></label>

        <div className="uai-itemcard-details-body">
          <ul>
            <li>It's slow to start talking, so it's a poor fit for snappy, interactive agents. The listed price also covers input audio only, not a full session, so real costs run higher.</li>
            <li>For low latency, look at the fast OpenAI or Qwen Flash tiers.</li>
          </ul>
        </div>
      </div>

      <div className="uai-itemcard-details">
        <input type="checkbox" id="speech-to-speech-qwen3-5-omni-plus-realtime-how-to-access" className="uai-itemcard-details-toggle" />

        <label htmlFor="speech-to-speech-qwen3-5-omni-plus-realtime-how-to-access"><span>How to access</span><span className="uai-itemcard-details-chevron" /></label>

        <div className="uai-itemcard-details-body">
          <ul>
            <li><strong>API</strong> — Accessible via <a href="https://www.alibabacloud.com/help/en/model-studio/realtime" target="_blank" rel="noreferrer" className="underline underline-offset-2">Alibaba Cloud Model Studio</a>.</li>
          </ul>
        </div>
      </div>
    </div>
  </div>
</div>

<div className="uai-itemcard" role="article">
  <div className="uai-itemcard-head">
    <span className="uai-itemcard-icon">
      <img src="https://mintcdn.com/usefulai/fTqNrv6I1z_YE97n/images/icons/144/stepfun.ai.png?fit=max&auto=format&n=fTqNrv6I1z_YE97n&q=85&s=903823390370a6da980e0c4246f8f6b4" alt="" noZoom loading="lazy" width="144" height="144" data-path="images/icons/144/stepfun.ai.png" />
    </span>

    <div className="uai-itemcard-identity">
      <div className="uai-itemcard-row uai-itemcard-row--title">
        ## [Step-Audio R1.1](https://huggingface.co/stepfun-ai/Step-Audio-R1.1)

        <span className="uai-itemcard-byline">StepFun</span>
      </div>

      <div className="uai-itemcard-row">
        <span className="uai-itemcard-note uai-itemcard-note--blue">Open-weight speech reasoning</span>
      </div>
    </div>

    <div className="uai-itemcard-end">
      <a href="https://huggingface.co/stepfun-ai/Step-Audio-R1.1" target="_blank" rel="noreferrer" aria-label="View Step-Audio R1.1 on Hugging Face" className="uai-itemcard-cta uai-itemcard-cta--blue no-underline">View on Hugging Face</a>
    </div>
  </div>

  <div className="uai-itemcard-body">
    <div className="uai-itemcard-summary">
      The standout open-weight pick - strong speech reasoning under an Apache-2.0 license, with first-party app and API routes if you'd rather not self-host.
    </div>

    <div className="uai-itemcard-facts" aria-label="Step-Audio R1.1 facts">
      <span>Score <strong>81</strong></span>
      <span>Price <strong>{"$0.06/hr"}</strong></span>
      <span>License <span className="uai-badge uai-badge--emerald">Open weight</span></span>
      <span>Time to first audio <strong>1.51s</strong></span>
    </div>

    <div className="uai-itemcard-details-group">
      <div className="uai-itemcard-details">
        <input type="checkbox" id="speech-to-speech-step-audio-r1-1-strengths" className="uai-itemcard-details-toggle" />

        <label htmlFor="speech-to-speech-step-audio-r1-1-strengths"><span>Strengths</span><span className="uai-itemcard-details-chevron" /></label>

        <div className="uai-itemcard-details-body">
          <ul>
            <li>It's the rare open-weight model that competes with hosted leaders on reasoning, and the permissive license lets you deploy it however your compliance needs dictate.</li>
            <li>First-party app and API routes mean you can use it immediately without standing up your own infrastructure.</li>
          </ul>
        </div>
      </div>

      <div className="uai-itemcard-details">
        <input type="checkbox" id="speech-to-speech-step-audio-r1-1-tradeoffs" className="uai-itemcard-details-toggle" />

        <label htmlFor="speech-to-speech-step-audio-r1-1-tradeoffs"><span>Tradeoffs</span><span className="uai-itemcard-details-chevron" /></label>

        <div className="uai-itemcard-details-body">
          <ul>
            <li>"Open weight" here doesn't mean easy - the official self-hosting path is infrastructure-heavy and multi-GPU, not a local-machine setup.</li>
            <li>If you want a voice you truly run yourself, PersonaPlex fits better; if you just want it hosted, the API route is the practical choice.</li>
          </ul>
        </div>
      </div>

      <div className="uai-itemcard-details">
        <input type="checkbox" id="speech-to-speech-step-audio-r1-1-how-to-access" className="uai-itemcard-details-toggle" />

        <label htmlFor="speech-to-speech-step-audio-r1-1-how-to-access"><span>How to access</span><span className="uai-itemcard-details-chevron" /></label>

        <div className="uai-itemcard-details-body">
          <ul>
            <li><strong>App</strong> — Available in <a href="https://www.stepfun.com/studio/audio" target="_blank" rel="noreferrer" className="underline underline-offset-2">StepFun Audio Studio</a>.</li>
            <li><strong>API</strong> — Accessible via the <a href="https://platform.stepfun.com/" target="_blank" rel="noreferrer" className="underline underline-offset-2">StepFun Open Platform</a>.</li>
            <li><strong>Run locally</strong> — Open weights are available from <a href="https://huggingface.co/stepfun-ai/Step-Audio-R1.1" target="_blank" rel="noreferrer" className="underline underline-offset-2">Hugging Face</a>, but in practice this needs self-hosting infrastructure, not a local machine.</li>
          </ul>
        </div>
      </div>
    </div>
  </div>
</div>

<div className="uai-itemcard" role="article">
  <div className="uai-itemcard-head">
    <span className="uai-itemcard-icon">
      <img src="https://mintcdn.com/usefulai/O2NMryn9hetXwNok/images/icons/144/aws.amazon.com.png?fit=max&auto=format&n=O2NMryn9hetXwNok&q=85&s=71c98e62e87ab4587dc368301e4ac9e0" alt="" noZoom loading="lazy" width="144" height="144" data-path="images/icons/144/aws.amazon.com.png" />
    </span>

    <div className="uai-itemcard-identity">
      <div className="uai-itemcard-row uai-itemcard-row--title">
        ## [Amazon Nova 2 Sonic](https://docs.aws.amazon.com/bedrock/latest/userguide/model-card-amazon-nova-2-sonic.html)

        <span className="uai-itemcard-byline">Amazon</span>
      </div>

      <div className="uai-itemcard-row">
        <span className="uai-itemcard-note uai-itemcard-note--blue">Enterprise voice agents</span>
      </div>
    </div>

    <div className="uai-itemcard-end">
      <a href="https://docs.aws.amazon.com/bedrock/latest/userguide/model-card-amazon-nova-2-sonic.html" target="_blank" rel="noreferrer" aria-label="Visit Amazon Nova 2 Sonic" className="uai-itemcard-cta uai-itemcard-cta--blue no-underline">Visit AWS</a>
    </div>
  </div>

  <div className="uai-itemcard-body">
    <div className="uai-itemcard-summary">
      A solid, mid-tier voice model built for production agents - streaming, tools, retrieval, and interruptions - though it trails the top scorers on raw quality.
    </div>

    <div className="uai-itemcard-facts" aria-label="Amazon Nova 2 Sonic facts">
      <span>Score <strong>74</strong></span>
      <span>Price <strong>{"$0.27/hr"}</strong></span>
      <span>License <span className="uai-badge uai-badge--zinc">Proprietary</span></span>
      <span>Time to first audio <strong>1.14s</strong></span>
    </div>

    <div className="uai-itemcard-details-group">
      <div className="uai-itemcard-details">
        <input type="checkbox" id="speech-to-speech-amazon-nova-2-sonic-strengths" className="uai-itemcard-details-toggle" />

        <label htmlFor="speech-to-speech-amazon-nova-2-sonic-strengths"><span>Strengths</span><span className="uai-itemcard-details-chevron" /></label>

        <div className="uai-itemcard-details-body">
          <ul>
            <li>It's built for real voice-agent work: low-latency streaming, tool use, retrieval, interruption handling, and multilingual support, all geared for production from the start.</li>
            <li>For teams that want a dependable, feature-complete agent model rather than the highest benchmark score, it delivers.</li>
          </ul>
        </div>
      </div>

      <div className="uai-itemcard-details">
        <input type="checkbox" id="speech-to-speech-amazon-nova-2-sonic-tradeoffs" className="uai-itemcard-details-toggle" />

        <label htmlFor="speech-to-speech-amazon-nova-2-sonic-tradeoffs"><span>Tradeoffs</span><span className="uai-itemcard-details-chevron" /></label>

        <div className="uai-itemcard-details-body">
          <ul>
            <li>On pure quality it sits mid-pack, behind GPT-Realtime-2, Grok Voice Think Fast 1.0, and the Gemini and Qwen reasoning models.</li>
            <li>The listed price covers input audio only, and long calls need a session-continuation pattern that adds engineering work.</li>
          </ul>
        </div>
      </div>

      <div className="uai-itemcard-details">
        <input type="checkbox" id="speech-to-speech-amazon-nova-2-sonic-how-to-access" className="uai-itemcard-details-toggle" />

        <label htmlFor="speech-to-speech-amazon-nova-2-sonic-how-to-access"><span>How to access</span><span className="uai-itemcard-details-chevron" /></label>

        <div className="uai-itemcard-details-body">
          <ul>
            <li><strong>API</strong> — Accessible via <a href="https://docs.aws.amazon.com/bedrock/latest/userguide/model-card-amazon-nova-2-sonic.html" target="_blank" rel="noreferrer" className="underline underline-offset-2">Amazon Bedrock</a>.</li>
          </ul>
        </div>
      </div>
    </div>
  </div>
</div>

<div className="uai-itemcard" role="article">
  <div className="uai-itemcard-head">
    <span className="uai-itemcard-icon">
      <img src="https://mintcdn.com/usefulai/sV7VJe4pqO2Le0pu/images/icons/144/x.ai.png?fit=max&auto=format&n=sV7VJe4pqO2Le0pu&q=85&s=421f0adc6c1e753bf2ad0cd1661abbc7" alt="" noZoom loading="lazy" width="144" height="144" data-path="images/icons/144/x.ai.png" />
    </span>

    <div className="uai-itemcard-identity">
      <div className="uai-itemcard-row uai-itemcard-row--title">
        ## [Grok Voice Agent](https://docs.x.ai/developers/model-capabilities/audio/voice-agent)

        <span className="uai-itemcard-byline">xAI</span>
      </div>

      <div className="uai-itemcard-row">
        <span className="uai-itemcard-note uai-itemcard-note--blue">Fast everyday voice agents</span>
      </div>
    </div>

    <div className="uai-itemcard-end">
      <a href="https://docs.x.ai/developers/model-capabilities/audio/voice-agent" target="_blank" rel="noreferrer" aria-label="Visit Grok Voice Agent" className="uai-itemcard-cta uai-itemcard-cta--blue no-underline">Visit xAI</a>
    </div>
  </div>

  <div className="uai-itemcard-body">
    <div className="uai-itemcard-summary">
      xAI's quick, practical voice agent - sub-second responses and an easy build path, but clearly a step below its Think Fast sibling on quality.
    </div>

    <div className="uai-itemcard-facts" aria-label="Grok Voice Agent facts">
      <span>Score <strong>71</strong></span>
      <span>Price <strong>{"$3.00/hr"}</strong></span>
      <span>License <span className="uai-badge uai-badge--zinc">Proprietary</span></span>
      <span>Time to first audio <strong>0.78s</strong></span>
    </div>

    <div className="uai-itemcard-details-group">
      <div className="uai-itemcard-details">
        <input type="checkbox" id="speech-to-speech-grok-voice-agent-strengths" className="uai-itemcard-details-toggle" />

        <label htmlFor="speech-to-speech-grok-voice-agent-strengths"><span>Strengths</span><span className="uai-itemcard-details-chevron" /></label>

        <div className="uai-itemcard-details-body">
          <ul>
            <li>It answers fast, near the quickest here, and comes with a straightforward builder and API, so you can stand up a responsive voice agent without much fuss.</li>
            <li>For everyday, latency-sensitive assistants that don't need frontier reasoning, it's a reasonable pick.</li>
          </ul>
        </div>
      </div>

      <div className="uai-itemcard-details">
        <input type="checkbox" id="speech-to-speech-grok-voice-agent-tradeoffs" className="uai-itemcard-details-toggle" />

        <label htmlFor="speech-to-speech-grok-voice-agent-tradeoffs"><span>Tradeoffs</span><span className="uai-itemcard-details-chevron" /></label>

        <div className="uai-itemcard-details-body">
          <ul>
            <li>It's meaningfully weaker than Grok Voice Think Fast 1.0 on both reasoning and tool use, so don't mix the two up.</li>
            <li>If your agent does real work mid-call, step up to Think Fast; if you only need speed, other fast tiers compete on price.</li>
          </ul>
        </div>
      </div>

      <div className="uai-itemcard-details">
        <input type="checkbox" id="speech-to-speech-grok-voice-agent-how-to-access" className="uai-itemcard-details-toggle" />

        <label htmlFor="speech-to-speech-grok-voice-agent-how-to-access"><span>How to access</span><span className="uai-itemcard-details-chevron" /></label>

        <div className="uai-itemcard-details-body">
          <ul>
            <li><strong>App</strong> — Available in <a href="https://x.ai/voice" target="_blank" rel="noreferrer" className="underline underline-offset-2">xAI Voice Agent Builder</a>.</li>
            <li><strong>API</strong> — Accessible via the <a href="https://docs.x.ai/developers/model-capabilities/audio/voice-agent" target="_blank" rel="noreferrer" className="underline underline-offset-2">xAI Voice Agent API</a>.</li>
          </ul>
        </div>
      </div>
    </div>
  </div>
</div>

<div className="uai-itemcard" role="article">
  <div className="uai-itemcard-head">
    <span className="uai-itemcard-icon">
      <img src="https://mintcdn.com/usefulai/Ez-pJDkPpztLE7Cr/images/icons/144/deepslate.eu.png?fit=max&auto=format&n=Ez-pJDkPpztLE7Cr&q=85&s=edc64b92879abd432eeb801a21a53822" alt="" noZoom loading="lazy" width="144" height="144" data-path="images/icons/144/deepslate.eu.png" />
    </span>

    <div className="uai-itemcard-identity">
      <div className="uai-itemcard-row uai-itemcard-row--title">
        ## [Deepslate Opal](https://docs.deepslate.eu/opal)

        <span className="uai-itemcard-byline">Deepslate</span>
      </div>

      <div className="uai-itemcard-row">
        <span className="uai-itemcard-note uai-itemcard-note--blue">Lowest-latency voice responses</span>
      </div>
    </div>

    <div className="uai-itemcard-end">
      <a href="https://docs.deepslate.eu/opal" target="_blank" rel="noreferrer" aria-label="Visit Deepslate Opal" className="uai-itemcard-cta uai-itemcard-cta--blue no-underline">Visit Deepslate</a>
    </div>
  </div>

  <div className="uai-itemcard-body">
    <div className="uai-itemcard-summary">
      The fastest model here by a clear margin on first response, with EU hosting and flexible integration routes, but only mid-tier on quality.
    </div>

    <div className="uai-itemcard-facts" aria-label="Deepslate Opal facts">
      <span>Score <strong>68</strong></span>
      <span>Price <strong>{"$6.48/hr"}</strong></span>
      <span>License <span className="uai-badge uai-badge--zinc">Proprietary</span></span>
      <span>Time to first audio <strong>0.44s</strong></span>
    </div>

    <div className="uai-itemcard-details-group">
      <div className="uai-itemcard-details">
        <input type="checkbox" id="speech-to-speech-deepslate-opal-strengths" className="uai-itemcard-details-toggle" />

        <label htmlFor="speech-to-speech-deepslate-opal-strengths"><span>Strengths</span><span className="uai-itemcard-details-chevron" /></label>

        <div className="uai-itemcard-details-body">
          <ul>
            <li>Nothing else starts talking as quickly, so conversations feel genuinely instant - the closest to human turn-taking in this group.</li>
            <li>REST, WebSocket, and SIP routes plus EU-based hosting make it easy to slot into telephony and privacy-sensitive setups.</li>
          </ul>
        </div>
      </div>

      <div className="uai-itemcard-details">
        <input type="checkbox" id="speech-to-speech-deepslate-opal-tradeoffs" className="uai-itemcard-details-toggle" />

        <label htmlFor="speech-to-speech-deepslate-opal-tradeoffs"><span>Tradeoffs</span><span className="uai-itemcard-details-chevron" /></label>

        <div className="uai-itemcard-details-body">
          <ul>
            <li>That speed comes with only middling reasoning and weaker agentic performance, so it's not the one for complex, tool-heavy tasks.</li>
            <li>The ecosystem is smaller and public pricing is thin. For more capability at similar latency, weigh the fast OpenAI tiers.</li>
          </ul>
        </div>
      </div>

      <div className="uai-itemcard-details">
        <input type="checkbox" id="speech-to-speech-deepslate-opal-how-to-access" className="uai-itemcard-details-toggle" />

        <label htmlFor="speech-to-speech-deepslate-opal-how-to-access"><span>How to access</span><span className="uai-itemcard-details-chevron" /></label>

        <div className="uai-itemcard-details-body">
          <ul>
            <li><strong>App</strong> — Available in <a href="https://docs.deepslate.eu/opal" target="_blank" rel="noreferrer" className="underline underline-offset-2">Deepslate Assistants and Agents</a>.</li>
            <li><strong>API</strong> — Accessible via <a href="https://deepslate.eu/" target="_blank" rel="noreferrer" className="underline underline-offset-2">Deepslate REST, WebSocket, and SIP integrations</a>.</li>
          </ul>
        </div>
      </div>
    </div>
  </div>
</div>

<div className="uai-itemcard" role="article">
  <div className="uai-itemcard-head">
    <span className="uai-itemcard-icon">
      <img src="https://mintcdn.com/usefulai/Te6KzZ86-OxPuEC2/images/icons/144/openai.com.png?fit=max&auto=format&n=Te6KzZ86-OxPuEC2&q=85&s=745b8837f7535bc53cd70fc2f7024d58" alt="" noZoom loading="lazy" width="144" height="144" data-path="images/icons/144/openai.com.png" />
    </span>

    <div className="uai-itemcard-identity">
      <div className="uai-itemcard-row uai-itemcard-row--title">
        ## [GPT-Realtime mini](https://developers.openai.com/api/docs/models/gpt-realtime-mini)

        <span className="uai-itemcard-byline">OpenAI</span>
      </div>

      <div className="uai-itemcard-row">
        <span className="uai-itemcard-note uai-itemcard-note--blue">Fast, low-cost realtime chat</span>
      </div>
    </div>

    <div className="uai-itemcard-end">
      <a href="https://developers.openai.com/api/docs/models/gpt-realtime-mini" target="_blank" rel="noreferrer" aria-label="Visit GPT-Realtime mini" className="uai-itemcard-cta uai-itemcard-cta--blue no-underline">Visit OpenAI</a>
    </div>
  </div>

  <div className="uai-itemcard-body">
    <div className="uai-itemcard-summary">
      The budget-friendly, low-latency OpenAI realtime option - great at natural conversation, but a real step down in reasoning and tool use.
    </div>

    <div className="uai-itemcard-facts" aria-label="GPT-Realtime mini facts">
      <span>Score <strong>58</strong></span>
      <span>Price <strong>{"$3.04/hr"}</strong></span>
      <span>License <span className="uai-badge uai-badge--zinc">Proprietary</span></span>
      <span>Time to first audio <strong>0.81s</strong></span>
    </div>

    <div className="uai-itemcard-details-group">
      <div className="uai-itemcard-details">
        <input type="checkbox" id="speech-to-speech-gpt-realtime-mini-strengths" className="uai-itemcard-details-toggle" />

        <label htmlFor="speech-to-speech-gpt-realtime-mini-strengths"><span>Strengths</span><span className="uai-itemcard-details-chevron" /></label>

        <div className="uai-itemcard-details-body">
          <ul>
            <li>It's quick to respond and handles everyday back-and-forth smoothly at a lower cost than the flagship.</li>
            <li>For high-volume, lightweight voice - simple Q\&A, routing, casual assistants - it's an efficient workhorse that keeps conversations feeling natural.</li>
          </ul>
        </div>
      </div>

      <div className="uai-itemcard-details">
        <input type="checkbox" id="speech-to-speech-gpt-realtime-mini-tradeoffs" className="uai-itemcard-details-toggle" />

        <label htmlFor="speech-to-speech-gpt-realtime-mini-tradeoffs"><span>Tradeoffs</span><span className="uai-itemcard-details-chevron" /></label>

        <div className="uai-itemcard-details-body">
          <ul>
            <li>Push it toward multi-step reasoning or serious tool use and it falls well short of GPT-Realtime-2 and the reasoning-led models.</li>
            <li>Note it's the older mini - GPT-Realtime-2.1 mini is a newer, distinct option worth testing before you commit.</li>
          </ul>
        </div>
      </div>

      <div className="uai-itemcard-details">
        <input type="checkbox" id="speech-to-speech-gpt-realtime-mini-how-to-access" className="uai-itemcard-details-toggle" />

        <label htmlFor="speech-to-speech-gpt-realtime-mini-how-to-access"><span>How to access</span><span className="uai-itemcard-details-chevron" /></label>

        <div className="uai-itemcard-details-body">
          <ul>
            <li><strong>API</strong> — Accessible via the <a href="https://developers.openai.com/api/docs/guides/realtime" target="_blank" rel="noreferrer" className="underline underline-offset-2">OpenAI Realtime API</a>.</li>
          </ul>
        </div>
      </div>
    </div>
  </div>
</div>

<div className="uai-itemcard" role="article">
  <div className="uai-itemcard-head">
    <span className="uai-itemcard-icon">
      <img src="https://mintcdn.com/usefulai/Te6KzZ86-OxPuEC2/images/icons/144/qwen.ai.png?fit=max&auto=format&n=Te6KzZ86-OxPuEC2&q=85&s=77ec239e207f3d7865895d11af129b39" alt="" noZoom loading="lazy" width="144" height="144" data-path="images/icons/144/qwen.ai.png" />
    </span>

    <div className="uai-itemcard-identity">
      <div className="uai-itemcard-row uai-itemcard-row--title">
        ## [Qwen3.5 Omni Flash Realtime](https://www.alibabacloud.com/help/en/model-studio/realtime)

        <span className="uai-itemcard-byline">Alibaba Cloud</span>
      </div>

      <div className="uai-itemcard-row">
        <span className="uai-itemcard-note uai-itemcard-note--blue">Cheap high-volume voice agents</span>
      </div>
    </div>

    <div className="uai-itemcard-end">
      <a href="https://www.alibabacloud.com/help/en/model-studio/realtime" target="_blank" rel="noreferrer" aria-label="Visit Qwen3.5 Omni Flash Realtime" className="uai-itemcard-cta uai-itemcard-cta--blue no-underline">Visit Alibaba Cloud</a>
    </div>
  </div>

  <div className="uai-itemcard-body">
    <div className="uai-itemcard-summary">
      The speed-and-value play from the Qwen line - very cheap, quick to respond, and multilingual, but noticeably weaker at reasoning than Omni Plus.
    </div>

    <div className="uai-itemcard-facts" aria-label="Qwen3.5 Omni Flash Realtime facts">
      <span>Score <strong>53</strong></span>
      <span>Price <strong>{"$0.16/hr"}</strong></span>
      <span>License <span className="uai-badge uai-badge--zinc">Proprietary</span></span>
      <span>Time to first audio <strong>0.79s</strong></span>
    </div>

    <div className="uai-itemcard-details-group">
      <div className="uai-itemcard-details">
        <input type="checkbox" id="speech-to-speech-qwen3-5-omni-flash-realtime-strengths" className="uai-itemcard-details-toggle" />

        <label htmlFor="speech-to-speech-qwen3-5-omni-flash-realtime-strengths"><span>Strengths</span><span className="uai-itemcard-details-chevron" /></label>

        <div className="uai-itemcard-details-body">
          <ul>
            <li>It's among the cheapest models here and starts talking fast, with broad language coverage.</li>
            <li>For high-volume, cost-sensitive voice where you need many concurrent sessions more than deep reasoning, it stretches a budget further than almost anything else on this list.</li>
          </ul>
        </div>
      </div>

      <div className="uai-itemcard-details">
        <input type="checkbox" id="speech-to-speech-qwen3-5-omni-flash-realtime-tradeoffs" className="uai-itemcard-details-toggle" />

        <label htmlFor="speech-to-speech-qwen3-5-omni-flash-realtime-tradeoffs"><span>Tradeoffs</span><span className="uai-itemcard-details-chevron" /></label>

        <div className="uai-itemcard-details-body">
          <ul>
            <li>Its reasoning is well behind Qwen3.5 Omni Plus Realtime, so it's the wrong tool for complex, multi-step conversations.</li>
            <li>Treat it as the speed-and-volume option; when answers have to be right, step up to Omni Plus or a top-tier model.</li>
          </ul>
        </div>
      </div>

      <div className="uai-itemcard-details">
        <input type="checkbox" id="speech-to-speech-qwen3-5-omni-flash-realtime-how-to-access" className="uai-itemcard-details-toggle" />

        <label htmlFor="speech-to-speech-qwen3-5-omni-flash-realtime-how-to-access"><span>How to access</span><span className="uai-itemcard-details-chevron" /></label>

        <div className="uai-itemcard-details-body">
          <ul>
            <li><strong>API</strong> — Accessible via <a href="https://www.alibabacloud.com/help/en/model-studio/realtime" target="_blank" rel="noreferrer" className="underline underline-offset-2">Alibaba Cloud Model Studio</a>.</li>
          </ul>
        </div>
      </div>
    </div>
  </div>
</div>

<div className="uai-itemcard" role="article">
  <div className="uai-itemcard-head">
    <span className="uai-itemcard-icon">
      <img src="https://mintcdn.com/usefulai/Te6KzZ86-OxPuEC2/images/icons/144/nvidia.com.png?fit=max&auto=format&n=Te6KzZ86-OxPuEC2&q=85&s=da3a69cf52131d76f5aa096d31f1a718" alt="" noZoom loading="lazy" width="144" height="144" data-path="images/icons/144/nvidia.com.png" />
    </span>

    <div className="uai-itemcard-identity">
      <div className="uai-itemcard-row uai-itemcard-row--title">
        ## [Nemotron Voicechat](https://build.nvidia.com/nvidia/nemotron-voicechat/modelcard)

        <span className="uai-itemcard-byline">NVIDIA</span>
      </div>

      <div className="uai-itemcard-row">
        <span className="uai-itemcard-note uai-itemcard-note--blue">Full-duplex enterprise evaluation</span>
      </div>
    </div>

    <div className="uai-itemcard-end">
      <a href="https://build.nvidia.com/nvidia/nemotron-voicechat/modelcard" target="_blank" rel="noreferrer" aria-label="Visit Nemotron Voicechat" className="uai-itemcard-cta uai-itemcard-cta--blue no-underline">Visit NVIDIA</a>
    </div>
  </div>

  <div className="uai-itemcard-body">
    <div className="uai-itemcard-summary">
      A full-duplex enterprise model you can trial, but it's early-access evaluation software - not something to build a product on yet.
    </div>

    <div className="uai-itemcard-facts" aria-label="Nemotron Voicechat facts">
      <span>Score <strong>38</strong></span>
      <span>Price <strong>{"n/a"}</strong></span>
      <span>License <span className="uai-badge uai-badge--zinc">Proprietary</span></span>
      <span>Time to first audio <strong>n/a</strong></span>
    </div>

    <div className="uai-itemcard-details-group">
      <div className="uai-itemcard-details">
        <input type="checkbox" id="speech-to-speech-nemotron-voicechat-strengths" className="uai-itemcard-details-toggle" />

        <label htmlFor="speech-to-speech-nemotron-voicechat-strengths"><span>Strengths</span><span className="uai-itemcard-details-chevron" /></label>

        <div className="uai-itemcard-details-body">
          <ul>
            <li>The full-duplex design - listening and speaking at once - is its most interesting trait, and a trial endpoint lets you evaluate it directly.</li>
            <li>For teams exploring where always-on, interruptible voice could go, it's worth a look.</li>
          </ul>
        </div>
      </div>

      <div className="uai-itemcard-details">
        <input type="checkbox" id="speech-to-speech-nemotron-voicechat-tradeoffs" className="uai-itemcard-details-toggle" />

        <label htmlFor="speech-to-speech-nemotron-voicechat-tradeoffs"><span>Tradeoffs</span><span className="uai-itemcard-details-chevron" /></label>

        <div className="uai-itemcard-details-body">
          <ul>
            <li>It's proprietary early-access under an evaluation license, not a normal release, and real deployment expects H100-class infrastructure.</li>
            <li>Overall quality also lands near the bottom here. For something you can actually ship today, almost everything above it is a safer bet.</li>
          </ul>
        </div>
      </div>

      <div className="uai-itemcard-details">
        <input type="checkbox" id="speech-to-speech-nemotron-voicechat-how-to-access" className="uai-itemcard-details-toggle" />

        <label htmlFor="speech-to-speech-nemotron-voicechat-how-to-access"><span>How to access</span><span className="uai-itemcard-details-chevron" /></label>

        <div className="uai-itemcard-details-body">
          <ul>
            <li><strong>API</strong> — Accessible via the <a href="https://build.nvidia.com/nvidia/nemotron-voicechat/modelcard" target="_blank" rel="noreferrer" className="underline underline-offset-2">NVIDIA Build trial endpoint</a>.</li>
            <li><strong>Run locally</strong> — Qualified model and NIM access are available through <a href="https://developer.nvidia.com/nemotron-voicechat-early-access" target="_blank" rel="noreferrer" className="underline underline-offset-2">NVIDIA Early Access</a>, but in practice this needs self-hosting infrastructure, not a local machine.</li>
          </ul>
        </div>
      </div>
    </div>
  </div>
</div>

<div className="uai-itemcard" role="article">
  <div className="uai-itemcard-head">
    <span className="uai-itemcard-icon">
      <img src="https://mintcdn.com/usefulai/Te6KzZ86-OxPuEC2/images/icons/144/nvidia.com.png?fit=max&auto=format&n=Te6KzZ86-OxPuEC2&q=85&s=da3a69cf52131d76f5aa096d31f1a718" alt="" noZoom loading="lazy" width="144" height="144" data-path="images/icons/144/nvidia.com.png" />
    </span>

    <div className="uai-itemcard-identity">
      <div className="uai-itemcard-row uai-itemcard-row--title">
        ## [PersonaPlex](https://huggingface.co/nvidia/personaplex-7b-v1)

        <span className="uai-itemcard-byline">NVIDIA</span>
      </div>

      <div className="uai-itemcard-row">
        <span className="uai-itemcard-note uai-itemcard-note--blue">Controllable local voice personas</span>
      </div>
    </div>

    <div className="uai-itemcard-end">
      <a href="https://huggingface.co/nvidia/personaplex-7b-v1" target="_blank" rel="noreferrer" aria-label="View PersonaPlex on Hugging Face" className="uai-itemcard-cta uai-itemcard-cta--blue no-underline">View on Hugging Face</a>
    </div>
  </div>

  <div className="uai-itemcard-body">
    <div className="uai-itemcard-summary">
      The one genuinely local, open-weight pick with real persona and voice control - if you've got high-end hardware and don't need strong reasoning.
    </div>

    <div className="uai-itemcard-facts" aria-label="PersonaPlex facts">
      <span>Score <strong>33</strong></span>
      <span>Price <strong>{"n/a"}</strong></span>
      <span>License <span className="uai-badge uai-badge--emerald">Open weight</span></span>
      <span>Time to first audio <strong>n/a</strong></span>
    </div>

    <div className="uai-itemcard-details-group">
      <div className="uai-itemcard-details">
        <input type="checkbox" id="speech-to-speech-personaplex-strengths" className="uai-itemcard-details-toggle" />

        <label htmlFor="speech-to-speech-personaplex-strengths"><span>Strengths</span><span className="uai-itemcard-details-chevron" /></label>

        <div className="uai-itemcard-details-body">
          <ul>
            <li>It's open weight, full-duplex, and unusually good at conversational dynamics, with persona and voice conditioning you can actually steer.</li>
            <li>If you want a private, customizable voice you run yourself and you have the GPU for it, nothing else here offers this mix.</li>
          </ul>
        </div>
      </div>

      <div className="uai-itemcard-details">
        <input type="checkbox" id="speech-to-speech-personaplex-tradeoffs" className="uai-itemcard-details-toggle" />

        <label htmlFor="speech-to-speech-personaplex-tradeoffs"><span>Tradeoffs</span><span className="uai-itemcard-details-chevron" /></label>

        <div className="uai-itemcard-details-body">
          <ul>
            <li>Reasoning is weak, so it's not for agents that need to think problems through, and official guidance targets A100/H100-class hardware, so "local" means a high-end rig, not a laptop.</li>
            <li>For capability, any hosted leader is far ahead.</li>
          </ul>
        </div>
      </div>

      <div className="uai-itemcard-details">
        <input type="checkbox" id="speech-to-speech-personaplex-how-to-access" className="uai-itemcard-details-toggle" />

        <label htmlFor="speech-to-speech-personaplex-how-to-access"><span>How to access</span><span className="uai-itemcard-details-chevron" /></label>

        <div className="uai-itemcard-details-body">
          <ul>
            <li><strong>Run locally</strong> — If you have a high-end machine, you can run it with <a href="https://github.com/NVIDIA/PersonaPlex" target="_blank" rel="noreferrer" className="underline underline-offset-2">PersonaPlex</a> after downloading weights from <a href="https://huggingface.co/nvidia/personaplex-7b-v1" target="_blank" rel="noreferrer" className="underline underline-offset-2">Hugging Face</a>.</li>
          </ul>
        </div>
      </div>
    </div>
  </div>
</div>

***

## How to Choose

When choosing between these models, consider:

* **Access:** First decide whether you want an app, an API, or a model you run yourself, because that changes cost, privacy, latency, and setup work. Most main picks are hosted services. Step-Audio R1.1 supports infrastructure-heavy self-hosting, PersonaPlex is the main high-end local option, and the weaker Moshi also runs on a typical machine.
* **Quality:** We use the Artificial Analysis Speech to Speech benchmark suite as the main score - an equal-weighted look at speech reasoning, conversational dynamics, and agentic voice performance. It measures how good the model is, not how fast it responds.
* **Price:** We compare USD per hour of input audio. We use Artificial Analysis's calculated hourly cost where available; otherwise the value is the listed input-audio rate, so output charges may still apply.
* **Time to First Audio:** This is how quickly audio starts, averaged across benchmark runs - not total call latency. Lower feels more human. A strong model can still start slowly (Qwen3.5 Omni Plus Realtime and Gemini 3.1 Flash Live both do), which matters a lot for snappy, interactive agents.

***

## Other Models We Considered

<div className="not-prose my-4 flex flex-col gap-1.5 uai-article-prose text-zinc-700 dark:text-zinc-300">
  <span className="flex items-baseline gap-2.5">
    <span aria-hidden="true" className="relative -top-0.5 inline-block h-1.5 w-1.5 shrink-0 rounded-full bg-zinc-300 dark:bg-zinc-600" />

    <span><img src="https://mintcdn.com/usefulai/C5xOaAf4Os-Vu41o/images/icons/48/openai.com.png?fit=max&auto=format&n=C5xOaAf4Os-Vu41o&q=85&s=21ac965dc6ee5751127e6d435043629b" alt="" noZoom className="relative -top-px mr-1 inline h-4 w-4 rounded-sm object-contain" width="48" height="48" data-path="images/icons/48/openai.com.png" /><a href="https://developers.openai.com/api/docs/models/gpt-realtime-2.1" target="_blank" rel="noreferrer" className="font-medium text-zinc-950 underline underline-offset-2 dark:text-white">GPT-Realtime-2.1</a> <span className="uai-ink-muted">(OpenAI)</span> — Current full-size OpenAI API model; test it beside GPT-Realtime-2.</span>
  </span>

  <span className="flex items-baseline gap-2.5">
    <span aria-hidden="true" className="relative -top-0.5 inline-block h-1.5 w-1.5 shrink-0 rounded-full bg-zinc-300 dark:bg-zinc-600" />

    <span><img src="https://mintcdn.com/usefulai/C5xOaAf4Os-Vu41o/images/icons/48/openai.com.png?fit=max&auto=format&n=C5xOaAf4Os-Vu41o&q=85&s=21ac965dc6ee5751127e6d435043629b" alt="" noZoom className="relative -top-px mr-1 inline h-4 w-4 rounded-sm object-contain" width="48" height="48" data-path="images/icons/48/openai.com.png" /><a href="https://developers.openai.com/api/docs/models/gpt-realtime-2.1-mini" target="_blank" rel="noreferrer" className="font-medium text-zinc-950 underline underline-offset-2 dark:text-white">GPT-Realtime-2.1 mini</a> <span className="uai-ink-muted">(OpenAI)</span> — Newer low-cost realtime option for lighter voice workloads.</span>
  </span>

  <span className="flex items-baseline gap-2.5">
    <span aria-hidden="true" className="relative -top-0.5 inline-block h-1.5 w-1.5 shrink-0 rounded-full bg-zinc-300 dark:bg-zinc-600" />

    <span><img src="https://mintcdn.com/usefulai/C5xOaAf4Os-Vu41o/images/icons/48/openai.com.png?fit=max&auto=format&n=C5xOaAf4Os-Vu41o&q=85&s=21ac965dc6ee5751127e6d435043629b" alt="" noZoom className="relative -top-px mr-1 inline h-4 w-4 rounded-sm object-contain" width="48" height="48" data-path="images/icons/48/openai.com.png" /><a href="https://openai.com/index/introducing-gpt-live/" target="_blank" rel="noreferrer" className="font-medium text-zinc-950 underline underline-offset-2 dark:text-white">GPT-Live-1</a> <span className="uai-ink-muted">(OpenAI)</span> — Powers ChatGPT Voice for paid users, with no API to build on yet.</span>
  </span>

  <span className="flex items-baseline gap-2.5">
    <span aria-hidden="true" className="relative -top-0.5 inline-block h-1.5 w-1.5 shrink-0 rounded-full bg-zinc-300 dark:bg-zinc-600" />

    <span><img src="https://mintcdn.com/usefulai/C5xOaAf4Os-Vu41o/images/icons/48/openai.com.png?fit=max&auto=format&n=C5xOaAf4Os-Vu41o&q=85&s=21ac965dc6ee5751127e6d435043629b" alt="" noZoom className="relative -top-px mr-1 inline h-4 w-4 rounded-sm object-contain" width="48" height="48" data-path="images/icons/48/openai.com.png" /><a href="https://openai.com/index/introducing-gpt-live/" target="_blank" rel="noreferrer" className="font-medium text-zinc-950 underline underline-offset-2 dark:text-white">GPT-Live-1 mini</a> <span className="uai-ink-muted">(OpenAI)</span> — The free ChatGPT Voice model, also without an API yet.</span>
  </span>

  <span className="flex items-baseline gap-2.5">
    <span aria-hidden="true" className="relative -top-0.5 inline-block h-1.5 w-1.5 shrink-0 rounded-full bg-zinc-300 dark:bg-zinc-600" />

    <span><img src="https://mintcdn.com/usefulai/82PG1Up2qz4DPkMj/images/icons/48/google.com.png?fit=max&auto=format&n=82PG1Up2qz4DPkMj&q=85&s=d44aa6f953cf72e889c25d7f615750e5" alt="" noZoom className="relative -top-px mr-1 inline h-4 w-4 rounded-sm object-contain" width="48" height="48" data-path="images/icons/48/google.com.png" /><a href="https://ai.google.dev/gemini-api/docs/live" target="_blank" rel="noreferrer" className="font-medium text-zinc-950 underline underline-offset-2 dark:text-white">Gemini 2.5 Flash Native Audio Dialog Thinking</a> <span className="uai-ink-muted">(Google)</span> — Stronger reasoning than newer Flash, but much slower to respond.</span>
  </span>

  <span className="flex items-baseline gap-2.5">
    <span aria-hidden="true" className="relative -top-0.5 inline-block h-1.5 w-1.5 shrink-0 rounded-full bg-zinc-300 dark:bg-zinc-600" />

    <span><img src="https://mintcdn.com/usefulai/82PG1Up2qz4DPkMj/images/icons/48/google.com.png?fit=max&auto=format&n=82PG1Up2qz4DPkMj&q=85&s=d44aa6f953cf72e889c25d7f615750e5" alt="" noZoom className="relative -top-px mr-1 inline h-4 w-4 rounded-sm object-contain" width="48" height="48" data-path="images/icons/48/google.com.png" /><a href="https://ai.google.dev/gemini-api/docs/live" target="_blank" rel="noreferrer" className="font-medium text-zinc-950 underline underline-offset-2 dark:text-white">Gemini 2.5 Flash Native Audio Dialog</a> <span className="uai-ink-muted">(Google)</span> — Fast starts, but weaker reasoning than current options.</span>
  </span>

  <span className="flex items-baseline gap-2.5">
    <span aria-hidden="true" className="relative -top-0.5 inline-block h-1.5 w-1.5 shrink-0 rounded-full bg-zinc-300 dark:bg-zinc-600" />

    <span><img src="https://mintcdn.com/usefulai/C5xOaAf4Os-Vu41o/images/icons/48/qwen.ai.png?fit=max&auto=format&n=C5xOaAf4Os-Vu41o&q=85&s=dd7a3821c277e432380635e28c70dc83" alt="" noZoom className="relative -top-px mr-1 inline h-4 w-4 rounded-sm object-contain" width="48" height="48" data-path="images/icons/48/qwen.ai.png" /><a href="https://www.alibabacloud.com/help/en/model-studio/realtime" target="_blank" rel="noreferrer" className="font-medium text-zinc-950 underline underline-offset-2 dark:text-white">Qwen3 Omni Realtime</a> <span className="uai-ink-muted">(Alibaba Cloud)</span> — Earlier Qwen realtime model, now behind the 3.5 versions.</span>
  </span>

  <span className="flex items-baseline gap-2.5">
    <span aria-hidden="true" className="relative -top-0.5 inline-block h-1.5 w-1.5 shrink-0 rounded-full bg-zinc-300 dark:bg-zinc-600" />

    <span><img src="https://mintcdn.com/usefulai/C5xOaAf4Os-Vu41o/images/icons/48/qwen.ai.png?fit=max&auto=format&n=C5xOaAf4Os-Vu41o&q=85&s=dd7a3821c277e432380635e28c70dc83" alt="" noZoom className="relative -top-px mr-1 inline h-4 w-4 rounded-sm object-contain" width="48" height="48" data-path="images/icons/48/qwen.ai.png" /><a href="https://www.alibabacloud.com/help/en/model-studio/qwen-omni" target="_blank" rel="noreferrer" className="font-medium text-zinc-950 underline underline-offset-2 dark:text-white">Qwen3 Omni Flash</a> <span className="uai-ink-muted">(Alibaba Cloud)</span> — Older, slower-starting Qwen option with weaker overall value.</span>
  </span>

  <span className="flex items-baseline gap-2.5">
    <span aria-hidden="true" className="relative -top-0.5 inline-block h-1.5 w-1.5 shrink-0 rounded-full bg-zinc-300 dark:bg-zinc-600" />

    <span><img src="https://mintcdn.com/usefulai/C5xOaAf4Os-Vu41o/images/icons/48/openai.com.png?fit=max&auto=format&n=C5xOaAf4Os-Vu41o&q=85&s=21ac965dc6ee5751127e6d435043629b" alt="" noZoom className="relative -top-px mr-1 inline h-4 w-4 rounded-sm object-contain" width="48" height="48" data-path="images/icons/48/openai.com.png" /><a href="https://developers.openai.com/api/docs/models/gpt-4o-realtime-preview" target="_blank" rel="noreferrer" className="font-medium text-zinc-950 underline underline-offset-2 dark:text-white">GPT-4o Realtime</a> <span className="uai-ink-muted">(OpenAI)</span> — Legacy realtime model for existing builds, not new ones.</span>
  </span>

  <span className="flex items-baseline gap-2.5">
    <span aria-hidden="true" className="relative -top-0.5 inline-block h-1.5 w-1.5 shrink-0 rounded-full bg-zinc-300 dark:bg-zinc-600" />

    <span><img src="https://mintcdn.com/usefulai/C5xOaAf4Os-Vu41o/images/icons/48/openai.com.png?fit=max&auto=format&n=C5xOaAf4Os-Vu41o&q=85&s=21ac965dc6ee5751127e6d435043629b" alt="" noZoom className="relative -top-px mr-1 inline h-4 w-4 rounded-sm object-contain" width="48" height="48" data-path="images/icons/48/openai.com.png" /><a href="https://developers.openai.com/api/docs/models/gpt-4o-mini-realtime-preview" target="_blank" rel="noreferrer" className="font-medium text-zinc-950 underline underline-offset-2 dark:text-white">GPT-4o mini Realtime</a> <span className="uai-ink-muted">(OpenAI)</span> — Legacy mini model; newer Realtime mini tiers are easier picks.</span>
  </span>

  <span className="flex items-baseline gap-2.5">
    <span aria-hidden="true" className="relative -top-0.5 inline-block h-1.5 w-1.5 shrink-0 rounded-full bg-zinc-300 dark:bg-zinc-600" />

    <span><img src="https://mintcdn.com/usefulai/C5xOaAf4Os-Vu41o/images/icons/48/huggingface.co.png?fit=max&auto=format&n=C5xOaAf4Os-Vu41o&q=85&s=760541dee73285f3eb28c4ced26d9a3c" alt="" noZoom className="relative -top-px mr-1 inline h-4 w-4 rounded-sm object-contain" width="48" height="48" data-path="images/icons/48/huggingface.co.png" /><a href="https://github.com/kyutai-labs/moshi" target="_blank" rel="noreferrer" className="font-medium text-zinc-950 underline underline-offset-2 dark:text-white">Moshi</a> <span className="uai-ink-muted">(Kyutai)</span> — Runs on a normal machine, but far behind the hosted leaders on quality.</span>
  </span>
</div>

***

## Frequently Asked Questions

<AccordionGroup>
  <Accordion title={"What is the best speech-to-speech model right now?"}>
    GPT-Realtime-2. It's the top scorer and the most reliable at complex, tool-driven conversations, so it's the default recommendation for demanding production voice agents. The main reason not to use it is cost on very high-volume, simple traffic.
  </Accordion>

  <Accordion title={"What is the best speech-to-speech model for most teams?"}>
    For most new builds, GPT-Realtime-2 if you want peak quality, or Gemini 3.1 Flash Live if you want strong reasoning at a much lower hourly price and can accept a slower start. For high-volume lightweight voice, GPT-Realtime mini or Qwen3.5 Omni Flash Realtime keep costs down.
  </Accordion>

  <Accordion title={"What is the best open-weight speech-to-speech model?"}>
    Step-Audio R1.1. It competes with hosted leaders on reasoning under an Apache-2.0 license, and you can use it through first-party app and API routes. Just know that self-hosting it is infrastructure-heavy, not a local-machine task.
  </Accordion>

  <Accordion title={"What is the best speech-to-speech model you can run locally?"}>
    PersonaPlex is our main local pick, and it needs a high-end GPU (A100/H100-class), not a laptop. Moshi runs on a typical machine but is far weaker. For serious quality, a hosted model is still the better route.
  </Accordion>

  <Accordion title={"Which speech-to-speech model has the lowest latency?"}>
    Deepslate Opal starts talking faster than anything else on this list, which makes conversations feel close to instant. It's only mid-tier on reasoning, though, so it's best where responsiveness matters more than deep capability.
  </Accordion>

  <Accordion title={"Do the benchmark scores match real-world use?"}>
    For quality, mostly yes - the score tracks how well a model reasons and holds a conversation. But it says nothing about speed. Always check Time to First Audio too, since a high-scoring model like Qwen3.5 Omni Plus Realtime can still feel sluggish in a live call.
  </Accordion>

  <Accordion title={"Is GPT-Realtime-2 better than Grok Voice Think Fast 1.0?"}>
    They're close. GPT-Realtime-2 edges ahead on overall quality and the hardest reasoning, while Grok Voice Think Fast 1.0 posts the strongest agentic, tool-using results. If your agent mainly takes actions and calls tools, test Think Fast; for the toughest reasoning, GPT-Realtime-2.
  </Accordion>

  <Accordion title={"What should I use instead of GPT-Realtime or GPT-4o Realtime?"}>
    GPT-Realtime-2 for most cases - it's more capable than both and much cheaper than GPT-Realtime. If you need lower latency or lower cost within the same realtime family, look at GPT-Realtime-1.5 for speed or GPT-Realtime mini for budget.
  </Accordion>
</AccordionGroup>
