Skip to main content
Updated July 12, 2026
Agent LLMs don’t just chat - they plan, call tools, and run multi-step tasks on their own. The catch: a high benchmark score can still hide tool hallucination, the failure that quietly derails unattended runs. We compared 15 models on agent-specific benchmarks.

Best LLMs for Agents


Hardest long-horizon agent work
The most capable agent model in this comparison, built for the longest autonomous runs where weaker models lose the thread - and priced to match.
Score 100%Price License ProprietaryTool hallucination +1.24%
  • It holds a plan together across long, multi-step tasks better than anything else here, staying coherent over runs that last hours and checking its own work.
  • It’s near the top at avoiding calls to tools that don’t exist. If the task is genuinely hard, this is the ceiling.
  • You pay the highest price on this list, so it’s overkill for the routine tool loops that Opus 4.8 or Sonnet 5 handle for far less.
  • Reach for it only when a task genuinely needs the extra ceiling.
All-around agent default
The default pick for serious agent work: it makes efficient tool decisions, recovers when a tool fails, and flags its mistakes rather than hiding them.
Score 85%Price License ProprietaryTool hallucination +0.70%
  • The best all-around agent here for browser and computer-use work, and unusually good at knowing when not to reach for a tool at all.
  • It recovers when a tool fails mid-task and, unlike earlier Claude models, flags its own flawed output rather than shipping it quietly.
  • It costs Opus-tier money, so high-volume, simple tool loops are cheaper to run elsewhere.
  • On pure terminal-style coding, GPT-5.5 has a slight edge, and for the absolute ceiling on the hardest runs, Fable 5 sits clearly above it.
Near-frontier agents at scale
Most of Opus 4.8’s agent reliability at a lower price - the one to run when volume matters more than peak capability.
Score 81%Price License ProprietaryTool hallucination +1.11%
  • It plans multi-step work, drives browsers and terminals, and stays on convention through clean, sequential changes.
  • It’s strong on brownfield code, tracing a failure to its root cause instead of patching symptoms, and it behaves well in long agent loops while keeping tool hallucination low.
  • Tool use is reliable on common APIs but slips when it must infer what an unusual tool does, and it recovers from mid-task failures less gracefully than Opus 4.8.
  • For the hardest reasoning or exotic tool surfaces, step up to Opus 4.8 or GPT-5.5.
Agentic coding and tool use
OpenAI’s strongest agentic coder, notably precise at picking the right tool and argument across large tool surfaces and long-running loops.
Score 80%Price License ProprietaryTool hallucination +1.24%
  • It plans well across messy, multi-part tasks and is precise about tool selection when the tool list is long - the setting where weaker models call the wrong function or invent arguments.
  • It’s also strong at avoiding nonexistent tool calls, which keeps long autonomous runs on track.
  • At high reasoning effort it runs slower, so it’s not the pick for cheap, high-volume loops.
  • It’s coder-first, too, so for the hardest long-horizon or computer-use work, Opus 4.8 and Fable 5 stay more reliable.
Best open-weight agents
The strongest open-weight agent model here by a clear margin, built coding-first with a long context and top-tier tool-call discipline.
Score 73%Price License Open weightTool hallucination +1.24%
  • The highest-scoring open-weight model here, priced well below the proprietary frontier, with a context long enough for repository-scale work. It’s tuned for tool-augmented, multi-step engineering and among the best here at avoiding nonexistent tool calls.
  • Open weights let you host it wherever cost or compliance dictates.
  • It’s text-only, so it won’t drive screenshot or GUI agents that need to see the screen - Gemini 3.5 Flash or Qwen3.6 27B fit there.
  • And despite open weights, it’s far too large for a local machine, so in practice you’re calling a hosted API.
  • App — Available in Z.ai.
  • API — Accessible via Z.ai API.
  • Run locally — Open weights are available from Z.ai, but in practice this needs self-hosting infrastructure, not a local machine.
Low-cost coding agents
xAI’s first model built ground-up for coding and agent work, aggressively priced and marketed as Opus-class - though results land mid-pack, not at the top.
Score 65%Price License ProprietaryTool hallucination Unavailable
  • Built from the ground up for coding and tool-driven tasks, learning from real coding-session data, and priced well below the proprietary frontier.
  • It’s notably token-efficient, and function calling, live web search, and code execution are built in, so it slots into agent loops with little scaffolding.
  • The Opus-class billing outruns the evidence - it sits below the top Claude models and GPT-5.5 overall, and its tool-hallucination reliability isn’t measured yet.
  • For higher-scoring open weights at a similar price, GLM-5.2 is the stronger buy.
  • App — Available in Grok.
  • API — Accessible via xAI API.
High-speed multimodal agents
The fastest capable agent here and the most multimodal, though a Flash-tier ceiling and a real tool-hallucination weakness hold it back from heavy autonomy.
Score 53%Price License ProprietaryTool hallucination -1.09%
  • The speed pick: it returns tokens far faster than anything else here, and it takes text, images, video, audio, and PDFs, so it’s the natural choice for high-throughput, multimodal, and screen-driven agents.
  • Tool orchestration is a genuine strength at this tier.
  • It’s a Flash-tier model, so it trails the top picks on the hardest reasoning and longest runs. It’s also more prone than average to calling tools that don’t exist, so supervise it on high-stakes automation.
  • For deep autonomy, reach for Opus 4.8 or GPT-5.5.
Cheapest capable agent
Frontier-adjacent agentic coding at a rounding-error price, and the best capability-per-dollar on this entire list.
Score 51%Price License Open weightTool hallucination +0.99%
  • You get open-weight agentic coding that holds up against far pricier models, with a long context and solid discipline about not inventing tool calls, all at a tiny fraction of frontier cost.
  • For cost-sensitive, high-volume agent work where you still want real capability, nothing here matches its value.
  • It’s a mid-pack scorer, so it won’t match Opus 4.8 or GPT-5.5 on the hardest long-horizon reasoning. And despite open weights, the full model is a server-cluster deployment, not a local one.
  • If you want cheaper still, DeepSeek V4 Flash undercuts it.
Low-cost multimodal agents
A cheap, open-weight generalist that pairs multimodal input with a long context, aimed at cost-sensitive agent and coding loops.
Score 46%Price License Open weightTool hallucination +0.99%
  • One of the few open-weight models here that takes images and video as well as text, with a long context and low per-task cost.
  • It’s built for autonomous task decomposition and multi-step tool use, and it’s solid at avoiding nonexistent tool calls - a reasonable low-cost base for multimodal agents.
  • It lands mid-pack, so it’s not the model for the hardest reasoning or longest autonomous runs.
  • Open weights don’t buy you local use - it’s a datacenter-class deployment - and cheaper open models like DeepSeek V4 Pro score higher, so its main draw is native multimodality.
  • App — Available in MiniMax.
  • API — Accessible via MiniMax API.
  • Run locally — Open weights are available from Hugging Face, but in practice this needs self-hosting infrastructure, not a local machine.
Long-horizon autonomous execution
An agent-first proprietary model built for very long autonomous runs, with strong tool discipline but a price that’s hard to justify against cheaper open weights.
Score 45%Price License ProprietaryTool hallucination +1.02%
  • Purpose-built for long-horizon autonomy - it sustains very long chains of sequential tool calls with state management and dead-end recovery, and it’s strong at not inventing tools along the way.
  • A long context and native tool support round it out for extended, unattended runs.
  • For its score it’s expensive, and it’s closed, so there’s no self-hosting or fine-tuning. Open-weight GLM-5.2 scores higher for less, and DeepSeek V4 Pro delivers similar-tier capability at a fraction of the price.
  • Long-autonomy is its main reason to choose it.
Open-weight agent specialist
A purpose-built open-weight agent model with respectable coding numbers, but from an obscure vendor with thin, API-only access.
Score 44%Price License Open weightTool hallucination Unavailable
  • Built specifically for agent work - planning, coding, tool use, and iterating on environment feedback - rather than general chat, and it’s competitive on coding for an open-weight model.
  • It also takes image input, and permissive licensing gives you full freedom to host and adapt it.
  • There’s no first-party app and no published task price, so your only real route is a third-party host.
  • It’s heavy to self-host, and better-known open weights like GLM-5.2 and DeepSeek V4 Pro score higher with far more support behind them.
  • API — Accessible via OpenRouter.
  • Run locally — Open weights are available from Hugging Face, but in practice this needs self-hosting infrastructure, not a local machine.
Open-weight coding agents
A code-specialized open-weight model tuned for long, end-to-end programming agents, with best-in-class discipline about calling only tools that exist.
Score 43%Price License Open weightTool hallucination +1.24%
  • Purpose-tuned for code and agentic tool use, and among the very best here at not hallucinating tool calls - exactly what you want in an unattended coding loop.
  • It’s notably token-efficient across multi-turn runs and priced well below the proprietary options.
  • It’s narrow - strong on code and tool use, weaker on broad reasoning - and its context is shorter than the frontier models here.
  • The full model is far too large to run locally, so you’re on a host. GLM-5.2 is the stronger all-round open-weight agent.
  • App — Available in Kimi.
  • API — Accessible via Kimi API.
  • Run locally — Open weights are available from Hugging Face, but in practice this needs self-hosting infrastructure, not a local machine.
Multimodal tool-use agents
Meta’s first Superintelligence Labs model leans hard into tool use, but it’s a limited-access preview and weak at coding.
Score 40%Price License ProprietaryTool hallucination Unavailable
  • Tool use is where it looks strongest - it handles native tools, MCP servers, and custom skills it hasn’t seen before, and tops scaled tool-use benchmarks.
  • It’s natively multimodal across text, images, video, and documents, and it manages its own context and delegates to subagents on longer tasks.
  • Access is the dealbreaker: the API has been a limited preview, so you can’t reliably build on it yet. Coding trails the field, and closed weights rule out self-hosting.
  • For dependable tool-use agents you can deploy today, Opus 4.8 or GPT-5.5 are safer.
  • App — Available in Meta AI.
  • API — Private API preview for select users via Meta.
Single-machine local agents
The rare capable agent model you can actually run on one high-end machine, with vision on board - the pick when local control beats peak score.
Score 38%Price License Open weightTool hallucination Unavailable
  • The most self-host-friendly model here: a dense 27B that fits on a single high-end GPU or a top-spec Apple-silicon Mac, so you get offline use, privacy, and no per-token cost.
  • It’s also one of the few open-weight picks that can see images, useful for local GUI or screenshot agents.
  • It’s the smallest model here, so its ceiling sits below the frontier - expect it to handle scoped tool tasks, not long-horizon runs.
  • If you don’t need local control, cloud open weights like GLM-5.2 or DeepSeek V4 Pro are far more capable for the money.
Cheapest high-volume agents
The cheapest model here by far, built for fast, high-volume tool loops where per-task cost matters more than peak capability.
Score 37%Price License Open weightTool hallucination -0.60%
  • Effectively free per task, with a long context and a smaller active footprint that keeps tool loops fast and cheap.
  • If your agent runs a lot of simple, well-scoped calls at high volume, this is the most economical way to do it.
  • It has the lowest capability score here and a negative tool-hallucination signal, a touch more prone than average to calling nonexistent tools - so keep it to simple, scoped work.
  • It’s a server deployment, not a laptop. Step up to DeepSeek V4 Pro for real capability.

How to Choose

When choosing between these models, consider:
  • Access: First decide whether you want the model in an app, called through an API, or running locally, because those paths change cost, privacy, latency, and setup work. Only Qwen3.6 27B here is a realistic single-machine option; the other open-weight picks need hosted or server-grade infrastructure, and the proprietary models rely on hosted apps or APIs.
  • Quality: We use a combined Agent Arena and Artificial Analysis score as the main number, blending Agent Arena’s Net Improvement signal with Artificial Analysis’s Agentic Index into one normalized figure where higher is better.
  • Price: We use cost per agentic task, drawn from Artificial Analysis where published, for the cleanest cross-model comparison. Some models don’t publish a comparable task cost, so we mark those unavailable.
  • Tool Hallucination: A causal signal from Agent Arena for whether a model avoids calling tools that don’t exist. Positive means fewer hallucinated tool calls than the average model, negative means more. It’s not a raw error rate, so weigh it alongside recovery behavior and your own tests for anything you won’t be watching.

Other Models We Considered

GPT-5.4 mini (OpenAI) — A cheaper OpenAI option, but GPT-5.5 is the stronger pick.MiMo-V2.5-Pro (Xiaomi) — Low-cost open weights, but the top budget picks beat it.Gemini 3.1 Pro (Google) — A familiar Gemini baseline, now behind Gemini 3.5 Flash.Qwen3.7 Plus (Alibaba) — A cheaper Qwen tier, but weaker than Qwen3.7 Max.Step 3.7 Flash (StepFun) — A capable open-weight option, but only a secondary agent pick.Nemotron 3 Ultra (NVIDIA) — Self-hostable, but weaker agent results hold it back.Mistral Medium 3.5 (Mistral) — Recognizable, but less convincing for agent work here.Ring-2.6-1T (InclusionAI) — A huge open model, but low score and thin access.Gemma 4 31B (Google) — Runs locally, but much weaker for agents.Llama 4 Maverick (Meta) — A familiar open model, but not a serious agent pick.

Frequently Asked Questions

Claude Fable 5 has the highest ceiling for the hardest, longest autonomous runs. But Opus 4.8 is the better default for most work - nearly as capable, cheaper, and unusually disciplined about tool calls and flagging its own mistakes.
Opus 4.8. It’s the most reliable all-rounder for tool use, computer use, and long tasks. If you run agents at high volume and want to spend less, Sonnet 5 gives you most of that reliability at a lower per-task cost.
GLM-5.2 is the strongest open-weight agent model here and the clearest value against the proprietary frontier. If cost is the priority, DeepSeek V4 Pro delivers similar-tier capability for far less. Both need server-grade infrastructure to self-host.
DeepSeek V4 Flash is effectively free per task and fine for simple, high-volume tool loops. DeepSeek V4 Pro costs a little more and is far more capable, so it’s usually the smarter cheap pick.
Qwen3.6 27B. It’s a dense 27B model that runs on a single high-end GPU or a top-spec Apple-silicon Mac, and it can read images too. Everything more capable here is either proprietary or too large to run outside a server cluster.
They’re close. GPT-5.5 has a slight edge on terminal-style coding and precise tool selection across large tool lists. Opus 4.8 is stronger on computer use, error recovery, and catching its own mistakes, which makes it the safer choice for unsupervised runs.
Roughly, for capability. But a high score doesn’t guarantee reliable tool use - some strong models still invent tool calls, which quietly derails unattended agents. That’s why we track tool hallucination separately; weight it heavily for anything you won’t be watching.
Reliability under autonomy, not just raw score. Decide your access route first, then weigh tool-call discipline and error recovery for unsupervised work, and match cost to your task volume. Peak capability matters least if the model drifts the moment you look away.