Skip to main content
Updated July 12, 2026
LLMs for coding write, debug, and refactor code - distinct from the tools like Claude Code or Cursor that wrap them. Choosing one means trading capability against price and how much you can run yourself. We ranked 15 on blind web-dev preference and agentic benchmarks.

Best LLMs for Coding


Frontier autonomous coding
The most capable coding model in this comparison, and it shows most on long, autonomous, repo-spanning work where lesser models drift - at frontier prices.
Score 99%Price License ProprietaryContext 1M
  • Best-in-class at sustained agentic coding, staying coherent across a long session and carrying a repo-wide migration through in one sitting. Strong vision too, so screenshot-to-code and figure-heavy work land better than on rivals.
  • When the task is genuinely hard and the ceiling matters, this is the pick.
  • It’s the priciest model here by a wide margin, so it’s overkill for routine edits and quick loops. For most daily work, Sonnet 5 or GPT-5.6 Sol give you most of the capability for far less.
  • Reserve Fable 5 for problems that need it.
Token-efficient agentic coding
OpenAI’s strongest agentic coder holds context across large, messy systems and is unusually token-efficient, making it the frontier pick that’s easiest to actually afford.
Score 99%Price License ProprietaryContext 1.05M
  • Excellent at reasoning through ambiguous failures and checking its own work across big systems, and it does it with fewer tokens than rivals - so the effective cost per finished task runs lower than the sticker price suggests.
  • A safe frontier default for heavy agent work.
  • It sits neck-and-neck with Fable 5 at the top, so the choice often comes down to which house style you prefer.
  • It’s still a premium model, and for lighter work GPT-5.6 Terra or Sonnet 5 cover the basics for less.

Grok 4.5

Value frontier-adjacent coding
The value standout near the top - close to frontier coding quality at a fraction of the price, with shorter context as the trade-off.
Score 89%Price License ProprietaryContext 500K
  • Punches well above its price, landing near the strongest proprietary coders while costing a fraction of them, and it’s fast and token-efficient.
  • If you want frontier-adjacent quality without frontier billing, and your work fits a mid-size context, this is one of the best deals on the list.
  • Its context window is the smallest among the leaders, so very large repo-spanning sessions can outgrow it - reach for Opus 4.8 or a 1M-context model there.
  • On the very hardest problems it trails Fable 5 and GPT-5.6 Sol.
Reliable heavy engineering
Near-top agentic coding with a reliability edge - it flags flawed code more readily than most, which matters when it’s committing to your repo unattended.
Score 89%Price License ProprietaryContext 1M
  • Anthropic tuned it to catch its own mistakes and flag flawed code far more often than the prior Opus, which matters when the model is committing to your repo unattended.
  • A large context and steady long-horizon behavior make it a safe default for heavy engineering work.
  • It’s expensive for daily use, and on the hardest tasks Fable 5 and GPT-5.6 Sol edge ahead.
  • If you need maximum reliability on unattended agent runs, it’s the safer step up from Sonnet 5; otherwise Sonnet 5 delivers most of the quality for less.
Best open-weight coding
The highest-scoring open-weight model here and the pick if you want frontier-adjacent coding without proprietary lock-in - priced like a budget option, with huge context.
Score 88%Price License Open weightContext 1M
  • Open weights let you route it through whichever host is cheapest or fits your compliance needs, and it beats every other open model here on coding while staying near budget pricing.
  • For serious open-weight engineering, or anyone avoiding proprietary lock-in, this is the one to beat.
  • Its weights are open, but it’s too large to run on your own hardware in practice - so you’re really calling a hosted API like any proprietary option.
  • On the hardest problems it lands just below Opus 4.8 and the frontier pair.
  • App — Available in Z.ai.
  • API — Accessible via Z.ai API and OpenRouter.
  • Run locally — Open weights are available from Hugging Face, but in practice this needs self-hosting infrastructure, not a local machine.
High-quality daily driver
The default daily-driver pick - most of the frontier’s coding quality at friendlier pricing and pace for everyday work.
Score 86%Price License ProprietaryContext 1M
  • The sweet spot of quality, speed, and price for most engineering work - close enough to Opus that you rarely feel the gap on routine tasks, with a large context and Anthropic’s reliable, cautious editing behavior.
  • For most developers, this is the one to standardize on.
  • On the hardest, longest-horizon problems it trails Opus 4.8 and the frontier pair, so escalate the genuinely difficult work.
  • If you need maximum reliability on unattended agent runs, Opus 4.8 is the safer step up; for lighter loads, cheaper models suffice.
Low-cost high-capability coding
Meta’s coder matches strong mid-pack quality at a low price, but it runs on a public-preview API - promising rather than production-ready today.
Score 86%Price License ProprietaryContext 1M
  • Strong coding quality for the price, competitive with pricier mid-tier proprietary models while undercutting them, and paired with a large context.
  • If the preview holds up and pricing sticks after general availability, it’s a genuinely appealing low-cost option for everyday coding.
  • The preview status is the catch: terms, limits, and pricing can shift before general availability, so it’s risky to build production workflows on it today.
  • For a stable low-cost pick now, GLM-5.2 or a proven proprietary model is safer.
Mid-tier general coding
Alibaba’s proprietary flagship is a competent all-rounder with a big context, but it’s boxed in by open-weight models that match it for less.
Score 81%Price License ProprietaryContext 1M
  • A solid general-purpose coder with a large context window, capable across everyday generation, edits, and mid-complexity refactors.
  • It holds its own in the middle of the pack and is a reasonable proprietary option if you want a big context without paying frontier prices.
  • The problem is its neighbors: GLM-5.2 scores higher at a lower price with open weights, and Gemini 3.5 Flash matches its score with more speed.
  • It’s competent but hard to single out when cheaper, stronger options sit right next to it.
Fast high-volume coding
Google’s speed-first coder - built for fast, high-volume work where throughput and latency matter more than topping the hardest reasoning tasks.
Score 81%Price License ProprietaryContext 1.05M
  • Fast and responsive with a very large context, which makes it a strong fit for high-volume coding loops, quick iterations, and tasks where you value low latency.
  • When you’re running many calls and want snappy turnarounds rather than the absolute top answer, Flash earns its place.
  • As a Flash-tier model it trails the top coders on the hardest reasoning and multi-step agent work, so reach for Opus 4.8, Sonnet 5, or GPT-5.6 Sol for deep debugging or tricky refactors.
  • And at its price, some stronger models sit uncomfortably close.
Multimodal coding and reasoning
Google’s Pro-tier preview brings strong multimodal range and a big context, but on our coding spine it lands below the cheaper, faster Gemini 3.5 Flash.
Score 76%Price License ProprietaryContext 1.05M
  • Broad, capable reasoning with strong multimodal handling and a very large context, so it’s comfortable on mixed tasks that pair code with images, diagrams, or long documents.
  • If your work is genuinely multimodal, its range is a real draw.
  • For pure coding it’s hard to justify: it scores below Gemini 3.5 Flash while costing more, and it’s still a preview.
  • Flash is the better pick between the two; for peak coding quality, the frontier models are well ahead.
Deliberate mid-tier coding
OpenAI’s mid-tier GPT-5.6 coder - a deliberate, high-effort option that sits below Sol on our coding spine while costing more than the stronger value picks.
Score 74%Price License ProprietaryContext 1.05M
  • A capable coder for mid-complexity work, with a very large context and a deliberate, self-checking reasoning style that suits carefully-worked problems over fast loops.
  • It handles everyday generation and refactors cleanly when you don’t need a top-of-table score.
  • It’s caught in the middle: GPT-5.6 Sol is far stronger near the top, while cheaper models match or beat Terra’s coding for less.
  • Its evidence also leans on a single benchmark component, so treat its standing as less settled than the frontier models’.
Cheap code-tuned tasks
Moonshot’s code-specific open-weight model is cheap and purpose-built for programming, with a context that covers most single-repo work rather than sprawling monorepos.
Score 73%Price License Open weightContext 262K
  • Purpose-tuned for code and priced low, a sensible budget option for straightforward generation and edits. Its context comfortably covers most single-repo tasks, and open weights give you routing and compliance flexibility if you can host it.
  • Good value for focused coding work.
  • Its context is smaller than the 1M-token leaders, so big cross-repo sessions won’t fit, and it trails GLM-5.2 on quality.
  • For stronger open-weight coding, GLM-5.2 is worth the step up; for the cheapest capable option, DeepSeek V4 Pro undercuts it.
  • App — Available in Kimi Code.
  • API — Accessible via Moonshot API and OpenRouter.
  • Run locally — Open weights are available from Hugging Face, but in practice this needs self-hosting infrastructure, not a local machine.
Cheapest capable coding
The value champion here - unusually cheap for its coding quality, with a huge context, though you reach it through an API, not an app.
Score 71%Price License Open weightContext 1.05M
  • By far the cheapest capable coder here, and it pairs that with a very large context - so for high-volume, cost-sensitive coding it’s hard to beat on price per useful output.
  • Open weights add routing and compliance flexibility for teams that can host it.
  • It’s too large to run locally despite open weights, so you’re on a hosted API in practice, and there’s no first-party app to wire it in for you.
  • On quality it sits below the leaders - a value play, not a frontier one.
  • API — Accessible via DeepSeek API and OpenRouter.
  • Run locally — Open weights are available from Hugging Face, but in practice this needs self-hosting infrastructure, not a local machine.
Local coding, strong hardware
A pick you can run yourself - offline on a high-end machine after quantization, trading a real quality drop for privacy and no per-token cost.
Score 53%Price License Open weightContext 262K
  • One of only two models here you can realistically run on your own hardware.
  • On a high-end machine with quantization you get offline use, privacy, and no per-token cost - good for private, low-stakes coding help, learning, and experimentation without sending code to a provider.
  • Its score is near the bottom, so expect struggles past simple, well-scoped tasks - it’s not a serious agent or refactoring model.
  • And “local” still means a high-memory machine, not an average laptop. If you can use the cloud, options above it are more capable.
Self-hosted local coding
The pick if you want to actually self-host a coding model and have a high-end GPU, accepting a big quality drop for control and privacy.
Score 49%Price License Open weightContext 262K
  • The other model here you can run on your own hardware.
  • With a high-end GPU and quantization you get full control, offline use, and privacy at no per-token cost - a fit for private experimentation and learning when keeping code off external servers matters most.
  • It has the lowest score here, handling only simple, well-scoped tasks, not agent or refactoring work - and that standing rests on a single benchmark.
  • If you can use the cloud, nearly everything above is more capable; for local use, Gemma 4 31B scores higher.

How to Choose

When choosing between these models, consider:
  • Access: First decide whether you’ll use the model in an app, call it through an API, or run it locally. That choice drives cost, privacy, latency, and setup work more than small score differences do. For proprietary models, local isn’t an option; only Gemma 4 31B and Qwen3.5 27B are realistic self-run picks, and both need a high-memory machine.
  • Quality: Our score is a normalized average of Code Arena’s WebDev Overall (blind human preference on web-app output) and the Artificial Analysis Coding Index (Terminal-Bench and SciCode, usually at high reasoning effort). Treat it as a comparison spine across models, not universal coding truth - a model can top it and still lose on your specific stack.
  • Price: We use blended API cost per 1M tokens at a 3:1 input-to-output ratio for the cleanest comparison. App subscriptions and self-hosting change the real math, so read this as a relative yardstick.
  • Context Window: This is the maximum input a model accepts, not a promise it stays sharp across the whole window. Long-session reliability varies, so a bigger number helps but doesn’t guarantee coherence on giant repos.
One thing worth clearing up: the model is not the tool. Claude Code, Codex, Cursor, and Copilot are harnesses that run these models, and the same model can feel different depending on the harness around it. This list ranks the models themselves, not the coding tools that wrap them.

Other Models We Considered

GPT-5.5 (OpenAI) — Still a strong coder, but GPT-5.6 Sol is the better current pick.GPT-5.6 Luna (OpenAI) — The cheaper GPT-5.6 tier - handy for fast loops, weaker on hard work.Claude Opus 4.7 (Anthropic) — Nearly as good as Opus 4.8, but the newer version wins.GPT-5.4 (OpenAI) — A recognizable older baseline, now clearly behind GPT-5.6.Seed 2.1 Pro (ByteDance) — Promising preview coder, but too little confirmed to rank.GPT-5.3 Codex (OpenAI) — A useful model-versus-harness reminder, now superseded.MiMo-V2.5-Pro (Xiaomi) — Cheap open-weight for long coding runs, but self-hosting only.MiniMax-M3 (MiniMax) — Low-cost open-weight option, weaker than the best value picks.Qwen3-Coder Next (Alibaba) — A coder-family Qwen, now behind newer, cheaper coders.Devstral 2 (Mistral) — A familiar Mistral coder, now weak and superseded by Medium 3.5.

Frequently Asked Questions

Claude Fable 5 and GPT-5.6 Sol are the two strongest, sitting together at the top of our score. Fable 5 has the highest ceiling on hard, long-horizon work; Sol matches it while using fewer tokens, which makes it cheaper to run at scale. For most people, though, Claude Sonnet 5 is the smarter default - most of that quality at a fraction of the cost.
Claude Sonnet 5. It lands close to the frontier on everyday coding, runs faster and cheaper than the top models, and is reliable enough to standardize on. Step up to Opus 4.8 or Fable 5 only when a task is genuinely hard.
At the very top they’re close: Fable 5 and GPT-5.6 Sol trade the lead depending on the task, so it’s more house style than a clear winner. Sol is notably token-efficient; Fable 5 has a slight edge on the hardest problems. Below them, Sonnet 5 and Opus 4.8 are strong Claude value picks, while GPT-5.6 Terra sits mid-pack.
GLM-5.2. It’s the highest-scoring open-weight model here and beats every other open option on coding, at near-budget pricing. Just know that “open weight” doesn’t mean “runs on your laptop” - it’s too large for that, so in practice you’ll call it through a host.
Gemma 4 31B, with Qwen3.5 27B as the other option. Both run offline, but only on a high-end, high-memory machine after quantization, and both drop a lot of quality versus the cloud models. They’re good for private, low-stakes coding and learning - not serious agent work.
The model is the underlying intelligence; the tool is the harness that feeds it your files, runs commands, and applies edits. Claude Code and Codex are harnesses that run Claude and GPT models. The same model can feel different across harnesses, which is why we rank the models here, not the tools.
Roughly, at the top. Our score blends blind human preference on web apps with agentic coding tests, which tracks real quality better than any single number. But it’s a comparison spine, not a guarantee - a model can top the table and still stumble on your language, framework, or codebase. Trust the ranking to narrow the field, then test your top two on your own work.