Hardest long-horizon agent work
The most capable agent model in this comparison, built for the longest autonomous runs where weaker models lose the thread - and priced to match.
Score 100%Price License ProprietaryTool hallucination +1.24%
- It holds a plan together across long, multi-step tasks better than anything else here, staying coherent over runs that last hours and checking its own work.
- It’s near the top at avoiding calls to tools that don’t exist. If the task is genuinely hard, this is the ceiling.
- You pay the highest price on this list, so it’s overkill for the routine tool loops that Opus 4.8 or Sonnet 5 handle for far less.
- Reach for it only when a task genuinely needs the extra ceiling.
- App — Available in Claude.
- API — Accessible via Anthropic API.