Skip to main content
Updated July 12, 2026
Vision LLMs read images, screenshots, charts, and PDFs, then answer in text. The hard part is matching one to your job: a cheap high-volume reader and a frontier document-reasoner sit far apart on price, speed, and accuracy. These 17 picks cover both ends of that range.

Best Vision LLMs


High-accuracy document and diagram reading
The strongest complete benchmark performer here for dense documents, diagrams, and charts, if your work rewards precision over price.
Score 96%Price License ProprietaryVision latency 4.32s
  • Opus 4.7 reads cluttered PDFs, nested tables, and technical figures with a care that cheaper models miss, and it stays reliable across long, multi-page documents.
  • When a misread number is expensive, in finance, legal, or analytics, this is the safe pick.
  • You pay premium rates, and it isn’t the fastest to first response. For high-volume extraction where small errors are tolerable, Gemini 3.5 Flash and GPT-5.5 cost far less.
  • Opus 4.8 is the newer sibling if you want the current flagship instead.
Cheap high-volume image and document work
Google’s low-cost workhorse ties for the highest complete score here while costing a fraction of the frontier models, making it the default for volume.
Score 96%Price License ProprietaryVision latency 11.24s
  • Flash pairs near-top vision accuracy with pricing built for scale, so batch document parsing and screen reading stay affordable.
  • It handles layout-heavy documents better than its price suggests, which makes it the sensible default for high-volume pipelines.
  • Time to first token is slow for a Flash model, so it’s less suited to snappy interactive use than GPT-5.5. On the hardest single-document reasoning, Opus 4.7 and Gemini 3.1 Pro pull ahead.
  • Pick it for volume, not peak accuracy.
Multimodal reasoning and tool use
The benchmarked Muse Spark release is a strong natively multimodal reasoner, but its direct route is Meta AI rather than a public API or self-hosting.
Score 95%Price License ProprietaryVision latency Not disclosed
  • Built from the ground up to reason across images, audio, and tools in one model, Muse Spark is a capable option for multimodal help inside Meta AI.
  • The benchmarked version has no public API, local route, image price, or comparable latency. Meta’s newer Muse Spark 1.1 has a public-preview API but is not the model scored here.
  • Choose Gemini 3.5 Flash or GPT-5.5 if you need a benchmarked API model.
Deep visual reasoning and analysis
The Pro-tier Gemini for tasks that need careful visual reasoning rather than fast extraction, deeper than Flash and priced close to it.
Score 95%Price License ProprietaryVision latency 16.42s
  • Gemini 3.1 Pro is designed for multi-step visual reasoning - reading a chart, connecting it to surrounding text, and drawing a conclusion - while staying inexpensive.
  • It’s a strong middle ground when accuracy matters but frontier prices don’t fit.
  • It’s a preview model and slow to first token, so it’s poor for latency-sensitive or high-volume work where Flash is faster and cheaper.
  • On the very hardest documents, Opus 4.7 still edges it. Confirm preview stability before you depend on it.
Fast document and chart extraction
OpenAI’s fastest strong vision model returns a quick first response with excellent document, chart, and layout reading, ideal for interactive tools.
Score 95%Price License ProprietaryVision latency 1.87s
  • GPT-5.5 returns a first token faster than the other highlighted frontier models while remaining strong on documents, charts, and screenshots.
  • That combination makes it the standout for interactive workflows where responsiveness is the point.
  • It costs substantially more than the Gemini tier for a similar overall score, so the premium only makes sense when its much faster first response matters.
  • For batch processing, Gemini 3.5 Flash is the better-value option.
Reliable agentic visual workflows
Anthropic’s current flagship is tuned to be more honest and reliable than 4.7, making it the pick when a vision agent runs unattended.
Score 94%Price License ProprietaryVision latency Not disclosed
  • Opus 4.8 is Anthropic’s current flagship for PDFs, diagrams, messy layouts, and agentic work.
  • It is the better default than 4.7 when current model support and unattended workflows matter more than the older version’s stronger complete benchmark result.
  • It’s among the priciest models here, so for straightforward extraction it’s overkill; Gemini 3.5 Flash and Qwen3.7 Plus do that job for far less.
  • Reach for 4.8 when reliability under autonomy, not cost, is what you’re optimizing for.
Long-context multimodal reasoning
xAI’s flagship pairs strong multimodal reasoning with a very large context window, so it can hold many images and long documents at once.
Score 94%Price License ProprietaryVision latency 6.57s
  • Grok 4.5 keeps many images and long documents in one context, which suits multi-image comparisons and large visual workloads.
  • It’s priced below the top Claude and GPT tiers for that capability.
  • Independent vision benchmarking is still thin, so treat its standing as less settled than Gemini’s or Claude’s. For document precision, Opus 4.7 and GPT-5.5 have a longer track record.
  • It’s at its best when context size is the binding constraint.
  • App — Available in Grok.
  • API — Accessible via xAI API.
Low-cost GUI and screen agents
A cheap, fast multimodal agent model built to read screens and drive interfaces, with strong value for GUI automation and screenshot work.
Score 92%Price License ProprietaryVision latency 3.25s
  • Qwen3.7 Plus reads screens and images and is tuned for agentic GUI and CLI tasks, all at a fraction of frontier pricing.
  • If you’re building screen-reading or app-navigating agents at scale, the cost-to-capability ratio here is hard to beat.
  • It’s proprietary and API-only, with no first-party app or local route, and on the hardest document reasoning it sits below Opus 4.7 and Gemini 3.1 Pro.
  • Great for high-volume agent work, weaker for peak-accuracy analysis.

Kimi K2.6

Open-weight agentic vision work
The strongest open-weight pick here for agentic multimodal work, with a hosted app and API if you’d rather not run it yourself.
Score 91%Price License Open weightVision latency 3.26s
  • Kimi K2.6 brings capable image understanding to a genuinely open-weight model, and it holds up on long-context, agent-style tasks.
  • You get a hosted app and API for convenience, plus the option to inspect or self-host the weights when you need control.
  • It’s a trillion-parameter model, so “open weight” doesn’t mean local; realistic self-hosting needs serious infrastructure, not a workstation.
  • For pure document accuracy, Opus 4.7 and Gemini still lead. Choose it when open weights genuinely matter to you.
  • App — Available in Kimi.
  • API — Accessible via Kimi Platform and OpenRouter.
  • Run locally — Open weights are available from Hugging Face, but in practice this needs self-hosting infrastructure, not a local machine.
Balanced everyday vision work
Anthropic’s balanced daily driver delivers fast, reliable vision that covers most everyday document and image tasks without paying Opus prices.
Score 89%Price License ProprietaryVision latency 2.57s
  • Sonnet 5 handles the bulk of real vision work - reading documents, screenshots, and charts - quickly and dependably, with the same careful behavior as the Opus line.
  • For most teams it hits the sweet spot of speed, accuracy, and cost.
  • On the hardest, densest documents it gives up ground to Opus 4.7 and 4.8, and cheaper models like Gemini 3.5 Flash undercut it on price.
  • Step up to Opus when precision is critical, step down when volume rules.
Cheapest capable open-weight vision
About the cheapest way to get solid open-weight vision through an API, and a strong value if raw cost is your main driver.
Score 88%Price License Open weightVision latency 3.22s
  • MiniMax-M3 delivers competent image and document understanding at rock-bottom hosted pricing, and its open weights let you route it through whichever host is cheapest.
  • For high-volume, cost-sensitive vision where you don’t need frontier accuracy, it’s a smart budget option.
  • Despite open weights, it’s too large for practical local use, so you’re on a hosted API anyway. It trails Kimi K2.6 and the proprietary leaders on hard reasoning.
  • Pick it for price, and look elsewhere for peak accuracy.
  • API — Accessible via MiniMax API and OpenRouter.
  • Run locally — Open weights are available from Hugging Face, but in practice this needs self-hosting infrastructure, not a local machine.
Local vision on a high-end GPU
The best genuinely self-hostable vision model here: with a strong GPU you get capable image understanding without a mandatory metered API fee.
Score 86%Price License Open weightVision latency 2.39s
  • Gemma 4 31B runs locally on a high-end machine, giving you private, offline vision without a model-usage fee. It’s also currently free through a hosted route if you’d rather not manage hardware.
  • That makes it useful for privacy-sensitive work, although local hardware still has a cost.
  • You need a serious GPU and enough memory to run it well, and it trails the proprietary leaders on the hardest documents.
  • If you can use the cloud, Gemini 3.5 Flash is stronger and still cheap. Choose it for control and privacy.
  • API — Accessible via OpenRouter.
  • Run locally — If you have a high-end machine, you can run it with Ollama after downloading weights from Hugging Face.
Vision-driven coding and UI work
A native multimodal model tuned to turn what it sees - screenshots, design drafts, layouts - into working code and UI actions.
Score 82%Price License ProprietaryVision latency Not disclosed
  • GLM-5V Turbo is built to fuse visual perception with code, so screenshot-to-code, design-to-UI, and layout-driven agent tasks land better than on general vision models.
  • If your vision work ends in code or interface actions, this is a purpose-built option.
  • It’s narrower than the generalist leaders and weaker on open-ended document reasoning, where Opus 4.7 and Gemini 3.1 Pro do more, and its latency isn’t published.
  • Reach for it for vision-to-code specifically, not broad visual analysis.
Frontier reasoning on complex documents
Anthropic’s new premium model targets deeply nested diagrams and tables, but at the highest price here it’s overkill for routine vision.
Score 75%Price License ProprietaryVision latency Not disclosed
  • Fable 5 excels at the hardest, most document-heavy reasoning - untangling diagrams, charts, and tables buried inside long PDFs - and it can carry demanding, long-horizon analysis further than lighter models.
  • When a problem genuinely needs frontier reasoning over visuals, it delivers.
  • The price is the dealbreaker for everyday vision; it’s the most expensive model here by a wide margin, and its standardized vision-benchmark coverage is thin.
  • For most document work, Opus 4.7 and Sonnet 5 give you most of the value for far less.
Frontier reasoning on hard visuals
OpenAI’s newest frontier model brings heavy reasoning to visual problems, but very high latency makes it a deliberate choice, not an interactive one.
Score 74%Price License ProprietaryVision latency 26.97s
  • GPT-5.6 Sol applies top-tier reasoning to hard visual and document problems, and when a task rewards slow, careful analysis over speed, that depth shows.
  • It’s a serious option for complex, high-stakes visual reasoning where you can afford to wait for the answer.
  • First-token latency is the highest here by far and it’s expensive, so it’s wrong for interactive or high-volume vision. As a new release its vision-benchmark standing is still thin.
  • For fast document work, GPT-5.5 is far quicker and cheaper.
Mid-range self-hosted vision
A mid-tier open-weight model you can self-host on strong hardware or call cheaply through an API, with decent rather than leading vision.
Score 68%Price License Open weightVision latency 3.02s
  • Qwen3.6 27B gives you open weights and a real self-hosting path on a high-end machine, plus cheap hosted access if you prefer.
  • For private, moderate-stakes vision work where you want control without the largest models’ footprint, it’s a reasonable middle option.
  • Accuracy sits well behind the leaders, so it’s not for demanding analysis.
  • Gemma 4 31B is a stronger open-weight pick at a similar size, making Qwen3.6 27B hard to choose unless its deployment profile fits your constraints better.
Vision on a typical laptop
The one model here that genuinely runs on a normal laptop: small and limited, but private and practical for light vision.
Score 61%Price License Open weightVision latency 0.70s
  • Qwen3.5 4B is small enough to run on a typical machine, giving you offline, private image understanding without a model-usage fee.
  • For simple captioning, basic document reading, and on-device prototyping, it’s a genuinely useful small model; actual speed depends on your hardware and quantization.
  • It has the lowest accuracy here, so it struggles with anything complex or detail-critical; don’t trust it on dense documents.
  • For real analysis, almost everything above it is far stronger. Use it for light, local, low-stakes tasks only.

How to Choose

When choosing between these models, consider:
  • Access: Decide first whether you’ll use the model in an app, call it through an API, or run it locally, because that single choice drives cost, privacy, latency, and setup work more than small score differences do. Qwen3.5 4B runs on a typical laptop; Gemma 4 31B and Qwen3.6 27B need high-end local hardware. Open-weight leaders like Kimi K2.6 and MiniMax-M3 need self-hosting infrastructure, not a workstation.
  • Quality: We use a vision score that blends Arena’s vision arena (human preference, style-controlled) with Artificial Analysis’s MMMU-Pro visual reasoning, normalized to a percentage. Two caveats matter. Claude Opus 4.8 and Grok 4.5 carry observed-only scores that aren’t directly comparable to the fully benchmarked models above them, and the newest premium models, Claude Fable 5 and GPT-5.6 Sol, rank lower than their reputations suggest mainly because standardized vision coverage lags their release.
  • Price: We compare USD per 1,000 one-megapixel images at 1024x1024, image input only. It’s the cleanest way to line up costs, though your real bill also depends on the text tokens each request generates.
  • Vision Latency: Time to first token for one image plus roughly 1,000 input tokens, where lower is better. It captures responsiveness, not throughput. Gemini 3.5 Flash is slow to first token but built for high-volume batches, so match the metric to how you’ll actually use the model.

Other Models We Considered

Qwen3.5 397B A17B (Alibaba) — Tops the Qwen3.5 line on quality, but far too large for local use.Gemini 3 Pro (Google) — Excellent in its day, now retired in favor of Gemini 3.1 Pro.Qwen3-VL 235B A22B (Alibaba) — Popular incumbent with mature tooling, now superseded by newer Qwen models.Moondream 3.1 9B A2B (Moondream) — Tiny local specialist for captioning, detection, and pointing on edge hardware.GPT-4o (OpenAI) — The multimodal baseline everyone knew, but the app and API route is retired.Qwen2.5-VL 72B (Alibaba) — Familiar predecessor still in existing deployments, since surpassed by newer models.LLaVA-OneVision 72B (LLaVA contributors) — A recognizable open baseline, now well behind current vision models.

Frequently Asked Questions

For peak document and diagram accuracy, Claude Opus 4.7 leads. For the best mix of accuracy and price, Gemini 3.5 Flash is the default recommendation; choose GPT-5.5 when first-response latency matters more.
Gemini 3.5 Flash covers the widest range of everyday image and document work cheaply and well. If you want Anthropic’s careful reading at a moderate price, Claude Sonnet 5 is the close alternative.
Gemma 4 31B has no model-usage fee when run locally on a high-end machine and is currently free through a hosted route. Qwen3.5 4B costs almost nothing and runs on a normal laptop, while MiniMax-M3 is the cheapest capable paid hosted option.
Qwen3.5 4B is the only pick here that runs on a typical laptop. If you have a high-end GPU, Gemma 4 31B is the stronger local choice.
Kimi K2.6 is the strongest open-weight model on this list, though it’s large enough that most people will use it hosted rather than self-hosted.
It depends on the job. Flash is much cheaper and is the better-value batch option; GPT-5.5 returns a much faster first response. For interactive tools, GPT-5.5; for high-volume pipelines, Flash.
Roughly. They track document reading and visual reasoning well, but they don’t capture your exact images, latency needs, or task mix. Test the top two or three candidates on your own inputs before committing.
GPT-5.5 is the direct upgrade for fast, high-detail document reading. If cost and volume matter more, Gemini 3.5 Flash is the better move.