Recommendation for OCR & documents
OCR & Documents
Our top recommendation for OCR & Documents, based on the public evidence we track, is Anthropic: Claude Opus 4.6.[1][2] Use for high-stakes OCR where human-rated image comprehension quality matters, as it sits in the top decile of vision arena rankings. Meta: Muse Spark 1.2 is the next-ranked alternative. Deploy when you need near-top-tier vision performance without the top-ranked model's infrastructure, given its 1292 Elo in the vision arena.
About this recommendation
- Updated
- Sep 4, 2026
- Evidence through
- Sep 4, 2026
- Sources
- 6
- Revision
- v56
Decision audit
Why this result
Inspect the inputs and the computed order behind the recommendation.
Models screened
20
live candidates
Evaluation feeds
5
task-weighted
Winner coverage
37%
intended feed weight
Largest provider share
1 of 5
Anthropic
Sources evaluated
The task sets these weights before any model is scored.
| Evaluation feed | Weight | Winner result | Field measured |
|---|---|---|---|
| VLMEvalKit tasksunavailable | 40% | feed unavailable | 0/20 |
| LMArena Document | 20% | #2 | 11/20 |
| LMArena Vision | 15% | #4 | 17/20 |
| Structured-output evalunavailable | 15% | feed unavailable | 0/20 |
| OpenRouter usage | 10% | 85/100 | 20/20 |
Provider concentration
Each exact model is scored separately; provider identity is not a ranking input.
- Anthropic1 model
- deepseek1 model
- Google1 model
- Meta1 model
- OpenAI1 model
Decision table
Every published model is shown in computed order. Practitioner sources are distinct community threads, not the citations repeated in the prose below.
| Rank | Model | Relative score | Coverage | Practitioner evidence | Strongest measured reason |
|---|---|---|---|---|---|
| 01 | Claude Opus 4.6Anthropic | 54 | 37% | no linked practitioner threads | #2 LMArena Document · #4 LMArena Vision |
| 02 | Muse Spark 1.2Meta | 48 | 28% | no linked practitioner threads | #5 LMArena Vision |
| 03 | DeepSeek V4 Flash 0423deepseek | 47 | 11% | 2 threads · 2 families · 1 cautions | OpenRouter usage 99/100 normalized |
| 04 | Gemini 3.5 Flash LiteGoogle | 47 | 28% | 1 threads · 1 families · 1 cautions | #18 LMArena Vision |
| 05 | GPT-5.6 LunaOpenAI | 46 | 37% | no linked practitioner threads | #12 LMArena Document · #23 LMArena Vision |
Relative score combines normalized benchmark quality and signal coverage; independent practitioner evidence and freshness are bounded tie-breakers. It is an ordering score, not an absolute quality percentage. The writing model receives this order and cannot change it.
Claude Opus 4.6 places fourth among 68 vision models on LMArena's human preference benchmark for image understanding tasks.
Best when: Use for high-stakes OCR where human-rated image comprehension quality matters, as it sits in the top decile of vision arena rankings.
Tips
- Use for high-stakes OCR where human-rated image comprehension quality matters, as it sits in the top decile of vision arena rankings.
Muse Spark 1.2 ranks fifth on LMArena's vision arena, narrowly trailing the top tier in human preference for image tasks.
Best when: Deploy when you need near-top-tier vision performance without the top-ranked model's infrastructure, given its 1292 Elo in the vision arena.
Tips
- Deploy when you need near-top-tier vision performance without the top-ranked model's infrastructure, given its 1292 Elo in the vision arena.
DeepSeek V4 Flash 0423 has no vision capability in its hosted API and shows systematic OCR extraction errors on Treasury Bulletins in open evaluations.
Best when: Consider only after reviewing the cited caution.
Watch out for
- Do not use for any OCR task: the hosted API is text-only, and when vision was tested it produced wrong final numbers, wrong cells/years, and wrong formulas on OCR'd government documents.
Gemini 3.5 Flash Lite ranks twentieth on LMArena's vision arena and shows documented reliability issues with PDF OCR pipelines.
Best when: Consider only after reviewing the cited caution.
Watch out for
- Avoid for whole_pdf OCR: community reports frequent rejections, empty API responses, and failures when used in paperless-gpt-auto workflows.
GPT-5.6 Luna ranks twenty-fifth on LMArena's vision arena, placing it in the bottom half of vision models by human preference.
Best when: Consider only after reviewing the cited caution.
Watch out for
- Avoid for OCR quality-critical workflows: its 1254 Elo sits well below median in the vision arena, suggesting weaker human-perceived image understanding than most alternatives listed.
Frequently asked
- What is the top-ranked model for OCR & Documents?
- Anthropic: Claude Opus 4.6 ranks first in the current evidence-weighted comparison. Use for high-stakes OCR where human-rated image comprehension quality matters, as it sits in the top decile of vision arena rankings.[1]
- What is an alternative to Anthropic: Claude Opus 4.6?
- Meta: Muse Spark 1.2 is the next-ranked option. Deploy when you need near-top-tier vision performance without the top-ranked model's infrastructure, given its 1292 Elo in the vision arena.[2]
Sources
- 1
“Ranks #4 of 68 on LMArena's vision arena (Elo 1299), based on human preference on image-understanding tasks.”
LMArena vision arena · Benchmark · Aug 27, 2026 - 2
“Ranks #5 of 68 on LMArena's vision arena (Elo 1292), based on human preference on image-understanding tasks.”
LMArena vision arena · Benchmark · Aug 27, 2026 - 3
“# Daily ironclaw failure taxonomy — 2026-09-03 ## Suites analyzed - [officeqa (63 non-pass)](https://nearai.github.io/benchmarks/#/runs/ironclaw/officeqa/4ac1b895-896e-4af3-94b3-e25bda616fad) — All 63 non-passes are genuine model-quality errors by deepseek-v4-flash over OCR'd Treasury Bulletins; no ironclaw defect drives any failure. The dominant mode is a wrong final number from a healthy or self-recovering trajectory (wrong cell/year, wrong formula/denominator, or — UID0237 — computing the co…”
pranavraja99 · GitHub · Sep 3, 2026 - 4
“## Standing order Do **not** OCR-then-DeepSeek / text-only 2.4T. DeepSeek v4-flash was unscored because the hosted API is text-only. If a VL host for Qwen3.8-2.4T-A95B appears, run the same seed-777 JSONL + scorer as Luna/restem. Until then this is parked.”
caiotheodoro · GitHub · Aug 21, 2026 - 5
“1. Gemini tends to reject whole_pdf OCR pretty frequently, switching to "image" mode helps get more documents OCRed without gettng flagged. 2. Despite #1, gemini will often reject the OCR text when "paperless-gpt-auto" is used (for the second LLM pass) to try and classify the scan 3. The v0.27 logs show "googleai GenerateContent API returned a candidate with no content parts" for whatever the prescribed number of empty worker responses is set by env. 4. After the retries, the paperless-gpt work…”
phaset · GitHub · Jul 28, 2026 - 6
“Ranks #25 of 68 on LMArena's vision arena (Elo 1254), based on human preference on image-understanding tasks.”
LMArena vision arena · Benchmark · Aug 27, 2026
Rankings synthesized from community evidence and open benchmarks. See methodology. Not driven by vendor marketing.