Recommendation for OCR & documents

OCR & Documents

Our top recommendation for OCR & Documents, based on the public evidence we track, is Anthropic: Claude Opus 4.6.[1][2] Use for high-stakes OCR where human-rated image comprehension quality matters, as it sits in the top decile of vision arena rankings. Meta: Muse Spark 1.2 is the next-ranked alternative. Deploy when you need near-top-tier vision performance without the top-ranked model's infrastructure, given its 1292 Elo in the vision arena.

About this recommendation

Updated
Sep 4, 2026
Evidence through
Sep 4, 2026
Sources
6
Revision
v56

Decision audit

Why this result

Inspect the inputs and the computed order behind the recommendation.

Models screened

20

live candidates

Evaluation feeds

5

task-weighted

Winner coverage

37%

intended feed weight

Largest provider share

1 of 5

Anthropic

Provisional source breadth. 4 citation families and 0 practitioner families support the top result; 0 cautionary threads is retained. The largest citation family contributes 57%.

Sources evaluated

The task sets these weights before any model is scored.

winner: Claude Opus 4.6
Evaluation feedWeightWinner resultField measured
VLMEvalKit tasksunavailable
40%
feed unavailable0/20
LMArena Document
20%
#211/20
LMArena Vision
15%
#417/20
Structured-output evalunavailable
15%
feed unavailable0/20
OpenRouter usage
10%
85/10020/20

Provider concentration

Each exact model is scored separately; provider identity is not a ranking input.

Anthropic20%
  • Anthropic1 model
  • deepseek1 model
  • Google1 model
  • Meta1 model
  • OpenAI1 model

Decision table

Every published model is shown in computed order. Practitioner sources are distinct community threads, not the citations repeated in the prose below.

RankModelRelative scoreCoveragePractitioner evidenceStrongest measured reason
01Claude Opus 4.6Anthropic
54
37%no linked practitioner threads#2 LMArena Document · #4 LMArena Vision
02Muse Spark 1.2Meta
48
28%no linked practitioner threads#5 LMArena Vision
03DeepSeek V4 Flash 0423deepseek
47
11%2 threads · 2 families · 1 cautionsOpenRouter usage 99/100 normalized
04Gemini 3.5 Flash LiteGoogle
47
28%1 threads · 1 families · 1 cautions#18 LMArena Vision
05GPT-5.6 LunaOpenAI
46
37%no linked practitioner threads#12 LMArena Document · #23 LMArena Vision

Relative score combines normalized benchmark quality and signal coverage; independent practitioner evidence and freshness are bounded tie-breakers. It is an ordering score, not an absolute quality percentage. The writing model receives this order and cannot change it.

  1. Claude Opus 4.6 places fourth among 68 vision models on LMArena's human preference benchmark for image understanding tasks.

    Best when: Use for high-stakes OCR where human-rated image comprehension quality matters, as it sits in the top decile of vision arena rankings.

    Tips

    • Use for high-stakes OCR where human-rated image comprehension quality matters, as it sits in the top decile of vision arena rankings.
      Source 1
      Ranks #4 of 68 on LMArena's vision arena (Elo 1299), based on human preference on image-understanding tasks.
      LMArena vision arenaOpen original ↗
  2. Muse Spark 1.2 ranks fifth on LMArena's vision arena, narrowly trailing the top tier in human preference for image tasks.

    Best when: Deploy when you need near-top-tier vision performance without the top-ranked model's infrastructure, given its 1292 Elo in the vision arena.

    Tips

    • Deploy when you need near-top-tier vision performance without the top-ranked model's infrastructure, given its 1292 Elo in the vision arena.
      Source 2
      Ranks #5 of 68 on LMArena's vision arena (Elo 1292), based on human preference on image-understanding tasks.
      LMArena vision arenaOpen original ↗
  3. DeepSeek V4 Flash 0423 has no vision capability in its hosted API and shows systematic OCR extraction errors on Treasury Bulletins in open evaluations.

    Best when: Consider only after reviewing the cited caution.

    Watch out for

    • Do not use for any OCR task: the hosted API is text-only, and when vision was tested it produced wrong final numbers, wrong cells/years, and wrong formulas on OCR'd government documents.
      Source 3
      # Daily ironclaw failure taxonomy — 2026-09-03 ## Suites analyzed - [officeqa (63 non-pass)](https://nearai.github.io/benchmarks/#/runs/ironclaw/officeqa/4ac1b895-896e-4af3-94b3-e25bda616fad) — All 63 non-passes are genuine model-quality errors by deepseek-v4-flash over OCR'd Treasury Bulletins; no ironclaw defect drives any failure. The dominant mode is a wrong final number from a healthy or self-recovering trajectory (wrong cell/year, wrong formula/denominator, or — UID0237 — computing the co…
      pranavraja99Open original ↗
      Source 4
      ## Standing order Do **not** OCR-then-DeepSeek / text-only 2.4T. DeepSeek v4-flash was unscored because the hosted API is text-only. If a VL host for Qwen3.8-2.4T-A95B appears, run the same seed-777 JSONL + scorer as Luna/restem. Until then this is parked.
      caiotheodoroOpen original ↗
  4. Gemini 3.5 Flash Lite ranks twentieth on LMArena's vision arena and shows documented reliability issues with PDF OCR pipelines.

    Best when: Consider only after reviewing the cited caution.

    Watch out for

    • Avoid for whole_pdf OCR: community reports frequent rejections, empty API responses, and failures when used in paperless-gpt-auto workflows.
      Source 5
      1. Gemini tends to reject whole_pdf OCR pretty frequently, switching to "image" mode helps get more documents OCRed without gettng flagged. 2. Despite #1, gemini will often reject the OCR text when "paperless-gpt-auto" is used (for the second LLM pass) to try and classify the scan 3. The v0.27 logs show "googleai GenerateContent API returned a candidate with no content parts" for whatever the prescribed number of empty worker responses is set by env. 4. After the retries, the paperless-gpt work…
  5. GPT-5.6 Luna ranks twenty-fifth on LMArena's vision arena, placing it in the bottom half of vision models by human preference.

    Best when: Consider only after reviewing the cited caution.

    Watch out for

    • Avoid for OCR quality-critical workflows: its 1254 Elo sits well below median in the vision arena, suggesting weaker human-perceived image understanding than most alternatives listed.
      Source 6
      Ranks #25 of 68 on LMArena's vision arena (Elo 1254), based on human preference on image-understanding tasks.
      LMArena vision arenaOpen original ↗

Frequently asked

What is the top-ranked model for OCR & Documents?
Anthropic: Claude Opus 4.6 ranks first in the current evidence-weighted comparison. Use for high-stakes OCR where human-rated image comprehension quality matters, as it sits in the top decile of vision arena rankings.[1]
What is an alternative to Anthropic: Claude Opus 4.6?
Meta: Muse Spark 1.2 is the next-ranked option. Deploy when you need near-top-tier vision performance without the top-ranked model's infrastructure, given its 1292 Elo in the vision arena.[2]

Sources

  1. 1

    Ranks #4 of 68 on LMArena's vision arena (Elo 1299), based on human preference on image-understanding tasks.

    LMArena vision arena · Benchmark · Aug 27, 2026
  2. 2

    Ranks #5 of 68 on LMArena's vision arena (Elo 1292), based on human preference on image-understanding tasks.

    LMArena vision arena · Benchmark · Aug 27, 2026
  3. 3

    # Daily ironclaw failure taxonomy — 2026-09-03 ## Suites analyzed - [officeqa (63 non-pass)](https://nearai.github.io/benchmarks/#/runs/ironclaw/officeqa/4ac1b895-896e-4af3-94b3-e25bda616fad) — All 63 non-passes are genuine model-quality errors by deepseek-v4-flash over OCR'd Treasury Bulletins; no ironclaw defect drives any failure. The dominant mode is a wrong final number from a healthy or self-recovering trajectory (wrong cell/year, wrong formula/denominator, or — UID0237 — computing the co…

    pranavraja99 · GitHub · Sep 3, 2026
  4. 4

    ## Standing order Do **not** OCR-then-DeepSeek / text-only 2.4T. DeepSeek v4-flash was unscored because the hosted API is text-only. If a VL host for Qwen3.8-2.4T-A95B appears, run the same seed-777 JSONL + scorer as Luna/restem. Until then this is parked.

    caiotheodoro · GitHub · Aug 21, 2026
  5. 5

    1. Gemini tends to reject whole_pdf OCR pretty frequently, switching to "image" mode helps get more documents OCRed without gettng flagged. 2. Despite #1, gemini will often reject the OCR text when "paperless-gpt-auto" is used (for the second LLM pass) to try and classify the scan 3. The v0.27 logs show "googleai GenerateContent API returned a candidate with no content parts" for whatever the prescribed number of empty worker responses is set by env. 4. After the retries, the paperless-gpt work…

    phaset · GitHub · Jul 28, 2026
  6. 6

    Ranks #25 of 68 on LMArena's vision arena (Elo 1254), based on human preference on image-understanding tasks.

    LMArena vision arena · Benchmark · Aug 27, 2026

Rankings synthesized from community evidence and open benchmarks. See methodology. Not driven by vendor marketing.