Recommendation for Vision / Documents
Vision & Documents
Our top recommendation for Vision & Documents, based on the public evidence we track, is Anthropic: Claude Fable 5.[1][2] Use for document parsing workflows where human preference rankings suggest reliable extraction from formatted pages. Anthropic: Claude Opus 4.6 is the next-ranked alternative. Deploy for document-heavy pipelines where the #2 ranking on a 29-model blind preference benchmark signals superior layout and text extraction.
About this recommendation
- Updated
- Sep 4, 2026
- Evidence through
- Sep 4, 2026
- Sources
- 15
- Revision
- v57
Decision audit
Why this result
Inspect the inputs and the computed order behind the recommendation.
Models screened
20
live candidates
Evaluation feeds
5
task-weighted
Winner coverage
37%
intended feed weight
Largest provider share
3 of 8
Anthropic
Sources evaluated
The task sets these weights before any model is scored.
| Evaluation feed | Weight | Winner result | Field measured |
|---|---|---|---|
| VLMEvalKit tasksunavailable | 40% | feed unavailable | 0/20 |
| LMArena Document | 20% | #3 | 10/20 |
| LMArena Vision | 15% | #1 | 16/20 |
| Structured-output evalunavailable | 15% | feed unavailable | 0/20 |
| OpenRouter usage | 10% | 87/100 | 20/20 |
Provider concentration
Each exact model is scored separately; provider identity is not a ranking input.
- Anthropic3 models
- Google2 models
- deepseek1 model
- Meta1 model
- OpenAI1 model
Decision table
Every published model is shown in computed order. Practitioner sources are distinct community threads, not the citations repeated in the prose below.
| Rank | Model | Relative score | Coverage | Practitioner evidence | Strongest measured reason |
|---|---|---|---|---|---|
| 01 | Claude Fable 5Anthropic | 55 | 37% | no linked practitioner threads | #1 LMArena Vision · #3 LMArena Document |
| 02 | Claude Opus 4.6Anthropic | 54 | 37% | no linked practitioner threads | #2 LMArena Document · #4 LMArena Vision |
| 03 | DeepSeek V4 Flash Vision Expdeepseek | 52 | 11% | 10 threads · 8 families · 2 cautions | OpenRouter usage 90/100 normalized |
| 04 | GPT-5.6 SolOpenAI | 50 | 37% | 1 threads · 1 families · 1 cautions | #6 LMArena Document · #10 LMArena Vision |
| 05 | Claude Sonnet 4.6Anthropic | 50 | 37% | 1 threads · 1 families · 1 cautions | #5 LMArena Document · #14 LMArena Vision |
| 06 | Gemini 3.6 FlashGoogle | 49 | 28% | 1 threads · 1 families · 0 cautions | #7 LMArena Vision |
| 07 | Muse Spark 1.2Meta | 48 | 28% | 1 threads · 1 families · 0 cautions | #5 LMArena Vision |
| 08 | Gemini 3.5 Flash LiteGoogle | 47 | 28% | 2 threads · 2 families · 1 cautions | #18 LMArena Vision |
Relative score combines normalized benchmark quality and signal coverage; independent practitioner evidence and freshness are bounded tie-breakers. It is an ordering score, not an absolute quality percentage. The writing model receives this order and cannot change it.
Claude Fable 5 places third in blind human preference for document tasks on LMArena, indicating strong OCR and layout parsing capabilities.
Best when: Use for document parsing workflows where human preference rankings suggest reliable extraction from formatted pages.
Tips
- Use for document parsing workflows where human preference rankings suggest reliable extraction from formatted pages.
Claude Opus 4.6 ranks second on LMArena's document arena, the highest placement among all candidates for blind preference on document understanding.
Best when: Deploy for document-heavy pipelines where the #2 ranking on a 29-model blind preference benchmark signals superior layout and text extraction.
Tips
- Deploy for document-heavy pipelines where the #2 ranking on a 29-model blind preference benchmark signals superior layout and text extraction.
DeepSeek V4 Flash Vision Exp supports 1M context windows with verified 900K stability, 220 tps reasoning, and multiple image input methods including base64 data URLs and external URLs.
Best when: Process lengthy document batches or multi-page scans in a single call using the 900K+ stable context window.
Tips
- Process lengthy document batches or multi-page scans in a single call using the 900K+ stable context window.
- Feed images via base64 data URLs for sensitive documents, or external URLs for public assets, with explicit detail control (low/high).
Watch out for
- Expect OOM failures during FP8-to-FP4 MoE conversion on 2x GB10 (DGX Spark) setups, blocking local deployment.
GPT-5.6 Sol ranks eighth on LMArena's document arena but demonstrates specific competence in sheet music transcription where other models fail.
Best when: Try for musical notation OCR when other vision models trip on horizontal staff line alignment.
Tips
- Try for musical notation OCR when other vision models trip on horizontal staff line alignment.
Watch out for
- Accept lower blind preference ranking (#8 of 29) for general document tasks compared to top Claude variants.
Claude Sonnet 4.6 ranks sixth on LMArena's document arena and was evaluated in a controlled food photography study with 26,904 total queries at temperature 0.01.
Best when: Use for consistent visual analysis where low-temperature reproducibility matters, as demonstrated in large-scale food image studies.
Tips
- Use for consistent visual analysis where low-temperature reproducibility matters, as demonstrated in large-scale food image studies.
Watch out for
- Note the #6 ranking places it below Fable 5 and Opus 4.6 for document preference, suggesting tradeoffs in layout handling.
Gemini 3.6 Flash ranks ninth on LMArena's vision arena and serves as a quota-aware fallback in free-tier applications.
Best when: Implement with automatic fallback to 3.5 Flash Lite when quota-blocked, preserving receipt scan workflows without user intervention.
Tips
- Implement with automatic fallback to 3.5 Flash Lite when quota-blocked, preserving receipt scan workflows without user intervention.
Watch out for
- Expect mid-tier vision performance (#9 of 68) that may trail specialized document models on complex layouts.
Muse Spark 1.2 ranks fifth on LMArena's vision arena for general image understanding, though deployment requires explicit vision capability configuration.
Best when: Consider for general image understanding tasks where the #5 vision ranking suggests strong visual reasoning.
Tips
- Consider for general image understanding tasks where the #5 vision ranking suggests strong visual reasoning.
Watch out for
- Verify vision capability is explicitly enabled in your client configuration, as default setups may reject image inputs.
Gemini 3.5 Flash Lite ranks 20th on LMArena's vision arena and functions primarily as a fallback tier, with observed OCR rejection patterns requiring workarounds.
Best when: Keep as a backup when primary models hit quota limits, maintaining basic receipt and document scan availability.
Tips
- Keep as a backup when primary models hit quota limits, maintaining basic receipt and document scan availability.
Watch out for
- Switch from whole-PDF OCR to per-page image mode to reduce rejection rates, and expect occasional empty API responses requiring retry logic.
Frequently asked
- What is the top-ranked model for Vision & Documents?
- Anthropic: Claude Fable 5 ranks first in the current evidence-weighted comparison. Use for document parsing workflows where human preference rankings suggest reliable extraction from formatted pages.[1]
- What is an alternative to Anthropic: Claude Fable 5?
- Anthropic: Claude Opus 4.6 is the next-ranked option. Deploy for document-heavy pipelines where the #2 ranking on a 29-model blind preference benchmark signals superior layout and text extraction.[2]
Sources
- 1
“Ranks #3 of 29 on LMArena's document arena (score 1504), based on blind preference for document tasks.”
LMArena document arena · Benchmark · Jul 30, 2026 - 2
“Ranks #2 of 29 on LMArena's document arena (score 1510), based on blind preference for document tasks.”
LMArena document arena · Benchmark · Jul 30, 2026 - 3
“https://github.com/local-inference-lab/rtx6kpro/blob/master/models/ds4-jovian-judgement-r4.md - working and stable, 1M context achievable while 900K is rock solid. 220tps reasoning and prose, 300tps code, 6500tps prefill on Epyc 7702 platform, LMCache supported.”
nephepritou · Hugging Face · Sep 4, 2026 - 4
“Thank you for sharing the public recipe *DeepSeek-v4-Flash-DSpark-2x-DGX-Spark*. I adapted your approach to my hardware setup, revalidated the configuration, and confirmed that the 200K-token context window is reproducible and reliable. **Validation Results** Starting from your recipe (your 14 runtime patches and vision overlay are used as-is), I ran three test suites on a freshly wiped pair of nodes (containers and config removed, weights reused): - **Needle-in-a-Haystack**: 9/9 pass (3 depths…”
tenaiaiai · GitHub · Sep 3, 2026 - 5
“## Resolution 研究完成(来源:官方图像理解指南 https://api-docs.deepseek.com/zh-cn/guides/vision + Tool Calls / 思考模式指南;**无任何 live 调用**)。完整契约见 `docs/research/vision-api-contract.md`(分支 `research/vision-api-contract`)。核心结论: - **模型**:`deepseek-v4-flash-vision-exp`(experimental,额外接受图像输入)。OpenAI 兼容 `/chat/completions`。 - **传图**:`content` 必须是**块数组**(非纯字符串)。三种方式——base64 data URL(`image_url` 块)、外链 URL(≤8192 字符,单图 ≤32 MiB,60s 内下载)、Files API `file_id`。扫雷棋盘建议走 base64 data URL。 - **`detail`**:`low`(512×512,更快/省)/ `high` /…”
MiSmiler · GitHub · Aug 24, 2026 - 6
“The verified 2x GB10 cookbook cell for DeepSeek-V4-Flash-Vision-Exp never finishes loading on our pair. The kernel OOM-kills the scheduler during the FP8 to FP4 MoE conversion. Five launches, five identical deaths. ## Setup - Image `lmsysorg/sglang:dev-v4f-2dgx-v2` (`sha256:67873eb93b994736ab534111f79b5aa93d2575b973ebf9d276c40d548ff9afec`), sglang `0.0.0.dev1+g452239a74`, torch 2.13.0+cu130. - 2x ASUS Ascent GX10 (DGX Spark, GB10, 121 GB unified), ConnectX-7 direct attach, RoCE 200 Gb/s x2. Dri…”
sinmkd · GitHub · Sep 4, 2026 - 7
“So far I haven't seen a single model succeeding at transcribing sheet music, but I just tested it again with 5.6 Sol and it nailed the small test case. Fluently reading music requires multiple years of training for most people, but I feel like accurately following the horizontal lines trips up vision models in particular.”
kherud · Hacker News · Aug 17, 2026 - 8
“Ranks #8 of 29 on LMArena's document arena (score 1479), based on blind preference for document tasks.”
LMArena document arena · Benchmark · Jul 30, 2026 - 9
“The study used a temperature of 0.01. > "Thirteen food photographs were each submitted 495–561 times to four LLM vision APIs (GPT-5.4, Claude Sonnet 4.6, Gemini 2.5 Pro, Gemini 3.1 Pro Preview) using an identical structured prompt adapted from the iAPS automated insulin delivery system (26,904 total queries, temperature 0.01)"”
muwtyhg · Hacker News · Apr 29, 2026 - 10
“Ranks #6 of 29 on LMArena's document arena (score 1483), based on blind preference for document tasks.”
LMArena document arena · Benchmark · Jul 30, 2026 - 11
“## Who As a PairPocket user on the Gemini free tier, I want the app to prefer `gemini-3.6-flash` and automatically use `gemini-3.5-flash-lite` while 3.6 is quota-blocked, so that onboarding and receipt scans keep working without waiting on a failed 3.6 call every time. ## Why Free-tier limits are per project and per model row (RPM, TPM, RPD). RPD resets at midnight Pacific Time. RPM/TPM use a rolling window. Hard-coding quota numbers is unreliable; AI Studio shows live limits. Trying 3.6 on eve…”
minsikpaul92 · GitHub · Jul 28, 2026 - 12
“Ranks #9 of 68 on LMArena's vision arena (Elo 1285), based on human preference on image-understanding tasks.”
LMArena vision arena · Benchmark · Aug 27, 2026 - 13
“Ranks #5 of 68 on LMArena's vision arena (Elo 1292), based on human preference on image-understanding tasks.”
LMArena vision arena · Benchmark · Aug 27, 2026 - 14
“Resolved model commandcode/gpt-5.6-luna does not support image input. Configure a vision-capable model for mo... same for muse spark 1.2 and other vision models”
DiyarD · GitHub · Aug 9, 2026 - 15
“1. Gemini tends to reject whole_pdf OCR pretty frequently, switching to "image" mode helps get more documents OCRed without gettng flagged. 2. Despite #1, gemini will often reject the OCR text when "paperless-gpt-auto" is used (for the second LLM pass) to try and classify the scan 3. The v0.27 logs show "googleai GenerateContent API returned a candidate with no content parts" for whatever the prescribed number of empty worker responses is set by env. 4. After the retries, the paperless-gpt work…”
phaset · GitHub · Jul 28, 2026
Rankings synthesized from community evidence and open benchmarks. See methodology. Not driven by vendor marketing.