Recommendation for Documents & forms
Document Parsing
Our top recommendation for Document Parsing, based on the public evidence we track, is OpenAI: GPT-5.6 Sol.[1][2][3] Deploy for technical document parsing involving hex dumps, packet analysis, or code decompilation where precise offset extraction matters. Watch out: Watch for sloppy handling on structured data tasks like database migration and health bench evaluations, where it may miss details that other models catch. Anthropic: Claude Fable 5 is the next-ranked alternative. Choose for general document parsing where high blind preference rankings suggest reliable output quality on typical document tasks.
About this recommendation
- Updated
- Sep 4, 2026
- Evidence through
- Sep 4, 2026
- Sources
- 13
- Revision
- v58
Decision audit
Why this result
Inspect the inputs and the computed order behind the recommendation.
Models screened
21
live candidates
Evaluation feeds
5
task-weighted
Winner coverage
37%
intended feed weight
Largest provider share
3 of 8
Anthropic
Sources evaluated
The task sets these weights before any model is scored.
| Evaluation feed | Weight | Winner result | Field measured |
|---|---|---|---|
| VLMEvalKit tasksunavailable | 40% | feed unavailable | 0/21 |
| LMArena Document | 20% | #6 | 11/21 |
| LMArena Vision | 15% | #10 | 18/21 |
| Structured-output evalunavailable | 15% | feed unavailable | 0/21 |
| OpenRouter usage | 10% | 97/100 | 21/21 |
Provider concentration
Each exact model is scored separately; provider identity is not a ranking input.
- Anthropic3 models
- Google2 models
- deepseek1 model
- Meta1 model
- OpenAI1 model
Decision table
Every published model is shown in computed order. Practitioner sources are distinct community threads, not the citations repeated in the prose below.
| Rank | Model | Relative score | Coverage | Practitioner evidence | Strongest measured reason |
|---|---|---|---|---|---|
| 01 | GPT-5.6 SolOpenAI | 57 | 37% | 2 threads · 1 families · 0 cautions | #6 LMArena Document · #10 LMArena Vision |
| 02 | Claude Fable 5Anthropic | 55 | 37% | 1 threads · 1 families · 1 cautions | #1 LMArena Vision · #3 LMArena Document |
| 03 | Claude Opus 4.6Anthropic | 54 | 37% | no linked practitioner threads | #2 LMArena Document · #4 LMArena Vision |
| 04 | DeepSeek V4 Flash Vision Expdeepseek | 52 | 11% | 2 threads · 1 families · 0 cautions | OpenRouter usage 90/100 normalized |
| 05 | Gemini 2.5 FlashGoogle | 51 | 28% | 2 threads · 2 families · 0 cautions | #39 LMArena Vision |
| 06 | Claude Sonnet 4.6Anthropic | 50 | 37% | no linked practitioner threads | #5 LMArena Document · #14 LMArena Vision |
| 07 | Gemini 3.6 FlashGoogle | 49 | 28% | 1 threads · 1 families · 1 cautions | #7 LMArena Vision |
| 08 | Muse Spark 1.2Meta | 49 | 28% | 1 threads · 1 families · 0 cautions | #5 LMArena Vision |
Relative score combines normalized benchmark quality and signal coverage; independent practitioner evidence and freshness are bounded tie-breakers. It is an ordering score, not an absolute quality percentage. The writing model receives this order and cannot change it.
GPT-5.6 Sol ranks eighth on LMArena's document arena with a score of 1479, and has demonstrated practical extraction capabilities in technical hex decoding and decompilation tasks.
Best when: Deploy for technical document parsing involving hex dumps, packet analysis, or code decompilation where precise offset extraction matters.
Tips
- Deploy for technical document parsing involving hex dumps, packet analysis, or code decompilation where precise offset extraction matters.
Watch out for
- Watch for sloppy handling on structured data tasks like database migration and health bench evaluations, where it may miss details that other models catch.
Claude Fable 5 ranks third on LMArena's document arena with a score of 1504, showing strong blind preference for document tasks, though it refuses many questions in specialized life science and medical chemistry benchmarks.
Best when: Choose for general document parsing where high blind preference rankings suggest reliable output quality on typical document tasks.
Tips
- Choose for general document parsing where high blind preference rankings suggest reliable output quality on typical document tasks.
Watch out for
- Expect refusals on life science, gene analysis, and medical chemistry document parsing due to alignment constraints that trigger rejection of majority questions in those domains.
Claude Opus 4.6 achieves second place on LMArena's document arena with a score of 1510, the highest document-specific ranking among all candidates.
Best when: Prioritize for document parsing where maximum blind human preference matters, as it leads the document arena benchmark.
Tips
- Prioritize for document parsing where maximum blind human preference matters, as it leads the document arena benchmark.
DeepSeek V4 Flash Vision Exp delivers production visual qualification in 9-13 seconds per call with strict schema extraction and multipage support, backed by implemented infrastructure for timeout, retry, and fallback observability.
Best when: Use for high-throughput invoice and form processing where p50 latency under 15 seconds per call and strict schema compliance are required, with built-in retry and fallback mechanisms.
Tips
- Use for high-throughput invoice and form processing where p50 latency under 15 seconds per call and strict schema compliance are required, with built-in retry and fallback mechanisms.
Gemini 2.5 Flash supports structured JSON output with schema validation for release note parsing and document decomposition, though it faces planned deprecation with eventual price increases.
Best when: Deploy for structured document decomposition tasks requiring JSON Schema output, such as breaking release notes into categorized change units with impact and migration assessments.
Tips
- Deploy for structured document decomposition tasks requiring JSON Schema output, such as breaking release notes into categorized change units with impact and migration assessments.
Watch out for
- Plan migration before October 2026, as this model enters ELA with price protection ending January 2027 and potential region availability changes that could break cost assumptions.
Claude Sonnet 4.6 places sixth in LMArena's document arena with a score of 1483, indicating solid blind human preference for document tasks among 29 models.
Best when: Use for document parsing workflows where you need a balance of capability and cost within the Claude family, as it ranks competitively on blind preference for document tasks.
Tips
- Use for document parsing workflows where you need a balance of capability and cost within the Claude family, as it ranks competitively on blind preference for document tasks.
Gemini 3.6 Flash ranks ninth on LMArena's vision arena with an Elo of 1285.
Best when: Consider for general vision understanding tasks where its vision arena ranking suggests acceptable baseline performance.
Tips
- Consider for general vision understanding tasks where its vision arena ranking suggests acceptable baseline performance.
Watch out for
- Do not rely on strict prompt compliance for document parsing, as live evaluation shows consistent failure to follow explicit 'MUST include' mandates across all thinking budget configurations.
Muse Spark 1.2 ranks fifth on LMArena's vision arena with an Elo of 1292, and is deployed specifically for second-stage validation to catch omissions and weak statements missed by primary extraction.
Best when: Add as a secondary validation layer after initial document extraction to scan for missed list items, named properties, and weakened statements that primary models skip.
Tips
- Add as a secondary validation layer after initial document extraction to scan for missed list items, named properties, and weakened statements that primary models skip.
Frequently asked
- What is the top-ranked model for Document Parsing?
- OpenAI: GPT-5.6 Sol ranks first in the current evidence-weighted comparison. Deploy for technical document parsing involving hex dumps, packet analysis, or code decompilation where precise offset extraction matters.[1]
- What should I watch out for with OpenAI: GPT-5.6 Sol?
- Watch for sloppy handling on structured data tasks like database migration and health bench evaluations, where it may miss details that other models catch.[2]
- What is an alternative to OpenAI: GPT-5.6 Sol?
- Anthropic: Claude Fable 5 is the next-ranked option. Choose for general document parsing where high blind preference rankings suggest reliable output quality on typical document tasks.[3]
Sources
- 1
“It's good enough at decoding hex from some packet dumps. And I was doing that even with 5.5. And it was good at decompiling some code (with tools) and searching for offsets of buffers and commands. Found viable exploit that allowed me to rescue broken update system in devices I was maintaining for my company (it was broken by chatgpt forgetting -v in hexdump, heh).”
yetihehe · Hacker News · Aug 13, 2026 - 2
“Data Science Tasks (Internal) doesn't include time for Astra... same for Database Migration Tasks (Internal)... But does for gpt 5.6 sol.... which is funny. Same for HealthBench Professional and a few others. Clearly either OpenAI is very sloppy or GPT-6 Astra is also sloppy.”
datadrivenangel · Hacker News · Sep 4, 2026 - 3
“Ranks #3 of 29 on LMArena's document arena (score 1504), based on blind preference for document tasks.”
LMArena document arena · Benchmark · Jul 30, 2026 - 4
“"Claude Fable 5 and 5.1 are not included in LifeSciBench Gold v1, GeneBench Pro v13, and MedChemBench because they refuse the majority of questions in these evaluations.12" Sounds about right. Alignment is important, but also being able to do mundane tasks is important too.”
HDBaseT · Hacker News · Sep 3, 2026 - 5
“Ranks #2 of 29 on LMArena's document arena (score 1510), based on blind preference for document tasks.”
LMArena document arena · Benchmark · Jul 30, 2026 - 6
“## Objetivo Reducir el tiempo real de calificación, digitalización y generación de presentaciones sin sacrificar calidad, trazabilidad ni perder solicitudes en curso. ## Evidencia de producción (últimos 14 días) - Calificación completa: p50 157 s; p95 516 s. - Digitalización: p50 189 s. - Presentación reciente: 366 s. - DeepSeek V4 Flash Vision Exp en calificación visual: ~9–13 s por llamada exitosa. - Presentaciones con ese modelo: ~95 s por llamada y hasta tres llamadas por regeneración/revis…”
Andres-back · GitHub · Sep 4, 2026 - 7
“Implementar specs/020-deepseek-vision: extractor desacoplado, multipágina, schema estricto, timeout/retry/fallback observable y benchmarks directos/backend. La solicitud detallada del usuario constituye aprobación de spec y plan.”
Andres-back · GitHub · Aug 24, 2026 - 8
“## 概要 GitHub Releaseの原文をGemini APIで解析し、変更単位の日本語要約・Breaking Change・Impact・Migrationを構造化して保存する。 ## 対応内容 - [ ] Gemini API連携 - [ ] `gemini-2.5-flash` 利用 - [ ] JSON Schema / structured output - [ ] Release Notesを変更単位へ分解 - [ ] 日本語タイトル・要約 - [ ] category判定 - [ ] impact判定 - [ ] breaking判定 - [ ] migration判定 - [ ] Gemini API失敗時のリトライ - [ ] APIキーをGitHub Secretsで管理 - [ ] Phase 2の自動収集との統合 - [ ] テスト ## 出力 `changes`, `category`, `title`, `summary`, `impact`, `breaking`, `migration` を構造化JSONとして保存する。 ## 完了条件 - Re…”
haruki33 · GitHub · Aug 17, 2026 - 9
“## 概要 Vertex AI(Gemini Enterprise Agent Platform)の `gemini-2.5-flash` が 2026-10-20 に廃止(ELA 入り)予定のため、Gemini 3 系への移行先モデルを実測で選定し、ADR にまとめる。 Google からの通知では影響プロジェクトとして `documentaisample-488504` が名指しされており、本リポジトリの Gemini 抽出経路が対象。 - 2026-10-20: ELA 入り(呼び出しは継続可能・価格据え置き) - 2027-01-28: ELA 標準価格終了、大幅値上げ+リージョン提供状況の変更可能性 呼び出しが即停止するわけではないため障害リスクは低いが、ADR-0010 が記録したコスト前提(`thinkingBudget:0` での平均 $0.00175/枚)が失効するため、実測をやり直す必要がある。 ## 変更内容 本 Issue のスコープは **移行先モデルの実測比較と方針決定(ADR)まで**。既定モデルの切替コードは別 Issue で対応する。 1. 移行先候…”
git-berian · GitHub · Aug 7, 2026 - 10
“Ranks #6 of 29 on LMArena's document arena (score 1483), based on blind preference for document tasks.”
LMArena document arena · Benchmark · Jul 30, 2026 - 11
“Ranks #9 of 68 on LMArena's vision arena (Elo 1285), based on human preference on image-understanding tasks.”
LMArena vision arena · Benchmark · Aug 27, 2026 - 12
“## Summary A live 130-cell eval (`cmd/eval_thinking_matrix`, [issue #80](https://github.com/weitzer-org/sound-profile-builder/issues/80)) found that **~85-90% of Cabinet blocks violate `12_architect_v4.md` Rule 12's explicit "MUST include... do not drop" mic-placement mandate**, uniformly across all 3 models (`gemini-3.5-flash`, `gemini-3.6-flash`, `gemini-3.1-pro-preview`) and all 4 thinking budgets tested. This is a pure prompt-compliance bug, independent of the separate open question (tracke…”
benw307 · GitHub · Jul 26, 2026 - 13
“## Problem Statement 当前 Entry Module 单阶段提取在 5 案例上均值 ~36.2(Sol 等效),距 flawless 档(≥40)差 4 分。剩余扣分第一来源是漏提取与陈述弱化:Skein1 的函子公理清单、RT 的 handle-slide 命名性质等以 prose 清单/命名性质形式存在的内容被跳过;部分已提取条目的平衡张量积记号被弱化、或幻觉出原文没有的边界条件。连续 4 个提示词单变量实验(v1.32–1.35)证明在同一容量内加规则只会搬分数,净零。 ## Solution 在现有工作流(MinerU → v1.31 提取 → 确定性合并 → 产物)后新增第二阶段:用 muse-spark-1.2 对初版产物做校验补漏。阶段二输入为初版产物 + 源文本全文,仅做两件事:1) 扫描源文本中以清单/命名性质形式存在但被初版遗漏的条目,补 1–3 条;2) 逐条比对初版中被评委判为弱化的陈述与源文本原文,逐字订正。不改提示词,不改窗口。 ## User Stories 1. As a benchmark operator, I want the…”
cyw6130 · GitHub · Aug 21, 2026
Rankings synthesized from community evidence and open benchmarks. See methodology. Not driven by vendor marketing.