Recommendation for Documents & forms

Document Parsing

The best LLM for document parsing is Claude Fable 5, which holds the top spot on LMArena's vision arena and combines that with second-place text capabilities. Claude Opus 4.7 and Claude Opus 4.6 follow closely, taking second and third in the rankings with strong vision and text scores. Document parsing demands precise visual understanding for layout and OCR, plus text reasoning to extract structured data accurately. The vision arena rankings are the strongest signal here, since they reflect human preference on image-understanding tasks directly relevant to parsing PDFs, forms, and invoices. Anthropic models dominate the top of the vision leaderboard, with Gemini 3.5 Flash and GPT-5.4 providing the strongest non-Anthropic alternatives at ranks 4 and 5 respectively.

About this recommendation

Updated
Jul 17, 2026
Evidence through
Jul 17, 2026
Sources
16
Revision
v1
  1. Claude Fable 5 is the top choice for document parsing, leading the vision arena at #1 (Elo 1335) while sitting at #2 for text (Elo 1493). That combination makes it ideal for extracting structured data from documents where visual layout matters.

    Best when: You need the best visual document understanding for complex layouts, tables, and scanned forms.

    Tips

    • Ranks #1 of 51 on LMArena's vision arena, the highest score for image-understanding tasks.
      Source 1
      Ranks #1 of 51 on LMArena's vision arena (Elo 1335), based on human preference on image-understanding tasks.
      LMArena vision arenaOpen original ↗
    • Holds #2 of 49 on text arena (Elo 1493), ensuring strong reasoning over extracted content.
      Source 2
      Ranks #2 of 17 on LMArena's overall text arena (Elo 1504), based on blind human preference votes.
      LMArena text arenaOpen original ↗
  2. Claude Opus 4.7 takes second place with #2 vision (Elo 1318) and #3 text rankings, making it a strong all-around performer for document workloads where you want slightly more text emphasis.

    Best when: You want excellent vision capability balanced with near-top-tier text reasoning for complex extraction logic.

    Tips

    • Ranks #2 of 51 on LMArena's vision arena, just behind Claude Fable 5.
      Source 3
      Ranks #2 of 51 on LMArena's vision arena (Elo 1318), based on human preference on image-understanding tasks.
      LMArena vision arenaOpen original ↗
    • Holds #3 of 49 on text arena (Elo 1490), competitive with the top text models.
      Source 4
      Ranks #4 of 17 on LMArena's overall text arena (Elo 1499), based on blind human preference votes.
      LMArena text arenaOpen original ↗
  3. Claude Opus 4.6 leads the text arena at #1 and ranks #3 in vision, suited for parsing tasks where the extraction logic is complex but the visual layout is straightforward.

    Best when: Your document parsing requires heavy post-extraction reasoning and you still need solid vision capability.

    Tips

    • Ranks #1 of 49 on LMArena's text arena (Elo 1501), the strongest text model available.
      Source 5
      Ranks #1 of 17 on LMArena's overall text arena (Elo 1512), based on blind human preference votes.
      LMArena text arenaOpen original ↗
    • Sits at #3 of 51 on vision arena (Elo 1317), still within the top tier for image understanding.
      Source 6
      Ranks #3 of 51 on LMArena's vision arena (Elo 1317), based on human preference on image-understanding tasks.
      LMArena vision arenaOpen original ↗
  4. Gemini 3.5 Flash is the highest-ranked non-Anthropic option at #4 on vision (Elo 1309) and #5 on text, offering a strong alternative if you prefer Google's ecosystem or pricing.

    Best when: You want competitive document parsing performance on Google infrastructure or at lower cost.

    Tips

    • Ranks #4 of 51 on LMArena's vision arena (Elo 1309), close to the Anthropic leaders.
      Source 7
      Ranks #4 of 51 on LMArena's vision arena (Elo 1309), based on human preference on image-understanding tasks.
      LMArena vision arenaOpen original ↗
    • Holds #5 of 49 on text arena (Elo 1480), maintaining strong reasoning capability.
      Source 8
      Ranks #8 of 17 on LMArena's overall text arena (Elo 1490), based on blind human preference votes.
      LMArena text arenaOpen original ↗
  5. GPT-5.4 ranks #5 in vision and #9 in text, a solid pick if you are already invested in OpenAI's API and need document parsing alongside other GPT-based workflows.

    Best when: Your stack is built on OpenAI and you need document parsing integrated with existing GPT pipelines.

    Tips

    • Ranks #5 of 51 on LMArena's vision arena (Elo 1299), in the top tier for image understanding.
      Source 9
      Ranks #5 of 51 on LMArena's vision arena (Elo 1299), based on human preference on image-understanding tasks.
      LMArena vision arenaOpen original ↗
    • Sits at #9 of 49 on text arena (Elo 1470), with respectable reasoning capability.
      Source 10
      Ranks #3 of 17 on LMArena's overall text arena (Elo 1499), based on blind human preference votes.
      LMArena text arenaOpen original ↗
  6. GPT-5.5 ranks #6 in vision and #10 in text, slightly trailing GPT-5.4 on both leaderboards but potentially useful for specific OpenAI feature integrations.

    Best when: You need a GPT model with vision capability and prefer the feature set or latency profile of the 5.5 variant.

    Tips

    • Ranks #6 of 51 on LMArena's vision arena (Elo 1297), just behind GPT-5.4.
      Source 11
      Ranks #6 of 51 on LMArena's vision arena (Elo 1297), based on human preference on image-understanding tasks.
      LMArena vision arenaOpen original ↗
    • Holds #10 of 49 on text arena (Elo 1470), providing competent text reasoning.
      Source 12
      Ranks #6 of 17 on LMArena's overall text arena (Elo 1491), based on blind human preference votes.
      LMArena text arenaOpen original ↗
  7. Gemini 3.1 Pro Preview ranks #7 in vision and #6 in text, offering balanced performance but sitting behind Gemini 3.5 Flash on the vision leaderboard that matters most for document parsing.

    Best when: You want a Google model with slightly higher text ranking than 3.5 Flash and can accept marginally lower vision scores.

    Tips

    • Ranks #7 of 51 on LMArena's vision arena (Elo 1296), within the top 10 for image understanding.
      Source 13
      Ranks #7 of 51 on LMArena's vision arena (Elo 1296), based on human preference on image-understanding tasks.
      LMArena vision arenaOpen original ↗
    • Sits at #6 of 49 on text arena (Elo 1479), strong for reasoning tasks.
      Source 14
      Ranks #9 of 17 on LMArena's overall text arena (Elo 1489), based on blind human preference votes.
      LMArena text arenaOpen original ↗
  8. Claude Opus 4.8 ranks #8 in vision and #13 in text, trailing the other Anthropic Claude models on both metrics but still viable for document tasks if you need this specific model version.

    Best when: You have a specific requirement for Claude Opus 4.8's behavior or are limited to this model in your environment.

    Tips

    • Ranks #8 of 51 on LMArena's vision arena (Elo 1295), within the top 10.
      Source 15
      Ranks #8 of 51 on LMArena's vision arena (Elo 1295), based on human preference on image-understanding tasks.
      LMArena vision arenaOpen original ↗
    • Holds #13 of 49 on text arena (Elo 1462), acceptable for most extraction workflows.
      Source 16
      Ranks #17 of 17 on LMArena's overall text arena (Elo 1472), based on blind human preference votes.
      LMArena text arenaOpen original ↗

Frequently asked

Which model is best for OCR and layout understanding?
Claude Fable 5 ranks #1 on LMArena's vision arena (Elo 1335), making it the top choice for visual tasks including OCR and document layout analysis.
Do I need a model with strong vision capabilities for document parsing?
Yes, extracting structured data from PDFs and scanned documents requires strong image understanding. The vision arena rankings directly measure this capability through human preference on image tasks.
What's a good budget-friendly option for document parsing?
Gemini 3.5 Flash ranks #4 on vision arena (Elo 1309) and #5 on text, offering solid document parsing performance at a lower price point than the top Anthropic models.
How important are text rankings for document parsing?
Text rankings matter for reasoning over extracted content and formatting structured output. Claude Opus 4.6 leads text at #1, while Claude Fable 5 sits at #2, giving the edge to Fable for a balanced skillset.

Sources

  1. 1

    Ranks #1 of 51 on LMArena's vision arena (Elo 1335), based on human preference on image-understanding tasks.

    LMArena vision arena · Benchmark · Jul 12, 2026
  2. 2

    Ranks #2 of 17 on LMArena's overall text arena (Elo 1504), based on blind human preference votes.

    LMArena text arena · Benchmark · Jul 20, 2026
  3. 3

    Ranks #2 of 51 on LMArena's vision arena (Elo 1318), based on human preference on image-understanding tasks.

    LMArena vision arena · Benchmark · Jul 12, 2026
  4. 4

    Ranks #4 of 17 on LMArena's overall text arena (Elo 1499), based on blind human preference votes.

    LMArena text arena · Benchmark · Jul 20, 2026
  5. 5

    Ranks #1 of 17 on LMArena's overall text arena (Elo 1512), based on blind human preference votes.

    LMArena text arena · Benchmark · Jul 20, 2026
  6. 6

    Ranks #3 of 51 on LMArena's vision arena (Elo 1317), based on human preference on image-understanding tasks.

    LMArena vision arena · Benchmark · Jul 12, 2026
  7. 7

    Ranks #4 of 51 on LMArena's vision arena (Elo 1309), based on human preference on image-understanding tasks.

    LMArena vision arena · Benchmark · Jul 12, 2026
  8. 8

    Ranks #8 of 17 on LMArena's overall text arena (Elo 1490), based on blind human preference votes.

    LMArena text arena · Benchmark · Jul 20, 2026
  9. 9

    Ranks #5 of 51 on LMArena's vision arena (Elo 1299), based on human preference on image-understanding tasks.

    LMArena vision arena · Benchmark · Jul 12, 2026
  10. 10

    Ranks #3 of 17 on LMArena's overall text arena (Elo 1499), based on blind human preference votes.

    LMArena text arena · Benchmark · Jul 20, 2026
  11. 11

    Ranks #6 of 51 on LMArena's vision arena (Elo 1297), based on human preference on image-understanding tasks.

    LMArena vision arena · Benchmark · Jul 12, 2026
  12. 12

    Ranks #6 of 17 on LMArena's overall text arena (Elo 1491), based on blind human preference votes.

    LMArena text arena · Benchmark · Jul 20, 2026
  13. 13

    Ranks #7 of 51 on LMArena's vision arena (Elo 1296), based on human preference on image-understanding tasks.

    LMArena vision arena · Benchmark · Jul 12, 2026
  14. 14

    Ranks #9 of 17 on LMArena's overall text arena (Elo 1489), based on blind human preference votes.

    LMArena text arena · Benchmark · Jul 20, 2026
  15. 15

    Ranks #8 of 51 on LMArena's vision arena (Elo 1295), based on human preference on image-understanding tasks.

    LMArena vision arena · Benchmark · Jul 12, 2026
  16. 16

    Ranks #17 of 17 on LMArena's overall text arena (Elo 1472), based on blind human preference votes.

    LMArena text arena · Benchmark · Jul 20, 2026

Rankings synthesized from community evidence and open benchmarks. See methodology. Not driven by vendor marketing.