Recommendation for Vision / Documents
Vision & Documents
The best LLM for vision and document tasks is Claude Fable 5, which holds the top spot on LMArena's vision arena with an Elo of 1335. Anthropic dominates this category, claiming four of the top six positions, with Claude Opus 4.6 and Opus 4.7 rounding out the top three. Gemini 3.5 Flash posts a strong fourth-place finish, making it the highest-ranked non-Anthropic model and a compelling budget-conscious option. GPT-5.4 secures fifth place, while GPT-5.5 lands at sixth. Community reports add important caveats: even top-tier models can struggle with complex real-world document interpretation tasks like appliance manuals, so none of these are magic bullets. For most image understanding work, the Claude family offers the highest accuracy ceiling, while Gemini 3.5 Flash delivers solid performance at a lower price point.
About this recommendation
- Updated
- Jul 17, 2026
- Evidence through
- Jul 17, 2026
- Sources
- 13
- Revision
- v1
Claude Fable 5 sits at the top of LMArena's vision rankings, making it the strongest choice for image understanding tasks where accuracy matters most.
Best when: You need maximum accuracy for OCR, chart interpretation, or document parsing and cost is secondary.
Tips
- Holds the #1 position on LMArena's vision arena with the highest Elo score (1335) among 51 models.
- Top ranking based on human preference for image-understanding tasks suggests consistent real-world performance.
Claude Opus 4.7 takes second place in vision benchmarks, offering high-end performance just behind Fable 5 for demanding document and image workloads.
Best when: You want near-top-tier vision capabilities with Anthropic's Opus-level reasoning.
Tips
- Ranks #2 of 51 on LMArena's vision arena with a strong Elo of 1318.
- Human preference-based ranking reflects reliable performance on image-understanding tasks.
Claude Opus 4.6 narrowly trails its successor in third place, delivering excellent vision performance for document and screenshot analysis.
Best when: You need proven vision capabilities and Opus 4.7 isn't available or you prefer the older version's behavior.
Tips
- Ranks #3 of 51 on LMArena's vision arena with Elo 1317.
- Strong human preference scores for image-understanding tasks.
Gemini 3.5 Flash breaks the Anthropic sweep with a fourth-place ranking and praise for its price-to-performance ratio on image tasks.
Best when: You want strong vision performance at a lower cost, especially for image and video inputs.
Tips
- Ranks #4 of 51 on LMArena's vision arena with Elo 1309.
- Community feedback notes it performs well for image and video inputs relative to its price point.
- High ranking suggests solid capability across OCR and document parsing tasks.
Watch out for
- Community report of failing to interpret washing machine manual photos suggests limitations on complex real-world document tasks.
GPT-5.4 holds fifth place in vision rankings and has been used in research studies requiring reliable image analysis, making it a trustworthy option.
Best when: You prefer OpenAI's ecosystem or need a model validated in academic-style research contexts.
Tips
- Ranks #5 of 51 on LMArena's vision arena with Elo 1299.
- Used in a structured research study analyzing food photographs with 26,904 total queries, indicating reliability for systematic tasks.
- Selected alongside other top models for academic-grade vision API testing.
GPT-5.5 lands sixth on the vision leaderboard and gets community praise for visual consistency checks, though real-world document tasks remain hit-or-miss.
Best when: You need good-enough vision plus strong conversational ability and visual consistency checking.
Tips
- Ranks #6 of 51 on LMArena's vision arena with Elo 1297.
- Community feedback notes strengths in visual consistency and assessing whether graphics and visuals "look right."
- Solid mid-tier ranking supports general-purpose vision tasks.
Watch out for
- Community report indicates failure on a practical washing machine manual interpretation task.
Gemini 3.1 Pro Preview ranks seventh and appears in academic research alongside other leading vision models, showing real-world applicability.
Best when: You want a Gemini model for vision tasks that falls between Flash and Pro in capability.
Tips
- Ranks #7 of 51 on LMArena's vision arena with Elo 1296.
- Used in a research study analyzing food photographs with 26,904 queries, demonstrating reliability for structured image analysis.
- Community notes it works well for verifying visual consistency in outputs.
Watch out for
- Community feedback suggests it can make mistakes on its own outputs despite being useful for verification.
Claude Opus 4.8 sits eighth in rankings but demonstrates specific OCR strengths, including reading complex content without prompting guidance.
Best when: You need strong OCR capabilities on difficult documents where layout might otherwise confuse models.
Tips
- Ranks #8 of 51 on LMArena's vision arena with Elo 1295.
- Community report shows it successfully reads complex content with a single prompt and no instructions on how to interpret it.
Frequently asked
- Which LLM has the best image understanding?
- Claude Fable 5 ranks #1 on LMArena's vision arena with an Elo of 1335, making it the current leader for image understanding tasks based on human preference.
- Is Gemini 3.5 Flash good for vision tasks?
- Yes, Gemini 3.5 Flash ranks #4 on LMArena's vision arena (Elo 1309) and community feedback notes it performs well for image and video inputs relative to its price.
- How do GPT models compare to Claude for vision?
- GPT-5.4 ranks #5 and GPT-5.5 ranks #6 on LMArena's vision arena, trailing the top three Claude models but remaining competitive for most vision tasks.
- Can vision LLMs read appliance manuals with photos?
- Community reports indicate even top models like GPT-5.5, Gemini 3.5 Flash, and Claude Opus 4.8 can fail at practical tasks like interpreting washing machine manual photos, so expect mixed results on complex real-world documents.
Sources
- 1
“Ranks #1 of 51 on LMArena's vision arena (Elo 1335), based on human preference on image-understanding tasks.”
LMArena vision arena · Benchmark · Jul 12, 2026 - 2
“Ranks #2 of 51 on LMArena's vision arena (Elo 1318), based on human preference on image-understanding tasks.”
LMArena vision arena · Benchmark · Jul 12, 2026 - 3
“Ranks #3 of 51 on LMArena's vision arena (Elo 1317), based on human preference on image-understanding tasks.”
LMArena vision arena · Benchmark · Jul 12, 2026 - 4
“Ranks #4 of 51 on LMArena's vision arena (Elo 1309), based on human preference on image-understanding tasks.”
LMArena vision arena · Benchmark · Jul 12, 2026 - 5
“That's the point: for Gemini 3.5 Flash, its price does not correlate well with its performance. It's pretty good for image video inputs, though.”
minimaxir · Hacker News · Jul 8, 2026 - 6
“FYI: I just had three SOTA LLMs + NotebookLM all fail at the simple task of explaining to me where to put a powder detergent in my particular newly bought washing machine, despite having photos of the machine and ability to find the manual (in case of NotebookLM, it literally had the manual as its only source). After first failure (Gemini 3.5 Flash + NotebookLM), I run the other two (Opus 4.8 on Extra; GPT 5.5 on High) in parallel, and looking at their thought streams, I gave up and dug up the…”
TeMPOraL · Hacker News · Jun 30, 2026 - 7
“Ranks #5 of 51 on LMArena's vision arena (Elo 1299), based on human preference on image-understanding tasks.”
LMArena vision arena · Benchmark · Jul 12, 2026 - 8
“The study used a temperature of 0.01. > "Thirteen food photographs were each submitted 495–561 times to four LLM vision APIs (GPT-5.4, Claude Sonnet 4.6, Gemini 2.5 Pro, Gemini 3.1 Pro Preview) using an identical structured prompt adapted from the iAPS automated insulin delivery system (26,904 total queries, temperature 0.01)"”
muwtyhg · Hacker News · Apr 29, 2026 - 9
“Ranks #6 of 51 on LMArena's vision arena (Elo 1297), based on human preference on image-understanding tasks.”
LMArena vision arena · Benchmark · Jul 12, 2026 - 10
“For chatting and getting informations, and be corrected on things you're wrong without being reprimended by your own tool, GPT 5.5 5.6 is way better. Gemini 3.1 pro is surprisingly good at verifying your stuff, even though it's always making mistakes about its own stuff (don't ask him question, but ask him to verify your answer to the question). Same for graphics, visual consistency, anything around the "does the look make sense and is pleasing" really, which makes claude design such a (good) s…”
nolok · Hacker News · Jul 11, 2026 - 11
“Ranks #7 of 51 on LMArena's vision arena (Elo 1296), based on human preference on image-understanding tasks.”
LMArena vision arena · Benchmark · Jul 12, 2026 - 12
“Ranks #8 of 51 on LMArena's vision arena (Elo 1295), based on human preference on image-understanding tasks.”
LMArena vision arena · Benchmark · Jul 12, 2026 - 13
“Claude Opus 4.8 can read it with a single prompt and no instructions on how to read it. https: ibb.co WWMSXQkQ”
dhruvkb · Hacker News · Jul 11, 2026
Rankings synthesized from community evidence and open benchmarks. See methodology. Not driven by vendor marketing.