Recommendation for Image understanding

Image Understanding

The best LLM for image understanding is Claude Fable 5, which holds the top spot on LMArena's vision arena with an Elo of 1335. It leads a crowded field where Anthropic models dominate the top four positions. Claude Opus 4.7 and Claude Opus 4.6 take second and third place, separated by a single Elo point. Google's Gemini 3.5 Flash breaks up the Anthropic sweep at fourth place with an Elo of 1309, while GPT-5.4 and GPT-5.5 hold down fifth and sixth. The rankings are based on human preference data from LMArena's vision arena, which tests models across photo description, visual question answering, and reasoning tasks. The gap between first and sixth place is significant: 38 Elo points separate Claude Fable 5 from GPT-5.5. For builders, Claude Fable 5 is the one to try first, but Gemini 3.5 Flash is worth a look if you need a non-Anthropic option or are already in the Google ecosystem.

About this recommendation

Updated
Jul 17, 2026
Evidence through
Jul 17, 2026
Sources
11
Revision
v1
  1. Claude Fable 5 is the top choice for general image understanding. It holds the number one spot on LMArena's vision arena with an Elo of 1335, which puts it 17 points ahead of the next competitor. If you need the strongest performer for describing photos, answering visual questions, or reasoning about images, this is the model to reach for first.

    Best when: You want the absolute best image understanding performance and are willing to use Anthropic's API.

    Tips

    • Ranks first out of 51 models on LMArena's vision arena.
      Source 1
      Ranks #1 of 51 on LMArena's vision arena (Elo 1335), based on human preference on image-understanding tasks.
      LMArena vision arenaOpen original ↗
    • Holds the highest Elo score (1335) among all evaluated models, giving it a comfortable lead.
      Source 1
      Ranks #1 of 51 on LMArena's vision arena (Elo 1335), based on human preference on image-understanding tasks.
      LMArena vision arenaOpen original ↗
  2. Claude Opus 4.7 takes second place with an Elo of 1318, just one point ahead of its sibling Claude Opus 4.6. The narrow gap suggests similar capability, so if you are already using Opus-class models for other tasks, this slides naturally into an existing workflow.

    Best when: You want near-top-tier vision performance and are already invested in the Claude Opus family for other tasks.

    Tips

    • Ranks second out of 51 models on LMArena's vision arena.
      Source 2
      Ranks #2 of 51 on LMArena's vision arena (Elo 1318), based on human preference on image-understanding tasks.
      LMArena vision arenaOpen original ↗
    • Elo of 1318 keeps it competitive with the leader while staying ahead of most alternatives.
      Source 2
      Ranks #2 of 51 on LMArena's vision arena (Elo 1318), based on human preference on image-understanding tasks.
      LMArena vision arenaOpen original ↗
  3. Claude Opus 4.6 sits in third place with an Elo of 1317, effectively tied with Claude Opus 4.7. It is one of four Anthropic models in the top four, which says a lot about Anthropic's current strength in multimodal tasks.

    Best when: You need high-quality image understanding and may prefer a more established Opus model over newer variants.

    Tips

    • Ranks third out of 51 models on LMArena's vision arena.
      Source 3
      Ranks #3 of 51 on LMArena's vision arena (Elo 1317), based on human preference on image-understanding tasks.
      LMArena vision arenaOpen original ↗
  4. Gemini 3.5 Flash is the highest-ranked non-Anthropic model at number four, with an Elo of 1309. It is eight points behind Claude Opus 4.6 but only ten points ahead of GPT-5.4, placing it in a solid middle position among top-tier options.

    Best when: You want strong vision capability but prefer Google's ecosystem or need to avoid vendor lock-in with Anthropic.

    Tips

    • Ranks fourth out of 51 models on LMArena's vision arena.
      Source 4
      Ranks #4 of 51 on LMArena's vision arena (Elo 1309), based on human preference on image-understanding tasks.
      LMArena vision arenaOpen original ↗
    • Best-performing model from Google in this evaluation set.
      Source 4
      Ranks #4 of 51 on LMArena's vision arena (Elo 1309), based on human preference on image-understanding tasks.
      LMArena vision arenaOpen original ↗
  5. GPT-5.4 ranks fifth with an Elo of 1299, leading the OpenAI entries in this list. It has been used in academic research analyzing food photographs, which suggests it handles practical image analysis tasks well enough for repeated use.

    Best when: You are building on OpenAI's infrastructure and need proven vision capability backed by real-world research use.

    Tips

    • Ranks fifth out of 51 models on LMArena's vision arena.
      Source 5
      Ranks #5 of 51 on LMArena's vision arena (Elo 1299), based on human preference on image-understanding tasks.
      LMArena vision arenaOpen original ↗
    • Used in a food photography study with over 26,000 queries, showing it can handle production-scale tasks.
      Source 6
      The study used a temperature of 0.01. > "Thirteen food photographs were each submitted 495–561 times to four LLM vision APIs (GPT-5.4, Claude Sonnet 4.6, Gemini 2.5 Pro, Gemini 3.1 Pro Preview) using an identical structured prompt adapted from the iAPS automated insulin delivery system (26,904 total queries, temperature 0.01)"
  6. GPT-5.5 takes sixth place with an Elo of 1297, two points behind GPT-5.4. The difference is marginal, so choosing between them may come down to pricing, latency, or other factors not captured in the arena rankings.

    Best when: You want OpenAI's latest model and are willing to accept slightly lower vision benchmarks for potential improvements elsewhere.

    Tips

    • Ranks sixth out of 51 models on LMArena's vision arena.
      Source 7
      Ranks #6 of 51 on LMArena's vision arena (Elo 1297), based on human preference on image-understanding tasks.
      LMArena vision arenaOpen original ↗
    • Elo of 1297 keeps it competitive within the top tier.
      Source 7
      Ranks #6 of 51 on LMArena's vision arena (Elo 1297), based on human preference on image-understanding tasks.
      LMArena vision arenaOpen original ↗
  7. Gemini 3.1 Pro Preview ranks seventh with an Elo of 1296. Like other models in this tier, it has seen research use for food photography analysis, making it a reasonable choice if you are testing Google's offerings.

    Best when: You want a preview model from Google with strong vision capability and are comfortable with potential API stability concerns.

    Tips

    • Ranks seventh out of 51 models on LMArena's vision arena.
      Source 8
      Ranks #7 of 51 on LMArena's vision arena (Elo 1296), based on human preference on image-understanding tasks.
      LMArena vision arenaOpen original ↗
    • Used alongside other models in food photography research, indicating practical applicability.
      Source 6
      The study used a temperature of 0.01. > "Thirteen food photographs were each submitted 495–561 times to four LLM vision APIs (GPT-5.4, Claude Sonnet 4.6, Gemini 2.5 Pro, Gemini 3.1 Pro Preview) using an identical structured prompt adapted from the iAPS automated insulin delivery system (26,904 total queries, temperature 0.01)"
  8. MiniMax M3 sits at nineteenth place with an Elo of 1260, well behind the leaders. But community discussion highlights a different use case: it works well as part of a multi-model pipeline where one model handles vision and another handles reasoning. If you are building an orchestration system rather than calling a single API, this approach is worth considering.

    Best when: You are building a multi-model pipeline and want an affordable vision processor that can feed outputs to a larger reasoning model.

    Tips

    • Community users report success using it as a dedicated vision processor in multi-model setups.
      Source 9
      If you use a good harness or add the right tools and plugins both the image and web search issues mentioned are non-issues. oh-my-pi (omp.sh) handles images for text models out of the box - as long as you have any vision capable provider enabled, it will be used when you paste images to a text model. Rather than let it guess I configured it to use MiniMax M3 for this task (as well as other utility tasks like code exploration & library functions). opencode has plugins that do the same thing, but…
    • Open weights make it affordable for specialized roles in a model ensemble.
      Source 10
      That's right, but there are other recent open weights and relatively big LLMs that are multimodal, e.g. MiniMax-M3. With open weights LLMs, it is affordable to use many different models, each for whatever it is better. Moreover, for analyzing "UIs, photos, screenshots, etc." there are small models that can be run locally on smartphones or laptops, e.g. IBM granite-vision-4.1-4B, certain Google Gemma 4 variants and certain Qwen variants, whose output you can use as input for a big LLM, in order…

    Watch out for

    • Ranks nineteenth out of 51 on LMArena's vision arena, significantly behind the leaders.
      Source 11
      Ranks #19 of 51 on LMArena's vision arena (Elo 1260), based on human preference on image-understanding tasks.
      LMArena vision arenaOpen original ↗
    • Requires extra tooling or harness setup to use effectively as a vision component.
      Source 9
      If you use a good harness or add the right tools and plugins both the image and web search issues mentioned are non-issues. oh-my-pi (omp.sh) handles images for text models out of the box - as long as you have any vision capable provider enabled, it will be used when you paste images to a text model. Rather than let it guess I configured it to use MiniMax M3 for this task (as well as other utility tasks like code exploration & library functions). opencode has plugins that do the same thing, but…

Frequently asked

Which model is best at understanding images?
Claude Fable 5 ranks first on LMArena's vision arena with an Elo of 1335, making it the strongest performer based on human preference for image understanding tasks.
How do GPT and Claude compare for vision tasks?
Claude models currently dominate: Claude Fable 5 ranks first while Claude Opus 4.7 and 4.6 take second and third. GPT-5.4 and GPT-5.5 sit at fifth and sixth place, roughly 35-40 Elo points behind the leader.
Is Gemini good at image understanding?
Gemini 3.5 Flash ranks fourth overall with an Elo of 1309, making it the strongest non-Anthropic option and competitive with top-tier models.
Which vision model should I use for food photography analysis?
Research studies have tested GPT-5.4, Claude Sonnet 4.6, and Gemini 3.1 Pro Preview for food photography analysis, suggesting any of these are reasonable choices for that use case.

Sources

  1. 1

    Ranks #1 of 51 on LMArena's vision arena (Elo 1335), based on human preference on image-understanding tasks.

    LMArena vision arena · Benchmark · Jul 12, 2026
  2. 2

    Ranks #2 of 51 on LMArena's vision arena (Elo 1318), based on human preference on image-understanding tasks.

    LMArena vision arena · Benchmark · Jul 12, 2026
  3. 3

    Ranks #3 of 51 on LMArena's vision arena (Elo 1317), based on human preference on image-understanding tasks.

    LMArena vision arena · Benchmark · Jul 12, 2026
  4. 4

    Ranks #4 of 51 on LMArena's vision arena (Elo 1309), based on human preference on image-understanding tasks.

    LMArena vision arena · Benchmark · Jul 12, 2026
  5. 5

    Ranks #5 of 51 on LMArena's vision arena (Elo 1299), based on human preference on image-understanding tasks.

    LMArena vision arena · Benchmark · Jul 12, 2026
  6. 6

    The study used a temperature of 0.01. > "Thirteen food photographs were each submitted 495–561 times to four LLM vision APIs (GPT-5.4, Claude Sonnet 4.6, Gemini 2.5 Pro, Gemini 3.1 Pro Preview) using an identical structured prompt adapted from the iAPS automated insulin delivery system (26,904 total queries, temperature 0.01)"

    muwtyhg · Hacker News · Apr 29, 2026
  7. 7

    Ranks #6 of 51 on LMArena's vision arena (Elo 1297), based on human preference on image-understanding tasks.

    LMArena vision arena · Benchmark · Jul 12, 2026
  8. 8

    Ranks #7 of 51 on LMArena's vision arena (Elo 1296), based on human preference on image-understanding tasks.

    LMArena vision arena · Benchmark · Jul 12, 2026
  9. 9

    If you use a good harness or add the right tools and plugins both the image and web search issues mentioned are non-issues. oh-my-pi (omp.sh) handles images for text models out of the box - as long as you have any vision capable provider enabled, it will be used when you paste images to a text model. Rather than let it guess I configured it to use MiniMax M3 for this task (as well as other utility tasks like code exploration & library functions). opencode has plugins that do the same thing, but…

    jeremyjh · Hacker News · Jul 7, 2026
  10. 10

    That's right, but there are other recent open weights and relatively big LLMs that are multimodal, e.g. MiniMax-M3. With open weights LLMs, it is affordable to use many different models, each for whatever it is better. Moreover, for analyzing "UIs, photos, screenshots, etc." there are small models that can be run locally on smartphones or laptops, e.g. IBM granite-vision-4.1-4B, certain Google Gemma 4 variants and certain Qwen variants, whose output you can use as input for a big LLM, in order…

    adrian_b · Hacker News · Jun 17, 2026
  11. 11

    Ranks #19 of 51 on LMArena's vision arena (Elo 1260), based on human preference on image-understanding tasks.

    LMArena vision arena · Benchmark · Jul 12, 2026

Rankings synthesized from community evidence and open benchmarks. See methodology. Not driven by vendor marketing.