Recommendation for Image understanding

Image Understanding

Our top recommendation for Image Understanding, based on the public evidence we track, is Anthropic: Claude Fable 5.[1][2] Use when you need the top-rated model for general image description and visual QA based on direct human comparisons. OpenAI: GPT-5.6 Sol is the next-ranked alternative. Choose for tasks requiring detection of subtle visual details, such as faint reflections or fine-grained image analysis.

About this recommendation

Updated
Sep 4, 2026
Evidence through
Sep 4, 2026
Sources
19
Revision
v57

Decision audit

Why this result

Inspect the inputs and the computed order behind the recommendation.

Models screened

23

live candidates

Evaluation feeds

4

task-weighted

Winner coverage

64%

intended feed weight

Largest provider share

2 of 8

Anthropic

Provisional source breadth. 13 citation families and 0 practitioner families support the top result; 0 cautionary threads is retained. The largest citation family contributes 26%.

Sources evaluated

The task sets these weights before any model is scored.

winner: Claude Fable 5
Evaluation feedWeightWinner resultField measured
LMArena Vision
45%
#122/23
VLMEvalKit tasksunavailable
35%
feed unavailable0/23
LiveBench Reasoning
10%
#821/23
OpenRouter usage
10%
87/10023/23

Provider concentration

Each exact model is scored separately; provider identity is not a ranking input.

Anthropic25%
  • Anthropic2 models
  • OpenAI2 models
  • Google1 model
  • Meta1 model
  • minimax1 model
  • Qwen1 model

Decision table

Every published model is shown in computed order. Practitioner sources are distinct community threads, not the citations repeated in the prose below.

RankModelRelative scoreCoveragePractitioner evidenceStrongest measured reason
01Claude Fable 5Anthropic
70
64%no linked practitioner threads#1 LMArena Vision · #8 LiveBench Reasoning
02GPT-5.6 SolOpenAI
67
64%2 threads · 2 families · 1 cautions#2 LiveBench Reasoning · #10 LMArena Vision
03Muse Spark 1.2Meta
66
64%1 threads · 1 families · 0 cautions#5 LMArena Vision · #7 LiveBench Reasoning
04Gemini 3.6 FlashGoogle
65
64%3 threads · 3 families · 0 cautions#7 LMArena Vision · #24 LiveBench Reasoning
05Claude Sonnet 4.6Anthropic
63
64%1 threads · 1 families · 1 cautions#14 LMArena Vision · #25 LiveBench Reasoning
06GPT-5.6 LunaOpenAI
61
64%2 threads · 2 families · 1 cautions#22 LiveBench Reasoning · #23 LMArena Vision
07Qwen3.8 27BQwen
59
64%5 threads · 5 families · 1 cautions#25 LMArena Vision · #33 LiveBench Reasoning
08MiniMax M3minimax
56
64%4 threads · 3 families · 1 cautions#31 LMArena Vision · #42 LiveBench Reasoning

Relative score combines normalized benchmark quality and signal coverage; independent practitioner evidence and freshness are bounded tie-breakers. It is an ordering score, not an absolute quality percentage. The writing model receives this order and cannot change it.

  1. Leads the vision arena with the highest human preference score for image understanding tasks.

    Best when: Use when you need the top-rated model for general image description and visual QA based on direct human comparisons.

    Tips

    • Use when you need the top-rated model for general image description and visual QA based on direct human comparisons.
      Source 1
      Ranks #1 of 68 on LMArena's vision arena (Elo 1313), based on human preference on image-understanding tasks.
      LMArena vision arenaOpen original ↗
  2. Demonstrated capability on an extremely challenging reflection recognition task involving faint visual details in complex lighting.

    Best when: Choose for tasks requiring detection of subtle visual details, such as faint reflections or fine-grained image analysis.

    Tips

    • Choose for tasks requiring detection of subtle visual details, such as faint reflections or fine-grained image analysis.
      Source 2
      I agree. It did very well on an extremely challenging task. I asked it to recognize and draw the very faint reflection of what I was wearing, visible in only a tiny black part of a very brightly lit poster behind glass. In addition, the poster itself also happened to contain similar clothing. You can see the reference images and its output in my writeup here: https: medium.com @rviragh gpt-5-6-sol-very-good-image-reco... While a human can focus on the reflection easily, this is an enormous chal…

    Watch out for

    • Watch for context compaction issues that may preserve large image payloads and reduce available context window.
      Source 3
      ### What version of Codex CLI is running? 0.146.0 ### What subscription do you have? Plus ### Which model were you using? gpt-5-6 Sol High ### What platform is your computer? Microsoft Windows NT 10.0.26200.0 x64 ### What terminal emulator and version are you using (if applicable)? Windows Terminal/PowerShell ### Codex doctor report ### What issue are you seeing? Compaction repeatedly preserves large image payloads, reducing available context In the latest compaction window (242), the preserved…
  3. Ranks fifth in vision arena preferences but has reported configuration issues with image input recognition.

    Best when: Consider for image understanding tasks where you can verify vision capability is properly configured.

    Tips

    • Consider for image understanding tasks where you can verify vision capability is properly configured.
      Source 4
      Ranks #5 of 68 on LMArena's vision arena (Elo 1292), based on human preference on image-understanding tasks.
      LMArena vision arenaOpen original ↗

    Watch out for

    • Verify vision input is actually enabled; some deployments incorrectly report the model lacks image support.
      Source 5
      Resolved model commandcode/gpt-5.6-luna does not support image input. Configure a vision-capable model for mo... same for muse spark 1.2 and other vision models
  4. Supports Google's Agentic Video Understanding protocol for long-form video analysis with reduced context token usage.

    Best when: Use for video understanding workflows where you want to avoid client-side frame sampling and reduce token costs.

    Tips

    • Use for video understanding workflows where you want to avoid client-side frame sampling and reduce token costs.
      Source 6
      ### Area Provider adapters ### What are you trying to accomplish? I need to route Gemini 3.7 Flash / 3.6 Flash multimodal video understanding requests from client agents (e.g. Hermes or custom AI agents) through the OpenCodeX local proxy, allowing requests that utilize Google's new **Agentic Video Understanding** protocol (). ### What prevents this today? Google has recently released Agentic Video Understanding ([Developer Guide](https://aistudio.google.com/learn/agentic-video-understanding-wit…
      GoldenLoaf24hOpen original ↗
      Source 7
      ### Problem or Use Case Analyzing long-form video in Hermes currently relies on client-side frame sampling or multi-frame fan-out (as described in ), which causes high context token usage, heavy CPU bursts during Base64 encoding, and risks gateway payload timeouts on longer clips. Google AI Studio recently introduced **Agentic Video Understanding** for Gemini 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite ([Official Developer Guide](https://aistudio.google.com/learn/agentic-video-understanding-with-g…
      GoldenLoaf24hOpen original ↗
    • Select when you need a model that responds well to contextual prompts about visual content like portraits.
      Source 8
      Opus 5 clearly frogmaxxed. gemini-3.6-flash runs 2 and 3 responded best to the royal portrait context.
  5. Evaluated in a controlled study for food photograph analysis with structured prompts at low temperature.

    Best when: Consider for structured visual analysis tasks like nutritional assessment from food images with reproducible outputs.

    Tips

    • Consider for structured visual analysis tasks like nutritional assessment from food images with reproducible outputs.
      Source 9
      The study used a temperature of 0.01. > "Thirteen food photographs were each submitted 495–561 times to four LLM vision APIs (GPT-5.4, Claude Sonnet 4.6, Gemini 2.5 Pro, Gemini 3.1 Pro Preview) using an identical structured prompt adapted from the iAPS automated insulin delivery system (26,904 total queries, temperature 0.01)"

    Watch out for

    • Note its lower ranking (#16) in human preference compared to top vision models for open-ended image description.
      Source 10
      Ranks #16 of 68 on LMArena's vision arena (Elo 1275), based on human preference on image-understanding tasks.
      LMArena vision arenaOpen original ↗
  6. Configured for image generation workloads with adjustable reasoning effort, though vision input support varies by deployment.

    Best when: Use for image generation tasks where you need configurable reasoning effort levels.

    Tips

    • Use for image generation tasks where you need configurable reasoning effort levels.
      Source 11
      Discovered while working #2826 (default Codex imagegen to gpt-5.6-luna / low effort, configurable in settings). ## Problem / Goal Codex renders now carry a per-render reasoning `effort` (defaulting to `low`, overridable via `imageGen.codex.effort`). A media job's params snapshot could preserve the effort a failed job used, but the **Edit & Retry** flow has no way to inspect or change it, and the server has no way to clear it back to the default. #2826 deliberately left retry to fall back to the…

    Watch out for

    • Confirm your deployment actually supports image input; some configurations incorrectly block vision capabilities.
      Source 5
      Resolved model commandcode/gpt-5.6-luna does not support image input. Configure a vision-capable model for mo... same for muse spark 1.2 and other vision models
  7. A 27B dense open-weight model with vision tower that runs efficiently on consumer GPUs up to 262K context.

    Best when: Deploy locally for interactive vision tasks when you need low TTFT and long context on modest hardware.

    Tips

    • Deploy locally for interactive vision tasks when you need low TTFT and long context on modest hardware.
      Source 12
      Benchmark [Qwen3.8-27B](https://unsloth.ai/docs/models/qwen3.8) at Q4 on 2× RTX 5060 Ti with `llama-bench`, and use the run to settle how a 27B dense model should actually be split across two cards. ## Target | | | |---|---| | Model | [`unsloth/Qwen3.8-27B-GGUF`](https://huggingface.co/unsloth/Qwen3.8-27B-GGUF) — 27B **dense**, vision + hybrid reasoning, 256K ctx | | Quant | `UD-Q4_K_XL` (17.92 GB), Unsloth Dynamic V3.0 preview — the 4-bit Unsloth's guide recommends | | GPU | 2× RTX 5060 Ti (16…
      Source 13
      ## Why Two workload shapes with opposite requirements are currently served by one always-on GPU: - **Sprinkled interactive traffic** (Discord, chat, vision, classifier, pi turns). Needs a resident model and low TTFT. Served well today by the 4090 running Qwen3.8-27B at 262K context. - **Clumped batch work** (`model-bench` sweeps, bulk classification, embeddings backfills, autonomous queue jobs). Latency-insensitive, preemption-tolerant, and wants a much larger model than the 4090 can hold. The…
      Source 14
      Qwen3.8-27B is a 27B dense hybrid-attention model with a vision tower and an MTP draft head. It fits one H200 GPU at the full 262144 context, so it queues faster than any recipe that needs a whole node. | | Measured | | --- | --- | | One caller | 122.1 tok/s | | Concurrency 1024 | 3957.3 tok/s, still rising | | Context | 262144, the checkpoint maximum | | KV cache | 1,804,253 tokens, 6.88 full-length requests at once | | Hardware | 1 H200 GPU, TP1, no NCCL | Scope - New recipe at recipes/Qwen3.…
    • Use for high-throughput batch vision processing with concurrency up to 1024 on a single H200 GPU.
      Source 14
      Qwen3.8-27B is a 27B dense hybrid-attention model with a vision tower and an MTP draft head. It fits one H200 GPU at the full 262144 context, so it queues faster than any recipe that needs a whole node. | | Measured | | --- | --- | | One caller | 122.1 tok/s | | Concurrency 1024 | 3957.3 tok/s, still rising | | Context | 262144, the checkpoint maximum | | KV cache | 1,804,253 tokens, 6.88 full-length requests at once | | Hardware | 1 H200 GPU, TP1, no NCCL | Scope - New recipe at recipes/Qwen3.…

    Watch out for

    • Watch for screenshot handling bugs in certain client integrations like VSCode browser tools.
      Source 15
      Hello, vision is working fine for me, when I upload separated image file, however, when the VSCode browser tool is sending screenshot, getting this error: Launching command: `.\ninfer-serve.exe models\qwen3_8_27b_nvfp4.ninfer --model-id qwen3.8-27b --max-context 220000 --default-max-tokens 220000 --kv-dtype int8 --lm-head-draft --spec mtp --draft-tokens 3 --lm-head-draft --host 127.0.0.1 --port 8081 --cors --preserve-thinking --webui --max-pending-requests 50 --pending-timeout-ms 3000000 --visi…
  8. An open-weight vision model available through multiple providers, with tooling support for image handling in chat workflows.

    Best when: Integrate into agent workflows where you need a dedicated vision model for utility tasks like code exploration and image analysis.

    Tips

    • Integrate into agent workflows where you need a dedicated vision model for utility tasks like code exploration and image analysis.
      Source 16
      If you use a good harness or add the right tools and plugins both the image and web search issues mentioned are non-issues. oh-my-pi (omp.sh) handles images for text models out of the box - as long as you have any vision capable provider enabled, it will be used when you paste images to a text model. Rather than let it guess I configured it to use MiniMax M3 for this task (as well as other utility tasks like code exploration & library functions). opencode has plugins that do the same thing, but…
    • Consider as part of a multi-model pipeline where smaller vision models feed into larger LLMs for UI and screenshot analysis.
      Source 17
      That's right, but there are other recent open weights and relatively big LLMs that are multimodal, e.g. MiniMax-M3. With open weights LLMs, it is affordable to use many different models, each for whatever it is better. Moreover, for analyzing "UIs, photos, screenshots, etc." there are small models that can be run locally on smartphones or laptops, e.g. IBM granite-vision-4.1-4B, certain Google Gemma 4 variants and certain Qwen variants, whose output you can use as input for a big LLM, in order…

    Watch out for

    • Verify your client supports the correct protocol; some deployments only expose vision on Anthropic-compatible endpoints, not OpenAI format.
      Source 18
      ### Bug type Regression (previous MiniMax VL routing fix no longer covers the newly image-capable M3 catalog entry) ### Beta release blocker No ### Summary OpenClaw 2026.7.2-beta.5 (`ee929db`) marks `minimax/MiniMax-M3` as accepting `input: ["text", "image"]` while the provider uses MiniMax's Anthropic-compatible endpoint. OpenClaw therefore treats the active model as natively vision-capable and the `image` tool returns the image to M3 as `Loaded 1 image for direct visual inspection` instead of…
      Source 19
      ## Problem `dsh-vision-toolkit` (and upstream `vision_client.py`) hard-requires an OpenAI-compatible endpoint: it POSTs `{baseUrl}/chat/completions` with `image_url` content blocks. But some subscriptions expose their best vision models **only on the Anthropic protocol**: - **OpenCode Go** ($10/mo) serves `Qwen3.7 Plus` / `Qwen3.7 Max` / `Qwen3.8 Max` / `MiniMax M3·M2.7·M2.5` exclusively at `https://opencode.ai/zen/go/v1/messages` (Anthropic Messages format), while the cheaper vision models on…

Frequently asked

What is the top-ranked model for Image Understanding?
Anthropic: Claude Fable 5 ranks first in the current evidence-weighted comparison. Use when you need the top-rated model for general image description and visual QA based on direct human comparisons.[1]
What is an alternative to Anthropic: Claude Fable 5?
OpenAI: GPT-5.6 Sol is the next-ranked option. Choose for tasks requiring detection of subtle visual details, such as faint reflections or fine-grained image analysis.[2]

Sources

  1. 1

    Ranks #1 of 68 on LMArena's vision arena (Elo 1313), based on human preference on image-understanding tasks.

    LMArena vision arena · Benchmark · Aug 27, 2026
  2. 2

    I agree. It did very well on an extremely challenging task. I asked it to recognize and draw the very faint reflection of what I was wearing, visible in only a tiny black part of a very brightly lit poster behind glass. In addition, the poster itself also happened to contain similar clothing. You can see the reference images and its output in my writeup here: https: medium.com @rviragh gpt-5-6-sol-very-good-image-reco... While a human can focus on the reflection easily, this is an enormous chal…

    logicallee · Hacker News · Aug 17, 2026
  3. 3

    ### What version of Codex CLI is running? 0.146.0 ### What subscription do you have? Plus ### Which model were you using? gpt-5-6 Sol High ### What platform is your computer? Microsoft Windows NT 10.0.26200.0 x64 ### What terminal emulator and version are you using (if applicable)? Windows Terminal/PowerShell ### Codex doctor report ### What issue are you seeing? Compaction repeatedly preserves large image payloads, reducing available context In the latest compaction window (242), the preserved…

    vural2123 · GitHub · Aug 1, 2026
  4. 4

    Ranks #5 of 68 on LMArena's vision arena (Elo 1292), based on human preference on image-understanding tasks.

    LMArena vision arena · Benchmark · Aug 27, 2026
  5. 5

    Resolved model commandcode/gpt-5.6-luna does not support image input. Configure a vision-capable model for mo... same for muse spark 1.2 and other vision models

    DiyarD · GitHub · Aug 9, 2026
  6. 6

    ### Area Provider adapters ### What are you trying to accomplish? I need to route Gemini 3.7 Flash / 3.6 Flash multimodal video understanding requests from client agents (e.g. Hermes or custom AI agents) through the OpenCodeX local proxy, allowing requests that utilize Google's new **Agentic Video Understanding** protocol (). ### What prevents this today? Google has recently released Agentic Video Understanding ([Developer Guide](https://aistudio.google.com/learn/agentic-video-understanding-wit…

    GoldenLoaf24h · GitHub · Sep 2, 2026
  7. 7

    ### Problem or Use Case Analyzing long-form video in Hermes currently relies on client-side frame sampling or multi-frame fan-out (as described in ), which causes high context token usage, heavy CPU bursts during Base64 encoding, and risks gateway payload timeouts on longer clips. Google AI Studio recently introduced **Agentic Video Understanding** for Gemini 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite ([Official Developer Guide](https://aistudio.google.com/learn/agentic-video-understanding-with-g…

    GoldenLoaf24h · GitHub · Sep 2, 2026
  8. 8

    Opus 5 clearly frogmaxxed. gemini-3.6-flash runs 2 and 3 responded best to the royal portrait context.

    wren6991 · Hacker News · Aug 2, 2026
  9. 9

    The study used a temperature of 0.01. > "Thirteen food photographs were each submitted 495–561 times to four LLM vision APIs (GPT-5.4, Claude Sonnet 4.6, Gemini 2.5 Pro, Gemini 3.1 Pro Preview) using an identical structured prompt adapted from the iAPS automated insulin delivery system (26,904 total queries, temperature 0.01)"

    muwtyhg · Hacker News · Apr 29, 2026
  10. 10

    Ranks #16 of 68 on LMArena's vision arena (Elo 1275), based on human preference on image-understanding tasks.

    LMArena vision arena · Benchmark · Aug 27, 2026
  11. 11

    Discovered while working #2826 (default Codex imagegen to gpt-5.6-luna / low effort, configurable in settings). ## Problem / Goal Codex renders now carry a per-render reasoning `effort` (defaulting to `low`, overridable via `imageGen.codex.effort`). A media job's params snapshot could preserve the effort a failed job used, but the **Edit & Retry** flow has no way to inspect or change it, and the server has no way to clear it back to the default. #2826 deliberately left retry to fall back to the…

    atomantic · GitHub · Jul 21, 2026
  12. 12

    Benchmark [Qwen3.8-27B](https://unsloth.ai/docs/models/qwen3.8) at Q4 on 2× RTX 5060 Ti with `llama-bench`, and use the run to settle how a 27B dense model should actually be split across two cards. ## Target | | | |---|---| | Model | [`unsloth/Qwen3.8-27B-GGUF`](https://huggingface.co/unsloth/Qwen3.8-27B-GGUF) — 27B **dense**, vision + hybrid reasoning, 256K ctx | | Quant | `UD-Q4_K_XL` (17.92 GB), Unsloth Dynamic V3.0 preview — the 4-bit Unsloth's guide recommends | | GPU | 2× RTX 5060 Ti (16…

    changtimwu · GitHub · Aug 19, 2026
  13. 13

    ## Why Two workload shapes with opposite requirements are currently served by one always-on GPU: - **Sprinkled interactive traffic** (Discord, chat, vision, classifier, pi turns). Needs a resident model and low TTFT. Served well today by the 4090 running Qwen3.8-27B at 262K context. - **Clumped batch work** (`model-bench` sweeps, bulk classification, embeddings backfills, autonomous queue jobs). Latency-insensitive, preemption-tolerant, and wants a much larger model than the 4090 can hold. The…

    jomcgi · GitHub · Aug 26, 2026
  14. 14

    Qwen3.8-27B is a 27B dense hybrid-attention model with a vision tower and an MTP draft head. It fits one H200 GPU at the full 262144 context, so it queues faster than any recipe that needs a whole node. | | Measured | | --- | --- | | One caller | 122.1 tok/s | | Concurrency 1024 | 3957.3 tok/s, still rising | | Context | 262144, the checkpoint maximum | | KV cache | 1,804,253 tokens, 6.88 full-length requests at once | | Hardware | 1 H200 GPU, TP1, no NCCL | Scope - New recipe at recipes/Qwen3.…

    mmshad · GitHub · Aug 26, 2026
  15. 15

    Hello, vision is working fine for me, when I upload separated image file, however, when the VSCode browser tool is sending screenshot, getting this error: Launching command: `.\ninfer-serve.exe models\qwen3_8_27b_nvfp4.ninfer --model-id qwen3.8-27b --max-context 220000 --default-max-tokens 220000 --kv-dtype int8 --lm-head-draft --spec mtp --draft-tokens 3 --lm-head-draft --host 127.0.0.1 --port 8081 --cors --preserve-thinking --webui --max-pending-requests 50 --pending-timeout-ms 3000000 --visi…

    kexar · GitHub · Aug 24, 2026
  16. 16

    If you use a good harness or add the right tools and plugins both the image and web search issues mentioned are non-issues. oh-my-pi (omp.sh) handles images for text models out of the box - as long as you have any vision capable provider enabled, it will be used when you paste images to a text model. Rather than let it guess I configured it to use MiniMax M3 for this task (as well as other utility tasks like code exploration & library functions). opencode has plugins that do the same thing, but…

    jeremyjh · Hacker News · Jul 7, 2026
  17. 17

    That's right, but there are other recent open weights and relatively big LLMs that are multimodal, e.g. MiniMax-M3. With open weights LLMs, it is affordable to use many different models, each for whatever it is better. Moreover, for analyzing "UIs, photos, screenshots, etc." there are small models that can be run locally on smartphones or laptops, e.g. IBM granite-vision-4.1-4B, certain Google Gemma 4 variants and certain Qwen variants, whose output you can use as input for a big LLM, in order…

    adrian_b · Hacker News · Jun 17, 2026
  18. 18

    ### Bug type Regression (previous MiniMax VL routing fix no longer covers the newly image-capable M3 catalog entry) ### Beta release blocker No ### Summary OpenClaw 2026.7.2-beta.5 (`ee929db`) marks `minimax/MiniMax-M3` as accepting `input: ["text", "image"]` while the provider uses MiniMax's Anthropic-compatible endpoint. OpenClaw therefore treats the active model as natively vision-capable and the `image` tool returns the image to M3 as `Loaded 1 image for direct visual inspection` instead of…

    SweetSophia · GitHub · Jul 31, 2026
  19. 19

    ## Problem `dsh-vision-toolkit` (and upstream `vision_client.py`) hard-requires an OpenAI-compatible endpoint: it POSTs `{baseUrl}/chat/completions` with `image_url` content blocks. But some subscriptions expose their best vision models **only on the Anthropic protocol**: - **OpenCode Go** ($10/mo) serves `Qwen3.7 Plus` / `Qwen3.7 Max` / `Qwen3.8 Max` / `MiniMax M3·M2.7·M2.5` exclusively at `https://opencode.ai/zen/go/v1/messages` (Anthropic Messages format), while the cheaper vision models on…

    EliteOtaku · GitHub · Aug 14, 2026

Rankings synthesized from community evidence and open benchmarks. See methodology. Not driven by vendor marketing.