Claude Fable 5.1

New

anthropic/claude-fable-5.1-20260831

Claude Fable 5.1 is a text, image, and file input model that generates text outputs. It supports a 1,000,000 token context window, tool use, and reasoning.

  • tool use
  • structured outputs
  • reasoning

At a glance

provideranthropic
releasedAug 29, 2026
context1M
input / M$10

Specifications

What the API gives you

Context window
1M tokens
Max output
128K tokens
Input price
$10 / 1M tokens
Output price
$50 / 1M tokens
Cache read
$0.25 / 1M tokens
Modalities
text, image, file → text
Knowledge cutoff
Released
Aug 29, 2026

Measured evidence

Public benchmark record

Design Arena Coding
1348 Elo2026-09-03
Design Arena Coding
1343 Elo2026-09-04
Design Arena UI
1379 Elo2026-09-03
Design Arena UI
1357 Elo2026-09-04
Design Arena Website
1330 Elo2026-09-03
Design Arena Website
1327 Elo2026-09-04
Frontier-Bench
57.9%2026-09-01Claude Code, max effort
LiveBench Coding
86.4%2026-06-25
LiveBench Language
89.5%2026-06-25
LiveBench Math
97.0%2026-06-25
LiveBench Reasoning
91.7%2026-06-25
Terminal-Bench 2.1
57.9%2026-09-01max effort

LMArena data licensed CC-BY-4.0. Aider (Apache-2.0) and SWE-bench (MIT) are open. Design Arena uses blind human preference. Submitted scaffold plus model result. Compare the harness and effort settings, not only the model name.

Task evidence

How broad is the support?

Each status is exact to this model version and the selected task. Official release claims remain provisional.

  • Best LLM for Frontend & UIEstablished evidence

    70% of intended weighted task evidence across 2 independent sources. 5 of 5 configured evaluation feeds currently have results.

  • Best LLM for Autonomous AgentsEstablished evidence

    63% of intended weighted task evidence across 3 independent sources. 7 of 7 configured evaluation feeds currently have results.

  • Best LLM for Competition MathProvisional evidence

    65% of intended weighted task evidence across 1 independent source; broader independent confirmation is still missing from LMArena Instruction Following or LMArena Math. 5 of 5 configured evaluation feeds currently have results.

  • Best LLM for Math & ReasoningProvisional evidence

    65% of intended weighted task evidence across 1 independent source; broader independent confirmation is still missing from LMArena Instruction Following or LMArena Math. 5 of 5 configured evaluation feeds currently have results. Official release evidence is also tracked provisionally.

  • 60% of intended weighted task evidence across 1 independent source; broader independent confirmation is still missing from LMArena Instruction Following or LMArena Math. 5 of 5 configured evaluation feeds currently have results. Official release evidence is also tracked provisionally.

  • Best LLM for Agentic CodingProvisional evidence

    55% of intended weighted task evidence across 3 independent sources; broader independent confirmation is still missing from Design Arena Full-stack or Design Arena Web-apps. 8 of 8 configured evaluation feeds currently have results. Official release evidence is also tracked provisionally.

  • 55% of intended weighted task evidence across 1 independent source; broader independent confirmation is still missing from LMArena Instruction Following or LMArena Text. 5 of 5 configured evaluation feeds currently have results.

  • Best LLM for CodingProvisional evidence

    50% of intended weighted task evidence across 3 independent sources; broader independent confirmation is still missing from Aider Polyglot or LMArena WebDev. 8 of 8 configured evaluation feeds currently have results. Official release evidence is also tracked provisionally.

  • Best LLM for Code CompletionProvisional evidence

    50% of intended weighted task evidence across 1 independent source; broader independent confirmation is still missing from Aider Polyglot or LMArena WebDev. 4 of 4 configured evaluation feeds currently have results.

  • Best LLM for Creative WritingProvisional evidence

    50% of intended weighted task evidence across 1 independent source; broader independent confirmation is still missing from LMArena Creative Writing or LMArena Text. 5 of 5 configured evaluation feeds currently have results. Official release evidence is also tracked provisionally.

  • Best LLM for WritingProvisional evidence

    50% of intended weighted task evidence across 1 independent source; broader independent confirmation is still missing from LMArena Creative Writing or LMArena Text. 5 of 5 configured evaluation feeds currently have results. Official release evidence is also tracked provisionally.

  • Best LLM for Fast SummariesProvisional evidence

    50% of intended weighted task evidence across 1 independent source; broader independent confirmation is still missing from LMArena Text or Route reliability. 3 of 4 configured evaluation feeds currently have results.

  • Best Value LLMProvisional evidence

    43% of intended weighted task evidence across 1 independent source; broader independent confirmation is still missing from LMArena Text or Route reliability. 3 of 4 configured evaluation feeds currently have results. Official release evidence is also tracked provisionally.

  • Best Cheap LLM for High VolumeProvisional evidence

    42% of intended weighted task evidence across 1 independent source; broader independent confirmation is still missing from LMArena Text or Route reliability. 3 of 4 configured evaluation feeds currently have results.

  • Cheapest Capable LLMProvisional evidence

    42% of intended weighted task evidence across 1 independent source; broader independent confirmation is still missing from LMArena Text or Route reliability. 3 of 4 configured evaluation feeds currently have results.

  • 42% of intended weighted task evidence across 1 independent source; broader independent confirmation is still missing from LMArena Text or Route reliability. 3 of 4 configured evaluation feeds currently have results.

  • 40% of intended weighted task evidence across 1 independent source; broader independent confirmation is still missing from LMArena Text or WMT translation. 4 of 5 configured evaluation feeds currently have results.

  • 40% of intended weighted task evidence across 1 independent source; broader independent confirmation is still missing from LMArena Text or WMT translation. 4 of 5 configured evaluation feeds currently have results.

  • Best LLM for TranslationProvisional evidence

    40% of intended weighted task evidence across 1 independent source; broader independent confirmation is still missing from LMArena Text or WMT translation. 4 of 5 configured evaluation feeds currently have results.

  • Best LLM for Business WritingProvisional evidence

    40% of intended weighted task evidence across 1 independent source; broader independent confirmation is still missing from LMArena Creative Writing or LMArena Instruction Following. 5 of 5 configured evaluation feeds currently have results. Official release evidence is also tracked provisionally.

  • Best LLM for Marketing CopyProvisional evidence

    40% of intended weighted task evidence across 1 independent source; broader independent confirmation is still missing from LMArena Creative Writing or LMArena Instruction Following. 5 of 5 configured evaluation feeds currently have results.

  • Best LLM for Chat & RoleplayProvisional evidence

    40% of intended weighted task evidence across 1 independent source; broader independent confirmation is still missing from LMArena Creative Writing or LMArena Text. 5 of 5 configured evaluation feeds currently have results.

  • 40% of intended weighted task evidence across 1 independent source; broader independent confirmation is still missing from LMArena Creative Writing or LMArena Text. 5 of 5 configured evaluation feeds currently have results.

  • Best LLM for AI AgentsProvisional evidence

    33% of intended weighted task evidence across 2 independent sources; broader independent confirmation is still missing from Berkeley Function Calling or LMArena Agent. 7 of 7 configured evaluation feeds currently have results. Official release evidence is also tracked provisionally.

  • 31% of intended weighted task evidence across 1 independent source; broader independent confirmation is still missing from Route reliability or Structured-output eval. 2 of 4 configured evaluation feeds currently have results.

  • 30% of intended weighted task evidence across 1 independent source; broader independent confirmation is still missing from FACTS Grounding or LMArena Document. 4 of 6 configured evaluation feeds currently have results.

  • Best LLM for SummarizationProvisional evidence

    30% of intended weighted task evidence across 1 independent source; broader independent confirmation is still missing from FACTS Grounding or LMArena Document. 4 of 6 configured evaluation feeds currently have results.

  • Best Local LLM for CodingProvisional evidence

    30% of intended weighted task evidence across 1 independent source; broader independent confirmation is still missing from Aider Polyglot or LMArena WebDev. 5 of 5 configured evaluation feeds currently have results.

  • 28% of intended weighted task evidence across 1 independent source; broader independent confirmation is still missing from Berkeley Function Calling or LMArena Text. 4 of 5 configured evaluation feeds currently have results.

  • 28% of intended weighted task evidence across 1 independent source; broader independent confirmation is still missing from Berkeley Function Calling or LMArena Text. 4 of 5 configured evaluation feeds currently have results. Official release evidence is also tracked provisionally.

  • 25% of intended weighted task evidence across 1 independent source; broader independent confirmation is still missing from Aider Polyglot or LMArena WebDev. 6 of 6 configured evaluation feeds currently have results. Official release evidence is also tracked provisionally.

  • Best LLM for Hermes AgentProvisional evidence

    24% of intended weighted task evidence across 1 independent source; broader independent confirmation is still missing from Berkeley Function Calling or LMArena Agent. 5 of 5 configured evaluation feeds currently have results.

  • Best LLM for OpenHandsProvisional evidence

    24% of intended weighted task evidence across 1 independent source; broader independent confirmation is still missing from Berkeley Function Calling or LMArena Agent. 5 of 5 configured evaluation feeds currently have results.

  • Best LLM for GooseProvisional evidence

    24% of intended weighted task evidence across 1 independent source; broader independent confirmation is still missing from Berkeley Function Calling or LMArena Agent. 5 of 5 configured evaluation feeds currently have results.

  • Best LLM for OpenClawProvisional evidence

    24% of intended weighted task evidence across 1 independent source; broader independent confirmation is still missing from Berkeley Function Calling or LMArena Agent. 5 of 5 configured evaluation feeds currently have results.

  • 20% of intended weighted task evidence across 1 independent source; broader independent confirmation is still missing from LMArena Vision or VLMEvalKit tasks. 3 of 4 configured evaluation feeds currently have results.

  • 10% of intended weighted task evidence across 0 independent sources; broader independent confirmation is still missing from Berkeley Function Calling or LMArena Agent. 4 of 5 configured evaluation feeds currently have results. Official release evidence is also tracked provisionally.

  • 10% of intended weighted task evidence across 0 independent sources; broader independent confirmation is still missing from Berkeley Function Calling or LMArena Agent. 4 of 4 configured evaluation feeds currently have results.

  • Best LLM for OCR & DocumentsProvisional evidence

    10% of intended weighted task evidence across 0 independent sources; broader independent confirmation is still missing from LMArena Document or LMArena Vision. 3 of 5 configured evaluation feeds currently have results.

  • 10% of intended weighted task evidence across 0 independent sources; broader independent confirmation is still missing from LMArena Document or LMArena Vision. 3 of 5 configured evaluation feeds currently have results.

  • Best LLM for Grounded RAGProvisional evidence

    10% of intended weighted task evidence across 0 independent sources; broader independent confirmation is still missing from CRAG or FACTS Grounding. 2 of 5 configured evaluation feeds currently have results. Official release evidence is also tracked provisionally.

  • Best LLM for Charts & DiagramsProvisional evidence

    10% of intended weighted task evidence across 0 independent sources; broader independent confirmation is still missing from LMArena Document or LMArena Vision. 3 of 5 configured evaluation feeds currently have results. Official release evidence is also tracked provisionally.

  • Best LLM for Document ParsingProvisional evidence

    10% of intended weighted task evidence across 0 independent sources; broader independent confirmation is still missing from LMArena Document or LMArena Vision. 3 of 5 configured evaluation feeds currently have results. Official release evidence is also tracked provisionally.

  • Best LLM for Massive ContextProvisional evidence

    6% of intended weighted task evidence across 0 independent sources; broader independent confirmation is still missing from LMArena Document or LMArena Long Query. 3 of 4 configured evaluation feeds currently have results.

  • 6% of intended weighted task evidence across 0 independent sources; broader independent confirmation is still missing from FACTS Grounding or LMArena Document. 4 of 6 configured evaluation feeds currently have results.

Official release evidence · provisional

What the launch material actually says

Verbatim claims from the model provider. These help on launch day but do not count as independent confirmation.

Recommendation coverage

Where this model appears

Not currently included in a published recommendation. The task evidence panel above says exactly what is confirmed and what is still missing.

Catalog data via OpenRouter and models.dev. Last checked within the past 6 hours.