Recommendation for Competition math

Competition Math

Our top recommendation for Competition Math, based on the public evidence we track, is Anthropic: Claude Fable 5.[1] Use for olympiad-style problems where LiveBench Mathematics scores matter, as it hits 95.99% (#3 of 51) on objectively scored competition tasks. OpenAI: GPT-5.6 Sol is the next-ranked alternative.

About this recommendation

Updated
Sep 4, 2026
Evidence through
Sep 4, 2026
Sources
9
Revision
v57

Decision audit

Why this result

Inspect the inputs and the computed order behind the recommendation.

Models screened

21

live candidates

Evaluation feeds

5

task-weighted

Winner coverage

100%

intended feed weight

Largest provider share

2 of 7

Anthropic

Provisional source breadth. 4 citation families and 0 practitioner families support the top result; 0 cautionary threads is retained. The largest citation family contributes 40%.

Sources evaluated

The task sets these weights before any model is scored.

winner: Claude Fable 5
Evaluation feedWeightWinner resultField measured
LiveBench Mathematics
35%
#321/21
LiveBench Reasoning
25%
#821/21
LMArena Math
25%
#320/21
LMArena Instruction Following
10%
#320/21
OpenRouter usage
5%
87/10021/21

Provider concentration

Each exact model is scored separately; provider identity is not a ranking input.

Anthropic29%
  • Anthropic2 models
  • OpenAI2 models
  • Google1 model
  • Qwen1 model
  • xAI1 model

Decision table

Every published model is shown in computed order. Practitioner sources are distinct community threads, not the citations repeated in the prose below.

RankModelRelative scoreCoveragePractitioner evidenceStrongest measured reason
01Claude Fable 5Anthropic
86
100%no linked practitioner threads#3 LiveBench Mathematics · #3 LMArena Instruction Following
02GPT-5.6 SolOpenAI
81
100%no linked practitioner threads#2 LiveBench Mathematics · #2 LiveBench Reasoning
03Claude Opus 4.6Anthropic
80
100%no linked practitioner threads#1 LMArena Instruction Following · #4 LMArena Math
04GPT-5.6 LunaOpenAI
79
100%1 threads · 1 families · 0 cautions#22 LiveBench Reasoning · #29 LiveBench Mathematics
05Qwen3.8 27BQwen
75
100%1 threads · 1 families · 0 cautions#30 LMArena Instruction Following · #32 LMArena Math
06Gemini 3.6 FlashGoogle
74
100%no linked practitioner threads#7 LMArena Math · #18 LMArena Instruction Following
07Grok 4.6xAI
72
100%no linked practitioner threads#6 LiveBench Reasoning · #12 LiveBench Mathematics

Relative score combines normalized benchmark quality and signal coverage; independent practitioner evidence and freshness are bounded tie-breakers. It is an ordering score, not an absolute quality percentage. The writing model receives this order and cannot change it.

  1. Claude Fable 5 delivers strong competition math performance with top-tier benchmark scores and human preference rankings in mathematical reasoning.

    Best when: Use for olympiad-style problems where LiveBench Mathematics scores matter, as it hits 95.99% (#3 of 51) on objectively scored competition tasks.

    Tips

    • Use for olympiad-style problems where LiveBench Mathematics scores matter, as it hits 95.99% (#3 of 51) on objectively scored competition tasks.
      Source 1
      Scores 95.99% on LiveBench Mathematics (#3 of 51), using objectively scored competition and olympiad-style tasks.
      LiveBench MathematicsOpen original ↗
    • Deploy when human preference alignment for math outputs is critical, ranking #3 on LMArena's maths category with Elo 1525.
      Source 2
      Ranks #3 of 139 on LMArena's maths category (Elo 1525), based on blind human preference votes for maths prompts.
      LMArena maths categoryOpen original ↗
  2. OpenAI: GPT-5.6 Sol scores 96.2% on LiveBench Mathematics (#2 of 51), using objectively scored competition and olympiad-style tasks.

    Best when: Consider only after reviewing the cited caution.

  3. Claude Opus 4.6 ranks highly in human preference for math but shows a significant gap in objective competition task performance compared to its sibling Fable 5.

    Best when: Use when human-rated math output quality matters, as it holds #4 rank on LMArena's maths category with Elo 1516.

    Tips

    • Use when human-rated math output quality matters, as it holds #4 rank on LMArena's maths category with Elo 1516.
      Source 3
      Ranks #4 of 139 on LMArena's maths category (Elo 1516), based on blind human preference votes for maths prompts.
      LMArena maths categoryOpen original ↗

    Watch out for

    • Avoid for pure competition accuracy, as its 89.32% on LiveBench Mathematics (#25 of 51) trails top performers by over 6 percentage points.
      Source 4
      Scores 89.32% on LiveBench Mathematics (#25 of 51), using objectively scored competition and olympiad-style tasks.
      LiveBench MathematicsOpen original ↗
  4. GPT-5.6 Luna demonstrates verified end-to-end capability solving a hard multi-step competition problem through tool-augmented reasoning with full judge agreement.

    Best when: Use for complex derivations requiring extended reasoning, as it completed a full slow-mode solve of problem 1.2.1-8 in 4m50s with both truth and audit judges passing.

    Tips

    • Use for complex derivations requiring extended reasoning, as it completed a full slow-mode solve of problem 1.2.1-8 in 4m50s with both truth and audit judges passing.
      Source 5
      ## End to end proven over the chatgpt-tool proxy Full slow-mode solve of 1.2.1-8 through server2's chatgpt-tool proxy on the Codex lane, free ChatGPT account, `gpt-5.6-luna`. Result: **PASS**, both judges agreeing. | | | | --- | --- | | verdict | PASS | | truth judge | PASS, score 7 | | audit judge | PASS | | evaluation | TRUE, complete, self-contained, human-readable, verifiable | | wall clock | 4m50s | | requests | 7 | | tokens | 29,272 in / 5,376 out, 34,648 total | | actual cost | $0, free…
    • Deploy via chatgpt-tool proxy when you need verifiable, self-contained outputs that third-party judges rate as complete and human-readable.
      Source 5
      ## End to end proven over the chatgpt-tool proxy Full slow-mode solve of 1.2.1-8 through server2's chatgpt-tool proxy on the Codex lane, free ChatGPT account, `gpt-5.6-luna`. Result: **PASS**, both judges agreeing. | | | | --- | --- | | verdict | PASS | | truth judge | PASS, score 7 | | audit judge | PASS | | evaluation | TRUE, complete, self-contained, human-readable, verifiable | | wall clock | 4m50s | | requests | 7 | | tokens | 29,272 in / 5,376 out, 34,648 total | | actual cost | $0, free…

    Watch out for

    • Expect high token consumption on hard problems, with 34,648 total tokens used for a single solve, which may impact cost at scale.
      Source 5
      ## End to end proven over the chatgpt-tool proxy Full slow-mode solve of 1.2.1-8 through server2's chatgpt-tool proxy on the Codex lane, free ChatGPT account, `gpt-5.6-luna`. Result: **PASS**, both judges agreeing. | | | | --- | --- | | verdict | PASS | | truth judge | PASS, score 7 | | audit judge | PASS | | evaluation | TRUE, complete, self-contained, human-readable, verifiable | | wall clock | 4m50s | | requests | 7 | | tokens | 29,272 in / 5,376 out, 34,648 total | | actual cost | $0, free…
  5. Qwen3.8 27B remains unverified for competition math with no direct benchmark scores, held in open status pending adoption criteria.

    Best when: Consider only after reviewing the cited caution.

    Watch out for

    • Do not rely on for competition math until verified, as the model is explicitly not adopted for this use case with benchmarks left unaudited.
      Source 6
      **Standing issue — intentionally left open. Revisit periodically; close only if the model is ruled out for good.** ## What this tracks Whether [Qwen3.8-27B](https://huggingface.co/Qwen/Qwen3.8-27B) should replace Gemma 4 26B-A4B as the local model, and what would have to change upstream for that to make sense. ## Why it was not adopted on release (2026-08-15) The benchmarks are excellent — GPQA Diamond 89.2, LiveCodeBench v6 90.3, Terminal-Bench 2.1 73.0, SWE-bench Pro 61.7, Apache 2.0, 262K na…
  6. Gemini 3.6 Flash shows competitive human preference rankings in math but lacks objective benchmark evidence for competition task performance.

    Best when: Consider for math workflows where speed and human preference alignment intersect, ranking #7 on LMArena's maths category with Elo 1505.

    Tips

    • Consider for math workflows where speed and human preference alignment intersect, ranking #7 on LMArena's maths category with Elo 1505.
      Source 7
      Ranks #7 of 139 on LMArena's maths category (Elo 1505), based on blind human preference votes for maths prompts.
      LMArena maths categoryOpen original ↗
  7. Grok 4.6 delivers mid-tier performance on objective competition math with a notable gap between its specialized benchmark score and general text arena ranking.

    Best when: Use for competition math where 92.57% accuracy on LiveBench Mathematics (#15 of 51) suffices, particularly if already integrated into xAI infrastructure.

    Tips

    • Use for competition math where 92.57% accuracy on LiveBench Mathematics (#15 of 51) suffices, particularly if already integrated into xAI infrastructure.
      Source 8
      Scores 92.57% on LiveBench Mathematics (#15 of 51), using objectively scored competition and olympiad-style tasks.
      LiveBench MathematicsOpen original ↗

    Watch out for

    • Expect weaker general reasoning quality, as its #29 rank on overall text arena (Elo 1461) suggests broader capability limitations.
      Source 9
      Ranks #29 of 144 on LMArena's overall text arena (Elo 1461), based on blind human preference votes.
      LMArena text arenaOpen original ↗

Frequently asked

What is the top-ranked model for Competition Math?
Anthropic: Claude Fable 5 ranks first in the current evidence-weighted comparison. Use for olympiad-style problems where LiveBench Mathematics scores matter, as it hits 95.99% (#3 of 51) on objectively scored competition tasks.[1]

Sources

  1. 1

    Scores 95.99% on LiveBench Mathematics (#3 of 51), using objectively scored competition and olympiad-style tasks.

    LiveBench Mathematics · Benchmark · Jun 25, 2026
  2. 2

    Ranks #3 of 139 on LMArena's maths category (Elo 1525), based on blind human preference votes for maths prompts.

    LMArena maths category · Benchmark · Sep 1, 2026
  3. 3

    Ranks #4 of 139 on LMArena's maths category (Elo 1516), based on blind human preference votes for maths prompts.

    LMArena maths category · Benchmark · Sep 1, 2026
  4. 4

    Scores 89.32% on LiveBench Mathematics (#25 of 51), using objectively scored competition and olympiad-style tasks.

    LiveBench Mathematics · Benchmark · Jun 25, 2026
  5. 5

    ## End to end proven over the chatgpt-tool proxy Full slow-mode solve of 1.2.1-8 through server2's chatgpt-tool proxy on the Codex lane, free ChatGPT account, `gpt-5.6-luna`. Result: **PASS**, both judges agreeing. | | | | --- | --- | | verdict | PASS | | truth judge | PASS, score 7 | | audit judge | PASS | | evaluation | TRUE, complete, self-contained, human-readable, verifiable | | wall clock | 4m50s | | requests | 7 | | tokens | 29,272 in / 5,376 out, 34,648 total | | actual cost | $0, free…

    tamnd · GitHub · Jul 25, 2026
  6. 6

    **Standing issue — intentionally left open. Revisit periodically; close only if the model is ruled out for good.** ## What this tracks Whether [Qwen3.8-27B](https://huggingface.co/Qwen/Qwen3.8-27B) should replace Gemma 4 26B-A4B as the local model, and what would have to change upstream for that to make sense. ## Why it was not adopted on release (2026-08-15) The benchmarks are excellent — GPQA Diamond 89.2, LiveCodeBench v6 90.3, Terminal-Bench 2.1 73.0, SWE-bench Pro 61.7, Apache 2.0, 262K na…

    nbramia · GitHub · Aug 15, 2026
  7. 7

    Ranks #7 of 139 on LMArena's maths category (Elo 1505), based on blind human preference votes for maths prompts.

    LMArena maths category · Benchmark · Sep 1, 2026
  8. 8

    Scores 92.57% on LiveBench Mathematics (#15 of 51), using objectively scored competition and olympiad-style tasks.

    LiveBench Mathematics · Benchmark · Jun 25, 2026
  9. 9

    Ranks #29 of 144 on LMArena's overall text arena (Elo 1461), based on blind human preference votes.

    LMArena text arena · Benchmark · Sep 1, 2026

Rankings synthesized from community evidence and open benchmarks. See methodology. Not driven by vendor marketing.