Recommendation for Competition math
Competition Math
Our top recommendation for Competition Math, based on the public evidence we track, is Anthropic: Claude Fable 5.[1] Use for olympiad-style problems where LiveBench Mathematics scores matter, as it hits 95.99% (#3 of 51) on objectively scored competition tasks. OpenAI: GPT-5.6 Sol is the next-ranked alternative.
About this recommendation
- Updated
- Sep 4, 2026
- Evidence through
- Sep 4, 2026
- Sources
- 9
- Revision
- v57
Decision audit
Why this result
Inspect the inputs and the computed order behind the recommendation.
Models screened
21
live candidates
Evaluation feeds
5
task-weighted
Winner coverage
100%
intended feed weight
Largest provider share
2 of 7
Anthropic
Sources evaluated
The task sets these weights before any model is scored.
| Evaluation feed | Weight | Winner result | Field measured |
|---|---|---|---|
| LiveBench Mathematics | 35% | #3 | 21/21 |
| LiveBench Reasoning | 25% | #8 | 21/21 |
| LMArena Math | 25% | #3 | 20/21 |
| LMArena Instruction Following | 10% | #3 | 20/21 |
| OpenRouter usage | 5% | 87/100 | 21/21 |
Provider concentration
Each exact model is scored separately; provider identity is not a ranking input.
- Anthropic2 models
- OpenAI2 models
- Google1 model
- Qwen1 model
- xAI1 model
Decision table
Every published model is shown in computed order. Practitioner sources are distinct community threads, not the citations repeated in the prose below.
| Rank | Model | Relative score | Coverage | Practitioner evidence | Strongest measured reason |
|---|---|---|---|---|---|
| 01 | Claude Fable 5Anthropic | 86 | 100% | no linked practitioner threads | #3 LiveBench Mathematics · #3 LMArena Instruction Following |
| 02 | GPT-5.6 SolOpenAI | 81 | 100% | no linked practitioner threads | #2 LiveBench Mathematics · #2 LiveBench Reasoning |
| 03 | Claude Opus 4.6Anthropic | 80 | 100% | no linked practitioner threads | #1 LMArena Instruction Following · #4 LMArena Math |
| 04 | GPT-5.6 LunaOpenAI | 79 | 100% | 1 threads · 1 families · 0 cautions | #22 LiveBench Reasoning · #29 LiveBench Mathematics |
| 05 | Qwen3.8 27BQwen | 75 | 100% | 1 threads · 1 families · 0 cautions | #30 LMArena Instruction Following · #32 LMArena Math |
| 06 | Gemini 3.6 FlashGoogle | 74 | 100% | no linked practitioner threads | #7 LMArena Math · #18 LMArena Instruction Following |
| 07 | Grok 4.6xAI | 72 | 100% | no linked practitioner threads | #6 LiveBench Reasoning · #12 LiveBench Mathematics |
Relative score combines normalized benchmark quality and signal coverage; independent practitioner evidence and freshness are bounded tie-breakers. It is an ordering score, not an absolute quality percentage. The writing model receives this order and cannot change it.
Claude Fable 5 delivers strong competition math performance with top-tier benchmark scores and human preference rankings in mathematical reasoning.
Best when: Use for olympiad-style problems where LiveBench Mathematics scores matter, as it hits 95.99% (#3 of 51) on objectively scored competition tasks.
Tips
- Use for olympiad-style problems where LiveBench Mathematics scores matter, as it hits 95.99% (#3 of 51) on objectively scored competition tasks.
- Deploy when human preference alignment for math outputs is critical, ranking #3 on LMArena's maths category with Elo 1525.
OpenAI: GPT-5.6 Sol scores 96.2% on LiveBench Mathematics (#2 of 51), using objectively scored competition and olympiad-style tasks.
Best when: Consider only after reviewing the cited caution.
Claude Opus 4.6 ranks highly in human preference for math but shows a significant gap in objective competition task performance compared to its sibling Fable 5.
Best when: Use when human-rated math output quality matters, as it holds #4 rank on LMArena's maths category with Elo 1516.
Tips
- Use when human-rated math output quality matters, as it holds #4 rank on LMArena's maths category with Elo 1516.
Watch out for
- Avoid for pure competition accuracy, as its 89.32% on LiveBench Mathematics (#25 of 51) trails top performers by over 6 percentage points.
GPT-5.6 Luna demonstrates verified end-to-end capability solving a hard multi-step competition problem through tool-augmented reasoning with full judge agreement.
Best when: Use for complex derivations requiring extended reasoning, as it completed a full slow-mode solve of problem 1.2.1-8 in 4m50s with both truth and audit judges passing.
Tips
- Use for complex derivations requiring extended reasoning, as it completed a full slow-mode solve of problem 1.2.1-8 in 4m50s with both truth and audit judges passing.
- Deploy via chatgpt-tool proxy when you need verifiable, self-contained outputs that third-party judges rate as complete and human-readable.
Watch out for
- Expect high token consumption on hard problems, with 34,648 total tokens used for a single solve, which may impact cost at scale.
Qwen3.8 27B remains unverified for competition math with no direct benchmark scores, held in open status pending adoption criteria.
Best when: Consider only after reviewing the cited caution.
Watch out for
- Do not rely on for competition math until verified, as the model is explicitly not adopted for this use case with benchmarks left unaudited.
Gemini 3.6 Flash shows competitive human preference rankings in math but lacks objective benchmark evidence for competition task performance.
Best when: Consider for math workflows where speed and human preference alignment intersect, ranking #7 on LMArena's maths category with Elo 1505.
Tips
- Consider for math workflows where speed and human preference alignment intersect, ranking #7 on LMArena's maths category with Elo 1505.
Grok 4.6 delivers mid-tier performance on objective competition math with a notable gap between its specialized benchmark score and general text arena ranking.
Best when: Use for competition math where 92.57% accuracy on LiveBench Mathematics (#15 of 51) suffices, particularly if already integrated into xAI infrastructure.
Tips
- Use for competition math where 92.57% accuracy on LiveBench Mathematics (#15 of 51) suffices, particularly if already integrated into xAI infrastructure.
Watch out for
- Expect weaker general reasoning quality, as its #29 rank on overall text arena (Elo 1461) suggests broader capability limitations.
Frequently asked
- What is the top-ranked model for Competition Math?
- Anthropic: Claude Fable 5 ranks first in the current evidence-weighted comparison. Use for olympiad-style problems where LiveBench Mathematics scores matter, as it hits 95.99% (#3 of 51) on objectively scored competition tasks.[1]
Sources
- 1
“Scores 95.99% on LiveBench Mathematics (#3 of 51), using objectively scored competition and olympiad-style tasks.”
LiveBench Mathematics · Benchmark · Jun 25, 2026 - 2
“Ranks #3 of 139 on LMArena's maths category (Elo 1525), based on blind human preference votes for maths prompts.”
LMArena maths category · Benchmark · Sep 1, 2026 - 3
“Ranks #4 of 139 on LMArena's maths category (Elo 1516), based on blind human preference votes for maths prompts.”
LMArena maths category · Benchmark · Sep 1, 2026 - 4
“Scores 89.32% on LiveBench Mathematics (#25 of 51), using objectively scored competition and olympiad-style tasks.”
LiveBench Mathematics · Benchmark · Jun 25, 2026 - 5
“## End to end proven over the chatgpt-tool proxy Full slow-mode solve of 1.2.1-8 through server2's chatgpt-tool proxy on the Codex lane, free ChatGPT account, `gpt-5.6-luna`. Result: **PASS**, both judges agreeing. | | | | --- | --- | | verdict | PASS | | truth judge | PASS, score 7 | | audit judge | PASS | | evaluation | TRUE, complete, self-contained, human-readable, verifiable | | wall clock | 4m50s | | requests | 7 | | tokens | 29,272 in / 5,376 out, 34,648 total | | actual cost | $0, free…”
tamnd · GitHub · Jul 25, 2026 - 6
“**Standing issue — intentionally left open. Revisit periodically; close only if the model is ruled out for good.** ## What this tracks Whether [Qwen3.8-27B](https://huggingface.co/Qwen/Qwen3.8-27B) should replace Gemma 4 26B-A4B as the local model, and what would have to change upstream for that to make sense. ## Why it was not adopted on release (2026-08-15) The benchmarks are excellent — GPQA Diamond 89.2, LiveCodeBench v6 90.3, Terminal-Bench 2.1 73.0, SWE-bench Pro 61.7, Apache 2.0, 262K na…”
nbramia · GitHub · Aug 15, 2026 - 7
“Ranks #7 of 139 on LMArena's maths category (Elo 1505), based on blind human preference votes for maths prompts.”
LMArena maths category · Benchmark · Sep 1, 2026 - 8
“Scores 92.57% on LiveBench Mathematics (#15 of 51), using objectively scored competition and olympiad-style tasks.”
LiveBench Mathematics · Benchmark · Jun 25, 2026 - 9
“Ranks #29 of 144 on LMArena's overall text arena (Elo 1461), based on blind human preference votes.”
LMArena text arena · Benchmark · Sep 1, 2026
Rankings synthesized from community evidence and open benchmarks. See methodology. Not driven by vendor marketing.