Recommendation for Math / Reasoning

Math & Reasoning

Our top recommendation for Math & Reasoning, based on the public evidence we track, is Anthropic: Claude Fable 5.[1][2][3] Use for competition mathematics and olympiad-style problems where LiveBench's 95.99% score indicates strong formal reasoning. Watch out: Watch token consumption closely on mathematical analysis tasks, as one user reported 60% higher session burn compared to Fable 5.1 for redoing similar work. Anthropic: Claude Opus 4.6 is the next-ranked alternative. Deploy for simulation and orbital mechanics coding when you lack deep math or physics background, as it reportedly writes correct simulation math code without issues.

About this recommendation

Updated
Sep 4, 2026
Evidence through
Sep 4, 2026
Sources
14
Revision
v57

Decision audit

Why this result

Inspect the inputs and the computed order behind the recommendation.

Models screened

21

live candidates

Evaluation feeds

5

task-weighted

Winner coverage

100%

intended feed weight

Largest provider share

2 of 6

Anthropic

Provisional source breadth. 5 citation families and 1 practitioner families support the top result; 0 cautionary threads is retained. The largest citation family contributes 31%.

Sources evaluated

The task sets these weights before any model is scored.

winner: Claude Fable 5
Evaluation feedWeightWinner resultField measured
LiveBench Mathematics
35%
#321/21
LiveBench Reasoning
25%
#821/21
LMArena Math
25%
#320/21
LMArena Instruction Following
10%
#320/21
OpenRouter usage
5%
87/10021/21

Provider concentration

Each exact model is scored separately; provider identity is not a ranking input.

Anthropic33%
  • Anthropic2 models
  • Google1 model
  • OpenAI1 model
  • Qwen1 model
  • xAI1 model

Decision table

Every published model is shown in computed order. Practitioner sources are distinct community threads, not the citations repeated in the prose below.

RankModelRelative scoreCoveragePractitioner evidenceStrongest measured reason
01Claude Fable 5Anthropic
86
100%1 threads · 1 families · 0 cautions#3 LiveBench Mathematics · #3 LMArena Instruction Following
02Claude Opus 4.6Anthropic
82
100%2 threads · 1 families · 0 cautions#1 LMArena Instruction Following · #4 LMArena Math
03GPT-5.6 SolOpenAI
81
100%no linked practitioner threads#2 LiveBench Mathematics · #2 LiveBench Reasoning
04Qwen3.8 27BQwen
75
100%4 threads · 2 families · 0 cautions#30 LMArena Instruction Following · #32 LMArena Math
05Gemini 3.6 FlashGoogle
74
100%no linked practitioner threads#7 LMArena Math · #18 LMArena Instruction Following
06Grok 4.6xAI
72
100%no linked practitioner threads#6 LiveBench Reasoning · #12 LiveBench Mathematics

Relative score combines normalized benchmark quality and signal coverage; independent practitioner evidence and freshness are bounded tie-breakers. It is an ordering score, not an absolute quality percentage. The writing model receives this order and cannot change it.

  1. Claude Fable 5 scores 95.99% on LiveBench Mathematics, placing third among 51 models on objectively scored competition and olympiad-style tasks.

    Best when: Use for competition mathematics and olympiad-style problems where LiveBench's 95.99% score indicates strong formal reasoning.

    Tips

    • Use for competition mathematics and olympiad-style problems where LiveBench's 95.99% score indicates strong formal reasoning.
      Source 1
      Scores 95.99% on LiveBench Mathematics (#3 of 51), using objectively scored competition and olympiad-style tasks.
      LiveBench MathematicsOpen original ↗
    • Consider for math-heavy workflows where human preference rankings matter, as it holds #3 position on LMArena's maths category with Elo 1525.
      Source 4
      Ranks #3 of 139 on LMArena's maths category (Elo 1525), based on blind human preference votes for maths prompts.
      LMArena maths categoryOpen original ↗

    Watch out for

    • Watch token consumption closely on mathematical analysis tasks, as one user reported 60% higher session burn compared to Fable 5.1 for redoing similar work.
      Source 2
      The writing style has significantly improved, however the token burn rate for tasks I have been working on seems to have skyrocketed. It definitely appears more capable (though I am unclear how much of that is just me liking the English it writes now vs actually more performant). I was using Fable 5 for some mathematical analysis assistance and redoing a part of it with 5.1 burned 60% of my session at a much faster rate.
  2. Claude Opus 4.6 ranks #4 on LMArena's maths category and handles simulation and game-related math code for users without strong physics backgrounds.

    Best when: Deploy for simulation and orbital mechanics coding when you lack deep math or physics background, as it reportedly writes correct simulation math code without issues.

    Tips

    • Deploy for simulation and orbital mechanics coding when you lack deep math or physics background, as it reportedly writes correct simulation math code without issues.
      Source 3
      > It's a pretty high moat getting into stuff like simulation software I'm currently working on a simulation game about space and orbital mechanics. I have a lot of software experience, I know how to build large projects and architect my code, and I know how to to test the end result to ensure I'm getting what I want. But I also don't have a strong math or physics background. In my experience, Claude (Opus 4.6+) has had no issues writing any simulation or game related math code. And the key thin…
    • Use for reasoning tasks where LiveBench's ground-truth scoring matters, though its 88.67% places it at #14 of 51, below top performers.
      Source 5
      Scores 88.67% on LiveBench Reasoning (#14 of 51), an objective, ground-truth-scored evaluation refreshed with new questions.
      LiveBench ReasoningOpen original ↗

    Watch out for

    • Verify reasoning transparency settings before relying on chain-of-thought visibility, as Anthropic has restricted reasoning display and users report uncertainty about whether reasoning remains exposed.
      Source 6
      The initial motivation for this was likely to thwart any competition. Already Anthropic has accused some companies of organized distillation efforts at a massive scale. Back when I used antigravity, it used to show the reasoning intact - at least for Gemini Pro 3.1, and likely for Claude Opus 4.6 (not 100% certain about it). I have some recollection of stopping the models mid-turn when they started going astray. As a power user, I find reasoning fascinating to read and genuinely useful at times…
  3. OpenAI: GPT-5.6 Sol scores 96.2% on LiveBench Mathematics (#2 of 51), using objectively scored competition and olympiad-style tasks.

    Best when: Consider only after reviewing the cited caution.

  4. Qwen3.8 27B runs locally on single RTX 5090 with verified functional correctness on math, tool calling, and 60K context retrieval, scoring 89.2 on GPQA Diamond and 90.3 on LiveCodeBench v6.

    Best when: Deploy locally for hard problems when you need Apache 2.0 licensing and 262K context, with GPQA Diamond 89.2 and LiveCodeBench v6 90.3 showing competitive reasoning.

    Tips

    • Deploy locally for hard problems when you need Apache 2.0 licensing and 262K context, with GPQA Diamond 89.2 and LiveCodeBench v6 90.3 showing competitive reasoning.
      Source 7
      **Standing issue — intentionally left open. Revisit periodically; close only if the model is ruled out for good.** ## What this tracks Whether [Qwen3.8-27B](https://huggingface.co/Qwen/Qwen3.8-27B) should replace Gemma 4 26B-A4B as the local model, and what would have to change upstream for that to make sense. ## Why it was not adopted on release (2026-08-15) The benchmarks are excellent — GPQA Diamond 89.2, LiveCodeBench v6 90.3, Terminal-Bench 2.1 73.0, SWE-bench Pro 61.7, Apache 2.0, 262K na…
    • Use NVFP4 quantization for ~1.6x faster decode on math tasks, accepting a small quality tax (92.5 vs 93.2 q_avg) versus GGUF Q4_K_M.
      Source 8
      Hardware: 1xRTX 5090 32 GB (SM 12.0), driver 595.84, CUDA 13.2 toolkit Engine: vLLM 0.27.1 (pip) + FlashInfer 0.6.17, torch 2.13.0+cu130 Model: Qwen3.8-27B, compressed-tensors NVFP4 (W4A16, group 16), MTP speculative head in BF16 ### Quality functional correctness verified (math, tool calling, 60K context retrieval, multi-turn); NVFP4 carries a small quality tax vs GGUF Q4_K_M on evals (independent measurement: q_avg 92.5 vs 93.2, HumanEval 89.6 vs 94.5) — the trade for ~1.6x faster decode than…

    Watch out for

    • Benchmark your specific configuration before trusting results, as one tester found 45 different settings configurations produced varying outcomes on 16 hard problems.
      Source 9
      A while back I posted a comparison of Qwen3.8-27B against Nemotron 3.5 Lightning and Muse Glimmer on 16 hard problems. Several people asked what settings I was running, and it turned out I did not have a good answer. So I went back and benchmarked the settings themselves: 45 configurations of this model alone. ![chart-ctx-fullprompt-latency](https://cdn-uploads.huggingface.co/production/uploads/6314e2e71994578f07ed8f89/_H2J3BUONs6AsrR2ruV4a.png) Sharing the findings here because a few of them c…
      Source 10
      ![chart-qwen-hero](https://cdn-uploads.huggingface.co/production/uploads/6314e2e71994578f07ed8f89/y0d7Q_aB0gLg27LarUwan.png) I ran Qwen3.8-27B against two other models in the same size class on a set of 16 hard problems, and wanted to share the numbers here because a few of them surprised me. ### Setup - One RTX 5090, 32GB - Ollama 0.32.12, tag `qwen3.8:latest` (Q4_K_M GGUF, 27.3B) - 64K context, temperature 0.2, generation cap 32,768 tokens - Every model loaded alone on the GPU, with all other…
  5. Gemini 3.6 Flash ranks #7 on LMArena's maths category with Elo 1505 but scores 85.15% on LiveBench Reasoning, placing #26 of 51.

    Best when: Consider for math tasks where human preference rankings guide selection, as its #7 position on LMArena maths suggests decent subjective performance.

    Tips

    • Consider for math tasks where human preference rankings guide selection, as its #7 position on LMArena maths suggests decent subjective performance.
      Source 11
      Ranks #7 of 139 on LMArena's maths category (Elo 1505), based on blind human preference votes for maths prompts.
      LMArena maths categoryOpen original ↗

    Watch out for

    • Verify output on ground-truth reasoning benchmarks before critical use, as its 85.15% on LiveBench Reasoning trails significantly behind top performers.
      Source 12
      Scores 85.15% on LiveBench Reasoning (#26 of 51), an objective, ground-truth-scored evaluation refreshed with new questions.
      LiveBench ReasoningOpen original ↗
  6. Grok 4.6 scores 90.51% on LiveBench Reasoning, placing #6 of 51, despite a lower overall LMArena ranking of #29.

    Best when: Use for reasoning tasks where objective ground-truth scoring matters more than human preference, as its #6 LiveBench Reasoning rank outperforms its overall LMArena standing.

    Tips

    • Use for reasoning tasks where objective ground-truth scoring matters more than human preference, as its #6 LiveBench Reasoning rank outperforms its overall LMArena standing.
      Source 13
      Scores 90.51% on LiveBench Reasoning (#6 of 51), an objective, ground-truth-scored evaluation refreshed with new questions.
      LiveBench ReasoningOpen original ↗

    Watch out for

    • Do not rely on general text arena rankings for math-specific selection, as its #29 overall position with Elo 1461 understates its reasoning capability.
      Source 14
      Ranks #29 of 144 on LMArena's overall text arena (Elo 1461), based on blind human preference votes.
      LMArena text arenaOpen original ↗

Frequently asked

What is the top-ranked model for Math & Reasoning?
Anthropic: Claude Fable 5 ranks first in the current evidence-weighted comparison. Use for competition mathematics and olympiad-style problems where LiveBench's 95.99% score indicates strong formal reasoning.[1]
What should I watch out for with Anthropic: Claude Fable 5?
Watch token consumption closely on mathematical analysis tasks, as one user reported 60% higher session burn compared to Fable 5.1 for redoing similar work.[2]
What is an alternative to Anthropic: Claude Fable 5?
Anthropic: Claude Opus 4.6 is the next-ranked option. Deploy for simulation and orbital mechanics coding when you lack deep math or physics background, as it reportedly writes correct simulation math code without issues.[3]

Sources

  1. 1

    Scores 95.99% on LiveBench Mathematics (#3 of 51), using objectively scored competition and olympiad-style tasks.

    LiveBench Mathematics · Benchmark · Jun 25, 2026
  2. 2

    The writing style has significantly improved, however the token burn rate for tasks I have been working on seems to have skyrocketed. It definitely appears more capable (though I am unclear how much of that is just me liking the English it writes now vs actually more performant). I was using Fable 5 for some mathematical analysis assistance and redoing a part of it with 5.1 burned 60% of my session at a much faster rate.

    ddahlen · Hacker News · Sep 2, 2026
  3. 3

    > It's a pretty high moat getting into stuff like simulation software I'm currently working on a simulation game about space and orbital mechanics. I have a lot of software experience, I know how to build large projects and architect my code, and I know how to to test the end result to ensure I'm getting what I want. But I also don't have a strong math or physics background. In my experience, Claude (Opus 4.6+) has had no issues writing any simulation or game related math code. And the key thin…

    pyth0 · Hacker News · Jun 7, 2026
  4. 4

    Ranks #3 of 139 on LMArena's maths category (Elo 1525), based on blind human preference votes for maths prompts.

    LMArena maths category · Benchmark · Sep 1, 2026
  5. 5

    Scores 88.67% on LiveBench Reasoning (#14 of 51), an objective, ground-truth-scored evaluation refreshed with new questions.

    LiveBench Reasoning · Benchmark · Jun 25, 2026
  6. 6

    The initial motivation for this was likely to thwart any competition. Already Anthropic has accused some companies of organized distillation efforts at a massive scale. Back when I used antigravity, it used to show the reasoning intact - at least for Gemini Pro 3.1, and likely for Claude Opus 4.6 (not 100% certain about it). I have some recollection of stopping the models mid-turn when they started going astray. As a power user, I find reasoning fascinating to read and genuinely useful at times…

    sheepscreek · Hacker News · Jun 23, 2026
  7. 7

    **Standing issue — intentionally left open. Revisit periodically; close only if the model is ruled out for good.** ## What this tracks Whether [Qwen3.8-27B](https://huggingface.co/Qwen/Qwen3.8-27B) should replace Gemma 4 26B-A4B as the local model, and what would have to change upstream for that to make sense. ## Why it was not adopted on release (2026-08-15) The benchmarks are excellent — GPQA Diamond 89.2, LiveCodeBench v6 90.3, Terminal-Bench 2.1 73.0, SWE-bench Pro 61.7, Apache 2.0, 262K na…

    nbramia · GitHub · Aug 15, 2026
  8. 8

    Hardware: 1xRTX 5090 32 GB (SM 12.0), driver 595.84, CUDA 13.2 toolkit Engine: vLLM 0.27.1 (pip) + FlashInfer 0.6.17, torch 2.13.0+cu130 Model: Qwen3.8-27B, compressed-tensors NVFP4 (W4A16, group 16), MTP speculative head in BF16 ### Quality functional correctness verified (math, tool calling, 60K context retrieval, multi-turn); NVFP4 carries a small quality tax vs GGUF Q4_K_M on evals (independent measurement: q_avg 92.5 vs 93.2, HumanEval 89.6 vs 94.5) — the trade for ~1.6x faster decode than…

    marco9899 · Hugging Face · Aug 18, 2026
  9. 9

    A while back I posted a comparison of Qwen3.8-27B against Nemotron 3.5 Lightning and Muse Glimmer on 16 hard problems. Several people asked what settings I was running, and it turned out I did not have a good answer. So I went back and benchmarked the settings themselves: 45 configurations of this model alone. ![chart-ctx-fullprompt-latency](https://cdn-uploads.huggingface.co/production/uploads/6314e2e71994578f07ed8f89/_H2J3BUONs6AsrR2ruV4a.png) Sharing the findings here because a few of them c…

    laxmimerit · Hugging Face · Aug 16, 2026
  10. 10

    ![chart-qwen-hero](https://cdn-uploads.huggingface.co/production/uploads/6314e2e71994578f07ed8f89/y0d7Q_aB0gLg27LarUwan.png) I ran Qwen3.8-27B against two other models in the same size class on a set of 16 hard problems, and wanted to share the numbers here because a few of them surprised me. ### Setup - One RTX 5090, 32GB - Ollama 0.32.12, tag `qwen3.8:latest` (Q4_K_M GGUF, 27.3B) - 64K context, temperature 0.2, generation cap 32,768 tokens - Every model loaded alone on the GPU, with all other…

    laxmimerit · Hugging Face · Aug 16, 2026
  11. 11

    Ranks #7 of 139 on LMArena's maths category (Elo 1505), based on blind human preference votes for maths prompts.

    LMArena maths category · Benchmark · Sep 1, 2026
  12. 12

    Scores 85.15% on LiveBench Reasoning (#26 of 51), an objective, ground-truth-scored evaluation refreshed with new questions.

    LiveBench Reasoning · Benchmark · Jun 25, 2026
  13. 13

    Scores 90.51% on LiveBench Reasoning (#6 of 51), an objective, ground-truth-scored evaluation refreshed with new questions.

    LiveBench Reasoning · Benchmark · Jun 25, 2026
  14. 14

    Ranks #29 of 144 on LMArena's overall text arena (Elo 1461), based on blind human preference votes.

    LMArena text arena · Benchmark · Sep 1, 2026

Rankings synthesized from community evidence and open benchmarks. See methodology. Not driven by vendor marketing.