Recommendation for Math / Reasoning

Math & Reasoning

The best LLM for math and reasoning is Claude Fable 5, which users describe as "so far ahead" of other frontier models for linear algebra and technical problem solving that the gap is "obscene." The best LLM for math and reasoning depends a bit on your specific workload: Fable 5 leads for deep technical insight, Claude Opus 4.6 holds the top LMArena Elo (1501) and handles simulation math well, while Claude Opus 4.7 and Gemini 3.1 Pro both excel at word reasoning tasks like the NYTimes Connections test. Avoid GPT-5.5 for now. Multiple reports describe it stalling, getting stuck on logic proofs, and degrading over time, which is a significant problem for multi-step reasoning work. The evidence shows a clear tier split. Anthropic's models dominate the top spots, with Fable 5, Opus 4.6, and Opus 4.7 all ranking in the top 3 for human preference. More importantly, there's direct user testimony about their reasoning quality: Fable 5 handles "delicate corner cases" in linear algebra without prompting, Opus 4.6 writes correct simulation math code, and both Opus 4.7 and Gemini 3.1 Pro consistently solve complex word puzzles. GPT-5.5's performance collapse is the surprise here. Users report it went from "one shotting stuff" to struggling badly on the same tasks.

About this recommendation

Updated
Jul 17, 2026
Evidence through
Jul 17, 2026
Sources
14
Revision
v1
  1. Fable 5 is the top pick for serious mathematical work. The LMArena ranking puts it at #2 overall, but what matters more is the direct user comparison describing it as dramatically better than other frontier models for linear algebra, specifically calling out better insights and better handling of delicate corner cases without needing explicit prompting.

    Best when: You need deep mathematical insight, especially for linear algebra or work involving subtle edge cases.

    Tips

    • Ranks #2 of 49 on LMArena with a 1493 Elo, confirming strong human preference for its outputs.
      Source 1
      Ranks #2 of 17 on LMArena's overall text arena (Elo 1504), based on blind human preference votes.
      LMArena text arenaOpen original ↗

    Watch out for

    • Limited direct evidence available compared to other models, so broader reasoning performance is less documented.
      Source 1
      Ranks #2 of 17 on LMArena's overall text arena (Elo 1504), based on blind human preference votes.
      LMArena text arenaOpen original ↗
  2. Opus 4.6 holds the highest LMArena Elo at 1501 and has solid evidence for mathematical reasoning tasks. A user working on an orbital mechanics simulation noted it "has had no issues writing any simulation or game related math code" even without a strong math background, which speaks to its ability to handle interdependent logic correctly.

    Best when: You want the highest-ranked overall model with proven math code generation for simulations or physics.

    Tips

    • Top LMArena ranking at #1 with Elo 1501 based on blind human preference votes.
      Source 2
      Ranks #1 of 17 on LMArena's overall text arena (Elo 1512), based on blind human preference votes.
      LMArena text arenaOpen original ↗
    • Handles simulation and game-related math code well, including orbital mechanics, for users without strong math backgrounds.
      Source 3
      > It's a pretty high moat getting into stuff like simulation software I'm currently working on a simulation game about space and orbital mechanics. I have a lot of software experience, I know how to build large projects and architect my code, and I know how to to test the end result to ensure I'm getting what I want. But I also don't have a strong math or physics background. In my experience, Claude (Opus 4.6+) has had no issues writing any simulation or game related math code. And the key thin…
    • Reasoning traces are visible and genuinely useful for power users monitoring the model's process.
      Source 4
      The initial motivation for this was likely to thwart any competition. Already Anthropic has accused some companies of organized distillation efforts at a massive scale. Back when I used antigravity, it used to show the reasoning intact - at least for Gemini Pro 3.1, and likely for Claude Opus 4.6 (not 100% certain about it). I have some recollection of stopping the models mid-turn when they started going astray. As a power user, I find reasoning fascinating to read and genuinely useful at times…

    Watch out for

    • Less direct evidence on pure mathematical reasoning compared to Fable 5.
      Source 3
      > It's a pretty high moat getting into stuff like simulation software I'm currently working on a simulation game about space and orbital mechanics. I have a lot of software experience, I know how to build large projects and architect my code, and I know how to to test the end result to ensure I'm getting what I want. But I also don't have a strong math or physics background. In my experience, Claude (Opus 4.6+) has had no issues writing any simulation or game related math code. And the key thin…
  3. Opus 4.7 ties with Gemini 3.1 Pro as the best documented performer on word reasoning tasks, nailing the NYTimes Connections puzzle as a standard test. It ranks #3 on LMArena with 1490 Elo, making it a strong all-around choice, though one user noted it can be trigger-happy with refusals on certain framings.

    Best when: You need strong word reasoning and puzzle-solving, or want a top-tier model with slightly different characteristics than Opus 4.6.

    Tips

    • Ranks #3 of 49 on LMArena with 1490 Elo.
      Source 5
      Ranks #4 of 17 on LMArena's overall text arena (Elo 1499), based on blind human preference votes.
      LMArena text arenaOpen original ↗

    Watch out for

    • Can refuse requests based on framing, even when the underlying request is benign.
      Source 6
      I was talking to a Claude Opus 4.7 chatbot about low-discrepancy sequences and made the mistake of framing a thought experiment as "deceiving" a vendor who owns a scrambler by making a hard to scramble low-discrepancy sequence generator. Admittedly that sounds sketchy but the conversation up to that point--without the explicit framing--was about the same subject just shorn of metaphor. With the framing I got a refusal message that likened my request to Dual_EC_DRBG!
  4. Gemini 3.1 Pro matches Opus 4.7 on the word reasoning test, consistently solving the NYTimes Connections puzzle. Benchmark evidence puts it marginally ahead of GPT-5.4 on reasoning tasks according to one comparison, and its reasoning traces are visible and useful for power users.

    Best when: You want Google's top reasoning model with proven puzzle-solving and visible thinking tokens.

    Tips

    • Consistently solves the NYTimes Connections puzzle, matching Opus 4.7.
      Source 7
      I have a standard test to look at the reasoning capabilities of a model - solve today's NYTimes connections problem. Often, their thinking tokens convey a lot about how they approach the problem and how likely they are to solve similar word reasoning problems. Claude 4.7 and Gemini 3.1 Pro have nailed all so far, GPT 5.5 failed miserably. Of the chinese models, Kimi-K-2.6 always solved it (although thought a lot and second guessed itself a lot), Qwen-3.6-Plus often gave wrong answers and GLM-5.…
    • Reasoning traces are visible and genuinely useful, allowing users to stop mid-turn if the model goes astray.
      Source 4
      The initial motivation for this was likely to thwart any competition. Already Anthropic has accused some companies of organized distillation efforts at a massive scale. Back when I used antigravity, it used to show the reasoning intact - at least for Gemini Pro 3.1, and likely for Claude Opus 4.6 (not 100% certain about it). I have some recollection of stopping the models mid-turn when they started going astray. As a power user, I find reasoning fascinating to read and genuinely useful at times…
    • Ranks #6 of 49 on LMArena with 1479 Elo.
      Source 8
      Ranks #9 of 17 on LMArena's overall text arena (Elo 1489), based on blind human preference votes.
      LMArena text arenaOpen original ↗

    Watch out for

    • Lower LMArena ranking than the top Anthropic models.
      Source 8
      Ranks #9 of 17 on LMArena's overall text arena (Elo 1489), based on blind human preference votes.
      LMArena text arenaOpen original ↗
  5. GPT-5.4 edges out Gemini 3.1 Pro on standard reasoning benchmarks according to comparative analysis, making it the strongest OpenAI option for reasoning tasks. However, it lacks the direct user testimonials that Fable 5 and the Opus models have.

    Best when: You're committed to OpenAI's ecosystem but want better reasoning than GPT-5.5 offers.

    Tips

    • Performs marginally better than Gemini 3.1 Pro on standard reasoning benchmarks.
      Source 9
      Two key quotes: • Reasoning: Through the expansion of reasoning tokens, DeepSeek-V4-Pro-Max demonstrates superior performance relative to GPT-5.2 and Gemini-3.0-Pro on standard reasoning benchmarks. Nevertheless, its performance falls marginally short of GPT-5.4 and Gemini3.1-Pro, suggesting a developmental trajectory that trails state-of-the-art frontier models by approximately 3 to 6 months. Furthermore, DeepSeek-V4-Flash-Max achieves comparable performance to GPT-5.2 and Gemini-3.0-Pro, esta…
      creamyhorrorOpen original ↗
    • Ranks #9 on LMArena with 1470 Elo, tied with GPT-5.5.
      Source 10
      Ranks #3 of 17 on LMArena's overall text arena (Elo 1499), based on blind human preference votes.
      LMArena text arenaOpen original ↗

    Watch out for

    • No direct user testimony on math or reasoning quality.
      Source 10
      Ranks #3 of 17 on LMArena's overall text arena (Elo 1499), based on blind human preference votes.
      LMArena text arenaOpen original ↗
      Source 9
      Two key quotes: • Reasoning: Through the expansion of reasoning tokens, DeepSeek-V4-Pro-Max demonstrates superior performance relative to GPT-5.2 and Gemini-3.0-Pro on standard reasoning benchmarks. Nevertheless, its performance falls marginally short of GPT-5.4 and Gemini3.1-Pro, suggesting a developmental trajectory that trails state-of-the-art frontier models by approximately 3 to 6 months. Furthermore, DeepSeek-V4-Flash-Max achieves comparable performance to GPT-5.2 and Gemini-3.0-Pro, esta…
      creamyhorrorOpen original ↗
  6. GPT-5.5 stands out for the wrong reasons: users report it routinely stalls, gets stuck on logic proofs, and has degraded significantly since launch. It failed the NYTimes Connections test badly, and one user had to build tooling just to make it keep working. There's evidence it used to be strong on logic, which makes the current state more frustrating for math work.

    Best when: Avoid for math and reasoning until OpenAI addresses the reported issues.

    Tips

    • At launch, could "one shot" complex tasks and was better than Opus on certain logic work according to one user.
      Source 11
      "But LLMs are prone to hallucinations which can really impact a string of interdependent logic like a proof. So I’m assuming it would respond with something that’s not complete nonsense to this proof most of the time." Unfortunately in my experience that's not really the case. For me, very often GPT 5.5 (which was a good deal better than Opus at this kind of task) would just get stuck for long periods when working in a logic like Iris. It wouldn't necessarily outright prove nonsense, but it wou…
      Source 12
      this explains so much why gpt 5.5 has been so bad lately it was really puzzling why it struggled so much where when it first came out it was one shotting stuff totally amazing, i tried the prompt that will tell you if your plan is degraded: codex exec --json --skip-git-repo-check --ephemeral -s read-only --disable memories -m gpt-5.5 -c model_reasoning_effort=high "Do not use external tools. A black bag contains candies with counts: round apple 7, round peach 9, round watermelon 8; star apple 7…
    • Ranks #10 on LMArena with 1470 Elo, showing human preference at scale.
      Source 13
      Ranks #6 of 17 on LMArena's overall text arena (Elo 1491), based on blind human preference votes.
      LMArena text arenaOpen original ↗

    Watch out for

    • Users report ongoing issues with stalling and needing quota monitoring to keep tasks running.
      Source 14
      I mean you can just ask them to do exactly that. Especially with GPT (5.5), I've been having a lot of issues with it just repeatedly stalling out. I had to build a quota monitoring skill so that it'd keep plowing forward until either the task was finished (in some way) or the quota budget was exhausted. I also had issues with the compaction. Codex seems to compact... weirdly, resulting in the agent becoming a newborn after each compaction event. Telling it to use a notes file is basically essen…
      perching_aixOpen original ↗
    • Gets stuck for long periods on logic like Iris proofs rather than making progress.
      Source 11
      "But LLMs are prone to hallucinations which can really impact a string of interdependent logic like a proof. So I’m assuming it would respond with something that’s not complete nonsense to this proof most of the time." Unfortunately in my experience that's not really the case. For me, very often GPT 5.5 (which was a good deal better than Opus at this kind of task) would just get stuck for long periods when working in a logic like Iris. It wouldn't necessarily outright prove nonsense, but it wou…
    • Performance has degraded noticeably since launch, struggling on tasks it previously handled.
      Source 12
      this explains so much why gpt 5.5 has been so bad lately it was really puzzling why it struggled so much where when it first came out it was one shotting stuff totally amazing, i tried the prompt that will tell you if your plan is degraded: codex exec --json --skip-git-repo-check --ephemeral -s read-only --disable memories -m gpt-5.5 -c model_reasoning_effort=high "Do not use external tools. A black bag contains candies with counts: round apple 7, round peach 9, round watermelon 8; star apple 7…

Frequently asked

Which model is best for linear algebra and technical math?
Claude Fable 5. User reports describe it as massively better than GPT-5.5 or Opus at linear algebra routines, with better insights and better handling of corner cases.
Is GPT-5.5 good for math reasoning?
No. Multiple users report issues with stalling, getting stuck on logic proofs, and quality degradation where it struggles on tasks it previously handled well.
Which models handle word reasoning puzzles well?
Claude Opus 4.7 and Gemini 3.1 Pro both consistently solve the NYTimes Connections puzzle, while GPT-5.5 fails it.
What's the highest-ranked model overall?
Claude Opus 4.6 ranks #1 on LMArena with an Elo of 1501, followed by Fable 5 at #2 with 1493.

Sources

  1. 1

    Ranks #2 of 17 on LMArena's overall text arena (Elo 1504), based on blind human preference votes.

    LMArena text arena · Benchmark · Jul 20, 2026
  2. 2

    Ranks #1 of 17 on LMArena's overall text arena (Elo 1512), based on blind human preference votes.

    LMArena text arena · Benchmark · Jul 20, 2026
  3. 3

    > It's a pretty high moat getting into stuff like simulation software I'm currently working on a simulation game about space and orbital mechanics. I have a lot of software experience, I know how to build large projects and architect my code, and I know how to to test the end result to ensure I'm getting what I want. But I also don't have a strong math or physics background. In my experience, Claude (Opus 4.6+) has had no issues writing any simulation or game related math code. And the key thin…

    pyth0 · Hacker News · Jun 7, 2026
  4. 4

    The initial motivation for this was likely to thwart any competition. Already Anthropic has accused some companies of organized distillation efforts at a massive scale. Back when I used antigravity, it used to show the reasoning intact - at least for Gemini Pro 3.1, and likely for Claude Opus 4.6 (not 100% certain about it). I have some recollection of stopping the models mid-turn when they started going astray. As a power user, I find reasoning fascinating to read and genuinely useful at times…

    sheepscreek · Hacker News · Jun 23, 2026
  5. 5

    Ranks #4 of 17 on LMArena's overall text arena (Elo 1499), based on blind human preference votes.

    LMArena text arena · Benchmark · Jul 20, 2026
  6. 6

    I was talking to a Claude Opus 4.7 chatbot about low-discrepancy sequences and made the mistake of framing a thought experiment as "deceiving" a vendor who owns a scrambler by making a hard to scramble low-discrepancy sequence generator. Admittedly that sounds sketchy but the conversation up to that point--without the explicit framing--was about the same subject just shorn of metaphor. With the framing I got a refusal message that likened my request to Dual_EC_DRBG!

    adampunk · Hacker News · Jun 7, 2026
  7. 7

    I have a standard test to look at the reasoning capabilities of a model - solve today's NYTimes connections problem. Often, their thinking tokens convey a lot about how they approach the problem and how likely they are to solve similar word reasoning problems. Claude 4.7 and Gemini 3.1 Pro have nailed all so far, GPT 5.5 failed miserably. Of the chinese models, Kimi-K-2.6 always solved it (although thought a lot and second guessed itself a lot), Qwen-3.6-Plus often gave wrong answers and GLM-5.…

    LeSavant · Hacker News · May 2, 2026
  8. 8

    Ranks #9 of 17 on LMArena's overall text arena (Elo 1489), based on blind human preference votes.

    LMArena text arena · Benchmark · Jul 20, 2026
  9. 9

    Two key quotes: • Reasoning: Through the expansion of reasoning tokens, DeepSeek-V4-Pro-Max demonstrates superior performance relative to GPT-5.2 and Gemini-3.0-Pro on standard reasoning benchmarks. Nevertheless, its performance falls marginally short of GPT-5.4 and Gemini3.1-Pro, suggesting a developmental trajectory that trails state-of-the-art frontier models by approximately 3 to 6 months. Furthermore, DeepSeek-V4-Flash-Max achieves comparable performance to GPT-5.2 and Gemini-3.0-Pro, esta…

    creamyhorror · Hacker News · Apr 24, 2026
  10. 10

    Ranks #3 of 17 on LMArena's overall text arena (Elo 1499), based on blind human preference votes.

    LMArena text arena · Benchmark · Jul 20, 2026
  11. 11

    "But LLMs are prone to hallucinations which can really impact a string of interdependent logic like a proof. So I’m assuming it would respond with something that’s not complete nonsense to this proof most of the time." Unfortunately in my experience that's not really the case. For me, very often GPT 5.5 (which was a good deal better than Opus at this kind of task) would just get stuck for long periods when working in a logic like Iris. It wouldn't necessarily outright prove nonsense, but it wou…

    Jweb_Guru · Hacker News · Jul 10, 2026
  12. 12

    this explains so much why gpt 5.5 has been so bad lately it was really puzzling why it struggled so much where when it first came out it was one shotting stuff totally amazing, i tried the prompt that will tell you if your plan is degraded: codex exec --json --skip-git-repo-check --ephemeral -s read-only --disable memories -m gpt-5.5 -c model_reasoning_effort=high "Do not use external tools. A black bag contains candies with counts: round apple 7, round peach 9, round watermelon 8; star apple 7…

    zuzululu · Hacker News · Jul 5, 2026
  13. 13

    Ranks #6 of 17 on LMArena's overall text arena (Elo 1491), based on blind human preference votes.

    LMArena text arena · Benchmark · Jul 20, 2026
  14. 14

    I mean you can just ask them to do exactly that. Especially with GPT (5.5), I've been having a lot of issues with it just repeatedly stalling out. I had to build a quota monitoring skill so that it'd keep plowing forward until either the task was finished (in some way) or the quota budget was exhausted. I also had issues with the compaction. Codex seems to compact... weirdly, resulting in the agent becoming a newborn after each compaction event. Telling it to use a notes file is basically essen…

    perching_aix · Hacker News · Jul 10, 2026

Rankings synthesized from community evidence and open benchmarks. See methodology. Not driven by vendor marketing.