Recommendation for Competition math

Competition Math

The best LLM for competition math among the candidates is Claude Opus 4.6, which holds the #1 ranking on LMArena's overall text arena. However, all available evidence consists of general preference rankings rather than math-specific benchmarks, so this ranking is provisional and assumes arena performance correlates with multi-step reasoning ability. The top tier includes Claude Fable 5 (#2) and Claude Opus 4.7 (#3), all showing strong Elo scores above 1490 (e2, e4, e5). These models likely handle complex tasks well given their standings, but builders should verify performance on actual olympiad-style problems. The mid-tier candidates, including Muse Spark 1.1 and Gemini 3.5 Flash, rank #4 through #8 and may offer a balance of capability and other factors like speed or cost. Without direct evidence of mathematical reasoning, this ranking defaults to overall arena performance as the strongest available proxy.

About this recommendation

Updated
Jul 17, 2026
Evidence through
Jul 17, 2026
Sources
7
Revision
v1
  1. Claude Opus 4.6 claims the top spot with a narrow Elo lead over the field. Competition math rewards sustained logical reasoning, and this model's #1 ranking suggests strong performance on difficult, multi-part prompts where voters likely evaluated answer quality (e2).

    Best when: You want the strongest available candidate based on human preference data and are willing to test it against actual olympiad problems.

    Tips

    • Highest-ranked model in the set at #1 with an Elo of 1501.
      Source 1
      Ranks #1 of 17 on LMArena's overall text arena (Elo 1512), based on blind human preference votes.
      LMArena text arenaOpen original ↗

    Watch out for

    • Arena ranking reflects general preference rather than specific mathematical reasoning.
      Source 1
      Ranks #1 of 17 on LMArena's overall text arena (Elo 1512), based on blind human preference votes.
      LMArena text arenaOpen original ↗
  2. Claude Fable 5 sits just behind the leader with a #2 ranking and an Elo of 1493. The eight-point gap to Opus 4.6 is small, and competitive math performance is likely similar between them (e4).

    Best when: You want near-top-tier performance and may prefer different pricing, latency, or availability characteristics.

    Tips

    • Second-highest ranking with Elo of 1493, very close to the leader.
      Source 2
      Ranks #2 of 17 on LMArena's overall text arena (Elo 1504), based on blind human preference votes.
      LMArena text arenaOpen original ↗

    Watch out for

    • Slightly trails Opus 4.6 in the rankings.
      Source 2
      Ranks #2 of 17 on LMArena's overall text arena (Elo 1504), based on blind human preference votes.
      LMArena text arenaOpen original ↗
  3. Claude Opus 4.7 rounds out the top three with an Elo of 1490. It sits only three points behind Fable 5 and eleven behind Opus 4.6, forming a clear top-tier cluster (e5).

    Best when: You want a top-3 model and may benefit from different tuning or behavior compared to other Claude variants.

    Tips

    • Ranks #3 overall with a strong Elo of 1490.
      Source 3
      Ranks #4 of 17 on LMArena's overall text arena (Elo 1499), based on blind human preference votes.
      LMArena text arenaOpen original ↗
    • Part of a top-tier cluster with single-digit ranks.
      Source 3
      Ranks #4 of 17 on LMArena's overall text arena (Elo 1499), based on blind human preference votes.
      LMArena text arenaOpen original ↗

    Watch out for

    • Third in the Claude lineup, trailing two siblings.
      Source 3
      Ranks #4 of 17 on LMArena's overall text arena (Elo 1499), based on blind human preference votes.
      LMArena text arenaOpen original ↗
  4. Muse Spark 1.1 leads the second tier at #4 with an Elo of 1481. It sits nine points behind Opus 4.7 but comfortably ahead of Gemini 3.5 Flash, positioning it as a strong alternative (e1).

    Best when: You want high performance outside the Claude family or need a different model provider.

    Tips

    • Ranks #4, leading the non-Claude contenders.
      Source 4
      Ranks #7 of 17 on LMArena's overall text arena (Elo 1491), based on blind human preference votes.
      LMArena text arenaOpen original ↗
    • Elo of 1481 remains highly competitive within the top five.
      Source 4
      Ranks #7 of 17 on LMArena's overall text arena (Elo 1491), based on blind human preference votes.
      LMArena text arenaOpen original ↗

    Watch out for

    • Trails the top-three Claude models by a clear margin.
      Source 4
      Ranks #7 of 17 on LMArena's overall text arena (Elo 1491), based on blind human preference votes.
      LMArena text arenaOpen original ↗
  5. Gemini 3.5 Flash ranks #5 with an Elo of 1480, just one point behind Muse Spark 1.1. The tight gap suggests comparable capability, and Flash models often prioritize speed (e6).

    Best when: You want a balance of ranking and potentially faster inference from a Google model.

    Tips

    • Ranks #5 with an Elo of 1480, very close to #4.
      Source 5
      Ranks #8 of 17 on LMArena's overall text arena (Elo 1490), based on blind human preference votes.
      LMArena text arenaOpen original ↗

    Watch out for

    • Outside the top tier by a small margin.
      Source 5
      Ranks #8 of 17 on LMArena's overall text arena (Elo 1490), based on blind human preference votes.
      LMArena text arenaOpen original ↗
    • No math-specific evidence confirms olympiad capability.
      Source 5
      Ranks #8 of 17 on LMArena's overall text arena (Elo 1490), based on blind human preference votes.
      LMArena text arenaOpen original ↗
  6. Gemini 3.1 Pro Preview takes #6 with an Elo of 1479, one point behind Gemini 3.5 Flash. The naming suggests a preview release, which may affect production readiness (e7).

    Best when: You want to test Google's Pro-tier capabilities and can tolerate preview-stage behavior.

    Tips

    • Top-10 ranking at #6 with Elo 1479.
      Source 6
      Ranks #9 of 17 on LMArena's overall text arena (Elo 1489), based on blind human preference votes.
      LMArena text arenaOpen original ↗

    Watch out for

    • Slightly trails Gemini 3.5 Flash.
      Source 6
      Ranks #9 of 17 on LMArena's overall text arena (Elo 1489), based on blind human preference votes.
      LMArena text arenaOpen original ↗
    • Preview designation may indicate lower stability.
      Source 6
      Ranks #9 of 17 on LMArena's overall text arena (Elo 1489), based on blind human preference votes.
      LMArena text arenaOpen original ↗
  7. Qwen3.7 Max ranks #8 with an Elo of 1476. Qwen models have a reputation for strong mathematical reasoning in benchmark circles, though the provided evidence does not confirm this (e8). It lands in the middle of the competitive pack.

    Best when: You want to test a Qwen variant specifically, possibly based on prior positive experience.

    Tips

    • Ranks #8 with Elo 1476 in the top-10 cluster.
      Source 7
      Ranks #11 of 17 on LMArena's overall text arena (Elo 1485), based on blind human preference votes.
      LMArena text arenaOpen original ↗

    Watch out for

    • Trails the top tier by roughly 20 Elo points.
      Source 7
      Ranks #11 of 17 on LMArena's overall text arena (Elo 1485), based on blind human preference votes.
      LMArena text arenaOpen original ↗

Frequently asked

Which model ranks highest overall in the arena?
Claude Opus 4.6 ranks #1 with an Elo of 1501, making it the top general-purpose text model among the candidates (e2).
Are there math-specific benchmarks available for these models?
No, the provided evidence only includes LMArena overall text rankings based on blind human preference votes. No math-specific benchmarks like MATH or GSM8K are listed.
Which Claude models are in the top tier?
Claude Opus 4.6 (#1), Claude Fable 5 (#2), and Claude Opus 4.7 (#3) all rank in the top three positions (e2, e4, e5).
Where does GPT-5.4 rank compared to other models?
GPT-5.4 ranks #9 with an Elo of 1470, tied with GPT-5.5 and outside the top-tier cluster (e10, e11).

Sources

  1. 1

    Ranks #1 of 17 on LMArena's overall text arena (Elo 1512), based on blind human preference votes.

    LMArena text arena · Benchmark · Jul 20, 2026
  2. 2

    Ranks #2 of 17 on LMArena's overall text arena (Elo 1504), based on blind human preference votes.

    LMArena text arena · Benchmark · Jul 20, 2026
  3. 3

    Ranks #4 of 17 on LMArena's overall text arena (Elo 1499), based on blind human preference votes.

    LMArena text arena · Benchmark · Jul 20, 2026
  4. 4

    Ranks #7 of 17 on LMArena's overall text arena (Elo 1491), based on blind human preference votes.

    LMArena text arena · Benchmark · Jul 20, 2026
  5. 5

    Ranks #8 of 17 on LMArena's overall text arena (Elo 1490), based on blind human preference votes.

    LMArena text arena · Benchmark · Jul 20, 2026
  6. 6

    Ranks #9 of 17 on LMArena's overall text arena (Elo 1489), based on blind human preference votes.

    LMArena text arena · Benchmark · Jul 20, 2026
  7. 7

    Ranks #11 of 17 on LMArena's overall text arena (Elo 1485), based on blind human preference votes.

    LMArena text arena · Benchmark · Jul 20, 2026

Rankings synthesized from community evidence and open benchmarks. See methodology. Not driven by vendor marketing.