Recommendation for Logic & reasoning
Logic & Problem Solving
The best LLM for logic and problem solving is Claude Opus 4.6, which holds the top spot on LMArena with an Elo of 1501 and offers visible reasoning traces that users find genuinely useful for following the model's logic. The best LLM for logic and problem solving should combine strong benchmark performance with transparent reasoning, and Opus 4.6 delivers on both. Close alternatives include Gemini 3.1 Pro Preview, which nails word reasoning tests like the NYTimes Connections puzzle, and Claude Opus 4.7 and Claude Fable 5, both ranked in the top three overall. This use case demands models that can work through multi-step problems without losing the thread, and the evidence points to Anthropic's Opus line and Google's Gemini 3.1 Pro as the leaders here.
About this recommendation
- Updated
- Jul 17, 2026
- Evidence through
- Jul 17, 2026
- Sources
- 10
- Revision
- v1
Claude Opus 4.6 sits at the top of LMArena's rankings and offers the visible reasoning traces that make it easier to trust the model's logic on complex puzzles. Power users report finding these thinking tokens genuinely useful for following the problem-solving process and identifying when the model starts to go wrong. It's the safest pick for logic-heavy work.
Best when: You need top-tier general performance and want to audit the model's reasoning process step by step.
Tips
- Holds the #1 rank on LMArena overall with an Elo of 1501.
- Shows reasoning transparently, letting users follow the logic and intervene if needed.
Gemini 3.1 Pro Preview is a strong contender for pure reasoning tasks. It consistently solves the NYTimes Connections puzzle, a word reasoning test that trips up other models like GPT-5.5. Combined with a #6 LMArena ranking and visible reasoning that users appreciate, it's a solid choice for multi-step problem solving.
Best when: You prioritize raw reasoning ability over overall chat quality, especially for word puzzles.
Tips
- Ranks #6 overall on LMArena with an Elo of 1479.
- Consistently solves the NYTimes Connections word reasoning puzzle.
- Displays reasoning tokens that users find useful for understanding the problem-solving approach.
Claude Opus 4.7 combines a top-three LMArena ranking with proven performance on word reasoning benchmarks. It solves the NYTimes Connections puzzle alongside Gemini 3.1 Pro, making it another strong option for logic tasks where you need to see the model work through the problem.
Best when: You want Anthropic's strong reasoning pedigree with a slightly newer model than Opus 4.6.
Tips
- Ranks #3 of 49 on LMArena with an Elo of 1490.
Claude Fable 5 holds the #2 spot on LMArena with an Elo of 1493, putting it just behind Opus 4.6 in blind human preference voting. While there's no specific reasoning test data for it, its proximity to Opus 4.6 in the rankings suggests similar capabilities for logic-heavy use cases.
Best when: You want a top-ranked Anthropic model but prefer a different flavor than Opus.
Tips
- Ranks #2 of 49 on LMArena with an Elo of 1493.
GPT-5.4 lands at #9 on LMArena and earns a mention as marginally ahead of DeepSeek on reasoning benchmarks. It's a capable model for logic tasks, though the lack of community testing data on puzzles like Connections makes it harder to gauge real-world reasoning performance compared to Anthropic and Google's offerings.
Best when: You're already committed to the OpenAI ecosystem and need stronger reasoning than GPT-5.5.
Tips
- Ranks #9 of 49 on LMArena with an Elo of 1470.
- Outperforms DeepSeek-V4-Pro-Max and trails only Gemini 3.1 Pro on reasoning benchmarks.
GPT-5.5 has slipped to #10 on LMArena and users report it struggles with tasks it previously handled well. It fails the NYTimes Connections test outright, which is a red flag for logic and reasoning work. Users have noticed degraded performance where it once excelled.
Best when: You need it for compatibility reasons, but most logic tasks are better served by other models.
Tips
- Still maintains a reasonable #10 rank on LMArena with an Elo of 1470.
Watch out for
- Fails miserably on the NYTimes Connections word reasoning test.
- Users report it struggles with tasks it previously handled well, suggesting performance regression.
Frequently asked
- Which model is best for word reasoning puzzles like NYTimes Connections?
- Gemini 3.1 Pro Preview and Claude 4.7 both solve these puzzles consistently, while GPT-5.5 fails them.
- Does reasoning transparency matter for logic tasks?
- Yes, power users find the visible thinking tokens useful for understanding how models approach problems and for catching when reasoning goes astray.
- What's the highest-ranked model on LMArena for general text?
- Claude Opus 4.6 ranks first with an Elo of 1501 out of 49 models.
- Is GPT-5.5 reliable for complex reasoning?
- No, users report it struggles with tasks it previously handled well, and it fails miserably on word reasoning tests.
Sources
- 1
“Ranks #1 of 17 on LMArena's overall text arena (Elo 1512), based on blind human preference votes.”
LMArena text arena · Benchmark · Jul 20, 2026 - 2
“The initial motivation for this was likely to thwart any competition. Already Anthropic has accused some companies of organized distillation efforts at a massive scale. Back when I used antigravity, it used to show the reasoning intact - at least for Gemini Pro 3.1, and likely for Claude Opus 4.6 (not 100% certain about it). I have some recollection of stopping the models mid-turn when they started going astray. As a power user, I find reasoning fascinating to read and genuinely useful at times…”
sheepscreek · Hacker News · Jun 23, 2026 - 3
“Ranks #9 of 17 on LMArena's overall text arena (Elo 1489), based on blind human preference votes.”
LMArena text arena · Benchmark · Jul 20, 2026 - 4
“I have a standard test to look at the reasoning capabilities of a model - solve today's NYTimes connections problem. Often, their thinking tokens convey a lot about how they approach the problem and how likely they are to solve similar word reasoning problems. Claude 4.7 and Gemini 3.1 Pro have nailed all so far, GPT 5.5 failed miserably. Of the chinese models, Kimi-K-2.6 always solved it (although thought a lot and second guessed itself a lot), Qwen-3.6-Plus often gave wrong answers and GLM-5.…”
LeSavant · Hacker News · May 2, 2026 - 5
“Ranks #4 of 17 on LMArena's overall text arena (Elo 1499), based on blind human preference votes.”
LMArena text arena · Benchmark · Jul 20, 2026 - 6
“Ranks #2 of 17 on LMArena's overall text arena (Elo 1504), based on blind human preference votes.”
LMArena text arena · Benchmark · Jul 20, 2026 - 7
“Ranks #3 of 17 on LMArena's overall text arena (Elo 1499), based on blind human preference votes.”
LMArena text arena · Benchmark · Jul 20, 2026 - 8
“Two key quotes: • Reasoning: Through the expansion of reasoning tokens, DeepSeek-V4-Pro-Max demonstrates superior performance relative to GPT-5.2 and Gemini-3.0-Pro on standard reasoning benchmarks. Nevertheless, its performance falls marginally short of GPT-5.4 and Gemini3.1-Pro, suggesting a developmental trajectory that trails state-of-the-art frontier models by approximately 3 to 6 months. Furthermore, DeepSeek-V4-Flash-Max achieves comparable performance to GPT-5.2 and Gemini-3.0-Pro, esta…”
creamyhorror · Hacker News · Apr 24, 2026 - 9
“Ranks #6 of 17 on LMArena's overall text arena (Elo 1491), based on blind human preference votes.”
LMArena text arena · Benchmark · Jul 20, 2026 - 10
“this explains so much why gpt 5.5 has been so bad lately it was really puzzling why it struggled so much where when it first came out it was one shotting stuff totally amazing, i tried the prompt that will tell you if your plan is degraded: codex exec --json --skip-git-repo-check --ephemeral -s read-only --disable memories -m gpt-5.5 -c model_reasoning_effort=high "Do not use external tools. A black bag contains candies with counts: round apple 7, round peach 9, round watermelon 8; star apple 7…”
zuzululu · Hacker News · Jul 5, 2026
Rankings synthesized from community evidence and open benchmarks. See methodology. Not driven by vendor marketing.