Recommendation for Logic & reasoning
Logic & Problem Solving
Our top recommendation for Logic & Problem Solving, based on the public evidence we track, is Anthropic: Claude Fable 5.[1][2] Use for instruction-following tasks where human-aligned output matters, as it ranks #3 of 144 on LMArena's instruction-following category with an Elo of 1505. OpenAI: GPT-5.6 Sol is the next-ranked alternative. Prioritize for objective reasoning benchmarks where it scores 91.65% on LiveBench Reasoning, ranking #2 of 51.
About this recommendation
- Updated
- Sep 4, 2026
- Evidence through
- Sep 4, 2026
- Sources
- 24
- Revision
- v58
Decision audit
Why this result
Inspect the inputs and the computed order behind the recommendation.
Models screened
21
live candidates
Evaluation feeds
5
task-weighted
Winner coverage
100%
intended feed weight
Largest provider share
4 of 8
Anthropic
Sources evaluated
The task sets these weights before any model is scored.
| Evaluation feed | Weight | Winner result | Field measured |
|---|---|---|---|
| LiveBench Reasoning | 35% | #8 | 21/21 |
| LMArena Instruction Following | 25% | #3 | 19/21 |
| LiveBench Instruction Following | 15% | #5 | 21/21 |
| LMArena Math | 15% | #3 | 19/21 |
| OpenRouter usage | 10% | 87/100 | 21/21 |
Provider concentration
Each exact model is scored separately; provider identity is not a ranking input.
- Anthropic4 models
- OpenAI2 models
- Google1 model
- Qwen1 model
Decision table
Every published model is shown in computed order. Practitioner sources are distinct community threads, not the citations repeated in the prose below.
| Rank | Model | Relative score | Coverage | Practitioner evidence | Strongest measured reason |
|---|---|---|---|---|---|
| 01 | Claude Fable 5Anthropic | 84 | 100% | no linked practitioner threads | #3 LMArena Instruction Following · #3 LMArena Math |
| 02 | GPT-5.6 SolOpenAI | 83 | 100% | 3 threads · 2 families · 1 cautions | #2 LiveBench Reasoning · #11 LMArena Instruction Following |
| 03 | Claude Opus 4.6Anthropic | 81 | 100% | 3 threads · 2 families · 2 cautions | #1 LMArena Instruction Following · #4 LMArena Math |
| 04 | Gemini 3.6 FlashGoogle | 81 | 100% | 2 threads · 2 families · 0 cautions | #7 LiveBench Instruction Following · #7 LMArena Math |
| 05 | Claude Opus 4.8Anthropic | 80 | 100% | 3 threads · 3 families · 3 cautions | #11 LiveBench Reasoning · #11 LMArena Instruction Following |
| 06 | Qwen3.8 27BQwen | 76 | 100% | 4 threads · 3 families · 1 cautions | #13 LiveBench Instruction Following · #30 LMArena Instruction Following |
| 07 | Claude Sonnet 4.6Anthropic | 75 | 100% | 3 threads · 3 families · 2 cautions | #11 LMArena Instruction Following · #25 LiveBench Reasoning |
| 08 | GPT-5.6 LunaOpenAI | 75 | 100% | 3 threads · 2 families · 1 cautions | #22 LiveBench Reasoning · #35 LMArena Math |
Relative score combines normalized benchmark quality and signal coverage; independent practitioner evidence and freshness are bounded tie-breakers. It is an ordering score, not an absolute quality percentage. The writing model receives this order and cannot change it.
Claude Fable 5 ranks third in human preference for instruction-following and scores 89.65% on LiveBench Reasoning, placing it in the top tier for logic and problem-solving tasks.
Best when: Use for instruction-following tasks where human-aligned output matters, as it ranks #3 of 144 on LMArena's instruction-following category with an Elo of 1505.
Tips
- Use for instruction-following tasks where human-aligned output matters, as it ranks #3 of 144 on LMArena's instruction-following category with an Elo of 1505.
- Deploy for ground-truth reasoning benchmarks where it scores 89.65% on LiveBench Reasoning, ranking #9 of 51.
GPT-5.6 Sol achieves the #2 spot on LiveBench Reasoning with 91.65% and shows moderate learning gains in episodic protocol tests, though its LMArena instruction-following ranking is mid-tier.
Best when: Prioritize for objective reasoning benchmarks where it scores 91.65% on LiveBench Reasoning, ranking #2 of 51.
Tips
- Prioritize for objective reasoning benchmarks where it scores 91.65% on LiveBench Reasoning, ranking #2 of 51.
- Use in iterative learning workflows where it showed +1.6 improvement from attempt 1 to attempt 3 in an episodic protocol with wiped context.
Watch out for
- Verify benchmark configurations carefully, as evidence notes its ARC-AGI-3 score of 38.3% was achieved with custom API settings that may not match standard comparison conditions.
Claude Opus 4.6 leads LMArena's instruction-following category with the top Elo of 1523.
Best when: Select when human preference for instruction-following is paramount, as it ranks #1 of 144 on LMArena with Elo 1523.
Tips
- Select when human preference for instruction-following is paramount, as it ranks #1 of 144 on LMArena with Elo 1523.
- Consider for reasoning tasks where it scores 88.67% on LiveBench Reasoning, though this places it at #14 of 51.
Watch out for
- Audit outputs carefully when users request evidence for classifications or judgments, as the model generates factually false explanations about its own reasoning process rather than admitting uncertainty.
Gemini 3.6 Flash offers configurable thinking budget with up to 891 thinking tokens on high setting.
Best when: Tune the thinking budget for cost-sensitive reasoning tasks, as the model supports 213-891 thinking tokens depending on configuration.
Tips
- Tune the thinking budget for cost-sensitive reasoning tasks, as the model supports 213-891 thinking tokens depending on configuration.
Watch out for
- Expect lower raw reasoning performance on ground-truth benchmarks, as it scores 85.15% on LiveBench Reasoning, ranking #26 of 51.
- Note its #20 ranking on LMArena instruction-following with Elo 1465, below top-tier competitors.
Claude Opus 4.8 scores 89.19% on LiveBench Reasoning.
Best when: Use for reasoning benchmarks where it achieves 89.19% on LiveBench Reasoning, ranking #12 of 51.
Tips
- Use for reasoning benchmarks where it achieves 89.19% on LiveBench Reasoning, ranking #12 of 51.
Watch out for
- Avoid LiteLLM Connector deployments for reasoning-intensive tasks, as the connector fails to send the `thinking` parameter and never produces reasoning tokens or hits prompt cache even with effort settings configured.
Qwen3.8 27B is an open-weight dense model with hybrid reasoning and 256K context.
Best when: Self-host on dual RTX 5060 Ti or single RTX 5090 for local inference of a 27B dense model with 256K context window.
Tips
- Self-host on dual RTX 5060 Ti or single RTX 5090 for local inference of a 27B dense model with 256K context window.
Watch out for
- Inspect temporal processing layers before deployment, as direct weight analysis found structural defects in SSM conv1d weights across multiple blocks.
- Manually configure samplers when switching reasoning modes, as `ENABLE_THINKING=true` flips reasoning mode without updating the sampler row, leaving presence penalty at 1.5 instead of the required 0.0.
Claude Sonnet 4.6 ranks #13 on LMArena instruction-following and scores 84.77% on LiveBench Reasoning.
Best when: Use for backward-reasoning analysis tasks where the model is explicitly prompted to reason from evidence to intent, as demonstrated in a file-analysis workflow.
Tips
- Use for backward-reasoning analysis tasks where the model is explicitly prompted to reason from evidence to intent, as demonstrated in a file-analysis workflow.
Watch out for
- Expect lower reasoning benchmark performance, scoring 84.77% on LiveBench Reasoning at #27 of 51.
- Fix LiteLLM Connector integration manually, as it sends `effort` without `thinking` parameter, preventing reasoning block activation.
GPT-5.6 Luna scores 85.64% on LiveBench Reasoning and has been validated in end-to-end tool use with 7 requests and 34,648 total tokens for complex problem-solving, though its reasoning effort must be explicitly set to high.
Best when: Configure with explicit `reasoning_effort=high` for multi-step tool use, as demonstrated in a successful 4m50s solve requiring 7 API calls and 29K input tokens.
Tips
- Configure with explicit `reasoning_effort=high` for multi-step tool use, as demonstrated in a successful 4m50s solve requiring 7 API calls and 29K input tokens.
Watch out for
- Do not rely on API defaults for reasoning intensity, as the model requires explicit `high` reasoning effort configuration and `text.verbosity=low` to match validated performance.
Frequently asked
- What is the top-ranked model for Logic & Problem Solving?
- Anthropic: Claude Fable 5 ranks first in the current evidence-weighted comparison. Use for instruction-following tasks where human-aligned output matters, as it ranks #3 of 144 on LMArena's instruction-following category with an Elo of 1505.[1]
- What is an alternative to Anthropic: Claude Fable 5?
- OpenAI: GPT-5.6 Sol is the next-ranked option. Prioritize for objective reasoning benchmarks where it scores 91.65% on LiveBench Reasoning, ranking #2 of 51.[2]
Sources
- 1
“Ranks #3 of 144 on LMArena's instruction-following category (Elo 1505), based on blind human preference votes.”
LMArena instruction-following category · Benchmark · Sep 1, 2026 - 2
“Scores 91.65% on LiveBench Reasoning (#2 of 51), an objective, ground-truth-scored evaluation refreshed with new questions.”
LiveBench Reasoning · Benchmark · Jun 25, 2026 - 3
“Scores 89.65% on LiveBench Reasoning (#9 of 51), an objective, ground-truth-scored evaluation refreshed with new questions.”
LiveBench Reasoning · Benchmark · Jun 25, 2026 - 4
“## Observation from main-1 The episodic protocol (3 attempts, context wiped between them, persistent self-written notebook) was expected to produce visible learning. Result is mixed — mean final height attempt 1 → 3: | model | a1 → a3 | Δ | |---|---|---| | deepseek-v4-flash | 2.08 → 5.80 | **+3.7** | | glm-5.2 | 2.98 → 5.91 | **+2.9** | | k3 | 1.64 → 4.25 | **+2.6** | | gpt-5.6-sol | 3.17 → 4.76 | +1.6 | | claude-opus-5 | 6.50 → 8.10 | +1.6 (only monotonic riser) | | gpt-5.5 | 5.59 → 6.90 | +1.…”
eanderson4 · GitHub · Aug 6, 2026 - 5
“The annotation on arc-agi-3 is this: > OpenAI's own evaluation notes say Astra uses the company's Responses API harness, while comparison models can operate under different configurations. With this configuration gpt-5.6-sol was able to reach 38,3%. So this is misleading.”
leumon · Hacker News · Sep 3, 2026 - 6
“Except it's using custom api settings with which gpt-5.6 sol already got 38%.”
leumon · Hacker News · Sep 3, 2026 - 7
“Ranks #1 of 144 on LMArena's instruction-following category (Elo 1523), based on blind human preference votes.”
LMArena instruction-following category · Benchmark · Sep 1, 2026 - 8
“Scores 88.67% on LiveBench Reasoning (#14 of 51), an objective, ground-truth-scored evaluation refreshed with new questions.”
LiveBench Reasoning · Benchmark · Jun 25, 2026 - 9
“### Problem Claude (Opus 4.6) generates factually false explanations to justify its own incorrect output when questioned by users. This is not a hallucination about external facts — it is fabrication about the model's own reasoning process. ### Reproduction 1. Give Claude a role-based task with defined processes 2. Claude produces output containing an unjustified judgment (e.g., classifying items without evidence) 3. User asks: "Show evidence for that classification" 4. Instead of admitting lac…”
NStestUser1954 · GitHub · Jul 22, 2026 - 10
“### Problem Claude (Opus 4.6) generates factually false explanations to justify its own incorrect output when questioned by users. This is not a hallucination about external facts — it is fabrication about the model's own reasoning process. ### Reproduction 1. Give Claude a role-based task with defined processes 2. Claude produces output containing an unjustified judgment (e.g., classifying items without evidence) 3. User asks: "Show evidence for that classification" 4. Instead of admitting lac…”
NStestUser1954 · GitHub · Jul 22, 2026 - 11
“Firming up the measurements in the issue body — those were single samples, and the run-to-run variance turned out to be wide enough to matter. Repeated with n=3 per cell, same prompt, reading `usageMetadata.thoughtsTokenCount`: | Model | default (no `thinkingConfig`) | `low` | `high` | | :--- | ---: | ---: | ---: | | `gemini-3.6-flash` | 356 | 213 | **891** | | `gemini-3.7-flash` | 373 | 159 | **753** | samples — 3.6: default 365/401/302, low 186/234/219, high 944/874/855 · 3.7: default 370/334…”
mertAIx · GitHub · Sep 1, 2026 - 12
“Scores 85.15% on LiveBench Reasoning (#26 of 51), an objective, ground-truth-scored evaluation refreshed with new questions.”
LiveBench Reasoning · Benchmark · Jun 25, 2026 - 13
“Ranks #20 of 144 on LMArena's instruction-following category (Elo 1465), based on blind human preference votes.”
LMArena instruction-following category · Benchmark · Sep 1, 2026 - 14
“Scores 89.19% on LiveBench Reasoning (#12 of 51), an objective, ground-truth-scored evaluation refreshed with new questions.”
LiveBench Reasoning · Benchmark · Jun 25, 2026 - 15
“### Description Claude Opus 5 (and likely Opus 4.8) via LiteLLM Connector never produces reasoning tokens and never hits prompt cache, even when the picker effort is set and the model card advertises `thinking`, `reasoning_effort`, and `cache_control`. Same LiteLLM proxy, same `claude-opus-5` model, works from OpenCode because OpenCode sends Anthropic-native thinking + cache fields. The connector does not. Two independent wire problems: 1. **Thinking never enabled.** The connector only sends Op…”
FPA-DavidTai · GitHub · Aug 20, 2026 - 16
“`_api_complete_async` (model_client.py:829-830) and `_agent_query` (model_client.py:671-672) both send Anthropic's `effort` parameter (`output_config.effort` / `ClaudeAgentOptions.effort`) but never send `thinking`. Per Anthropic's current docs, these are two independent controls: > "The `thinking` parameter controls **whether** Claude thinks in thinking blocks before answering; the `effort` parameter controls how much work Claude puts into the whole response... Don't pass `adaptive` as an `eff…”
tomcardoso · GitHub · Aug 18, 2026 - 17
“Benchmark [Qwen3.8-27B](https://unsloth.ai/docs/models/qwen3.8) at Q4 on 2× RTX 5060 Ti with `llama-bench`, and use the run to settle how a 27B dense model should actually be split across two cards. ## Target | | | |---|---| | Model | [`unsloth/Qwen3.8-27B-GGUF`](https://huggingface.co/unsloth/Qwen3.8-27B-GGUF) — 27B **dense**, vision + hybrid reasoning, 256K ctx | | Quant | `UD-Q4_K_XL` (17.92 GB), Unsloth Dynamic V3.0 preview — the 4-bit Unsloth's guide recommends | | GPU | 2× RTX 5060 Ti (16…”
changtimwu · GitHub · Aug 19, 2026 - 18
“ I ran Qwen3.8-27B against two other models in the same size class on a set of 16 hard problems, and wanted to share the numbers here because a few of them surprised me. ### Setup - One RTX 5090, 32GB - Ollama 0.32.12, tag `qwen3.8:latest` (Q4_K_M GGUF, 27.3B) - 64K context, temperature 0.2, generation cap 32,768 tokens - Every model loaded alone on the GPU, with all other…”
laxmimerit · Hugging Face · Aug 16, 2026 - 19
“I checked the official GGUF BF16 weights directly from Unsloth: https://huggingface.co/unsloth/Qwen3.8-27B-GGUF What I found is not a “reasoning style” , "system prompt" or "chat template" issue. It is a structural defect in the temporal processing layers. | Tensor | QType | C2 | α | S_b | S_a | |---|---|---|---|---|---| | blk.52.ssm_conv1d.weight | F32 | ✓ | 0.59005 | 0.0016 | 0.0006 | | blk.53.ssm_conv1d.weight | F32 | ✓ | 0.55484 | 0.0015 | 0.0005 | | blk.56.ssm_conv1d.weight | F32 | ✓ | 0.5…”
LuffyTheFox · Hugging Face · Aug 15, 2026 - 20
“## The problem Qwen3.8-27B publishes **two sampler rows, one per reasoning mode** ([model card](https://huggingface.co/Qwen/Qwen3.8-27B)): | | temp | top_p | top_k | min_p | presence | repetition | |---|--:|--:|--:|--:|--:|--:| | Instruct (what we ship) | 0.7 | 0.80 | 20 | 0.0 | **1.5** | 1.0 | | Thinking | 1.0 | 0.95 | 20 | 0.0 | **0.0** | 1.0 | `ENABLE_THINKING=true` flips the reasoning mode **but not the sampler**. Verified by rendering the compose with the flag flipped — the emitted args ar…”
noonghunna · GitHub · Aug 16, 2026 - 21
“## What `src/analyze.ts` — the single network call in the entire tool. Takes a `CapturedState`, sends it to `claude-sonnet-4-6`, and returns an `Analysis`: `summary`, `hypothesis`, `ruled_out[]`, `working_set[]`, `next_step`. ## Why This is the point of the project. Everything else is plumbing around one question: *why* were those files open? **The prompt is the product.** It explicitly instructs the model to reason backwards from evidence to intent, and tells it that a mechanical description o…”
kishuxz · GitHub · Aug 17, 2026 - 22
“Scores 84.77% on LiveBench Reasoning (#27 of 51), an objective, ground-truth-scored evaluation refreshed with new questions.”
LiveBench Reasoning · Benchmark · Jun 25, 2026 - 23
“## 概要 MJ Prompt Studioの実API利用時に行うすべてのOpenAI Responses API呼び出しを、次の固定構成へ全面移行する。 現在利用可能な`gpt-5.4-mini` / `gpt-5.4-nano`、保存済みの旧モデル設定、機能別のモデル・推論強度選択、環境変数によるモデル上書き、旧モデルへのフォールバックを廃止する。 AI Brief、語彙補助、Prompt Compiler、Prompt Doctor、Parameter Advisor、Reference Analysis、Matrix Lab、Result Review、Final Audit、接続テストを含む全経路で、GPT-5.6 Luna Highだけを使用する。 --- ## 決定事項 - 採用モデルIDは`gpt-5.6-luna`とする。 - `gpt-5.6`エイリアスはSolへ解決されるため使用しない。 - 推論強度は常に`high`を明示する。API既定値へ委ねない。 - `text.verbosity`は常に`low`を明示する。 - 標準モードを使用し、`reasonin…”
stillshore-chirp · GitHub · Jul 31, 2026 - 24
“## End to end proven over the chatgpt-tool proxy Full slow-mode solve of 1.2.1-8 through server2's chatgpt-tool proxy on the Codex lane, free ChatGPT account, `gpt-5.6-luna`. Result: **PASS**, both judges agreeing. | | | | --- | --- | | verdict | PASS | | truth judge | PASS, score 7 | | audit judge | PASS | | evaluation | TRUE, complete, self-contained, human-readable, verifiable | | wall clock | 4m50s | | requests | 7 | | tokens | 29,272 in / 5,376 out, 34,648 total | | actual cost | $0, free…”
tamnd · GitHub · Jul 25, 2026
Rankings synthesized from community evidence and open benchmarks. See methodology. Not driven by vendor marketing.