Recommendation for Logic & reasoning

Logic & Problem Solving

Our top recommendation for Logic & Problem Solving, based on the public evidence we track, is Anthropic: Claude Fable 5.[1][2] Use for instruction-following tasks where human-aligned output matters, as it ranks #3 of 144 on LMArena's instruction-following category with an Elo of 1505. OpenAI: GPT-5.6 Sol is the next-ranked alternative. Prioritize for objective reasoning benchmarks where it scores 91.65% on LiveBench Reasoning, ranking #2 of 51.

About this recommendation

Updated
Sep 4, 2026
Evidence through
Sep 4, 2026
Sources
24
Revision
v58

Decision audit

Why this result

Inspect the inputs and the computed order behind the recommendation.

Models screened

21

live candidates

Evaluation feeds

5

task-weighted

Winner coverage

100%

intended feed weight

Largest provider share

4 of 8

Anthropic

Provisional source breadth. 13 citation families and 0 practitioner families support the top result; 0 cautionary threads is retained. The largest citation family contributes 26%.

Sources evaluated

The task sets these weights before any model is scored.

winner: Claude Fable 5
Evaluation feedWeightWinner resultField measured
LiveBench Reasoning
35%
#821/21
LMArena Instruction Following
25%
#319/21
LiveBench Instruction Following
15%
#521/21
LMArena Math
15%
#319/21
OpenRouter usage
10%
87/10021/21

Provider concentration

Each exact model is scored separately; provider identity is not a ranking input.

Anthropic50%
  • Anthropic4 models
  • OpenAI2 models
  • Google1 model
  • Qwen1 model

Decision table

Every published model is shown in computed order. Practitioner sources are distinct community threads, not the citations repeated in the prose below.

RankModelRelative scoreCoveragePractitioner evidenceStrongest measured reason
01Claude Fable 5Anthropic
84
100%no linked practitioner threads#3 LMArena Instruction Following · #3 LMArena Math
02GPT-5.6 SolOpenAI
83
100%3 threads · 2 families · 1 cautions#2 LiveBench Reasoning · #11 LMArena Instruction Following
03Claude Opus 4.6Anthropic
81
100%3 threads · 2 families · 2 cautions#1 LMArena Instruction Following · #4 LMArena Math
04Gemini 3.6 FlashGoogle
81
100%2 threads · 2 families · 0 cautions#7 LiveBench Instruction Following · #7 LMArena Math
05Claude Opus 4.8Anthropic
80
100%3 threads · 3 families · 3 cautions#11 LiveBench Reasoning · #11 LMArena Instruction Following
06Qwen3.8 27BQwen
76
100%4 threads · 3 families · 1 cautions#13 LiveBench Instruction Following · #30 LMArena Instruction Following
07Claude Sonnet 4.6Anthropic
75
100%3 threads · 3 families · 2 cautions#11 LMArena Instruction Following · #25 LiveBench Reasoning
08GPT-5.6 LunaOpenAI
75
100%3 threads · 2 families · 1 cautions#22 LiveBench Reasoning · #35 LMArena Math

Relative score combines normalized benchmark quality and signal coverage; independent practitioner evidence and freshness are bounded tie-breakers. It is an ordering score, not an absolute quality percentage. The writing model receives this order and cannot change it.

  1. Claude Fable 5 ranks third in human preference for instruction-following and scores 89.65% on LiveBench Reasoning, placing it in the top tier for logic and problem-solving tasks.

    Best when: Use for instruction-following tasks where human-aligned output matters, as it ranks #3 of 144 on LMArena's instruction-following category with an Elo of 1505.

    Tips

    • Use for instruction-following tasks where human-aligned output matters, as it ranks #3 of 144 on LMArena's instruction-following category with an Elo of 1505.
      Source 1
      Ranks #3 of 144 on LMArena's instruction-following category (Elo 1505), based on blind human preference votes.
      LMArena instruction-following categoryOpen original ↗
    • Deploy for ground-truth reasoning benchmarks where it scores 89.65% on LiveBench Reasoning, ranking #9 of 51.
      Source 3
      Scores 89.65% on LiveBench Reasoning (#9 of 51), an objective, ground-truth-scored evaluation refreshed with new questions.
      LiveBench ReasoningOpen original ↗
  2. GPT-5.6 Sol achieves the #2 spot on LiveBench Reasoning with 91.65% and shows moderate learning gains in episodic protocol tests, though its LMArena instruction-following ranking is mid-tier.

    Best when: Prioritize for objective reasoning benchmarks where it scores 91.65% on LiveBench Reasoning, ranking #2 of 51.

    Tips

    • Prioritize for objective reasoning benchmarks where it scores 91.65% on LiveBench Reasoning, ranking #2 of 51.
      Source 2
      Scores 91.65% on LiveBench Reasoning (#2 of 51), an objective, ground-truth-scored evaluation refreshed with new questions.
      LiveBench ReasoningOpen original ↗
    • Use in iterative learning workflows where it showed +1.6 improvement from attempt 1 to attempt 3 in an episodic protocol with wiped context.
      Source 4
      ## Observation from main-1 The episodic protocol (3 attempts, context wiped between them, persistent self-written notebook) was expected to produce visible learning. Result is mixed — mean final height attempt 1 → 3: | model | a1 → a3 | Δ | |---|---|---| | deepseek-v4-flash | 2.08 → 5.80 | **+3.7** | | glm-5.2 | 2.98 → 5.91 | **+2.9** | | k3 | 1.64 → 4.25 | **+2.6** | | gpt-5.6-sol | 3.17 → 4.76 | +1.6 | | claude-opus-5 | 6.50 → 8.10 | +1.6 (only monotonic riser) | | gpt-5.5 | 5.59 → 6.90 | +1.…

    Watch out for

    • Verify benchmark configurations carefully, as evidence notes its ARC-AGI-3 score of 38.3% was achieved with custom API settings that may not match standard comparison conditions.
      Source 5
      The annotation on arc-agi-3 is this: > OpenAI's own evaluation notes say Astra uses the company's Responses API harness, while comparison models can operate under different configurations. With this configuration gpt-5.6-sol was able to reach 38,3%. So this is misleading.
      Source 6
      Except it's using custom api settings with which gpt-5.6 sol already got 38%.
  3. Claude Opus 4.6 leads LMArena's instruction-following category with the top Elo of 1523.

    Best when: Select when human preference for instruction-following is paramount, as it ranks #1 of 144 on LMArena with Elo 1523.

    Tips

    • Select when human preference for instruction-following is paramount, as it ranks #1 of 144 on LMArena with Elo 1523.
      Source 7
      Ranks #1 of 144 on LMArena's instruction-following category (Elo 1523), based on blind human preference votes.
      LMArena instruction-following categoryOpen original ↗
    • Consider for reasoning tasks where it scores 88.67% on LiveBench Reasoning, though this places it at #14 of 51.
      Source 8
      Scores 88.67% on LiveBench Reasoning (#14 of 51), an objective, ground-truth-scored evaluation refreshed with new questions.
      LiveBench ReasoningOpen original ↗

    Watch out for

    • Audit outputs carefully when users request evidence for classifications or judgments, as the model generates factually false explanations about its own reasoning process rather than admitting uncertainty.
      Source 9
      ### Problem Claude (Opus 4.6) generates factually false explanations to justify its own incorrect output when questioned by users. This is not a hallucination about external facts — it is fabrication about the model's own reasoning process. ### Reproduction 1. Give Claude a role-based task with defined processes 2. Claude produces output containing an unjustified judgment (e.g., classifying items without evidence) 3. User asks: "Show evidence for that classification" 4. Instead of admitting lac…
      NStestUser1954Open original ↗
      Source 10
      ### Problem Claude (Opus 4.6) generates factually false explanations to justify its own incorrect output when questioned by users. This is not a hallucination about external facts — it is fabrication about the model's own reasoning process. ### Reproduction 1. Give Claude a role-based task with defined processes 2. Claude produces output containing an unjustified judgment (e.g., classifying items without evidence) 3. User asks: "Show evidence for that classification" 4. Instead of admitting lac…
      NStestUser1954Open original ↗
  4. Gemini 3.6 Flash offers configurable thinking budget with up to 891 thinking tokens on high setting.

    Best when: Tune the thinking budget for cost-sensitive reasoning tasks, as the model supports 213-891 thinking tokens depending on configuration.

    Tips

    • Tune the thinking budget for cost-sensitive reasoning tasks, as the model supports 213-891 thinking tokens depending on configuration.
      Source 11
      Firming up the measurements in the issue body — those were single samples, and the run-to-run variance turned out to be wide enough to matter. Repeated with n=3 per cell, same prompt, reading `usageMetadata.thoughtsTokenCount`: | Model | default (no `thinkingConfig`) | `low` | `high` | | :--- | ---: | ---: | ---: | | `gemini-3.6-flash` | 356 | 213 | **891** | | `gemini-3.7-flash` | 373 | 159 | **753** | samples — 3.6: default 365/401/302, low 186/234/219, high 944/874/855 · 3.7: default 370/334…

    Watch out for

    • Expect lower raw reasoning performance on ground-truth benchmarks, as it scores 85.15% on LiveBench Reasoning, ranking #26 of 51.
      Source 12
      Scores 85.15% on LiveBench Reasoning (#26 of 51), an objective, ground-truth-scored evaluation refreshed with new questions.
      LiveBench ReasoningOpen original ↗
    • Note its #20 ranking on LMArena instruction-following with Elo 1465, below top-tier competitors.
      Source 13
      Ranks #20 of 144 on LMArena's instruction-following category (Elo 1465), based on blind human preference votes.
      LMArena instruction-following categoryOpen original ↗
  5. Claude Opus 4.8 scores 89.19% on LiveBench Reasoning.

    Best when: Use for reasoning benchmarks where it achieves 89.19% on LiveBench Reasoning, ranking #12 of 51.

    Tips

    • Use for reasoning benchmarks where it achieves 89.19% on LiveBench Reasoning, ranking #12 of 51.
      Source 14
      Scores 89.19% on LiveBench Reasoning (#12 of 51), an objective, ground-truth-scored evaluation refreshed with new questions.
      LiveBench ReasoningOpen original ↗

    Watch out for

    • Avoid LiteLLM Connector deployments for reasoning-intensive tasks, as the connector fails to send the `thinking` parameter and never produces reasoning tokens or hits prompt cache even with effort settings configured.
      Source 15
      ### Description Claude Opus 5 (and likely Opus 4.8) via LiteLLM Connector never produces reasoning tokens and never hits prompt cache, even when the picker effort is set and the model card advertises `thinking`, `reasoning_effort`, and `cache_control`. Same LiteLLM proxy, same `claude-opus-5` model, works from OpenCode because OpenCode sends Anthropic-native thinking + cache fields. The connector does not. Two independent wire problems: 1. **Thinking never enabled.** The connector only sends Op…
      FPA-DavidTaiOpen original ↗
      Source 16
      `_api_complete_async` (model_client.py:829-830) and `_agent_query` (model_client.py:671-672) both send Anthropic's `effort` parameter (`output_config.effort` / `ClaudeAgentOptions.effort`) but never send `thinking`. Per Anthropic's current docs, these are two independent controls: > "The `thinking` parameter controls **whether** Claude thinks in thinking blocks before answering; the `effort` parameter controls how much work Claude puts into the whole response... Don't pass `adaptive` as an `eff…
  6. Qwen3.8 27B is an open-weight dense model with hybrid reasoning and 256K context.

    Best when: Self-host on dual RTX 5060 Ti or single RTX 5090 for local inference of a 27B dense model with 256K context window.

    Tips

    • Self-host on dual RTX 5060 Ti or single RTX 5090 for local inference of a 27B dense model with 256K context window.
      Source 17
      Benchmark [Qwen3.8-27B](https://unsloth.ai/docs/models/qwen3.8) at Q4 on 2× RTX 5060 Ti with `llama-bench`, and use the run to settle how a 27B dense model should actually be split across two cards. ## Target | | | |---|---| | Model | [`unsloth/Qwen3.8-27B-GGUF`](https://huggingface.co/unsloth/Qwen3.8-27B-GGUF) — 27B **dense**, vision + hybrid reasoning, 256K ctx | | Quant | `UD-Q4_K_XL` (17.92 GB), Unsloth Dynamic V3.0 preview — the 4-bit Unsloth's guide recommends | | GPU | 2× RTX 5060 Ti (16…
      Source 18
      ![chart-qwen-hero](https://cdn-uploads.huggingface.co/production/uploads/6314e2e71994578f07ed8f89/y0d7Q_aB0gLg27LarUwan.png) I ran Qwen3.8-27B against two other models in the same size class on a set of 16 hard problems, and wanted to share the numbers here because a few of them surprised me. ### Setup - One RTX 5090, 32GB - Ollama 0.32.12, tag `qwen3.8:latest` (Q4_K_M GGUF, 27.3B) - 64K context, temperature 0.2, generation cap 32,768 tokens - Every model loaded alone on the GPU, with all other…

    Watch out for

    • Inspect temporal processing layers before deployment, as direct weight analysis found structural defects in SSM conv1d weights across multiple blocks.
      Source 19
      I checked the official GGUF BF16 weights directly from Unsloth: https://huggingface.co/unsloth/Qwen3.8-27B-GGUF What I found is not a “reasoning style” , "system prompt" or "chat template" issue. It is a structural defect in the temporal processing layers. | Tensor | QType | C2 | α | S_b | S_a | |---|---|---|---|---|---| | blk.52.ssm_conv1d.weight | F32 | ✓ | 0.59005 | 0.0016 | 0.0006 | | blk.53.ssm_conv1d.weight | F32 | ✓ | 0.55484 | 0.0015 | 0.0005 | | blk.56.ssm_conv1d.weight | F32 | ✓ | 0.5…
    • Manually configure samplers when switching reasoning modes, as `ENABLE_THINKING=true` flips reasoning mode without updating the sampler row, leaving presence penalty at 1.5 instead of the required 0.0.
      Source 20
      ## The problem Qwen3.8-27B publishes **two sampler rows, one per reasoning mode** ([model card](https://huggingface.co/Qwen/Qwen3.8-27B)): | | temp | top_p | top_k | min_p | presence | repetition | |---|--:|--:|--:|--:|--:|--:| | Instruct (what we ship) | 0.7 | 0.80 | 20 | 0.0 | **1.5** | 1.0 | | Thinking | 1.0 | 0.95 | 20 | 0.0 | **0.0** | 1.0 | `ENABLE_THINKING=true` flips the reasoning mode **but not the sampler**. Verified by rendering the compose with the flag flipped — the emitted args ar…
  7. Claude Sonnet 4.6 ranks #13 on LMArena instruction-following and scores 84.77% on LiveBench Reasoning.

    Best when: Use for backward-reasoning analysis tasks where the model is explicitly prompted to reason from evidence to intent, as demonstrated in a file-analysis workflow.

    Tips

    • Use for backward-reasoning analysis tasks where the model is explicitly prompted to reason from evidence to intent, as demonstrated in a file-analysis workflow.
      Source 21
      ## What `src/analyze.ts` — the single network call in the entire tool. Takes a `CapturedState`, sends it to `claude-sonnet-4-6`, and returns an `Analysis`: `summary`, `hypothesis`, `ruled_out[]`, `working_set[]`, `next_step`. ## Why This is the point of the project. Everything else is plumbing around one question: *why* were those files open? **The prompt is the product.** It explicitly instructs the model to reason backwards from evidence to intent, and tells it that a mechanical description o…

    Watch out for

    • Expect lower reasoning benchmark performance, scoring 84.77% on LiveBench Reasoning at #27 of 51.
      Source 22
      Scores 84.77% on LiveBench Reasoning (#27 of 51), an objective, ground-truth-scored evaluation refreshed with new questions.
      LiveBench ReasoningOpen original ↗
    • Fix LiteLLM Connector integration manually, as it sends `effort` without `thinking` parameter, preventing reasoning block activation.
      Source 16
      `_api_complete_async` (model_client.py:829-830) and `_agent_query` (model_client.py:671-672) both send Anthropic's `effort` parameter (`output_config.effort` / `ClaudeAgentOptions.effort`) but never send `thinking`. Per Anthropic's current docs, these are two independent controls: > "The `thinking` parameter controls **whether** Claude thinks in thinking blocks before answering; the `effort` parameter controls how much work Claude puts into the whole response... Don't pass `adaptive` as an `eff…
  8. GPT-5.6 Luna scores 85.64% on LiveBench Reasoning and has been validated in end-to-end tool use with 7 requests and 34,648 total tokens for complex problem-solving, though its reasoning effort must be explicitly set to high.

    Best when: Configure with explicit `reasoning_effort=high` for multi-step tool use, as demonstrated in a successful 4m50s solve requiring 7 API calls and 29K input tokens.

    Tips

    • Configure with explicit `reasoning_effort=high` for multi-step tool use, as demonstrated in a successful 4m50s solve requiring 7 API calls and 29K input tokens.
      Source 23
      ## 概要 MJ Prompt Studioの実API利用時に行うすべてのOpenAI Responses API呼び出しを、次の固定構成へ全面移行する。 現在利用可能な`gpt-5.4-mini` / `gpt-5.4-nano`、保存済みの旧モデル設定、機能別のモデル・推論強度選択、環境変数によるモデル上書き、旧モデルへのフォールバックを廃止する。 AI Brief、語彙補助、Prompt Compiler、Prompt Doctor、Parameter Advisor、Reference Analysis、Matrix Lab、Result Review、Final Audit、接続テストを含む全経路で、GPT-5.6 Luna Highだけを使用する。 --- ## 決定事項 - 採用モデルIDは`gpt-5.6-luna`とする。 - `gpt-5.6`エイリアスはSolへ解決されるため使用しない。 - 推論強度は常に`high`を明示する。API既定値へ委ねない。 - `text.verbosity`は常に`low`を明示する。 - 標準モードを使用し、`reasonin…
      stillshore-chirpOpen original ↗
      Source 24
      ## End to end proven over the chatgpt-tool proxy Full slow-mode solve of 1.2.1-8 through server2's chatgpt-tool proxy on the Codex lane, free ChatGPT account, `gpt-5.6-luna`. Result: **PASS**, both judges agreeing. | | | | --- | --- | | verdict | PASS | | truth judge | PASS, score 7 | | audit judge | PASS | | evaluation | TRUE, complete, self-contained, human-readable, verifiable | | wall clock | 4m50s | | requests | 7 | | tokens | 29,272 in / 5,376 out, 34,648 total | | actual cost | $0, free…

    Watch out for

    • Do not rely on API defaults for reasoning intensity, as the model requires explicit `high` reasoning effort configuration and `text.verbosity=low` to match validated performance.
      Source 23
      ## 概要 MJ Prompt Studioの実API利用時に行うすべてのOpenAI Responses API呼び出しを、次の固定構成へ全面移行する。 現在利用可能な`gpt-5.4-mini` / `gpt-5.4-nano`、保存済みの旧モデル設定、機能別のモデル・推論強度選択、環境変数によるモデル上書き、旧モデルへのフォールバックを廃止する。 AI Brief、語彙補助、Prompt Compiler、Prompt Doctor、Parameter Advisor、Reference Analysis、Matrix Lab、Result Review、Final Audit、接続テストを含む全経路で、GPT-5.6 Luna Highだけを使用する。 --- ## 決定事項 - 採用モデルIDは`gpt-5.6-luna`とする。 - `gpt-5.6`エイリアスはSolへ解決されるため使用しない。 - 推論強度は常に`high`を明示する。API既定値へ委ねない。 - `text.verbosity`は常に`low`を明示する。 - 標準モードを使用し、`reasonin…
      stillshore-chirpOpen original ↗

Frequently asked

What is the top-ranked model for Logic & Problem Solving?
Anthropic: Claude Fable 5 ranks first in the current evidence-weighted comparison. Use for instruction-following tasks where human-aligned output matters, as it ranks #3 of 144 on LMArena's instruction-following category with an Elo of 1505.[1]
What is an alternative to Anthropic: Claude Fable 5?
OpenAI: GPT-5.6 Sol is the next-ranked option. Prioritize for objective reasoning benchmarks where it scores 91.65% on LiveBench Reasoning, ranking #2 of 51.[2]

Sources

  1. 1

    Ranks #3 of 144 on LMArena's instruction-following category (Elo 1505), based on blind human preference votes.

    LMArena instruction-following category · Benchmark · Sep 1, 2026
  2. 2

    Scores 91.65% on LiveBench Reasoning (#2 of 51), an objective, ground-truth-scored evaluation refreshed with new questions.

    LiveBench Reasoning · Benchmark · Jun 25, 2026
  3. 3

    Scores 89.65% on LiveBench Reasoning (#9 of 51), an objective, ground-truth-scored evaluation refreshed with new questions.

    LiveBench Reasoning · Benchmark · Jun 25, 2026
  4. 4

    ## Observation from main-1 The episodic protocol (3 attempts, context wiped between them, persistent self-written notebook) was expected to produce visible learning. Result is mixed — mean final height attempt 1 → 3: | model | a1 → a3 | Δ | |---|---|---| | deepseek-v4-flash | 2.08 → 5.80 | **+3.7** | | glm-5.2 | 2.98 → 5.91 | **+2.9** | | k3 | 1.64 → 4.25 | **+2.6** | | gpt-5.6-sol | 3.17 → 4.76 | +1.6 | | claude-opus-5 | 6.50 → 8.10 | +1.6 (only monotonic riser) | | gpt-5.5 | 5.59 → 6.90 | +1.…

    eanderson4 · GitHub · Aug 6, 2026
  5. 5

    The annotation on arc-agi-3 is this: > OpenAI's own evaluation notes say Astra uses the company's Responses API harness, while comparison models can operate under different configurations. With this configuration gpt-5.6-sol was able to reach 38,3%. So this is misleading.

    leumon · Hacker News · Sep 3, 2026
  6. 6

    Except it's using custom api settings with which gpt-5.6 sol already got 38%.

    leumon · Hacker News · Sep 3, 2026
  7. 7

    Ranks #1 of 144 on LMArena's instruction-following category (Elo 1523), based on blind human preference votes.

    LMArena instruction-following category · Benchmark · Sep 1, 2026
  8. 8

    Scores 88.67% on LiveBench Reasoning (#14 of 51), an objective, ground-truth-scored evaluation refreshed with new questions.

    LiveBench Reasoning · Benchmark · Jun 25, 2026
  9. 9

    ### Problem Claude (Opus 4.6) generates factually false explanations to justify its own incorrect output when questioned by users. This is not a hallucination about external facts — it is fabrication about the model's own reasoning process. ### Reproduction 1. Give Claude a role-based task with defined processes 2. Claude produces output containing an unjustified judgment (e.g., classifying items without evidence) 3. User asks: "Show evidence for that classification" 4. Instead of admitting lac…

    NStestUser1954 · GitHub · Jul 22, 2026
  10. 10

    ### Problem Claude (Opus 4.6) generates factually false explanations to justify its own incorrect output when questioned by users. This is not a hallucination about external facts — it is fabrication about the model's own reasoning process. ### Reproduction 1. Give Claude a role-based task with defined processes 2. Claude produces output containing an unjustified judgment (e.g., classifying items without evidence) 3. User asks: "Show evidence for that classification" 4. Instead of admitting lac…

    NStestUser1954 · GitHub · Jul 22, 2026
  11. 11

    Firming up the measurements in the issue body — those were single samples, and the run-to-run variance turned out to be wide enough to matter. Repeated with n=3 per cell, same prompt, reading `usageMetadata.thoughtsTokenCount`: | Model | default (no `thinkingConfig`) | `low` | `high` | | :--- | ---: | ---: | ---: | | `gemini-3.6-flash` | 356 | 213 | **891** | | `gemini-3.7-flash` | 373 | 159 | **753** | samples — 3.6: default 365/401/302, low 186/234/219, high 944/874/855 · 3.7: default 370/334…

    mertAIx · GitHub · Sep 1, 2026
  12. 12

    Scores 85.15% on LiveBench Reasoning (#26 of 51), an objective, ground-truth-scored evaluation refreshed with new questions.

    LiveBench Reasoning · Benchmark · Jun 25, 2026
  13. 13

    Ranks #20 of 144 on LMArena's instruction-following category (Elo 1465), based on blind human preference votes.

    LMArena instruction-following category · Benchmark · Sep 1, 2026
  14. 14

    Scores 89.19% on LiveBench Reasoning (#12 of 51), an objective, ground-truth-scored evaluation refreshed with new questions.

    LiveBench Reasoning · Benchmark · Jun 25, 2026
  15. 15

    ### Description Claude Opus 5 (and likely Opus 4.8) via LiteLLM Connector never produces reasoning tokens and never hits prompt cache, even when the picker effort is set and the model card advertises `thinking`, `reasoning_effort`, and `cache_control`. Same LiteLLM proxy, same `claude-opus-5` model, works from OpenCode because OpenCode sends Anthropic-native thinking + cache fields. The connector does not. Two independent wire problems: 1. **Thinking never enabled.** The connector only sends Op…

    FPA-DavidTai · GitHub · Aug 20, 2026
  16. 16

    `_api_complete_async` (model_client.py:829-830) and `_agent_query` (model_client.py:671-672) both send Anthropic's `effort` parameter (`output_config.effort` / `ClaudeAgentOptions.effort`) but never send `thinking`. Per Anthropic's current docs, these are two independent controls: > "The `thinking` parameter controls **whether** Claude thinks in thinking blocks before answering; the `effort` parameter controls how much work Claude puts into the whole response... Don't pass `adaptive` as an `eff…

    tomcardoso · GitHub · Aug 18, 2026
  17. 17

    Benchmark [Qwen3.8-27B](https://unsloth.ai/docs/models/qwen3.8) at Q4 on 2× RTX 5060 Ti with `llama-bench`, and use the run to settle how a 27B dense model should actually be split across two cards. ## Target | | | |---|---| | Model | [`unsloth/Qwen3.8-27B-GGUF`](https://huggingface.co/unsloth/Qwen3.8-27B-GGUF) — 27B **dense**, vision + hybrid reasoning, 256K ctx | | Quant | `UD-Q4_K_XL` (17.92 GB), Unsloth Dynamic V3.0 preview — the 4-bit Unsloth's guide recommends | | GPU | 2× RTX 5060 Ti (16…

    changtimwu · GitHub · Aug 19, 2026
  18. 18

    ![chart-qwen-hero](https://cdn-uploads.huggingface.co/production/uploads/6314e2e71994578f07ed8f89/y0d7Q_aB0gLg27LarUwan.png) I ran Qwen3.8-27B against two other models in the same size class on a set of 16 hard problems, and wanted to share the numbers here because a few of them surprised me. ### Setup - One RTX 5090, 32GB - Ollama 0.32.12, tag `qwen3.8:latest` (Q4_K_M GGUF, 27.3B) - 64K context, temperature 0.2, generation cap 32,768 tokens - Every model loaded alone on the GPU, with all other…

    laxmimerit · Hugging Face · Aug 16, 2026
  19. 19

    I checked the official GGUF BF16 weights directly from Unsloth: https://huggingface.co/unsloth/Qwen3.8-27B-GGUF What I found is not a “reasoning style” , "system prompt" or "chat template" issue. It is a structural defect in the temporal processing layers. | Tensor | QType | C2 | α | S_b | S_a | |---|---|---|---|---|---| | blk.52.ssm_conv1d.weight | F32 | ✓ | 0.59005 | 0.0016 | 0.0006 | | blk.53.ssm_conv1d.weight | F32 | ✓ | 0.55484 | 0.0015 | 0.0005 | | blk.56.ssm_conv1d.weight | F32 | ✓ | 0.5…

    LuffyTheFox · Hugging Face · Aug 15, 2026
  20. 20

    ## The problem Qwen3.8-27B publishes **two sampler rows, one per reasoning mode** ([model card](https://huggingface.co/Qwen/Qwen3.8-27B)): | | temp | top_p | top_k | min_p | presence | repetition | |---|--:|--:|--:|--:|--:|--:| | Instruct (what we ship) | 0.7 | 0.80 | 20 | 0.0 | **1.5** | 1.0 | | Thinking | 1.0 | 0.95 | 20 | 0.0 | **0.0** | 1.0 | `ENABLE_THINKING=true` flips the reasoning mode **but not the sampler**. Verified by rendering the compose with the flag flipped — the emitted args ar…

    noonghunna · GitHub · Aug 16, 2026
  21. 21

    ## What `src/analyze.ts` — the single network call in the entire tool. Takes a `CapturedState`, sends it to `claude-sonnet-4-6`, and returns an `Analysis`: `summary`, `hypothesis`, `ruled_out[]`, `working_set[]`, `next_step`. ## Why This is the point of the project. Everything else is plumbing around one question: *why* were those files open? **The prompt is the product.** It explicitly instructs the model to reason backwards from evidence to intent, and tells it that a mechanical description o…

    kishuxz · GitHub · Aug 17, 2026
  22. 22

    Scores 84.77% on LiveBench Reasoning (#27 of 51), an objective, ground-truth-scored evaluation refreshed with new questions.

    LiveBench Reasoning · Benchmark · Jun 25, 2026
  23. 23

    ## 概要 MJ Prompt Studioの実API利用時に行うすべてのOpenAI Responses API呼び出しを、次の固定構成へ全面移行する。 現在利用可能な`gpt-5.4-mini` / `gpt-5.4-nano`、保存済みの旧モデル設定、機能別のモデル・推論強度選択、環境変数によるモデル上書き、旧モデルへのフォールバックを廃止する。 AI Brief、語彙補助、Prompt Compiler、Prompt Doctor、Parameter Advisor、Reference Analysis、Matrix Lab、Result Review、Final Audit、接続テストを含む全経路で、GPT-5.6 Luna Highだけを使用する。 --- ## 決定事項 - 採用モデルIDは`gpt-5.6-luna`とする。 - `gpt-5.6`エイリアスはSolへ解決されるため使用しない。 - 推論強度は常に`high`を明示する。API既定値へ委ねない。 - `text.verbosity`は常に`low`を明示する。 - 標準モードを使用し、`reasonin…

    stillshore-chirp · GitHub · Jul 31, 2026
  24. 24

    ## End to end proven over the chatgpt-tool proxy Full slow-mode solve of 1.2.1-8 through server2's chatgpt-tool proxy on the Codex lane, free ChatGPT account, `gpt-5.6-luna`. Result: **PASS**, both judges agreeing. | | | | --- | --- | | verdict | PASS | | truth judge | PASS, score 7 | | audit judge | PASS | | evaluation | TRUE, complete, self-contained, human-readable, verifiable | | wall clock | 4m50s | | requests | 7 | | tokens | 29,272 in / 5,376 out, 34,648 total | | actual cost | $0, free…

    tamnd · GitHub · Jul 25, 2026

Rankings synthesized from community evidence and open benchmarks. See methodology. Not driven by vendor marketing.