Recommendation for Agentic coding

Agentic Coding

Our top recommendation for Agentic Coding, based on the public evidence we track, is Anthropic: Claude Fable 5.[1][2][3][4][5][6] Use for autonomous game builds and creative coding projects where it can generate complete applications from minimal prompts, as shown in the Raccoon Heist implementation. Watch out: Watch for silent model fallback where pinned claude-fable-5 sessions switch to claude-opus-5 server-side without recovery, breaking reproducibility. OpenAI: GPT-5.6 Sol is the next-ranked alternative. Select as the default implementation model when your agent framework supports explicit routing tiers, as it is preferred over Luna/Terra for context gathering and extensive changes.

About this recommendation

Updated
Sep 4, 2026
Evidence through
Sep 4, 2026
Sources
24
Revision
v63

Decision audit

Why this result

Inspect the inputs and the computed order behind the recommendation.

Models screened

21

live candidates

Evaluation feeds

8

task-weighted

Winner coverage

100%

intended feed weight

Largest provider share

2 of 7

Anthropic

Established source breadth. 20 citation families and 4 practitioner families support the top result; 1 cautionary thread is retained. The largest citation family contributes 16%.

Sources evaluated

The task sets these weights before any model is scored.

winner: Claude Fable 5
Evaluation feedWeightWinner resultField measured
Terminal-Bench 2.1
22%
#1113/21
SWE-rebench
18%
#114/21
Frontier-Bench
14%
#511/21
LiveBench Agentic Coding
14%
#620/21
Design Arena Full-stack
10%
#514/21
LMArena Agent
9%
#217/21
Design Arena Web-apps
8%
#314/21
OpenRouter usage
5%
87/10021/21

Provider concentration

Each exact model is scored separately; provider identity is not a ranking input.

Anthropic29%
  • Anthropic2 models
  • OpenAI2 models
  • deepseek1 model
  • Moonshot AI1 model
  • Z.ai1 model

Decision table

Every published model is shown in computed order. Practitioner sources are distinct community threads, not the citations repeated in the prose below.

RankModelRelative scoreCoveragePractitioner evidenceStrongest measured reason
01Claude Fable 5Anthropic
79
100%5 threads · 4 families · 1 cautions#1 SWE-rebench · #2 LMArena Agent
02GPT-5.6 SolOpenAI
76
84%21 threads · 14 families · 8 cautions#3 LMArena Agent · #4 Frontier-Bench
03GLM 5.2Z.ai
69
91%20 threads · 8 families · 12 cautions#8 LMArena Agent · #9 SWE-rebench
04GPT-5.6 LunaOpenAI
66
84%21 threads · 16 families · 6 cautions#4 Terminal-Bench 2.1 · #11 Frontier-Bench
05Claude Sonnet 4.6Anthropic
60
86%4 threads · 4 families · 3 cautions#12 Design Arena Full-stack · #13 SWE-rebench
06Kimi K2.6Moonshot AI
59
86%6 threads · 2 families · 0 cautions#8 Design Arena Web-apps · #21 LMArena Agent
07DeepSeek V4 Flash Vision Expdeepseek
53
26%3 threads · 3 families · 1 cautions#3 LiveBench Agentic Coding

Relative score combines normalized benchmark quality and signal coverage; independent practitioner evidence and freshness are bounded tie-breakers. It is an ordering score, not an absolute quality percentage. The writing model receives this order and cannot change it.

  1. Claude Fable 5 ranks #2 on LMArena's agentic arena and #1 on SWE-rebench with 64.5% resolution, demonstrating strong multi-step tool use and repository-wide issue fixing.

    Best when: Use for autonomous game builds and creative coding projects where it can generate complete applications from minimal prompts, as shown in the Raccoon Heist implementation.

    Tips

    • Use for autonomous game builds and creative coding projects where it can generate complete applications from minimal prompts, as shown in the Raccoon Heist implementation.
      Source 1
      <p><strong><a href="https://simonw.github.io/raccoon-heist-codex/">Moonlight &amp; Mayhem (Raccoon Heist by Codex + GPT-5.6 Sol Ultra)</a></strong></p> On Wednesday I wrote about <a href="https://simonwillison.net/2026/Aug/5/raccoon-heist/">One-shotting a Raccoon Heist game using Claude Fable 5</a>, where I had Claude Fable 5 build a full working game from a premise I generated with GPT-3 and DALL-E <a href="https://twitter.com/simonw/status/1555626060384911360">four years ago</a>.</p> <p>I dec…
      Source 2
      <p>Back in 2022 <a href="https://twitter.com/simonw/status/1555626060384911360">I tweeted</a> screenshots of a game concept generated by GPT-3 and some concept "art" created using DALL-E. Today, on the fourth anniversary of that tweet, I decided to see if Claude Fable 5 (running in <a href="https://code.claude.com/docs/en/claude-code-on-the-web">Claude Code for web</a>) could build the entire game from the content of that tweet. It did a pretty good job of it!</p> <p>You can <a href="https://si…

    Watch out for

    • Watch for silent model fallback where pinned `claude-fable-5` sessions switch to `claude-opus-5` server-side without recovery, breaking reproducibility.
      Source 3
      **Bug Description** # Bug report: `claude-fable-5` silently falls back to `claude-opus-5` (13 occurrences, one-way, never recovers) ## Environment | Item | Value | |---|---| | Claude Code | 2.1.220 | | OS | Windows 11 Pro, Build 26200 (x64) | | Node | v24.14.1 | | `settings.json` model | `claude-fable-5[1m]` | | `effortLevel` | `high` | | Concurrent sessions | 8 | ## Summary With `claude-fable-5[1m]` pinned in `settings.json`, sessions are repeatedly switched to `claude-opus-5` by a server-emit…
    • Expect Frontier-Bench performance to trail top-tier models at 34% resolution, suggesting limitations on the hardest diverse agent tasks.
      Source 7
      using Claude Code at max effort, resolves 34.05% ± 1.71 of Frontier-Bench tasks (#5 of 14 public submissions), measuring a submitted model-and-agent configuration on diverse, difficult agent work.
      Frontier-BenchOpen original ↗
  2. GPT-5.6 Sol ranks #3 on LMArena's agentic arena and is explicitly designated as the implementation lane in production agent routing protocols, prioritizing it for autonomous coding execution.

    Best when: Select as the default implementation model when your agent framework supports explicit routing tiers, as it is preferred over Luna/Terra for context gathering and extensive changes.

    Tips

    • Select as the default implementation model when your agent framework supports explicit routing tiers, as it is preferred over Luna/Terra for context gathering and extensive changes.
      Source 4
      it's very fast but it still doesnt come close to 5.6 sol, at least for me, in terms of gathering the context necessary to do extensive changes.
      Source 5
      > 🤖 SPARQ agent\n\nThe current registry routing contract still assigns general implementation to Opus 5 and computes review by inverting the implementer's provider. The maintainer's new default protocol is:\n\n- GPT-5.6 Sol is the implementation lane, including autonomous review fixes.\n- Claude Opus 5 is the review and soundness lane.\n- Capacity handling must remain explicit and fail closed; retired model aliases must stay unreachable.\n- PLAN and CLAIM routing must continue to agree exactly…
      Source 6
      > 🤖 SPARQ agent\n\nAdopt the maintainer's new default model protocol in Sparq's target routing catalog:\n\n- GPT-5.6 Sol leads implementation work.\n- Claude Opus 5 owns review and soundness.\n- Preserve the restricted security/soundness posture and explicit exhaustion behavior.\n- Keep target PLAN routing and registry CLAIM routing in exact agreement.\n\nUpdate the routing table and its regression tests; do not modify protected agent briefs.

    Watch out for

    • Monitor for transport errors and session termination in high-volume deployments, with reports showing 11 incidents of fetch failed/terminated in a single large session.
      Source 8
      ## 背景 Pi Web Desktop 中 openai-codex(ChatGPT 订阅 OAuth)会话频繁出现「输出一半停止 / fetch failed / terminated」,其他模型(aliyun/deepseek)偶发但次数少。 ## 实锤数据(2026-08-01,scripts/session-stops.mjs 修复后全量扫描 185 会话) - transport-error 共 37 条:openai-codex/GPT 19(gpt-5.6-sol×11 + gpt-5.6-terra×8)、aliyun 9、pi-router 8、deepseek 1。 - 最大异常会话:311 消息 / ~2.2M tokens,会话开头连续 6 次 fetch failed/terminated。 - compaction 仅 1 次且在非中断点,**不是主因**;主因是 transport error。 ## 工具修复(commit bfe020d) session-stops.mjs 原判定只在独立 type=error 事件时识别 transport-er…
    • Account for opaque automatic model selection in some Codex configurations where Sol may be retried after Lite models fail, complicating cost and latency prediction.
      Source 9
      ### Description Codex web search has a model override, `PI_CODEX_WEB_SEARCH_MODEL`, but no equivalent setting in `config.yml`. Automatic selection is also opaque: OMP tries `gpt-5.6-luna`, `gpt-5.6-terra`, `gpt-5.6-sol`, then `gpt-5.5` and older candidates. It retries when a Lite model returns no `web_search_call` or Codex rejects the model, but the result only reports the final model. Reasoning effort has the same problem. The hosted search request does not set an effort. Responses-Lite reques…
  3. GLM 5.2 ranks #5 on SWE-rebench with 57% resolution and #9 on LMArena's agentic arena, offering competitive open-weight performance for repository-level issue resolution.

    Best when: Use as a cost-effective open-weight alternative for implementation work after initial context discovery with larger closed models, reducing token spend on execution phases.

    Tips

    • Use as a cost-effective open-weight alternative for implementation work after initial context discovery with larger closed models, reducing token spend on execution phases.
      Source 10
      I have tested this exact scenario, and it works. Opus 5 had access to IDA over MCP, and I simply asked it HOW certain things were done in the target binary. Purely informational, educational, discovery, it was very helpful creating context documents. Then I took those over to GLM-5.2 to actually accomplish something.

    Watch out for

    • Watch for tool-call formatting loops where the model plans correct commands in thinking but emits truncated or malformed tool calls with leading artifacts like `cd && `.
      Source 11
      ## Environment - **Model:** glm-5.2 (GLM Coding Plan, Anthropic-compatible endpoint) - **Harness:** Claude Code CLI v2.1.215 (`ANTHROPIC_BASE_URL` pointed at Z.ai) - **Context:** agentic coding session, Japanese-language context, ~450 assistant turns at the time of failure - **Date observed:** 2026-07-21 (JST) ## Summary During error recovery, GLM-5.2 entered a loop where it **planned the correct Bash command verbatim inside `thinking`, then emitted a tool call with the leading `cd && ` segment…
    • Expect significantly lower Frontier-Bench performance at 4.6% resolution, indicating limitations on complex multi-step agent tasks.
      Source 12
      using Claude Code at max effort, resolves 4.59% ± 0.97 of Frontier-Bench tasks (#14 of 14 public submissions), measuring a submitted model-and-agent configuration on diverse, difficult agent work.
      Frontier-BenchOpen original ↗
    • Verify clinerules compliance manually, as the model may ignore explicit instruction files for tasks like commit message formatting.
      Source 13
      ### Cline Surface VSCode Extension ### Cline Version 4.1.2 ### Beta version - [ ] I am using a beta version of Cline ### What happened? 用户在clinerules目录创建了commit-messages.md规则文件,明确要求Git提交信息使用中文撰写,包括subject和body。但Cline生成的提交信息仍为英文,未遵循规则。 ### Steps to reproduce 1.创建clinerules/commit-messages.md文件 2.写入中文提交信息规则 3.请求Cline生成提交信息 4.观察到生成的提交信息为英文 ### Provider/Model cline / cline-free/glm-5.2 ### IDE / CLI Diagnostics _No response_ ### System Information Visual Studio Code: 1.131.0, Node.js: v24.18.0, Arc…
  4. GPT-5.6 Luna is positioned as a volume-tier model in OpenAI'sResponses API lineup, with automatic fallback ordering behind Sol in production routing systems.

    Best when: Use for high-throughput control-plane and architecture work when cost optimization outweighs reasoning depth, accepting the tradeoff of routing simpler tasks through the volume tier.

    Tips

    • Use for high-throughput control-plane and architecture work when cost optimization outweighs reasoning depth, accepting the tradeoff of routing simpler tasks through the volume tier.
      Source 14
      The live Symphony workflow currently hardcodes `/home/timwhite/.local/bin/codex-rotate --config model="gpt-5.6-luna" app-server` for every issue. The checked-in Gem model registry supports mechanical, code, tests, review, semantic, architecture, and root-cause capabilities plus Luna/Terra/Sol routes, but Symphony does not consume any per-issue classification or routing receipt. This sends control-plane and architecture work such as the fleet-gate freshness repair through the volume model and pr…

    Watch out for

    • Avoid for agentic tasks requiring structured JSON outputs via the Responses API, as the legacy `response_format` parameter causes HTTP 400 errors requiring migration to `text.format`.
      Source 15
      ## Summary The Responses API used by OpenCode Go for `gpt-5.6-luna`, `grok-4.5`, and `muse-spark-1.2-contributor` rejects the legacy `response_format` field with HTTP 400. The error message directs callers to the new nested location: > Unsupported parameter: 'response_format'. In the Responses API, this parameter has moved to 'text.format'. Try again with the new parameter. `OpenAICompatProvider` in `src/llm/openai_compat.rs` currently serializes the old shape for JSON-required roles, so every…
    • Watch for session migration failures when switching from Luna to external models, as compaction state dialect mismatches can crash EMP middleware.
      Source 16
      ## Summary EMP 0.6.0 fails before the upstream request when a ChatGPT/Codex task containing native compaction state switches from gpt-5.6-luna to OR/stealth/ox-alpha. This is related to, but distinct from, #3. Issue #3 covers the reverse external-to-native history path and GLM compatibility. ## Environment - EMP: 0.6.0, commit fa70e6a - Client: ChatGPT desktop task using EMP - Source dialect: codex_native - Destination dialect: portable_responses No credentials, prompts, session identifiers, lo…
      cnsunfisheggOpen original ↗
  5. Claude Sonnet 4.6 ranks #20 on Design Arena and #22 on LMArena's agentic arena, showing weaker preference scores than Fable and Opus variants while maintaining utility for structured planning tasks.

    Best when: Use for specification refinement and implementation planning, where iterative review of edge cases and step-by-step code generation with automatic tests proves effective.

    Tips

    • Use for specification refinement and implementation planning, where iterative review of edge cases and step-by-step code generation with automatic tests proves effective.
      Source 17
      I'm about to complete a new non trivial functionality in a project of a costumer of mine. I spent an hour writing the spec. Then I asked Claude (Sonnet 4.6) to check if I missed something. I did, the sort of minor issues one notice after starting writing code, edge cases etc. That made me think about more issues and after a few iterations we settled down on a spec. I asked Claude to make an implementation plan and we ended up with 9 steps. It wrote the code for a step with new automatic tests a…

    Watch out for

    • Watch for device routing regressions in generated configurations, as seen in FlagGems operator boxing where ~150 routes incorrectly flipped from `cuda` to `flagos_python`.
      Source 18
      ## Issue Type - [x] Bug Report ## AI Agent Information - **Agent**: Claude Code CLI - **Model**: Claude Sonnet 4.6 - **Session Context**: Code review of #150/#151 surfaced regressions introduced by regenerating the FlagGems configs against flag_gems 7fb49bad: ~150 routes flipped from `cuda` to `flagos_python`, and a measured subset of them fails on the flagos device. ## Summary Regenerating the FlagGems configs against flag_gems 7fb49bad (PR #150) flipped ~150 operators from the `cuda` boxing r…
    • Expect 422 validation errors from GitLab Duo on first tool calls, requiring provider-specific workarounds.
      Source 19
      ### OmniRoute Version 3.8.49 ### Installation Method npm (global) ### Operating System Linux ### OS Version Ubuntu 22.04.5 LTS (kernel 5.15.0-1106-nvidia) ### Node.js Version 24.18.0 ### Provider(s) Involved GitLab Duo ### Model(s) Involved gitlab-duo/claude-sonnet-4-6, gitlab-duo/claude-haiku-4-5, gitlab-duo/claude-sonnet-4-6-high ### Client Tool Claude Code v2.1.220 (launched via omniroute launch --profile) ### Description GitLab Duo returns 422 {"detail":"Validation error"} on the very first…
  6. Kimi K2.6 ranks #8 on Design Arena's web-app agent category and is positioned as a top open-weight option for one-shot coding reasoning, though with noted struggles in longer agentic contexts.

    Best when: Consider only after reviewing the cited caution.

    Watch out for

    • Avoid for long-horizon agentic tasks requiring extended context maintenance, as open-weight models typically struggle with longer contexts in agentic evaluations.
      Source 20
      Early benchmarks show tremendous improvement over Kimi K2 Thinking, which didn't perform well on our benchmarks (and we do use best available quantization). Kimi K2.6 is currently the top open weights model in one-shot coding reasoning, a little better than GLM 5.1, and still a strong contender against SOTA models from ~3 months ago (comparable to Gemini 3.1 Pro Preview). Agentic tests are still running, check back tomorrow. Open weights models typically struggle with longer contexts in agentic…
    • Question benchmark results showing K2.6 near the bottom relative to smaller models like Ornith 35B, as community reports suggest these rankings may not reflect practical utility.
      Source 21
      That benchmark ranks Kimi K2.6 and K2.7 Code near the bottom. Both are below Ornith 35B. It ranks Gemma 4 26B much higher than GLM-5.2. The results don't make much sense.
      juliangoldsmithOpen original ↗
  7. DeepSeek V4 Flash Vision Exp scores 65.1% on LiveBench Agentic Coding (#3 of 51), placing it competitively for executable JavaScript, TypeScript, and Python tasks.

    Best when: Consider for cost-sensitive agentic coding where LiveBench results indicate strong performance on executable code tasks.

    Tips

    • Consider for cost-sensitive agentic coding where LiveBench results indicate strong performance on executable code tasks.
      Source 22
      Scores 65.1% on LiveBench Agentic Coding (#3 of 51), covering executable JavaScript, TypeScript, and Python tasks.
      LiveBench Agentic CodingOpen original ↗

    Watch out for

    • Avoid as an auxiliary client in Responses-only gateway configurations, as Chat-Completions-only models cannot serve in that role without reverse translation support.
      Source 23
      ## Bug Description The auxiliary client only speaks the Responses surface, so Chat-Completions-only models cannot serve as auxiliaries even when the gateway serves them correctly. There is no Responses→Chat reverse translation for that direction (the forward Chat→Responses translation exists). ## Steps to Reproduce 1. Configure an auxiliary task with a Chat-only model (e.g. `opencode/deepseek-v4-flash-vision-exp`, verified Chat-served by the gateway). 2. Trigger the auxiliary task. 3. The reque…
    • Configure explicit maxTokens limits (65536 or below) to prevent 400 errors from hidden model mapping where the vision variant routes to a lower-limit endpoint.
      Source 24
      复跑确认了一个更窄的事实:未配置 `maxTokens` 的 `cpa/deepseek-v4-flash-vision-exp` 会在 `packages/coding-agent/src/config/custom-models.ts:112-136` 按同名 catalog reference 继承 `384000`,随后 `packages/ai/src/providers/anthropic.ts:3429-3442` 原样发送 `max_tokens: 384000`;记录的最小复现中,继承值和 wire 值均为 `384000`。但 400 同时表明 CPA 把请求 ID `deepseek-v4-flash-vision-exp` 映射到了另一个实际模型 `deepseek-v4-flash:0731`,后者的服务端上限是 `65536`;该隐藏映射及其上限不在 omp 的模型元数据中,客户端无法推导。`models.yml` 的 `maxTokens` 正是自定义 endpoint 声明该实际上限的入口;为该模型配置 `maxTokens: 65536`(保守地也可…

Frequently asked

What is the top-ranked model for Agentic Coding?
Anthropic: Claude Fable 5 ranks first in the current evidence-weighted comparison. Use for autonomous game builds and creative coding projects where it can generate complete applications from minimal prompts, as shown in the Raccoon Heist implementation.[1][2]
What should I watch out for with Anthropic: Claude Fable 5?
Watch for silent model fallback where pinned `claude-fable-5` sessions switch to `claude-opus-5` server-side without recovery, breaking reproducibility.[3]
What is an alternative to Anthropic: Claude Fable 5?
OpenAI: GPT-5.6 Sol is the next-ranked option. Select as the default implementation model when your agent framework supports explicit routing tiers, as it is preferred over Luna/Terra for context gathering and extensive changes.[4][5][6]

Sources

  1. 1

    <p><strong><a href="https://simonw.github.io/raccoon-heist-codex/">Moonlight &amp; Mayhem (Raccoon Heist by Codex + GPT-5.6 Sol Ultra)</a></strong></p> On Wednesday I wrote about <a href="https://simonwillison.net/2026/Aug/5/raccoon-heist/">One-shotting a Raccoon Heist game using Claude Fable 5</a>, where I had Claude Fable 5 build a full working game from a premise I generated with GPT-3 and DALL-E <a href="https://twitter.com/simonw/status/1555626060384911360">four years ago</a>.</p> <p>I dec…

    Engineering publication · Aug 7, 2026
  2. 2

    <p>Back in 2022 <a href="https://twitter.com/simonw/status/1555626060384911360">I tweeted</a> screenshots of a game concept generated by GPT-3 and some concept "art" created using DALL-E. Today, on the fourth anniversary of that tweet, I decided to see if Claude Fable 5 (running in <a href="https://code.claude.com/docs/en/claude-code-on-the-web">Claude Code for web</a>) could build the entire game from the content of that tweet. It did a pretty good job of it!</p> <p>You can <a href="https://si…

    Engineering publication · Aug 5, 2026
  3. 3

    **Bug Description** # Bug report: `claude-fable-5` silently falls back to `claude-opus-5` (13 occurrences, one-way, never recovers) ## Environment | Item | Value | |---|---| | Claude Code | 2.1.220 | | OS | Windows 11 Pro, Build 26200 (x64) | | Node | v24.14.1 | | `settings.json` model | `claude-fable-5[1m]` | | `effortLevel` | `high` | | Concurrent sessions | 8 | ## Summary With `claude-fable-5[1m]` pinned in `settings.json`, sessions are repeatedly switched to `claude-opus-5` by a server-emit…

    XrentX · GitHub · Aug 2, 2026
  4. 4

    it's very fast but it still doesnt come close to 5.6 sol, at least for me, in terms of gathering the context necessary to do extensive changes.

    prodigycorp · Hacker News · Sep 2, 2026
  5. 5

    > 🤖 SPARQ agent\n\nThe current registry routing contract still assigns general implementation to Opus 5 and computes review by inverting the implementer's provider. The maintainer's new default protocol is:\n\n- GPT-5.6 Sol is the implementation lane, including autonomous review fixes.\n- Claude Opus 5 is the review and soundness lane.\n- Capacity handling must remain explicit and fail closed; retired model aliases must stay unreachable.\n- PLAN and CLAIM routing must continue to agree exactly…

    jeswr · GitHub · Sep 2, 2026
  6. 6

    > 🤖 SPARQ agent\n\nAdopt the maintainer's new default model protocol in Sparq's target routing catalog:\n\n- GPT-5.6 Sol leads implementation work.\n- Claude Opus 5 owns review and soundness.\n- Preserve the restricted security/soundness posture and explicit exhaustion behavior.\n- Keep target PLAN routing and registry CLAIM routing in exact agreement.\n\nUpdate the routing table and its regression tests; do not modify protected agent briefs.

    jeswr · GitHub · Sep 2, 2026
  7. 7

    using Claude Code at max effort, resolves 34.05% ± 1.71 of Frontier-Bench tasks (#5 of 14 public submissions), measuring a submitted model-and-agent configuration on diverse, difficult agent work.

    Frontier-Bench · Benchmark · Jul 13, 2026
  8. 8

    ## 背景 Pi Web Desktop 中 openai-codex(ChatGPT 订阅 OAuth)会话频繁出现「输出一半停止 / fetch failed / terminated」,其他模型(aliyun/deepseek)偶发但次数少。 ## 实锤数据(2026-08-01,scripts/session-stops.mjs 修复后全量扫描 185 会话) - transport-error 共 37 条:openai-codex/GPT 19(gpt-5.6-sol×11 + gpt-5.6-terra×8)、aliyun 9、pi-router 8、deepseek 1。 - 最大异常会话:311 消息 / ~2.2M tokens,会话开头连续 6 次 fetch failed/terminated。 - compaction 仅 1 次且在非中断点,**不是主因**;主因是 transport error。 ## 工具修复(commit bfe020d) session-stops.mjs 原判定只在独立 type=error 事件时识别 transport-er…

    dust617 · GitHub · Aug 1, 2026
  9. 9

    ### Description Codex web search has a model override, `PI_CODEX_WEB_SEARCH_MODEL`, but no equivalent setting in `config.yml`. Automatic selection is also opaque: OMP tries `gpt-5.6-luna`, `gpt-5.6-terra`, `gpt-5.6-sol`, then `gpt-5.5` and older candidates. It retries when a Lite model returns no `web_search_call` or Codex rejects the model, but the result only reports the final model. Reasoning effort has the same problem. The hosted search request does not set an effort. Responses-Lite reques…

    mixmav · GitHub · Aug 2, 2026
  10. 10

    I have tested this exact scenario, and it works. Opus 5 had access to IDA over MCP, and I simply asked it HOW certain things were done in the target binary. Purely informational, educational, discovery, it was very helpful creating context documents. Then I took those over to GLM-5.2 to actually accomplish something.

    gfosco · Hacker News · Jul 28, 2026
  11. 11

    ## Environment - **Model:** glm-5.2 (GLM Coding Plan, Anthropic-compatible endpoint) - **Harness:** Claude Code CLI v2.1.215 (`ANTHROPIC_BASE_URL` pointed at Z.ai) - **Context:** agentic coding session, Japanese-language context, ~450 assistant turns at the time of failure - **Date observed:** 2026-07-21 (JST) ## Summary During error recovery, GLM-5.2 entered a loop where it **planned the correct Bash command verbatim inside `thinking`, then emitted a tool call with the leading `cd && ` segment…

    carrotRakko · GitHub · Jul 21, 2026
  12. 12

    using Claude Code at max effort, resolves 4.59% ± 0.97 of Frontier-Bench tasks (#14 of 14 public submissions), measuring a submitted model-and-agent configuration on diverse, difficult agent work.

    Frontier-Bench · Benchmark · Jul 19, 2026
  13. 13

    ### Cline Surface VSCode Extension ### Cline Version 4.1.2 ### Beta version - [ ] I am using a beta version of Cline ### What happened? 用户在clinerules目录创建了commit-messages.md规则文件,明确要求Git提交信息使用中文撰写,包括subject和body。但Cline生成的提交信息仍为英文,未遵循规则。 ### Steps to reproduce 1.创建clinerules/commit-messages.md文件 2.写入中文提交信息规则 3.请求Cline生成提交信息 4.观察到生成的提交信息为英文 ### Provider/Model cline / cline-free/glm-5.2 ### IDE / CLI Diagnostics _No response_ ### System Information Visual Studio Code: 1.131.0, Node.js: v24.18.0, Arc…

    wangbaic · GitHub · Jul 31, 2026
  14. 14

    The live Symphony workflow currently hardcodes `/home/timwhite/.local/bin/codex-rotate --config model="gpt-5.6-luna" app-server` for every issue. The checked-in Gem model registry supports mechanical, code, tests, review, semantic, architecture, and root-cause capabilities plus Luna/Terra/Sol routes, but Symphony does not consume any per-issue classification or routing receipt. This sends control-plane and architecture work such as the fleet-gate freshness repair through the volume model and pr…

    itstimwhite · GitHub · Aug 13, 2026
  15. 15

    ## Summary The Responses API used by OpenCode Go for `gpt-5.6-luna`, `grok-4.5`, and `muse-spark-1.2-contributor` rejects the legacy `response_format` field with HTTP 400. The error message directs callers to the new nested location: > Unsupported parameter: 'response_format'. In the Responses API, this parameter has moved to 'text.format'. Try again with the new parameter. `OpenAICompatProvider` in `src/llm/openai_compat.rs` currently serializes the old shape for JSON-required roles, so every…

    airvzxf · GitHub · Aug 25, 2026
  16. 16

    ## Summary EMP 0.6.0 fails before the upstream request when a ChatGPT/Codex task containing native compaction state switches from gpt-5.6-luna to OR/stealth/ox-alpha. This is related to, but distinct from, #3. Issue #3 covers the reverse external-to-native history path and GLM compatibility. ## Environment - EMP: 0.6.0, commit fa70e6a - Client: ChatGPT desktop task using EMP - Source dialect: codex_native - Destination dialect: portable_responses No credentials, prompts, session identifiers, lo…

    cnsunfishegg · GitHub · Aug 23, 2026
  17. 17

    I'm about to complete a new non trivial functionality in a project of a costumer of mine. I spent an hour writing the spec. Then I asked Claude (Sonnet 4.6) to check if I missed something. I did, the sort of minor issues one notice after starting writing code, edge cases etc. That made me think about more issues and after a few iterations we settled down on a spec. I asked Claude to make an implementation plan and we ended up with 9 steps. It wrote the code for a step with new automatic tests a…

    pmontra · Hacker News · Jun 6, 2026
  18. 18

    ## Issue Type - [x] Bug Report ## AI Agent Information - **Agent**: Claude Code CLI - **Model**: Claude Sonnet 4.6 - **Session Context**: Code review of #150/#151 surfaced regressions introduced by regenerating the FlagGems configs against flag_gems 7fb49bad: ~150 routes flipped from `cuda` to `flagos_python`, and a measured subset of them fails on the flagos device. ## Summary Regenerating the FlagGems configs against flag_gems 7fb49bad (PR #150) flipped ~150 operators from the `cuda` boxing r…

    Postroggy · GitHub · Aug 24, 2026
  19. 19

    ### OmniRoute Version 3.8.49 ### Installation Method npm (global) ### Operating System Linux ### OS Version Ubuntu 22.04.5 LTS (kernel 5.15.0-1106-nvidia) ### Node.js Version 24.18.0 ### Provider(s) Involved GitLab Duo ### Model(s) Involved gitlab-duo/claude-sonnet-4-6, gitlab-duo/claude-haiku-4-5, gitlab-duo/claude-sonnet-4-6-high ### Client Tool Claude Code v2.1.220 (launched via omniroute launch --profile) ### Description GitLab Duo returns 422 {"detail":"Validation error"} on the very first…

    jayparmar88 · GitHub · Jul 31, 2026
  20. 20

    Early benchmarks show tremendous improvement over Kimi K2 Thinking, which didn't perform well on our benchmarks (and we do use best available quantization). Kimi K2.6 is currently the top open weights model in one-shot coding reasoning, a little better than GLM 5.1, and still a strong contender against SOTA models from ~3 months ago (comparable to Gemini 3.1 Pro Preview). Agentic tests are still running, check back tomorrow. Open weights models typically struggle with longer contexts in agentic…

    gertlabs · Hacker News · Apr 20, 2026
  21. 21

    That benchmark ranks Kimi K2.6 and K2.7 Code near the bottom. Both are below Ornith 35B. It ranks Gemma 4 26B much higher than GLM-5.2. The results don't make much sense.

    juliangoldsmith · Hacker News · Jun 29, 2026
  22. 22

    Scores 65.1% on LiveBench Agentic Coding (#3 of 51), covering executable JavaScript, TypeScript, and Python tasks.

    LiveBench Agentic Coding · Benchmark · Jun 25, 2026
  23. 23

    ## Bug Description The auxiliary client only speaks the Responses surface, so Chat-Completions-only models cannot serve as auxiliaries even when the gateway serves them correctly. There is no Responses→Chat reverse translation for that direction (the forward Chat→Responses translation exists). ## Steps to Reproduce 1. Configure an auxiliary task with a Chat-only model (e.g. `opencode/deepseek-v4-flash-vision-exp`, verified Chat-served by the gateway). 2. Trigger the auxiliary task. 3. The reque…

    DarkArty07 · GitHub · Sep 4, 2026
  24. 24

    复跑确认了一个更窄的事实:未配置 `maxTokens` 的 `cpa/deepseek-v4-flash-vision-exp` 会在 `packages/coding-agent/src/config/custom-models.ts:112-136` 按同名 catalog reference 继承 `384000`,随后 `packages/ai/src/providers/anthropic.ts:3429-3442` 原样发送 `max_tokens: 384000`;记录的最小复现中,继承值和 wire 值均为 `384000`。但 400 同时表明 CPA 把请求 ID `deepseek-v4-flash-vision-exp` 映射到了另一个实际模型 `deepseek-v4-flash:0731`,后者的服务端上限是 `65536`;该隐藏映射及其上限不在 omp 的模型元数据中,客户端无法推导。`models.yml` 的 `maxTokens` 正是自定义 endpoint 声明该实际上限的入口;为该模型配置 `maxTokens: 65536`(保守地也可…

    roboomp · GitHub · Aug 25, 2026

Rankings synthesized from community evidence and open benchmarks. See methodology. Not driven by vendor marketing.