Recommendation for Goose

Goose

Our top recommendation for Goose, based on the public evidence we track, is OpenAI: GPT-5.6 Sol. Anthropic: Claude Fable 5 is the next-ranked alternative.

About this recommendation

Updated
Sep 6, 2026
Evidence through
Sep 6, 2026
Sources
1
Revision
v50

Decision audit

Why this result

Inspect the inputs and the computed order behind the recommendation.

Models screened

20

live candidates

Evaluation feeds

6

task-weighted

Winner coverage

79%

intended feed weight

Largest provider share

2 of 3

Anthropic

Provisional source breadth. 1 citation families and 0 practitioner families support the top result; 0 cautionary threads is retained. The largest citation family contributes 100%.

Sources evaluated

The task sets these weights before any model is scored.

winner: GPT-5.6 Sol
Evaluation feedWeightWinner resultField measured
Berkeley Function Calling
32%
not measured3/20
LMArena Agent
23%
#317/20
price weight
15%
99/10020/20
Terminal-Bench 2.1
15%
#1411/20
SWE-rebench
10%
#415/20
OpenRouter usage
5%
97/10020/20

Provider concentration

Each exact model is scored separately; provider identity is not a ranking input.

Anthropic67%
  • Anthropic2 models
  • OpenAI1 model

Decision table

Every published model is shown in computed order. Practitioner sources are distinct community threads, not the citations repeated in the prose below.

RankModelRelative scoreCoveragePractitioner evidenceStrongest measured reason
01GPT-5.6 SolOpenAI
73
79%no linked practitioner threads#3 LMArena Agent · #4 SWE-rebench
02Claude Fable 5Anthropic
72
79%no linked practitioner threads#1 SWE-rebench · #2 LMArena Agent
03Claude Opus 4.6Anthropic
66
73%no linked practitioner threads#15 LMArena Agent · #16 SWE-rebench

Relative score combines normalized benchmark quality and signal coverage; independent practitioner evidence and freshness are bounded tie-breakers. It is an ordering score, not an absolute quality percentage. The writing model receives this order and cannot change it.

  1. OpenAI: GPT-5.6 Sol ranks #3 of 47 on LMArena's agentic arena (score 9.3), measuring tool-use and multi-step task performance from human preference.

    Best when: Consider only after reviewing the cited caution.

  2. Anthropic: Claude Fable 5 ranks #2 of 47 on LMArena's agentic arena (score 10.2), measuring tool-use and multi-step task performance from human preference.

    Best when: Consider only after reviewing the cited caution.

  3. Mid-tier agentic preference ranking with no repository automation benchmark available for comparison.

    Best when: Consider only after reviewing the cited caution.

    Watch out for

    • Lack of SWE-rebench data makes it difficult to assess repository automation suitability relative to benchmarked alternatives.
      Source 1
      Ranks #16 of 47 on LMArena's agentic arena (score 4.6), measuring tool-use and multi-step task performance from human preference.
      LMArena agentic arenaOpen original ↗

Sources

  1. 1

    Ranks #16 of 47 on LMArena's agentic arena (score 4.6), measuring tool-use and multi-step task performance from human preference.

    LMArena agentic arena · Benchmark · Sep 5, 2026

Rankings synthesized from community evidence and open benchmarks. See methodology. Not driven by vendor marketing.