Recommendation for Goose
Goose
Our top recommendation for Goose, based on the public evidence we track, is OpenAI: GPT-5.6 Sol. Anthropic: Claude Fable 5 is the next-ranked alternative.
About this recommendation
- Updated
- Sep 6, 2026
- Evidence through
- Sep 6, 2026
- Sources
- 1
- Revision
- v50
Decision audit
Why this result
Inspect the inputs and the computed order behind the recommendation.
Models screened
20
live candidates
Evaluation feeds
6
task-weighted
Winner coverage
79%
intended feed weight
Largest provider share
2 of 3
Anthropic
Sources evaluated
The task sets these weights before any model is scored.
| Evaluation feed | Weight | Winner result | Field measured |
|---|---|---|---|
| Berkeley Function Calling | 32% | not measured | 3/20 |
| LMArena Agent | 23% | #3 | 17/20 |
| price weight | 15% | 99/100 | 20/20 |
| Terminal-Bench 2.1 | 15% | #14 | 11/20 |
| SWE-rebench | 10% | #4 | 15/20 |
| OpenRouter usage | 5% | 97/100 | 20/20 |
Provider concentration
Each exact model is scored separately; provider identity is not a ranking input.
- Anthropic2 models
- OpenAI1 model
Decision table
Every published model is shown in computed order. Practitioner sources are distinct community threads, not the citations repeated in the prose below.
| Rank | Model | Relative score | Coverage | Practitioner evidence | Strongest measured reason |
|---|---|---|---|---|---|
| 01 | GPT-5.6 SolOpenAI | 73 | 79% | no linked practitioner threads | #3 LMArena Agent · #4 SWE-rebench |
| 02 | Claude Fable 5Anthropic | 72 | 79% | no linked practitioner threads | #1 SWE-rebench · #2 LMArena Agent |
| 03 | Claude Opus 4.6Anthropic | 66 | 73% | no linked practitioner threads | #15 LMArena Agent · #16 SWE-rebench |
Relative score combines normalized benchmark quality and signal coverage; independent practitioner evidence and freshness are bounded tie-breakers. It is an ordering score, not an absolute quality percentage. The writing model receives this order and cannot change it.
OpenAI: GPT-5.6 Sol ranks #3 of 47 on LMArena's agentic arena (score 9.3), measuring tool-use and multi-step task performance from human preference.
Best when: Consider only after reviewing the cited caution.
Anthropic: Claude Fable 5 ranks #2 of 47 on LMArena's agentic arena (score 10.2), measuring tool-use and multi-step task performance from human preference.
Best when: Consider only after reviewing the cited caution.
Mid-tier agentic preference ranking with no repository automation benchmark available for comparison.
Best when: Consider only after reviewing the cited caution.
Watch out for
- Lack of SWE-rebench data makes it difficult to assess repository automation suitability relative to benchmarked alternatives.
Sources
- 1
“Ranks #16 of 47 on LMArena's agentic arena (score 4.6), measuring tool-use and multi-step task performance from human preference.”
LMArena agentic arena · Benchmark · Sep 5, 2026
Rankings synthesized from community evidence and open benchmarks. See methodology. Not driven by vendor marketing.