Recommendation for Goose
Goose
The best LLM for Goose in this evidence-weighted comparison is Anthropic: Claude Fable 5.[1][2] Reach for Anthropic: Claude Fable 5 as a default Goose model: its #1 agentic-arena standing tracks tool calling and multi-step control across a long task, which is what this workload leans on. OpenAI: GPT-5.6 Sol is the next-ranked alternative. Reach for OpenAI: GPT-5.6 Sol as a default Goose model: its #2 agentic-arena standing tracks tool calling and multi-step control across a long task, which is what this workload leans on.
About this recommendation
- Updated
- Jul 24, 2026
- Evidence through
- Jul 24, 2026
- Sources
- 6
- Revision
- v1
Anthropic: Claude Fable 5 ranks #1 of 32 on LMArena's agentic arena (score 12.7), measuring tool-use and multi-step task performance from human preference.
Best when: Reach for Anthropic: Claude Fable 5 as a default Goose model: its #1 agentic-arena standing tracks tool calling and multi-step control across a long task, which is what this workload leans on.
Tips
- Reach for Anthropic: Claude Fable 5 as a default Goose model: its #1 agentic-arena standing tracks tool calling and multi-step control across a long task, which is what this workload leans on.
OpenAI: GPT-5.6 Sol ranks #2 of 32 on LMArena's agentic arena (score 10.1), measuring tool-use and multi-step task performance from human preference.
Best when: Reach for OpenAI: GPT-5.6 Sol as a default Goose model: its #2 agentic-arena standing tracks tool calling and multi-step control across a long task, which is what this workload leans on.
Tips
- Reach for OpenAI: GPT-5.6 Sol as a default Goose model: its #2 agentic-arena standing tracks tool calling and multi-step control across a long task, which is what this workload leans on.
MoonshotAI: Kimi K3 ranks #4 of 32 on LMArena's agentic arena (score 9.7), measuring tool-use and multi-step task performance from human preference.
Best when: Reach for MoonshotAI: Kimi K3 as a default Goose model: its #3 agentic-arena standing tracks tool calling and multi-step control across a long task, which is what this workload leans on.
Tips
- Reach for MoonshotAI: Kimi K3 as a default Goose model: its #3 agentic-arena standing tracks tool calling and multi-step control across a long task, which is what this workload leans on.
Anthropic: Claude Opus 4.8 ranks #3 of 32 on LMArena's agentic arena (score 9.7), measuring tool-use and multi-step task performance from human preference.
Best when: Reach for Anthropic: Claude Opus 4.8 as a default Goose model: its #3 agentic-arena standing tracks tool calling and multi-step control across a long task, which is what this workload leans on.
Tips
- Reach for Anthropic: Claude Opus 4.8 as a default Goose model: its #3 agentic-arena standing tracks tool calling and multi-step control across a long task, which is what this workload leans on.
Anthropic: Claude Sonnet 5 ranks #5 of 32 on LMArena's agentic arena (score 8.7), measuring tool-use and multi-step task performance from human preference.
Best when: Reach for Anthropic: Claude Sonnet 5 as a default Goose model: its #5 agentic-arena standing tracks tool calling and multi-step control across a long task, which is what this workload leans on.
Tips
- Reach for Anthropic: Claude Sonnet 5 as a default Goose model: its #5 agentic-arena standing tracks tool calling and multi-step control across a long task, which is what this workload leans on.
Anthropic: Claude Opus 4.7 ranks #7 of 32 on LMArena's agentic arena (score 7.9), measuring tool-use and multi-step task performance from human preference.
Best when: Reach for Anthropic: Claude Opus 4.7 as a default Goose model: its #6 agentic-arena standing tracks tool calling and multi-step control across a long task, which is what this workload leans on.
Tips
- Reach for Anthropic: Claude Opus 4.7 as a default Goose model: its #6 agentic-arena standing tracks tool calling and multi-step control across a long task, which is what this workload leans on.
Frequently asked
- What is the top-ranked model for Goose?
- Anthropic: Claude Fable 5 ranks first in the current evidence-weighted comparison. Reach for Anthropic: Claude Fable 5 as a default Goose model: its #1 agentic-arena standing tracks tool calling and multi-step control across a long task, which is what this workload leans on.[1]
- What is an alternative to Anthropic: Claude Fable 5?
- OpenAI: GPT-5.6 Sol is the next-ranked option. Reach for OpenAI: GPT-5.6 Sol as a default Goose model: its #2 agentic-arena standing tracks tool calling and multi-step control across a long task, which is what this workload leans on.[2]
Sources
- 1
“Ranks #1 of 32 on LMArena's agentic arena (score 12.7), measuring tool-use and multi-step task performance from human preference.”
LMArena agentic arena · Benchmark · Jul 21, 2026 - 2
“Ranks #2 of 32 on LMArena's agentic arena (score 10.1), measuring tool-use and multi-step task performance from human preference.”
LMArena agentic arena · Benchmark · Jul 21, 2026 - 3
“Ranks #4 of 32 on LMArena's agentic arena (score 9.7), measuring tool-use and multi-step task performance from human preference.”
LMArena agentic arena · Benchmark · Jul 21, 2026 - 4
“Ranks #3 of 32 on LMArena's agentic arena (score 9.7), measuring tool-use and multi-step task performance from human preference.”
LMArena agentic arena · Benchmark · Jul 21, 2026 - 5
“Ranks #5 of 32 on LMArena's agentic arena (score 8.7), measuring tool-use and multi-step task performance from human preference.”
LMArena agentic arena · Benchmark · Jul 21, 2026 - 6
“Ranks #7 of 32 on LMArena's agentic arena (score 7.9), measuring tool-use and multi-step task performance from human preference.”
LMArena agentic arena · Benchmark · Jul 21, 2026
Rankings synthesized from community evidence and open benchmarks. See methodology. Not driven by vendor marketing.