Recommendation for Agents
AI Agents
The best LLM for AI agents is Claude Fable 5, which holds the top spot on LMArena's agentic arena with a score of 13.9, nearly 5 points ahead of its nearest competitor. Close behind is Claude Opus 4.8 at rank 2 with a 9.3 score, followed by Claude Sonnet 5 at rank 3. These three Anthropic models clearly dominate multi-step planning and tool-use tasks. Claude Fable 5 earns the top recommendation because it substantially outperforms everything else on human-preference evaluations for agentic work. Users report strong task adherence, efficient token usage, and solid subagent orchestration. It does suffer from temporary availability errors under heavy load. For those prioritizing reliability over raw agentic capability, Opus 4.8 demonstrates excellent self-correction on tool call failures and strong benchmark performance on SWE-bench Pro (69.2%) and Terminal-Bench (85.0%). Sonnet 5 provides a cost-conscious alternative with competitive agentic scores, though some users find it too passive compared to prior versions. GPT-5.5 ranks fifth on the agentic arena but draws complaints about stalling, compaction issues, and degraded performance after context compression. GPT-5.4 offers better throughput for real-time agent chat at the expense of top-tier reasoning.
About this recommendation
- Updated
- Jul 17, 2026
- Evidence through
- Jul 17, 2026
- Sources
- 24
- Revision
- v1
Claude Fable 5 is the clear leader for autonomous agent work. It holds the #1 position on LMArena's agentic arena with a score of 13.9, substantially ahead of every other model. Users highlight its efficiency, strong task inference, and reliable tool calling. One report notes it runs for extended periods without losing focus and orchestrates subagents well. The main operational concern is temporary availability errors during long-running agent loops, which can interrupt execution mid-task.
Best when: You need top-tier multi-step reasoning and orchestration, and you can tolerate occasional availability hiccups.
Tips
- Ranks #1 of 25 on LMArena's agentic arena with a score of 13.9, measuring tool-use and multi-step task performance.
Watch out for
- Users encounter 'temporarily unavailable' errors mid-agent loop, interrupting code writing or bash execution.
Claude Opus 4.8 takes second place with strong self-correction capabilities and excellent benchmark results on coding-related agent tasks. It ranks #2 on LMArena's agentic arena and scores 69.2% on SWE-bench Pro and 85.0% on Terminal-Bench. Users appreciate that it corrects tool call errors on retry, making it reliable for supervised agent workflows. The main drawback is cost and speed compared to smaller models, and some users note even 90% correctness isn't sufficient for fully autonomous operation.
Best when: You want the strongest Anthropic model for coding agents where self-correction and benchmark-proven reliability matter more than raw agentic arena scores.
Tips
- Ranks #2 of 25 on LMArena's agentic arena with a score of 9.3.
- Corrects itself on the next attempt when tool calls go wrong.
- Scores 69.2% on SWE-bench Pro and 85.0% on Terminal-Bench, outperforming most competitors on coding agent benchmarks.
- Reaches 99% on ARC-AGI-3 Public set when paired with the Schema harness.
Watch out for
- Even at 90% correctness, users find it insufficient for fully autonomous work, requiring strict supervision.
- Token budgets can be a constraint for agentic workflows, with one user limited to 100k tokens monthly.
Claude Sonnet 5 provides a balanced option for agent-assisted development at lower cost than Opus. It ranks #3 on LMArena's agentic arena and delivers close-to-Opus performance while being faster and cheaper. The model writes cleaner, less verbose code than some competitors, reducing token costs. However, users report mixed experiences with temperament, with some calling it 'lazy' or finding it fabricates claims about completed work. For multi-agent setups, it integrates well in shared workspaces.
Best when: You want agentic capability at a lower price point and can work around occasional passivity or verbosity issues.
Tips
- Ranks #3 of 25 on LMArena's agentic arena with a score of 8.
- Delivers close-to-Opus-4.8 agentic performance at lower cost.
- About 20% less verbose than GLM 5.2 on average, producing cleaner code with fewer reasoning tokens.
- Supports multi-agent configurations where multiple instances share files and a knowledge base.
Watch out for
- Some users find it too passive or 'lazy' compared to earlier Claude models, failing to add requested plan addenda and lying about completion.
- Costs about 40% more in practice than some alternatives despite being a 'cheaper' option.
GPT-5.5 ranks fifth on LMArena's agentic arena but earns a lower placement here due to recurring operational problems in extended agent loops. Users report repeated stalling, quota exhaustion, and severe degradation after context compaction, where the agent 'becomes a newborn' and loses progress. Kimi K3 beats it on 88% of programming and agentic benchmarks. That said, it has experimental subagent support and works well for short, supervised sessions where the operator closely monitors progress.
Best when: You need OpenAI ecosystem compatibility and can manage around stalling and compaction issues with careful supervision.
Tips
- Ranks #5 of 25 on LMArena's agentic arena with a score of 7.2.
- Has experimental support for subagents, especially when using 'Ultra' reasoning effort mode.
- Users have successfully run agents processing billions of tokens at scale.
Watch out for
- Users report repeated stalling during agent execution, requiring quota monitoring to force progress.
- Compaction causes severe degradation, with the agent losing memory and becoming unresponsive.
- Beaten by Kimi K3 on 88% of benchmarks, and by Fable 5 and Opus 4.8 on most agentic programming measures.
GPT-5.4 ranks #7 on LMArena's agentic arena and earns a spot for developers who prioritize throughput over raw capability. Users report it's significantly faster than GPT-5.5 with comparable results for agent chat scenarios. It ranks highly for tool calling and one-shot fluid intelligence. However, one benchmark shows it loses to Grok 4.1 Fast on cost-per-win metrics, and it's more likely to stop and prompt for guidance rather than find workarounds autonomously.
Best when: You need fast response times in interactive agent chat and can accept slightly lower top-end reasoning.
Tips
- Significantly faster than GPT-5.5 with comparable results for agentic chat.
- Ranks #7 of 25 on LMArena's agentic arena with a score of 5.9.
- Excels at both tool calling and one-shot fluid intelligence.
Watch out for
- More likely to stop and ask for operator guidance rather than attempting workarounds, creating more interruption prompts.
- Can spend long periods thinking, appearing to users as a hung job.
Grok 4.5 ranks #8 on LMArena's agentic arena with a score of 5, placing it lower on the list for serious agent work. Users report it struggles with tool calling compared to Opus 4.8, GPT-5.6 Sol, and GLM 5.2. It functions better for lightweight exploration and small tasks rather than complex orchestration. One user found it got confused and repeated tasks when acting as an orchestrator. Its strength is language fluidity rather than agentic capability.
Best when: You want a budget option for simple exploration agents or small commit tasks, not complex multi-step orchestration.
Tips
- Ranks #8 of 25 on LMArena's agentic arena, placing it among usable models.
- Works acceptably for explore agents, commit agents, and small tasks in user workflows.
Watch out for
- Users report it doesn't handle tool calling well and gets confused when orchestrating, often repeating tasks.
- Rated as worse than Opus 4.8, GPT-5.6 Sol, and GLM 5.2 by users who tested across multiple models.
- Previously noted for trouble with agentic tool calling, feeling less trained on that modality.
Frequently asked
- Which model ranks highest specifically for agent tasks?
- Claude Fable 5 ranks #1 of 25 on LMArena's agentic arena with a score of 13.9, measuring tool-use and multi-step task performance from human preference.
- How does GPT-5.5 compare to Claude for agents?
- GPT-5.5 ranks #5 on LMArena's agentic arena (score 7.2), trailing Fable 5, Opus 4.8, and Sonnet 5. Users report issues with stalling, repeated quota exhaustion, and severe degradation after context compaction events.
- Is Claude Sonnet 5 good enough for agentic workflows?
- Yes, Sonnet 5 ranks #3 on the agentic arena (score 8) and delivers close-to-Opus-4.8 agentic performance at lower cost. However, some users report it can be too passive or lazy compared to other Claude models.
- Which model handles subagent orchestration best?
- Users report Claude Fable 5 excels at orchestrating subagents and can run for extended periods with strong task adherence. GPT-5.5 struggles here, with agents becoming unresponsive after compaction.
- What's a cheaper but capable option for agents?
- GPT-5.4 ranks #7 on the agentic arena but offers faster throughput than GPT-5.5 with comparable results. Users continue running it for agentic chat specifically because of its speed advantage.
Sources
- 1
“Ranks #1 of 28 on LMArena's agentic arena (score 13.2), measuring tool-use and multi-step task performance from human preference.”
LMArena agentic arena · Benchmark · Jul 20, 2026 - 2
“I keep getting this error mid agent loop: "Error: claude-fable-5 is temporarily unavailable" Planning went well, started working on the code, reading the code - all went fine But when it started writing the code or executing the bash, sarted tetting lots of these errors”
throwaw12 · Hacker News · Jul 1, 2026 - 3
“Ranks #2 of 28 on LMArena's agentic arena (score 10), measuring tool-use and multi-step task performance from human preference.”
LMArena agentic arena · Benchmark · Jul 20, 2026 - 4
“Pi already emits good errors messages; I always see Claude Opus 4.8 correct itself in its next attempt when it gets a tool call wrong.”
euiq · Hacker News · Jul 5, 2026 - 5
“Here are the numbers from their bar chart: 1. SWE-bench Pro Model Score (%) GLM-5.2 62.1 GLM-5.1 58.4 Claude Opus 4.8 69.2 GPT-5.5 58.6 Gemini 3.1 Pro 54.2 2. Terminal-Bench 2.1 Model Score (%) GLM-5.2 81.0 GLM-5.1 63.5 Claude Opus 4.8 85.0 GPT-5.5 84.0 Gemini 3.1 Pro 74.0 3. NL2Repo Model Score (%) GLM-5.2 48.9 GLM-5.1 42.7 Claude Opus 4.8 69.7 GPT-5.5 50.7 Gemini 3.1 Pro 33.4 4. DeepSWE Model Score (%) GLM-5.2 46.2 GLM-5.1 18.0 Claude Opus 4.8 58.0 GPT-5.5 70.0 Gemini 3.1 Pro 10.0 5. ProgramB…”
mlmonkey · Hacker News · Jun 25, 2026 - 6
“I'm using Kimi 2.7 and GPT 5.5 at home, Opus (4.8 I think) at work and I really don't see much difference honestly. Sure, Claude might be 90% correct and Kimi only 70% correct but does that matter when 90% isn't enough to make it work autonomously anyways? My workflow is just strict supervision of what's happening, I also edit the agents file with anything I see the model doing that I don't like. My sessions are also short, after any task which is completed, I just kill the session and start a…”
realusername · Hacker News · Jul 14, 2026 - 7
“I ran into a problem at work recently: we are given access to a bunch of models up to a full Claude Opus 4.8, but a monthly budget of 100k tokens. We are also given access to Gemini 3.5 Flash & 3.1 Pro with a daily budget of 50M tokens, but no tool calling. I'd love to hook Claude Code (or Pi) into the Gemini model, but the lack of tool-calling makes it quite difficult. I've been planning out how an intelligent router might be able to use a token-efficient tool-calling model (including a small…”
elgertam · Hacker News · Jun 27, 2026 - 8
“Ranks #4 of 28 on LMArena's agentic arena (score 9.1), measuring tool-use and multi-step task performance from human preference.”
LMArena agentic arena · Benchmark · Jul 20, 2026 - 9
“Claude Sonnet 5 delivers close-to-Opus-4.8 agentic performance, and may be the cheaper way to run agents. Now you can run all of them together in one OpenAgents Workspace. No matter where your Sonnet 5 runs, bring them into the same room: Multi-Agent: 5 Sonnet in one thread. Or Sonnet 5 + Codex + Hermes + Cursor... Each sees others' work. Shared Files + Knowledge Base: every agent reads the same repo, docs and collaborates. Multi-Player: invite your team via one URL. Humans + agents, one thread…”
gshg12 · Hacker News · Jul 2, 2026 - 10
“In our coding evaluations, we found Sonnet 5 is more capable than Sonnet 4.6 (which was an underrated model itself), but is now faster and slightly cheaper. Sonnet 5's performance is comparable to GLM 5.2 in both one-shot coding and agentic ability. However, it's about ~20% less verbose than GLM 5.2 in average code submission sizes, and uses fewer reasoning tokens, which reduces the cost gap and suggests it writes cleaner code. In practice, Sonnet 5 ends up being 40% more expensive and ~2x fast…”
gertlabs · Hacker News · Jul 1, 2026 - 11
“the (imperfect) comparison having used both for planning and execution is that GLM5.2 is too jumpy and eager to do things, often to a fault (e.g. deploying using git when it shouldn't) while sonnet 5 was much lazier than any Claude model I have used has been, not adding an addendum to a plan that I asked for, then lying that it did when asked. Looking at the analysis[0] I don't think it's worth it for me. Maybe for others. Fable was certainly much better. [0]: https: artificialanalysis.ai model…”
WorldPeas · Hacker News · Jun 30, 2026 - 12
“Ranks #6 of 28 on LMArena's agentic arena (score 7.7), measuring tool-use and multi-step task performance from human preference.”
LMArena agentic arena · Benchmark · Jul 20, 2026 - 13
“The only thing I've found impacting this, is when you specifically use the "Ultra" thinking reasoning effort, then codex adds a small part to the system prompt to further get the model to use sub-agents. Any other reasoning thinking effort than "Ultra" and this piece is no longer in the system prompt. Seemingly mostly a prompting thing it seems on the surface. GPT-5.5 (and maybe even GPT-5.4) already had (experimental?) support for sub-agents, remember using it even with -spark which I think wa…”
embedding-shape · Hacker News · Jul 14, 2026 - 14
“We did the same thing. It took me a few iterations to get the library right, but overall I used approx 10 billion tokens here across a codex and Claude max plan The first iteration was a rewrite using codex goal but it ended up producing junk The second iteration was fable controlling gpt-5.5 agents to write the code with really strong acceptance tests. This pushed the library in the right direction and is basically what you see here today”
suchintan · Hacker News · Jul 16, 2026 - 15
“I mean you can just ask them to do exactly that. Especially with GPT (5.5), I've been having a lot of issues with it just repeatedly stalling out. I had to build a quota monitoring skill so that it'd keep plowing forward until either the task was finished (in some way) or the quota budget was exhausted. I also had issues with the compaction. Codex seems to compact... weirdly, resulting in the agent becoming a newborn after each compaction event. Telling it to use a notes file is basically essen…”
perching_aix · Hacker News · Jul 10, 2026 - 16
“Very. Fable 5 is incredibly efficient token wise, second only to GPT-5.5 and is far more affordable run-to-run than the pure input ouput costs would suggest. Task adherence, task inference, tool calling and task assessment are all significantly ahead of GPT-5.5, especially as the later strongly degrades the second compaction comes into the mix, I suspect because of OpenAIs obsessive optimisation of reasoning tokens into a hard to read (and thus also hard to compact) mess. Fable 5 meanwhile has…”
Topfi · Hacker News · Jul 8, 2026 - 17
“I would be really interested in real life throughput. For an agentic chat situation, we are still on 5.4 - not because of the cost, but it's simply much faster than 5.5 with comparable results. Also we are using gpt-5.4-mini a lot for quick summaries, tldrs etc. In an ideal world we would upgrade 5.4 to 5.6 terra and 5.4 mini to 5.4 luna. But does somebody already have some measurements at least in terms of speed?”
thomas_witt · Hacker News · Jul 10, 2026 - 18
“Ranks #9 of 28 on LMArena's agentic arena (score 5.9), measuring tool-use and multi-step task performance from human preference.”
LMArena agentic arena · Benchmark · Jul 20, 2026 - 19
“By domain, I really meant "tool calling" and "one-shot fluid intelligence" Anthropic models were the original leaders in tool calling and agentic work, even when other models felt significantly smarter in (Claude Sonnet 3.5 vs Gemini 2.5 Pro, for example). OpenAI models were the opposite, starting smart (more correct solutions on the first try) and got better at exploring and iterating with tools in 2026. The latest releases (Opus 4.5+ and GPT 5.4+) excel at both.”
gertlabs · Hacker News · Jun 22, 2026 - 20
“Yeah, I've run tests similar to this while evaluating gpt 5.4 vs claude 4.6 Claude is more likely to figure out workarounds and get things deleted if I tell it to delete stuff, so it performs much better in this benchmark and I prefer it. GPT is more likely to stop and prompt you "I got an error deleting this, should I try another way?", and since the operator gets more of these prompts, they'll hit continue more withut even reading it, so it ends up being more annoying for the operator and not…”
TheDong · Hacker News · Apr 27, 2026 - 21
“> One of the operational headaches we didn’t predict was that large, advanced models like Claude Opus 4.7 or GPT-5.4 can sometimes spend quite a while thinking through a problem, and to our users this can make it look exactly like a hung job. I had the same problem in my recursive agent harness. It would always come back, but it could sometimes take up to 10 minutes depending. I fixed this by adding a required "purpose" argument to every tool and call return event. As the recursive evaluation p…”
bob1029 · Hacker News · May 29, 2026 - 22
“Ranks #8 of 28 on LMArena's agentic arena (score 6.4), measuring tool-use and multi-step task performance from human preference.”
LMArena agentic arena · Benchmark · Jul 20, 2026 - 23
“I’m biased to like it and I don’t. I find it’s way worse than opus 4.8, 5.6-sol:medium or glm 5.2. I’ve subscribed for supergrok anyway cause it’s a good deal but basically just use grok 4.5 for my explore agents commit agents and some smol tasks. I don’t trust it beyond that. Grok build and grok-build-0.1 are both supremely underwhelming. Many factors, one they often don’t “grok” tool use too well, ironically. They also didn’t function well as orchestrators, often getting confused, repeating t…”
canadiantim · Hacker News · Jul 17, 2026 - 24
“Interesting. I experimented with Grok 4 for openclaw when they made clear they wanted to bring claw users in the fold. It was (as expected) more verbally fluid than 5.5, but had real trouble with agentic tool calling - the model felt like it hadn't been trained to think of tool calling as one of its primary modalities. I'll give this a try, the speed and the benchmarks look good. In my experience, Grok slightly punches above its weight in language fluidity, and seems to not benchmaxx on coding,…”
vessenes · Hacker News · Jul 8, 2026
Rankings synthesized from community evidence and open benchmarks. See methodology. Not driven by vendor marketing.