Recommendation for Tool calling
Tool & Function Calling
The best LLM for tool and function calling is Claude Fable 5, which ranks #1 on LMArena's agentic arena with a score of 13.9, followed closely by Claude Opus 4.8 at #2, with Claude Sonnet 5 rounding out the top tier at #3. These Anthropic models have established a clear lead in human preference evaluations specifically measuring tool-use and multi-step task performance. Claude Opus 4.8 has additional community validation showing it can self-correct when tool calls go wrong. Beyond the top three, Grok 4.5 sits at #8 and shows issues with understanding tool use as a core modality, despite decent benchmark fluidity. Gemini 3.5 Flash presents a mixed picture, boasting strong benchmark numbers but drawing significant community criticism for over-calling tools and struggling with instruction adherence in agentic contexts. GPT-5 appears in limited evidence suggesting it's larger and slower than efficiency-focused models while still producing broken tool calls.
About this recommendation
- Updated
- Jul 17, 2026
- Evidence through
- Jul 17, 2026
- Sources
- 13
- Revision
- v1
Claude Fable 5 claims the top spot with the highest agentic arena score at 13.9. It's the clear choice when you need the best validated performance for multi-step tool workflows.
Best when: You need the highest-ranked model for complex multi-turn agentic tasks and are already in the Anthropic ecosystem.
Tips
- Ranks #1 of 25 on LMArena's agentic arena with a score of 13.9, the highest recorded.
- Specifically evaluated on tool-use and multi-step task performance through human preference.
Claude Opus 4.8 ties together excellent arena ranking with real-world reliability. It self-corrects when tool calls fail, which matters more than raw scores in production.
Best when: You need a battle-tested model that can recover from errors in multi-step tool chains.
Tips
- Ranks #2 of 25 on LMArena's agentic arena with a strong score of 9.3.
- Demonstrates self-correction capabilities, fixing tool call mistakes on subsequent attempts.
- Validated for tool-use and multi-step task performance through human preference evaluation.
Claude Sonnet 5 offers strong agentic performance at rank #3 with a score of 8. It's a solid fallback if Fable or Opus aren't available.
Best when: You want reliable tool calling within the Anthropic family but don't need the absolute top tier.
Tips
- Ranks #3 of 25 on LMArena's agentic arena with a score of 8.
- Measured specifically for tool-use and multi-step task performance.
Grok 4.5 sits at #8 on the agentic arena, but users report it struggles to treat tool calling as a primary modality. It works for light tasks but falters as an orchestrator.
Best when: You need verbal fluidity for simple tasks and can tolerate some confusion in complex tool chains.
Tips
- Ranks #8 of 25 on LMArena's agentic arena, placing it in the upper half.
- Punches above its weight in language fluidity according to user experience.
Watch out for
- Users report it doesn't understand tool use well and often gets confused or repeats itself as an orchestrator.
- Doesn't seem trained to treat tool calling as a primary modality.
- Some users don't trust it beyond small tasks.
Gemini 3.5 Flash is a polarizing option. It ranks #16 on the agentic arena and drawss complaints about over-calling tools and poor instruction adherence, though some users defend its tool calling capabilities.
Best when: You want raw efficiency and can handle potentially degenerate query patterns or excessive tool calls.
Tips
- Rocks benchmark charts according to community feedback.
- One user claims it's better than Fable at tool calling.
- Likes to stay quiet and just make tool calls, unlike more verbose models.
Watch out for
- Ranks #16 of 25 on LMArena's agentic arena with a negative score of -1.1.
- Users report horrible instruction adherence and way too many tool calls.
- Can hit query budget limits 10x faster than previous versions with degenerate patterns like running SELECT * LIMIT 1.
- Described as terrible as an agent despite strong benchmarks.
GPT-5 has minimal direct evidence for tool calling. Early signals suggest it produces broken tool calls and struggles with agentic tasks compared to its raw problem-solving abilities.
Best when: You need raw problem-solving without tools or search, where user reports suggest it matches Opus.
Tips
- Matches Opus class models for raw problem solving without tools.
Watch out for
- Produces broken tool calls and generally struggles with agentic tasks.
- Larger and slower than efficiency-focused models, producing drastically more tokens to solve problems.
Frequently asked
- Which model ranks highest specifically for agentic tool use?
- Claude Fable 5 ranks #1 on LMArena's agentic arena with a score of 13.9, measuring tool-use and multi-step task performance from human preference.
- Does Claude Opus 4.8 handle tool call failures well?
- Yes, community feedback indicates Claude Opus 4.8 can correct itself on its next attempt when it gets a tool call wrong.
- Is Gemini 3.5 Flash good for agentic workflows despite its benchmarks?
- Community reports are split. Some say it's terrible as an agent with horrible instruction adherence and excessive tool calls, while others claim it's better than Fable at tool calling.
- What are the main issues with Grok 4.5 for function calling?
- Users report Grok 4.5 often doesn't understand tool use well and can get confused or repeat itself when acting as an orchestrator.
Sources
- 1
“Ranks #1 of 28 on LMArena's agentic arena (score 13.2), measuring tool-use and multi-step task performance from human preference.”
LMArena agentic arena · Benchmark · Jul 20, 2026 - 2
“Ranks #2 of 28 on LMArena's agentic arena (score 10), measuring tool-use and multi-step task performance from human preference.”
LMArena agentic arena · Benchmark · Jul 20, 2026 - 3
“Pi already emits good errors messages; I always see Claude Opus 4.8 correct itself in its next attempt when it gets a tool call wrong.”
euiq · Hacker News · Jul 5, 2026 - 4
“Ranks #4 of 28 on LMArena's agentic arena (score 9.1), measuring tool-use and multi-step task performance from human preference.”
LMArena agentic arena · Benchmark · Jul 20, 2026 - 5
“Ranks #8 of 28 on LMArena's agentic arena (score 6.4), measuring tool-use and multi-step task performance from human preference.”
LMArena agentic arena · Benchmark · Jul 20, 2026 - 6
“Interesting. I experimented with Grok 4 for openclaw when they made clear they wanted to bring claw users in the fold. It was (as expected) more verbally fluid than 5.5, but had real trouble with agentic tool calling - the model felt like it hadn't been trained to think of tool calling as one of its primary modalities. I'll give this a try, the speed and the benchmarks look good. In my experience, Grok slightly punches above its weight in language fluidity, and seems to not benchmaxx on coding,…”
vessenes · Hacker News · Jul 8, 2026 - 7
“I’m biased to like it and I don’t. I find it’s way worse than opus 4.8, 5.6-sol:medium or glm 5.2. I’ve subscribed for supergrok anyway cause it’s a good deal but basically just use grok 4.5 for my explore agents commit agents and some smol tasks. I don’t trust it beyond that. Grok build and grok-build-0.1 are both supremely underwhelming. Many factors, one they often don’t “grok” tool use too well, ironically. They also didn’t function well as orchestrators, often getting confused, repeating t…”
canadiantim · Hacker News · Jul 17, 2026 - 8
“I second that. Gemini 3.5 Flash rocks the benchmark charts but is terrible as an agent. Horrible instruction adherence and makes WAY too many tool calls”
speedping · Hacker News · Jul 15, 2026 - 9
“Gemini 3.5 flash is better than fable at tool calling. Tool calling is probably one of the easier things to do post training for.”
ai_fry_ur_brain · Hacker News · Jul 9, 2026 - 10
“Seems like I have much better results with Gemini <= 3.5 Flash being able to solve 10 Lake levels. Of course there may be methodology differences (like how many times the level is attempted from scratch). I wonder if the reasoning tokens are still being accidentally discarded between turns in those tests. It seems to be much more important to preserve those for Gemini, as it likes to stay quiet and and just make tool calls, unlike Claude that yaps a lot what it's planning to do.”
dezgeg · Hacker News · Jul 17, 2026 - 11
“Ranks #15 of 28 on LMArena's agentic arena (score -0.7), measuring tool-use and multi-step task performance from human preference.”
LMArena agentic arena · Benchmark · Jul 20, 2026 - 12
“I'm seeing this too. I have a SQL agent and my tests with 3.5 are resulting in hitting query budget limits that have never been hit before. On average, to answer the same question, 3.5 is spending 10x more on SQL queries vs gemini-3-flash-preview. The query patterns can be extremely degenerate too. E.g. the agent will hit the semantic layer tool to pull the schema, then run `SELECT * FROM table LIMIT 1`, which hits the query budget limit and fails. I've only really been looking this morning, so…”
data-ottawa · Hacker News · May 20, 2026 - 13
“I already felt that gemini 3 proved what is possible if you train a model for efficiency. If I had to guess the pro and flash variants are 5x to 10x smaller than opus and gpt-5 class models. They produce drastically lower amount of tokens to solve a problem, but they haven't seem to have put enough effort into refinining their reasoning and execution as they produce broken toolcalls and generally struggle with 'agentic' tasks, but for raw problem solving without tools or search they match opus…”
himata4113 · Hacker News · Apr 22, 2026
Rankings synthesized from community evidence and open benchmarks. See methodology. Not driven by vendor marketing.