Recommendation for Tool calling

Tool & Function Calling

The best LLM for tool and function calling is Claude Fable 5, which ranks #1 on LMArena's agentic arena with a score of 13.9, followed closely by Claude Opus 4.8 at #2, with Claude Sonnet 5 rounding out the top tier at #3. These Anthropic models have established a clear lead in human preference evaluations specifically measuring tool-use and multi-step task performance. Claude Opus 4.8 has additional community validation showing it can self-correct when tool calls go wrong. Beyond the top three, Grok 4.5 sits at #8 and shows issues with understanding tool use as a core modality, despite decent benchmark fluidity. Gemini 3.5 Flash presents a mixed picture, boasting strong benchmark numbers but drawing significant community criticism for over-calling tools and struggling with instruction adherence in agentic contexts. GPT-5 appears in limited evidence suggesting it's larger and slower than efficiency-focused models while still producing broken tool calls.

About this recommendation

Updated
Jul 17, 2026
Evidence through
Jul 17, 2026
Sources
13
Revision
v1
  1. Claude Fable 5 claims the top spot with the highest agentic arena score at 13.9. It's the clear choice when you need the best validated performance for multi-step tool workflows.

    Best when: You need the highest-ranked model for complex multi-turn agentic tasks and are already in the Anthropic ecosystem.

    Tips

    • Ranks #1 of 25 on LMArena's agentic arena with a score of 13.9, the highest recorded.
      Source 1
      Ranks #1 of 28 on LMArena's agentic arena (score 13.2), measuring tool-use and multi-step task performance from human preference.
      LMArena agentic arenaOpen original ↗
    • Specifically evaluated on tool-use and multi-step task performance through human preference.
      Source 1
      Ranks #1 of 28 on LMArena's agentic arena (score 13.2), measuring tool-use and multi-step task performance from human preference.
      LMArena agentic arenaOpen original ↗
  2. Claude Opus 4.8 ties together excellent arena ranking with real-world reliability. It self-corrects when tool calls fail, which matters more than raw scores in production.

    Best when: You need a battle-tested model that can recover from errors in multi-step tool chains.

    Tips

    • Ranks #2 of 25 on LMArena's agentic arena with a strong score of 9.3.
      Source 2
      Ranks #2 of 28 on LMArena's agentic arena (score 10), measuring tool-use and multi-step task performance from human preference.
      LMArena agentic arenaOpen original ↗
    • Demonstrates self-correction capabilities, fixing tool call mistakes on subsequent attempts.
      Source 3
      Pi already emits good errors messages; I always see Claude Opus 4.8 correct itself in its next attempt when it gets a tool call wrong.
    • Validated for tool-use and multi-step task performance through human preference evaluation.
      Source 2
      Ranks #2 of 28 on LMArena's agentic arena (score 10), measuring tool-use and multi-step task performance from human preference.
      LMArena agentic arenaOpen original ↗
  3. Claude Sonnet 5 offers strong agentic performance at rank #3 with a score of 8. It's a solid fallback if Fable or Opus aren't available.

    Best when: You want reliable tool calling within the Anthropic family but don't need the absolute top tier.

    Tips

    • Ranks #3 of 25 on LMArena's agentic arena with a score of 8.
      Source 4
      Ranks #4 of 28 on LMArena's agentic arena (score 9.1), measuring tool-use and multi-step task performance from human preference.
      LMArena agentic arenaOpen original ↗
    • Measured specifically for tool-use and multi-step task performance.
      Source 4
      Ranks #4 of 28 on LMArena's agentic arena (score 9.1), measuring tool-use and multi-step task performance from human preference.
      LMArena agentic arenaOpen original ↗
  4. Grok 4.5 sits at #8 on the agentic arena, but users report it struggles to treat tool calling as a primary modality. It works for light tasks but falters as an orchestrator.

    Best when: You need verbal fluidity for simple tasks and can tolerate some confusion in complex tool chains.

    Tips

    • Ranks #8 of 25 on LMArena's agentic arena, placing it in the upper half.
      Source 5
      Ranks #8 of 28 on LMArena's agentic arena (score 6.4), measuring tool-use and multi-step task performance from human preference.
      LMArena agentic arenaOpen original ↗
    • Punches above its weight in language fluidity according to user experience.
      Source 6
      Interesting. I experimented with Grok 4 for openclaw when they made clear they wanted to bring claw users in the fold. It was (as expected) more verbally fluid than 5.5, but had real trouble with agentic tool calling - the model felt like it hadn't been trained to think of tool calling as one of its primary modalities. I'll give this a try, the speed and the benchmarks look good. In my experience, Grok slightly punches above its weight in language fluidity, and seems to not benchmaxx on coding,…

    Watch out for

    • Users report it doesn't understand tool use well and often gets confused or repeats itself as an orchestrator.
      Source 7
      I’m biased to like it and I don’t. I find it’s way worse than opus 4.8, 5.6-sol:medium or glm 5.2. I’ve subscribed for supergrok anyway cause it’s a good deal but basically just use grok 4.5 for my explore agents commit agents and some smol tasks. I don’t trust it beyond that. Grok build and grok-build-0.1 are both supremely underwhelming. Many factors, one they often don’t “grok” tool use too well, ironically. They also didn’t function well as orchestrators, often getting confused, repeating t…
      Source 6
      Interesting. I experimented with Grok 4 for openclaw when they made clear they wanted to bring claw users in the fold. It was (as expected) more verbally fluid than 5.5, but had real trouble with agentic tool calling - the model felt like it hadn't been trained to think of tool calling as one of its primary modalities. I'll give this a try, the speed and the benchmarks look good. In my experience, Grok slightly punches above its weight in language fluidity, and seems to not benchmaxx on coding,…
    • Doesn't seem trained to treat tool calling as a primary modality.
      Source 6
      Interesting. I experimented with Grok 4 for openclaw when they made clear they wanted to bring claw users in the fold. It was (as expected) more verbally fluid than 5.5, but had real trouble with agentic tool calling - the model felt like it hadn't been trained to think of tool calling as one of its primary modalities. I'll give this a try, the speed and the benchmarks look good. In my experience, Grok slightly punches above its weight in language fluidity, and seems to not benchmaxx on coding,…
    • Some users don't trust it beyond small tasks.
      Source 7
      I’m biased to like it and I don’t. I find it’s way worse than opus 4.8, 5.6-sol:medium or glm 5.2. I’ve subscribed for supergrok anyway cause it’s a good deal but basically just use grok 4.5 for my explore agents commit agents and some smol tasks. I don’t trust it beyond that. Grok build and grok-build-0.1 are both supremely underwhelming. Many factors, one they often don’t “grok” tool use too well, ironically. They also didn’t function well as orchestrators, often getting confused, repeating t…
  5. Gemini 3.5 Flash is a polarizing option. It ranks #16 on the agentic arena and drawss complaints about over-calling tools and poor instruction adherence, though some users defend its tool calling capabilities.

    Best when: You want raw efficiency and can handle potentially degenerate query patterns or excessive tool calls.

    Tips

    • Rocks benchmark charts according to community feedback.
      Source 8
      I second that. Gemini 3.5 Flash rocks the benchmark charts but is terrible as an agent. Horrible instruction adherence and makes WAY too many tool calls
    • One user claims it's better than Fable at tool calling.
      Source 9
      Gemini 3.5 flash is better than fable at tool calling. Tool calling is probably one of the easier things to do post training for.
      ai_fry_ur_brainOpen original ↗
    • Likes to stay quiet and just make tool calls, unlike more verbose models.
      Source 10
      Seems like I have much better results with Gemini <= 3.5 Flash being able to solve 10 Lake levels. Of course there may be methodology differences (like how many times the level is attempted from scratch). I wonder if the reasoning tokens are still being accidentally discarded between turns in those tests. It seems to be much more important to preserve those for Gemini, as it likes to stay quiet and and just make tool calls, unlike Claude that yaps a lot what it's planning to do.

    Watch out for

    • Ranks #16 of 25 on LMArena's agentic arena with a negative score of -1.1.
      Source 11
      Ranks #15 of 28 on LMArena's agentic arena (score -0.7), measuring tool-use and multi-step task performance from human preference.
      LMArena agentic arenaOpen original ↗
    • Users report horrible instruction adherence and way too many tool calls.
      Source 8
      I second that. Gemini 3.5 Flash rocks the benchmark charts but is terrible as an agent. Horrible instruction adherence and makes WAY too many tool calls
    • Can hit query budget limits 10x faster than previous versions with degenerate patterns like running SELECT * LIMIT 1.
      Source 12
      I'm seeing this too. I have a SQL agent and my tests with 3.5 are resulting in hitting query budget limits that have never been hit before. On average, to answer the same question, 3.5 is spending 10x more on SQL queries vs gemini-3-flash-preview. The query patterns can be extremely degenerate too. E.g. the agent will hit the semantic layer tool to pull the schema, then run `SELECT * FROM table LIMIT 1`, which hits the query budget limit and fails. I've only really been looking this morning, so…
    • Described as terrible as an agent despite strong benchmarks.
      Source 8
      I second that. Gemini 3.5 Flash rocks the benchmark charts but is terrible as an agent. Horrible instruction adherence and makes WAY too many tool calls
  6. GPT-5 has minimal direct evidence for tool calling. Early signals suggest it produces broken tool calls and struggles with agentic tasks compared to its raw problem-solving abilities.

    Best when: You need raw problem-solving without tools or search, where user reports suggest it matches Opus.

    Tips

    • Matches Opus class models for raw problem solving without tools.
      Source 13
      I already felt that gemini 3 proved what is possible if you train a model for efficiency. If I had to guess the pro and flash variants are 5x to 10x smaller than opus and gpt-5 class models. They produce drastically lower amount of tokens to solve a problem, but they haven't seem to have put enough effort into refinining their reasoning and execution as they produce broken toolcalls and generally struggle with 'agentic' tasks, but for raw problem solving without tools or search they match opus…

    Watch out for

    • Produces broken tool calls and generally struggles with agentic tasks.
      Source 13
      I already felt that gemini 3 proved what is possible if you train a model for efficiency. If I had to guess the pro and flash variants are 5x to 10x smaller than opus and gpt-5 class models. They produce drastically lower amount of tokens to solve a problem, but they haven't seem to have put enough effort into refinining their reasoning and execution as they produce broken toolcalls and generally struggle with 'agentic' tasks, but for raw problem solving without tools or search they match opus…
    • Larger and slower than efficiency-focused models, producing drastically more tokens to solve problems.
      Source 13
      I already felt that gemini 3 proved what is possible if you train a model for efficiency. If I had to guess the pro and flash variants are 5x to 10x smaller than opus and gpt-5 class models. They produce drastically lower amount of tokens to solve a problem, but they haven't seem to have put enough effort into refinining their reasoning and execution as they produce broken toolcalls and generally struggle with 'agentic' tasks, but for raw problem solving without tools or search they match opus…

Frequently asked

Which model ranks highest specifically for agentic tool use?
Claude Fable 5 ranks #1 on LMArena's agentic arena with a score of 13.9, measuring tool-use and multi-step task performance from human preference.
Does Claude Opus 4.8 handle tool call failures well?
Yes, community feedback indicates Claude Opus 4.8 can correct itself on its next attempt when it gets a tool call wrong.
Is Gemini 3.5 Flash good for agentic workflows despite its benchmarks?
Community reports are split. Some say it's terrible as an agent with horrible instruction adherence and excessive tool calls, while others claim it's better than Fable at tool calling.
What are the main issues with Grok 4.5 for function calling?
Users report Grok 4.5 often doesn't understand tool use well and can get confused or repeat itself when acting as an orchestrator.

Sources

  1. 1

    Ranks #1 of 28 on LMArena's agentic arena (score 13.2), measuring tool-use and multi-step task performance from human preference.

    LMArena agentic arena · Benchmark · Jul 20, 2026
  2. 2

    Ranks #2 of 28 on LMArena's agentic arena (score 10), measuring tool-use and multi-step task performance from human preference.

    LMArena agentic arena · Benchmark · Jul 20, 2026
  3. 3

    Pi already emits good errors messages; I always see Claude Opus 4.8 correct itself in its next attempt when it gets a tool call wrong.

    euiq · Hacker News · Jul 5, 2026
  4. 4

    Ranks #4 of 28 on LMArena's agentic arena (score 9.1), measuring tool-use and multi-step task performance from human preference.

    LMArena agentic arena · Benchmark · Jul 20, 2026
  5. 5

    Ranks #8 of 28 on LMArena's agentic arena (score 6.4), measuring tool-use and multi-step task performance from human preference.

    LMArena agentic arena · Benchmark · Jul 20, 2026
  6. 6

    Interesting. I experimented with Grok 4 for openclaw when they made clear they wanted to bring claw users in the fold. It was (as expected) more verbally fluid than 5.5, but had real trouble with agentic tool calling - the model felt like it hadn't been trained to think of tool calling as one of its primary modalities. I'll give this a try, the speed and the benchmarks look good. In my experience, Grok slightly punches above its weight in language fluidity, and seems to not benchmaxx on coding,…

    vessenes · Hacker News · Jul 8, 2026
  7. 7

    I’m biased to like it and I don’t. I find it’s way worse than opus 4.8, 5.6-sol:medium or glm 5.2. I’ve subscribed for supergrok anyway cause it’s a good deal but basically just use grok 4.5 for my explore agents commit agents and some smol tasks. I don’t trust it beyond that. Grok build and grok-build-0.1 are both supremely underwhelming. Many factors, one they often don’t “grok” tool use too well, ironically. They also didn’t function well as orchestrators, often getting confused, repeating t…

    canadiantim · Hacker News · Jul 17, 2026
  8. 8

    I second that. Gemini 3.5 Flash rocks the benchmark charts but is terrible as an agent. Horrible instruction adherence and makes WAY too many tool calls

    speedping · Hacker News · Jul 15, 2026
  9. 9

    Gemini 3.5 flash is better than fable at tool calling. Tool calling is probably one of the easier things to do post training for.

    ai_fry_ur_brain · Hacker News · Jul 9, 2026
  10. 10

    Seems like I have much better results with Gemini <= 3.5 Flash being able to solve 10 Lake levels. Of course there may be methodology differences (like how many times the level is attempted from scratch). I wonder if the reasoning tokens are still being accidentally discarded between turns in those tests. It seems to be much more important to preserve those for Gemini, as it likes to stay quiet and and just make tool calls, unlike Claude that yaps a lot what it's planning to do.

    dezgeg · Hacker News · Jul 17, 2026
  11. 11

    Ranks #15 of 28 on LMArena's agentic arena (score -0.7), measuring tool-use and multi-step task performance from human preference.

    LMArena agentic arena · Benchmark · Jul 20, 2026
  12. 12

    I'm seeing this too. I have a SQL agent and my tests with 3.5 are resulting in hitting query budget limits that have never been hit before. On average, to answer the same question, 3.5 is spending 10x more on SQL queries vs gemini-3-flash-preview. The query patterns can be extremely degenerate too. E.g. the agent will hit the semantic layer tool to pull the schema, then run `SELECT * FROM table LIMIT 1`, which hits the query budget limit and fails. I've only really been looking this morning, so…

    data-ottawa · Hacker News · May 20, 2026
  13. 13

    I already felt that gemini 3 proved what is possible if you train a model for efficiency. If I had to guess the pro and flash variants are 5x to 10x smaller than opus and gpt-5 class models. They produce drastically lower amount of tokens to solve a problem, but they haven't seem to have put enough effort into refinining their reasoning and execution as they produce broken toolcalls and generally struggle with 'agentic' tasks, but for raw problem solving without tools or search they match opus…

    himata4113 · Hacker News · Apr 22, 2026

Rankings synthesized from community evidence and open benchmarks. See methodology. Not driven by vendor marketing.