FUURAA AI Knowledge Library · tool-use evaluation

How to evaluate AI tool-use and function-calling capability claims

Use six evidence gates to turn “can call tools and complete tasks” into an exact task, tools and authority, schemas and environment, selection and arguments, real execution, final state, all failures, recovery, human work, cost and a dated target-workflow boundary.

Published29 August 2026Evidence statusMethod synthesis grounded in primary tool-use and function-calling researchScopeLLM tool selection, function calling and stateful task execution research

A valid call is not completion; fluent narration is not state evidence

Tool capability must keep task, authority, tool surface, arguments, execution, final state, failures, recovery and side effects in one evidence chain.

A correct function name or valid JSON only shows that a call is well formed. The real decision must also establish tool suitability, argument provenance, execution, intended world-state change, rule compliance and whether failures stop or recover safely.

Applicability boundaryThis is a public research method, not a FUURAA or FUUVO product-capability claim or an assessment of any model, provider, agent, tool, API, dataset or leaderboard. It is not a guarantee of task completion, safety, authority or reliability.

Six rejectable evidence gates

Each gate requires minimum evidence and stops transfer or narrows the conclusion when material unknowns remain.

01

Freeze the task, authority and acceptance decision

Decision question
What exact user outcome may the system pursue, with which tools, credentials, side effects and consequences?
Minimum evidence
Verbatim claim and date; exact agent, model, prompt and orchestration versions; user task and starting state; eligible tools; read, write, spend, send and delete authority; approval points; success, safety, latency and cost thresholds.
Stop condition
Stop when a generic tool-use score or demonstration replaces a named task, authority boundary, target state and acceptance decision.
02

Version the tool surface and environment

Decision question
Which tool definitions, schemas, documentation, credentials, dependencies and world state were actually available?
Minimum evidence
Tool registry and hashes; names, descriptions, endpoints and schema versions; required and optional parameters; examples and error contracts; credentials, scopes, sandbox, rate limits, clocks, locale, network, dependencies and initial state snapshots.
Stop condition
Stop when tool availability, documentation version, access scope, initial state or simulator fidelity is unknown.
03

Separate discovery, selection and arguments

Decision question
Can the system decide whether a tool is needed, find the right one and construct semantically correct calls?
Minimum evidence
No-tool and unavailable-tool cases; candidate-set size and distractors; selection precision and recall; required, optional, nested and constrained arguments; value provenance; ordering, parallelism, dependencies and clarification behaviour.
Stop condition
Stop when valid JSON, a matching function name or one gold trajectory is treated as proof that the intended call is correct.
04

Execute, observe state and verify completion

Decision question
Did calls run, produce the intended state change and support the final answer or user-visible outcome?
Minimum evidence
Raw requests, responses and timestamps; tool and transport errors; before-and-after state; idempotency keys; return-value grounding; policy checks; partial completion; final-state or task-specific validators independent of the agent's own narration.
Stop condition
Stop when a syntactically accepted call, HTTP success, plausible response or fluent final message substitutes for verified state and outcome.
05

Count every call, failure, retry and intervention

Decision question
What happened across every eligible task and run, including loops, duplicate effects, recoveries and human work?
Minimum evidence
Complete task and call denominators; repeated trials and variance; wrong, omitted, malformed and unnecessary calls; refusals, timeouts and rate limits; retries, fallbacks, rollbacks and duplicate effects; human approvals, corrections, latency, tokens, API spend and total cost.
Stop condition
Stop when selected successes hide failed tasks, extra calls, unstable repeats, manual rescue, unrolled side effects or cost.
06

Transfer to the target workflow and expire the claim

Decision question
Does the controlled system remain useful under target users, policies, tools and consequences after anything changes?
Minimum evidence
Shadow or staged use; representative users and exceptions; least-privilege and approval controls; incident and rollback drills; task outcome and harm checks; monitoring; model, prompt, tool, schema, credential, policy and dependency change triggers; owner and expiry date.
Stop condition
Stop when a benchmark or sandbox result transfers to live tools without direct workflow, permission, state, monitoring and dated recheck evidence.

Minimum tool-use claim failure matrix

Check these conditions deliberately before transferring call format or selected trajectories into real task completion.

  • 01
    The system chooses the right tool family but an obsolete endpoint, version or similarly named function.

    Record the affected task, tool, authority, state, side effect, user and decision; preserve the narrowest conclusion that remains and specify whether to add tests, rerun, roll back, hand off, restrict authority or withdraw the claim.

  • 02
    A call satisfies the schema while a date, identifier, unit, enum, locale or nested argument is semantically wrong.

    Record the affected task, tool, authority, state, side effect, user and decision; preserve the narrowest conclusion that remains and specify whether to add tests, rerun, roll back, hand off, restrict authority or withdraw the claim.

  • 03
    The agent omits clarification, exceeds the user's authority or turns a read request into a write, send, spend or delete action.

    Record the affected task, tool, authority, state, side effect, user and decision; preserve the narrowest conclusion that remains and specify whether to add tests, rerun, roll back, hand off, restrict authority or withdraw the claim.

  • 04
    A needed tool is absent, yet the system hallucinates a function, parameter, result or unsupported completion.

    Record the affected task, tool, authority, state, side effect, user and decision; preserve the narrowest conclusion that remains and specify whether to add tests, rerun, roll back, hand off, restrict authority or withdraw the claim.

  • 05
    The API returns success while the intended record, recipient, quantity, permission or final state is wrong or unchanged.

    Record the affected task, tool, authority, state, side effect, user and decision; preserve the narrowest conclusion that remains and specify whether to add tests, rerun, roll back, hand off, restrict authority or withdraw the claim.

  • 06
    One step succeeds and the final response hides a later failure, missing dependency, stale result or partial task.

    Record the affected task, tool, authority, state, side effect, user and decision; preserve the narrowest conclusion that remains and specify whether to add tests, rerun, roll back, hand off, restrict authority or withdraw the claim.

  • 07
    Recovery repeats an irreversible call, creates duplicate side effects or continues a loop without a safe stop.

    Record the affected task, tool, authority, state, side effect, user and decision; preserve the narrowest conclusion that remains and specify whether to add tests, rerun, roll back, hand off, restrict authority or withdraw the claim.

  • 08
    A model, prompt, schema, documentation, credential, rate limit, policy or service change invalidates prior evidence.

    Record the affected task, tool, authority, state, side effect, user and decision; preserve the narrowest conclusion that remains and specify whether to add tests, rerun, roll back, hand off, restrict authority or withdraw the claim.

Minimum tool-use capability evaluation record

Let the next reviewer reconstruct the conclusion with the same task, tools, authority, environment, execution traces and complete runs.

  1. 01Claim, date, owner, user task, initial state, consequence and acceptance thresholds
  2. 02Agent, model, prompt, orchestrator, memory, runtime and deployment versions
  3. 03Tool registry, definitions, schemas, documentation, endpoints and dependency versions
  4. 04Credentials, scopes, sandbox, approval points, rate limits and policy rules
  5. 05Task set, tool/no-tool cases, distractors, languages, users, exceptions and exclusions
  6. 06Selected tools, arguments, value provenance, order, parallel calls and clarifications
  7. 07Raw requests, responses, timestamps, errors and before-and-after state
  8. 08Independent task, final-state, return-grounding, policy and side-effect validation
  9. 09All tasks, calls, failures, retries, rollbacks, interventions, latency and total cost
  10. 10Target-workflow evidence, monitoring, change triggers, residual limits and expiry

Common evidence states

Bind conclusions to the exact toolchain, task, authority, complete denominator, target workflow and date.

Supported

The exact system completes representative tasks within authority, final-state, reliability, side-effect, human-effort, latency and cost thresholds, with a current review date.

Conditional

Support holds only for named tools, schema versions, tasks, permissions, environments, user groups or operating controls.

Mixed

Selection, arguments, execution, final state, policy compliance, recovery, repeated-run reliability, latency or cost vary materially across slices.

Insufficient

Tool surface, authority, initial state, executable traces, independent outcome checks, complete denominator, target transfer or expiry is missing.

FUURAA analysisThe minimum decision unit for a tool-use and function-calling capability claim is exact agent, model, prompt and toolchain version × user task, authority, consequence and goal state × tool registry, schemas, documentation, credentials, sandbox and initial state × discovery, selection, arguments, order and execution × final state, return grounding, rule compliance and task completion × all calls, failures, retries, rollbacks, interventions, latency and cost × target-workflow boundary and cut-off date. Producing one well-formed call is the start of a candidate action, not evidence of completion.

Primary sources and non-transfer boundaries

These sources constrain call timing, API planning, arguments and documentation, tool retrieval, function selection and final-state reliability.

Sources rechecked 29 August 2026. Each retains its publication timing, role in this method and non-transfer boundary.

Published 9 February 2023Toolformer — Language Models Can Teach Themselves to Use Tools

Studies how a model can learn when to call simple APIs, which arguments to pass and how to incorporate returned results.

BoundaryA small fixed set of read-oriented tools and downstream language tasks does not establish permission handling, state-changing execution, recovery or production completion.

Open primary source ↗
Published 14 April 2023API-Bank — A Comprehensive Benchmark for Tool-Augmented LLMs

A Chinese-led study provides a runnable system for planning, retrieving and calling APIs across annotated tool-use dialogues.

BoundaryIts selected APIs, synthetic training data and dialogue tasks do not cover every schema, language, permission model, live service or irreversible action.

Open primary source ↗
Published 24 May 2023Gorilla — Large Language Model Connected with Massive APIs

Introduces APIBench and tests API-call generation, retrieval over documentation and adaptation to documentation changes.

BoundaryCorrect-looking calls against model-hub APIs do not prove successful execution, target-state correctness, safe authority or end-to-end task completion.

Open primary source ↗
Published 31 July 2023ToolLLM and ToolBench — Mastering 16,000+ Real-world APIs

A Chinese-led framework spans API collection, instruction generation, multi-tool solution paths, retrieval and automatic evaluation over many REST APIs.

BoundaryAPI count and an automatic evaluator do not guarantee live endpoint stability, response semantics, access control, side-effect safety or human acceptance.

Open primary source ↗
Released February 2024; updated 19 August 2024Berkeley Function-Calling Leaderboard

Tests relevance detection plus simple, multiple, parallel and executable function calls across several programming and API formats.

BoundaryFunction-selection, AST and execution categories do not by themselves establish a user's final state, policy compliance, repeated-run reliability or deployment controls.

Open primary source ↗
Published 17 June 2024τ-bench — Tool-Agent-User Interaction in Real-World Domains

Evaluates dynamic user interaction, domain rules, API tools, final database state and reliability across repeated trials.

BoundarySimulated users, two domains and annotated goal states do not represent every person, policy, tool surface, exception or real operational consequence.

Open primary source ↗

Continue checking

Move from tool calling into AI agents, RAG, coding capability or the full library.

Enter the AI Knowledge LibraryEvaluate AI agentsEvaluate RAGEvaluate coding capabilityEnter AI Evidence Atlas