Skip to content

Compare Tool-Call Models

Pick a catalog model for a ToolCall node using measured exact accuracy, not parameter count. Two models with the same overall score can be strong and weak in opposite places — the tables below are for choosing by job, not for crowning a winner.

Everything here is greedy decoding (temperature 0, seed 42) on the shipped tool-call catalog. Temperature 0.3 and 0.7 move the field by about a point; use greedy numbers unless you already know you will sample hotter in game.

How a case is scored

Exact accuracy: the call matches the expected tool name and every argument, or the model correctly calls nothing. No partial credit, no judge model. Constrained decoding holds the syntax; these numbers are about which tool and which arguments.

Which table to read

Your tools look like… Read
Game commands, NPC actuation, log queries, trading Tryll-authored suites
Generic APIs, several tools at once, "don't call anything" BFCL slice

The two corpora do not rank models the same way. NVIDIA Orchestrator 8B leads the authored suites and sits in a pack on BFCL. Llama 3.1 8B is usable on authored command parsing and collapses on BFCL parallel cases. Read the corpus that matches the shape of your tools.

Catalog models whose tool_call_support is unsupported are not in these tables — CreateAgent rejects them for a ToolCall node. See Not every model can call tools.


Tryll-authored suites

Seven game-shaped suites, 132 cases per model. Twenty catalog models; Mistral 7B did not complete this corpus (a constrained-decoding abort on one suite). Run authored_sweep_20260821_0300.

Model Exact accuracy
Ministral 3 14B Instruct (Q4_K_M) 96.2%
NVIDIA Orchestrator 8B (Q4_K_M) 96.2%
QVikhr 3 4B Instruction (Q4_K_M) 93.9%
Bonsai 27B (Q1_0 1-bit) 93.2%
Ministral 3 3B Instruct (Q4_K_M) 91.7%
Ministral 3 8B Instruct (Q4_K_M) 91.7%
Granite 4.0 H-Micro Instruct (Q4_K_M) 90.9%
Qwen2.5 7B Instruct (Q4_K_M) 88.6%
Bonsai 8B (Q2_0 2-bit) 87.1%
NVIDIA Nemotron 3 Nano 4B (Q4_K_M) 87.1%
Qwen 3.5 4B (Q4_K_M) 84.1%
Gemma 4 E4B Instruct (Q4_K_M) 81.1%
Qwen2.5 3B Instruct (Q4_K_M) 81.1%
Granite 4.0 H-Tiny Instruct (Q4_K_M) 80.3%
Llama 3.1 8B Instruct (Q4_K_M) 78.0%
Qwen 3.5 2B (Q4_K_M) 76.5%
Bonsai 8B (Q1_0 1-bit) 72.7%
Llama 3.2 3B Instruct (Q4_K_M) 50.8%
Qwen2.5 0.5B Instruct (Q4_K_M) 50.8%
Llama 3.2 1B Instruct (Q4_K_M) 36.4%

Shape matters more than the average

Pass rate per suite, pooled across three temperatures, for four models chosen because their profiles differ. Outer ring is 100%.

Radar of per-suite tool-call pass rate for NVIDIA Orchestrator 8B, Qwen2.5 7B, Bonsai 8B Q1_0, and Llama 3.2 1B

Bonsai 8B (Q1_0) is the argument for not shipping on the headline number alone. It is near-perfect on both command-parsing suites and near the floor on servant_room and merchant_bargain. Given a radio command it is a top-tier model; given a conversation to read it is not.

Pick a job, then a model

If you already know the kind of tool you are shipping, read that column, not the average. Figures are greedy, cases passed / cases in the suite.

Model Commands (20) Game log (20) NPC reaction (20)
NVIDIA Orchestrator 8B 20 / 20 19 / 20 20 / 20
Ministral 3 14B 20 / 20 19 / 20 19 / 20
Qwen 3.5 4B 18 / 20 19 / 20 16 / 20
Qwen2.5 7B 19 / 20 18 / 20 16 / 20
Bonsai 8B (Q1_0) 20 / 20 18 / 20 9 / 20
Llama 3.1 8B 17 / 20 12 / 20 14 / 20
Llama 3.2 1B 14 / 20 7 / 20 2 / 20
  • Commands is heist_command_per_action — several tools with different parameter shapes, player speech mapped onto a fixed vocabulary.
  • Game log is game_log_query — two tools, four filters, the player never names every slot. The hardest authored suite, and the closest to a companion that looks things up.
  • NPC reaction is npc_reaction — zero-argument tools. Nothing is scored except the decision (attack, trade, follow, leave, or stay quiet). This is the floor of the authored corpus.

A worked example: you want a companion that answers "where did I last see the orc?" by calling a query tool. Read Game log, not the overall column. Orchestrator, Ministral 14B, and Qwen 3.5 4B are 19 / 20 there; Llama 3.1 8B is 12 / 20 even though it is an 8B model.

For the schema and prompt shape that suite uses, see Design a tool-call prompt.


BFCL slice

300 generic single-turn cases from BFCL v4 (simple, multiple, parallel, irrelevance). All 21 catalog tool-call models, including Mistral 7B. Run bfcl_sweep_20260826_1204.

Model Exact accuracy
QVikhr 3 4B Instruction (Q4_K_M) 85.3%
Bonsai 8B (Q2_0 2-bit) 85.0%
Ministral 3 8B Instruct (Q4_K_M) 85.0%
NVIDIA Orchestrator 8B (Q4_K_M) 85.0%
Bonsai 8B (Q1_0 1-bit) 84.7%
Qwen 3.5 4B (Q4_K_M) 84.3%
Granite 4.0 H-Micro Instruct (Q4_K_M) 84.0%
Ministral 3 3B Instruct (Q4_K_M) 84.0%
Granite 4.0 H-Tiny Instruct (Q4_K_M) 83.3%
Ministral 3 14B Instruct (Q4_K_M) 83.3%
NVIDIA Nemotron 3 Nano 4B (Q4_K_M) 82.3%
Gemma 4 E4B Instruct (Q4_K_M) 80.0%
Qwen2.5 7B Instruct (Q4_K_M) 79.3%
Qwen2.5 3B Instruct (Q4_K_M) 77.3%
Qwen 3.5 2B (Q4_K_M) 72.7%
Bonsai 27B (Q1_0 1-bit) 69.0%
Mistral 7B Instruct (Q4_K_M) 62.0%
Qwen2.5 0.5B Instruct (Q4_K_M) 53.7%
Llama 3.1 8B Instruct (Q4_K_M) 46.3%
Llama 3.2 3B Instruct (Q4_K_M) 29.3%
Llama 3.2 1B Instruct (Q4_K_M) 10.0%

Eleven models sit between 82% and 85%. Below that the field drops off, and parallel is what does it: one utterance that should produce several calls. Llama 3.1 8B, Llama 3.2 3B, and Llama 3.2 1B score 0 / 54 parallel cases. Bonsai 27B is 10 / 54 there, which is why a 27B model ranks sixteenth on this corpus and near the top of the authored one.

If your game issues one tool per player line, BFCL parallel is the wrong number to optimise for. If a single line must open every door, it is the right one — and you need parallel_tool_calls = true on the node as well.


What these numbers do not tell you

  • VRAM and RAM. An 8B Q4 and a 27B 1-bit are different machines. See Estimate memory footprint.
  • Whether the template can render tools at all. That is a fail-fast at CreateAgent, not a score.
  • Multi-turn tool use. Both corpora are single-turn.
  • Your prompt and your schema. A model that scores 19 / 20 on the game-log suite still fails if your own tool has five required filters and no name for "unfiltered". Shape the schema first — Design a tool-call prompt.