Skip to content

Design a Tool-Call Prompt

How to write the tool schema and the ToolCall node's system_prompt so a small on-device model fills the right slots — and leaves the others alone.

The recipes come from iterating the Tryll-authored tool-call suites against the catalog (twenty models, greedy decoding). The finding that paid off was not cleverer prose. It was changing what the model is asked to say.

Tryll does not let you pick a tool-call format. The model's chat template renders the schema you declare. Your levers are the tool list, the parameter enums, required, the descriptions, and the system prompt. See Tool calling.

How claims are marked

  • Tested — measured on the authored suites; the effect repeated across the catalog.
  • Hint — real, but narrower: one suite, or a design rule rather than a score delta.
  • No effect / made it worse — we tried it; do not spend time here.

The failure that looks like a knowledge gap

Small models extract a filter the player did name very reliably. They fail at restraint. Asked "where did I see the orc?" with optional region / event_type / time_window, they invent a region rather than omit the argument — and they invent the first value in the enum.

On the game-log suite that pattern was the bulk of the error: specified filters were right about 97% of the time; an unspecified region was wrong 63% of the time, and Ironwood (then enum[0]) beat the other three regions 51 to 7. That is slot-filling pressure, not geography.

The same pressure shows up as a guessed door (left, the first door listed) when the player said "lock the door" without saying which.


Recipe 1 — name "unfiltered", put it first, require the slot

Tested. This is the change that moved the field.

Do not ask the model to omit a parameter. Give "the player did not narrow this down" a real value, list it first in the enum, and mark the parameter required. The thing a pressured model reaches for becomes the right answer.

Across the catalog, adding any this way took the game-log suite from 46% to 77% exact accuracy. The three fields that gained any dropped from 44–63% wrong (when they should have been unfiltered) to roughly the error rate the models already had on filters they were told to set.

{
  "name": "search_events",
  "description": "Find matching events in the activity log.",
  "parameters": {
    "type": "dict",
    "properties": {
      "entity": {
        "type": "string",
        "description": "Who or what the question is about. Pass \"any\" when the player did not name one."
      },
      "event_type": {
        "type": "string",
        "description": "Kind of logged event. Pass \"any\" when the player did not imply one.",
        "enum": ["any", "kill", "sighting", "dialogue", "loot", "quest"]
      },
      "region": {
        "type": "string",
        "description": "Where it happened. Pass \"any\" when the player did not name a place.",
        "enum": ["any", "Ironwood", "Riverwood", "Karhold", "Gutter Ward"]
      },
      "time_window": {
        "type": "string",
        "description": "When it happened. Pass \"any\" when the player did not give a time.",
        "enum": ["any", "today", "last_session", "last_week"]
      }
    },
    "required": ["entity", "event_type", "region", "time_window"]
  }
}

The system prompt has to say the same thing the schema does:

Every call sets all four filters. When the player did not narrow one down,
pass `any`. That is the correct value, not a fallback — `any` means "do not
filter on this". Never guess a region, a time window, or a subject the player
did not mention.

The same pattern applies to actuation: lead door and window enums with unspecified, require both parameters, and tell the servant that "unspecified" is the right answer when the master did not say which.

Your client then treats any / unspecified as "no filter / ask which one", instead of hoping the model leaves a JSON key out.


Recipe 2 — do not add a parameter you can live without

Tested. Made it worse.

"Where did I last see the orc?" is not expressible if the tool has no "most recent" knob. Adding a required sort argument so that line could be answered made every model we re-ran worse — not only on the new slot, but on the four existing filters. Counts on four models: 19→18, 19→15, 18→12, 15→9 out of 20.

Each extra required parameter adds to the pressure to fill every slot. Four filters was the ceiling on that suite, not an arbitrary choice. Rename tricks did not help either: swapping last_session so the word "last" would stop pulling a time window changed nothing. A query with no "last" in it still drew a time filter.

If a distinction cannot be said without a fifth argument, prefer a slightly wrong query over a schema the model cannot fill. Keep the tool list short for the same reason — see What makes tool calls reliable.


Recipe 3 — do not forbid the call you want

Tested.

A prompt that says "never call a tool with every parameter unspecified" will be obeyed — including on the line whose correct call is exactly operate_window(unspecified, unspecified) ("Do something about that window"). Half the field answered in character instead of calling, which is what we asked for and then scored as a miss.

Write the exception where you mean it: greetings and questions get no call; a vague order still gets a call, with unspecified filled in. Put that in the system prompt next to the enum, not as a blanket ban.


Recipe 4 — split tools on intent, not on wording

Hint.

Three tools that sound alike need a rule the model can apply in one step:

  • propose_price — they named a figure.
  • request_discount — they want a lower price and named no figure.
  • refuse_deal — they are walking away ("No deal.", "Forget it.").

"Not at that price, no." is not a refusal in that scheme — it is a discount request with no number. If you cannot say in one sentence why a line maps to exactly one tool, do not score it, and do not ship it as a distinct tool. Ambiguous splits train you to argue with the eval; they do not make the runtime call more reliable.

Parameterless classifier tools (attack_player, start_trading, end_dialogue) are a different job again. Models that extract arguments well still decline to fire a zero-argument tool and answer in character instead. That is capability, not missing instructions — every suite already tells the model when not to call. See Compare tool-call models for which catalog models hold up on that flavour.


Recipe 5 — descriptions are the prompt the model actually sees

Hint — this is also the concept-page advice; the eval did not contradict it.

The chat template inlines your description strings next to the JSON schema. "Pass any when the player did not name a place" on the region field does more than a long character biography in system_prompt. Keep the system prompt for routing (which tool, when to stay quiet) and put slot-filling rules on the slots.

Use temperature = 0 on the ToolCall node. Sampling is not a meaningful lever here; both corpora move about a point from greedy to 0.7.


What not to do

Temptation What happened
Leave optional args off the required list so the model can "just omit them" Models fill enum[0] instead.
Put the sentinel last in the enum (…, "unspecified") The guessed value is still enum[0].
Add sort / extra filters so every English distinction is expressible Accuracy fell on the new slot and the old ones.
Rename values to dodge a word in the utterance No change.
Ban "all unspecified" in the prompt to stop chit-chat from calling It also stops the vague order you wanted.

Abstention ("Hello", "Nice weather") is mostly model quality. A clear do-not-call line in the prompt is necessary and not sufficient; weaker models still grab a tool whose argument matches a noun in the sentence.