Tool calling is the mechanism that lets a model do more than generate text: instead of answering directly, it can emit a structured request to run a function, get the result back, and continue reasoning with that result in hand. The model never executes anything itself — it only ever produces a description of a call. Your code is what actually runs it.
What Tool Calling Actually Is
A tool definition is a name, a description, and a JSON schema for its inputs, sent alongside a request. When the model decides a tool would help answer the request, instead of (or in addition to) generating text, it emits a tool_use content block: a tool name and a JSON object matching that schema. The response stops there — the model is waiting for you to actually run the tool and hand back a result before it continues.
Defining a Tool
The schema is the entire interface the model has to work with — there’s no additional documentation it can consult mid-request. A vague description or an under-specified schema produces malformed or wrong calls even when the model’s underlying reasoning about what to do is correct.
{
"name": "get_order_status",
"description": "Look up the current status of a customer order by order ID. Returns shipping status, estimated delivery, and last known location if in transit.",
"input_schema": {
"type": "object",
"properties": {
"order_id": {
"type": "string",
"description": "The order ID, formatted like ORD-12345."
}
},
"required": ["order_id"]
}
}
Every field in that schema is load-bearing. The description tells the model when to reach for this tool versus a different one; the parameter description tells it what format to produce, which matters more than it sounds — a model given no format hint will guess, and guesses are where malformed inputs come from.
The Round Trip
The critical detail is that the tool result goes back as a new message in the conversation, not as a special side channel — from the model’s perspective, seeing a tool result is just more context it reasons over, exactly like it reasons over anything else in the conversation history.
claude-sonnet-5Tools available: get_order_status.
User: Where's my order ORD-88214?Output
[tool_use: get_order_status(order_id="ORD-88214")]
→ {"status": "in_transit", "eta": "2026-08-09", "location": "Regional hub, Denver CO"}
Your order ORD-88214 is in transit and expected to arrive by August 9. It's currently at the regional hub in Denver, CO.
Beyond the JSON schema itself, tool descriptions benefit from stating not just what a tool does but when to prefer it over a similar one. If an application has both search_orders and get_order_status, a description that only says “looks up an order” leaves the model to guess which is appropriate for a given phrasing of a request. Naming the disambiguating condition explicitly — “use this when you already have an exact order ID; use search_orders when you only have a customer name or date range” — resolves that ambiguity at definition time instead of leaving it to be guessed wrong at call time.
Parallel Tool Calls
A single model response can contain multiple tool_use blocks at once, when the calls don’t depend on each other’s results. Execute them concurrently and return every corresponding tool_result in one message, not split across several — returning them one at a time trains the model, over the course of a conversation, to stop batching independent calls together, which costs you latency on every subsequent multi-tool turn.
const toolUses = response.content.filter((b) => b.type === 'tool_use');
const results = await Promise.all(
toolUses.map(async (block) => ({
type: 'tool_result' as const,
tool_use_id: block.id,
content: await executeTool(block.name, block.input),
})),
);
messages.push({ role: 'user', content: results });
Designing Tools the Model Can Use Reliably
| Design choice | Why it matters |
|---|---|
| Narrow, single-purpose tools | Easier for the model to pick correctly among several options |
| Descriptive names over generic ones | get_order_status beats query, especially with many tools defined |
| Return structured errors, not silent failure | Model can reason about and recover from a clear error |
| Idempotent where possible | A retried call after a timeout doesn’t double-execute a side effect |
| Keep required parameters minimal | Fewer chances for the model to omit something needed |
Error Handling
A failed tool call should come back as a tool_result with an error flag, not be dropped from the conversation. If the model doesn’t see that a call failed, it has no reason to try a different approach — from its perspective, the call is still pending. Returning a clear, actionable error message (not a raw stack trace) often lets the model retry with corrected arguments or fall back to a different tool without any code-side intervention.
Distinguish, in the error content you return, between a validation failure the model caused (a malformed order ID it can fix by trying again) and an infrastructure failure it can’t do anything about (the downstream service is down). Collapsing both into a generic “tool failed” message pushes the model toward retrying a call that will keep failing for reasons entirely outside its control, burning turns and tokens on retries that were never going to succeed.
Takeaway
Tool calling is a structured request-response protocol layered on top of ordinary conversation turns, not a separate execution channel — every safety property has to be built into the tool implementation and the surrounding code, because the model only ever proposes a call. Design tools the way you’d design a small, well-documented API: narrow scope, clear errors, and idempotent where a retry is even remotely possible.