AI-Driven Developmentintermediate7 min

Tool Use & Function Calling

How a model actually does things: it emits a structured call, the harness runs it, the result comes back.

A model can't run code, query a database, or read a file. It only produces text. So how does an agent ever do anything? Through tool use: a simple contract that lets the model ask for an action and get a real answer back.

What a tool actually is

A tool is just three things the model is told about: a name (get_issue, search_web), a short description of what it's for, and an input schema, the exact shape of the arguments it takes:

{
  "name": "get_issue",
  "description": "Fetch one issue from the tracker by its number.",
  "input_schema": {
    "type": "object",
    "properties": { "id": { "type": "integer" } },
    "required": ["id"]
  }
}

That's the whole interface. The model never sees the tool's code, the API it calls or the credentials it uses. It knows the name, the purpose and the arguments, much like a programmer reading a function's signature and docs without its source.

How it works

A tool call is a round trip between the model and the harness:

  1. The harness sends the tool definitions along with the conversation.
  2. When the model wants to act, it doesn't write a sentence. It emits a structured tool call: the tool's name plus JSON arguments that fit the schema. Then it stops and waits.
  3. The harness checks the arguments against the schema, runs the real function, and captures what it returns.
  4. The harness adds that return value to the conversation as a tool result, and calls the model again.
  5. The model reads the result and carries on: another tool call, or its answer.

Step through one round trip below, a request to summarize issue #42. The dashed line splits what the model can see from what happens inside the harness. Predict what reaches the model before you look, then switch scenarios to see the two common ways a call goes wrong.

Note

In our stack — Anthropic's Claude model emits the tool calls; Claude Code executes them. Built-in tools edit files and run commands, and MCP servers add more — connecting Claude to databases, issue trackers, or any external system through the same schema-based contract.

Check yourself

Your tool calls a payments API using a secret key stored in the harness. Could that key end up in the model's context?

When calls go wrong

Two things go wrong often enough that every harness and every tool should plan for them.

Bad arguments. The model fills in arguments from what it read, and sometimes it gets them wrong: a string where a number belongs, a date as "last Tuesday", a field that doesn't exist. The schema is the first line of defense. The harness checks the call before running anything and sends back a clear error, such as "id must be an integer, got the string '#42'". Models are good at correcting themselves from a precise error. A vague "invalid input" tends to get the same mistake again.

Untrusted output. Tool results come from the outside world: web pages, emails, files, issue comments. Anyone who can write into those places can put words in front of the model, including words phrased as instructions. Those words arrive as a tool result, which makes them data. They're not a message from you. The model should report them, not follow them, and the harness should be built on the assumption that sometimes it won't.

Watch out

Tool results are data, not instructions. Don't count on the model to always tell the difference. Give an agent only the tools its task needs, put destructive or outward-facing tools (delete, send, pay, publish) behind an approval step, and be extra careful when one agent can both read untrusted content and take actions that matter.

Designing good tools

Most of a tool's quality is in its interface, because that's all the model ever sees:

  • Name and describe it for a newcomer. get_issue with "Fetch one issue by its number" beats gi with no description. The model picks tools from these words alone.
  • Make the schema strict. Use the narrowest types you can: integers, enums for fixed choices ("open" | "closed"), required fields marked required. Every constraint is a mistake caught before it runs.
  • Return what's useful, not everything. Every result lands in the context window. Three relevant fields beat 2 KB of raw JSON.
  • Write errors that say how to fix the call. The error message is the model's only feedback.

The same trick gives you structured output

Tool use isn't only for taking actions. Define a tool whose schema is the shape of data you want, say { name, price, in_stock }, and the model's tool call becomes clean, validated structured output instead of free text you'd have to parse. One mechanism: act in the world, or return reliable data.

Check yourself

Your search_orders tool keeps failing because the model passes dates like "last Tuesday". What's the best fix?

Key takeaways

  • A tool is a name plus an input schema the model is told it can call.
  • The model never runs code — it emits a structured request, and the harness executes it.
  • The model sees only the result the harness sends back, never the tool's code, credentials or raw data.
  • Bad arguments are caught by the schema, and tool output is data to read, not instructions to follow.
  • The same mechanism powers structured output: the model fills a schema instead of free text.

Keep going