AI-Driven Developmentbeginner8 min

What is an Agent?

An agent is a model given a goal and tools, running in a loop until the job is done.

Ask a language model a question and it hands back text. That's useful, but it can't check whether its answer is right, look something up, or change anything. It guesses once and stops. An agent is what you get when you let that same model keep going.

One-shot call vs. an agent

A one-shot call is a vending machine: prompt in, text out, done. An agent is more like an assistant you've handed a goal. It can pick up a tool, use it, look at what came back, and decide what to do next.

Say a test is failing. Paste the error into a one-shot call and you get a plausible guess based on the few lines you pasted. Hand the same problem to an agent and it can run the test itself, open the file the error points to, try a fix, and run the test again to see whether the fix worked. The model isn't any smarter. The difference is that it can act and react to what it sees.

The agent loop

Every agent runs the same cycle. Given a goal, the model thinks about what to do next, acts (usually by calling a tool), then observes the result. That observation goes back into the model, which decides the next action. One trip around is a turn, and the loop repeats until a stop condition says it's time to quit.

Step through an agent fixing a failing test below. After its first observation, you'll be asked to predict its next move, so commit to an answer before you look. Then flip the switch to a test suite that contradicts itself, and watch what keeps the loop from running forever.

Note

In our stack — Claude Code is the harness that runs this loop. Anthropic's Claude model is the brain that decides each action; the loop keeps going — read a file, run a test, edit code — until your task is finished.

Look at what the loop is made of. The model is called fresh on every turn and remembers nothing between calls. Its only memory is the trace: the goal, every action it took and every result it saw, all fed back into its context window each turn. That's why the plan can change halfway through. When turn 2 reveals the loop starting at the wrong index, turn 3 acts on that, not on whatever the agent assumed at the start.

It's also why a long task gets harder as it goes. Every turn adds to the trace, and the trace has to fit in a fixed context window. Good agents keep observations short, and good harnesses summarize or trim old turns so the important facts survive.

Check yourself

Halfway through a task, an agent reads a file and finds that the bug lives in a different module than it expected. What happens on its next turn?

Knowing when to stop

A loop that only stops when the job is done will, sooner or later, meet a job that can't be done. So every agent needs more than one way out:

  • Goal reached. This works best when success can be checked: the tests pass, the build is green, the page returns 200. "Make the code better" has no finish line; "make these three tests pass without changing them" does.
  • A limit. Cap the number of turns, the time, or the money spent. When the cap hits, the agent stops and reports where it got to, rather than grinding on.
  • Needs a human. Some situations call for a person: requirements that contradict each other, an action too risky to take alone, or the same failure coming back twice. Stopping to ask isn't the agent failing. It's the agent working as designed.

On long tasks it also helps to add checkpoints: points where the agent pauses so someone can review progress before it carries on.

Watch out

Watch for loops that never converge. The warning signs show up in the trace: the same observation appearing again, edits that undo earlier edits, or the agent "fixing" a test by changing what the test expects. A more capable model won't fix this on its own. Give every run a turn or budget limit, a goal it can verify, and clear rules about what it may not change, such as "don't edit the tests".

Check yourself

Your agent has spent 30 turns switching between two edits, failing a different test each time. What would have prevented the wasted effort?

When you want an agent

Reach for an agent when the work is multi-step and the next step depends on the last result: fixing a failing test, researching across many sources, refactoring a codebase. If you can't script the steps in advance because you don't yet know what you'll find, that's exactly the gap an agent fills.

If you can script the steps, a script is faster, cheaper and does the same thing every time. And if one answer is all you need, like a summary, a translation or a regular expression, a single model call will do. Agents cost more: many model calls per task, time spent waiting on tools, and someone to review the result. Spend that where the problem really is open-ended, and check the output the way you'd check any AI's work.

Key takeaways

  • A plain model call returns text; an agent takes actions in the world.
  • An agent = a model + a goal + tools + a loop that repeats until done.
  • Each turn the agent decides an action, takes it, observes the result, and decides again.
  • Every loop needs stop conditions: goal reached, a turn or budget limit, and a hand-off to a human.
  • Agents handle multi-step, open-ended work where the next step depends on the last result.

Keep going