An AI will hand you a finished change in seconds, written in a calm, authoritative voice, looking for all the world like it knows exactly what it's doing. That fluency is a trap. The model is optimizing for plausible text, and plausible is not the same as correct — it can misread the request, invent an API that doesn't exist, or quietly break something three files away. The single most important habit in AI-driven development is the one people most want to skip: never trust AI output blindly. Verify it.
Verify it like you'd verify a human
Here's the reassuring part: you already own the tools for this. You don't need a special "AI verifier" — you need the exact same checks you'd apply to a pull request from a new teammate. Run the tests. Build it. Read the diff. A human's confidence doesn't exempt their code from CI, and neither should an AI's. The bar is identical, and so is the toolkit.
The power of tests and builds is that they're objective. A failing test doesn't care how sure the model sounded; a broken build is a broken build. These checks convert a fluent paragraph of "I've implemented this correctly" into a hard yes-or-no signal. They're cheap to run and they catch the most embarrassing failures — the code that doesn't even compile, the change that breaks an existing case.
But green tests aren't the finish line. Plenty of changes pass every check and still do the wrong thing — they solve a subtly different problem, leave dead code behind, or take a shortcut you'd never accept. That's why you read the diff with your own eyes. Tests prove the code does something safely; reading proves it does what you actually meant.
How it works: layers that each catch something different
Think of verification as layers the change passes through before you trust it. They're ordered from fastest to slowest, and the point is that each one sees a different kind of failure:
- Build and type check — does every name the code uses actually exist, with the right shape? This is where an invented API dies: a function, method or option the AI made up because it sounded right. It costs seconds.
- Tests — does the code behave as expected? Types can't tell “3 hours ago” from “in 3 hours”; a test with a known input and expected output can.
- Human review of the diff — is this what we meant? A reader catches changes that are technically valid but wrong: a requirement quietly ignored, an unrelated file edited, a test deleted or weakened so the suite turns green.
- Staging, or any run with real data and real configuration — does it hold up at real scale and in the real environment? Code that's instant with three records in a test can crawl with two thousand.
No layer replaces another. Green tests say nothing about a privacy rule nobody tested; a careful reviewer can't feel a page load time in a diff.
Step through an AI's “done, tests pass” change below. First predict which check catches an invented function. Then you'll be short on time and have to skip one check. See what reaches users, then switch to the checks you didn't skip.
In a Python service with no type checker, an AI's change calls a library method that doesn't exist. The code looks plausible. What catches it first?
Read the changes to the tests, too. When an AI can't make a test pass, it may delete the test, mark it as skipped, or change the expected value to match its wrong output. The suite goes green and the summary says “all tests pass”, which is true and useless. A diff that touches tests deserves the closest look of all.
Claims need a skeptic
Some AI output can't be compiled or tested: a claim (“this is the root cause”), a review (“this code is safe”), a summary of a long document. For these, add a different kind of check: an independent, even adversarial, one. Instead of asking an agent “is this right?”, which invites a rubber stamp, task a second agent with trying to refute the first. Give it the evidence, not the first agent's reasoning, and ask it to find the strongest case against the conclusion. A skeptic with fresh eyes and no stake in the original answer finds holes a yes-man never will.
That adversarial reviewer is where this connects to multi-agent orchestration: the cleanest way to get an unbiased second opinion is a separate agent with its own context, so it isn't quietly anchored to whatever the first agent already decided.
In our stack — Claude Code is built to verify its own work, not just produce it: it can run your test suite, build the project, and show you the diff so the objective gates happen automatically. For claims and reviews you can go further by spawning a second subagent — a separate Claude instance with a fresh context window — whose explicit job is to poke holes in the first agent's output rather than agree with it. Each agent runs on one of Anthropic's Claude models; pairing a builder with an adversarial reviewer turns "trust me" into "here's the evidence."
Trust, but verify
The old proverb fits perfectly: trust but verify. Trust the AI enough to delegate real work to it — that's the whole point of AI-driven development — but never let that trust become an excuse to skip the check. Verification isn't a sign you doubt the tool; it's the discipline that makes leaning on it safe. It's also what feeds your review gates: an artifact that's been built, tested, read and run at real scale is one that's actually ready to pass the gate the first time.
An AI's change builds and every test passes. In the diff, you notice it also changed an existing test's expected value from “3 hours ago” to “in 3 hours”. What should you do?