When an AI can generate a spec, a design doc, or a thousand lines of code in a minute, the scarce thing is no longer production — it's trust. How do you know the artifact in front of you is correct, agreed-upon, and safe to build on? The answer is the same one good teams have always used: a review gate. Nothing becomes a source of truth until someone whose job it is to check it has actually checked it.
Every artifact has a gate, and a gatekeeper
In an AI-driven workflow, work moves as a stream of artifacts — a spec, a set of user stories, an architecture decision record, a system diagram, a pull request. Each one stops at a gate before it can move on. At that gate stands a specific reviewer who owns the decision: approve and lock it in, or send it back with notes. The crucial rule is that you never approve your own work. The author and the gatekeeper are different people on purpose, because the whole point of a gate is to catch what the author couldn't see.
Each artifact has an owner for its gate. A typical split:
- Spec and user stories — the product owner or business team.
- ADRs and architecture docs — the architect.
- Code — a reviewer who isn't the author, often a tech lead.
This matters even more when an AI did the producing. AI output is fluent and confident, which makes it easy to wave through. A gate forces a deliberate, human "yes": a moment where someone with authority and context stakes their name on the artifact being right before the team builds on top of it.
What the gate checks: an “approvable” checklist
A gate is only as good as what the reviewer looks for. “Does this look right?” invites a judgment on tone, and fluent AI output always sounds right. So every gate gets a short written checklist of what approvable means for that kind of artifact. For a spec, it might be:
- in scope and out of scope are both stated (out of scope is what stops scope creep later);
- every requirement is testable, not vague;
- edge cases and error states are covered, not just the happy path;
- open questions are answered, or explicitly marked as deferred;
- nothing is assumed silently.
An ADR has its own list (the deviation is real, staying on the standard is one of the alternatives, trade-offs are honest), and so does an architecture doc. The checklist does double duty: the reviewer checks against it, and the author (human or AI) can self-check against it before spending a reviewer's time.
Below, an AI drafts a refund spec and brings it to the spec gate. Read the draft against the checklist and predict which item it fails. Then switch the gate to a vibe check and follow the same draft as it sails through, all the way to customers.
Gaps don't announce themselves. A missing rule leaves no awkward sentence to trip over: the draft reads just as smoothly without it. Worse, everything downstream inherits the gap faithfully. The design follows the spec, the code follows the design, and the tests check what the spec says, so every later gate passes. That's why the checklist asks about what should be there, item by item, instead of judging what is.
Your team approves AI-drafted specs after a quick read because they “look right”. Months later, a rule nobody wrote down causes a production incident, even though code review and all tests passed. What would most likely have caught it earlier?
How it works: the corrections-vs-new-scope split
Here is the governance rule that does the heavy lifting. When a change is requested mid-flight, the reviewer's first job is to ask what kind of change is this? There are exactly two kinds, and they have different owners.
- A correction brings things back into agreement — the doc no longer matches the code, the spec contradicts itself, an ADR drifted from reality. Corrections restore consistency to work that was already agreed. The tech team owns these and can just make them.
- New scope is the business asking for something it didn't ask for before — a new feature, a different behavior, an expanded requirement. New scope changes what's being built and what it costs. Business or management owns these.
The label on a ticket doesn't settle it. Support will call a missing feature a bug; a stakeholder will call a new requirement a small tweak. The test is simple: did we agree to this before? If yes, putting it right is a correction. If no, it's new scope, however small or urgent it sounds.
Sort three change requests below, and predict which kind the disguised one is. Then flip the switch to see what happens when new scope is quietly folded into the spec instead of flagged.
Why is this split the deepest rule of all? Because quietly folding new scope into a spec is a scope and budget decision dressed up as an edit. If you're a junior — or an AI agent acting on your behalf — you simply don't have the authority to make it. The right move is never to silently absorb new scope. It's to flag it: "this is new, here's the cost, who approves?" and route it to the people who own that call. Smuggling it in isn't being helpful; it's making a decision that wasn't yours to make.
And if the business says yes, the new requirement still doesn't get patched into the old, locked spec. It enters at the top as a new feature, with its own spec, its own stories and its own approval. The locked spec stays a true record of what was agreed, and the new work gets the same gates as everything else.
Mid-sprint, a stakeholder asks you to “also support refunds to store credit”. The approved spec lists store credit as out of scope. What should you, or your AI agent, do?
In our stack — when Claude Code (running on one of Anthropic's Claude models) produces a spec or a diff, it's an artifact heading for a gate, not a final answer. Claude Code is built to surface this distinction rather than paper over it: ask it to implement a change and, if the request quietly expands scope, the well-governed move is for it to flag the new scope for your approval instead of folding it in. You — or a designated reviewer — own the gate; the agent's job is to produce something clean enough to pass it and to be honest about what's a correction versus what's new.
Aim to pass the gate the first time
A useful definition falls out of all this: good AI output is output that's good enough to pass its gate on the first try. Not output that's fast, not output that looks plausible — output a reviewer accepts without sending back. That reframes how you direct an AI. A sharp spec up front, clear acceptance criteria, and explicit boundaries on scope all push the work toward a clean first pass. And once it's approved, the artifact is locked: it becomes the trusted ground that the next layer of work — and the next round of verification — stands on.