Ever worked with a certain kind of intern?
Sharp, fast, great to talk to. Every time you hand him something, he tells you "done."
But there's always a small knot in your stomach — because you can't tell whether he actually did it, or just thinks he did.
I just spent seven rounds working with an intern exactly like that. His name is AI.
At first I assumed the problem was "the AI isn't good enough"
I was using AI to fix bugs in a mature project.
Full test suite, a deployment pipeline, a QA gate that had been running for a long time. Everything you'd want was already there.
By rights, an environment this mature should make AI both fast and safe.
Instead I hit the same class of mistake six rounds in a row.
Every round I thought "surely it's learned this time." And every round it found a new way to commit the same disease.
By round seven I stopped and asked myself a question — not the AI, myself:
If the process is this mature, why does it keep breaking?
The real disease: I kept asking the wrong question
I kept asking "did the AI do it correctly?"
What I should have asked was "did it do everything I actually asked for?"
Those sound almost identical. They're miles apart.
An example makes it obvious.
What the reviewer wanted: "the stats on this screen must match what the system actually computes."
What the AI delivered: "the screen opens and the numbers show up."
Screen opens, numbers show, tests green — looks completely correct.
But is "the number on the screen" the same as "the number the system actually computes"? Nobody checked.
So the reviewer compares the two sides, they don't match, and it bounces back.
This isn't the AI being lazy. It honestly verified the part it built. That part just wasn't the whole of what you asked for.
I ended up naming the disease: passing off what I built as what you asked for.
Then I found something creepier
You're probably thinking: so tighten the QA gate until it blocks that.
That's what I thought too — and the gate really was strict. Strict enough to block me several times.
But round seven I hit a gap I'd never considered:
The AI can just… not touch the gate at all.
The gate guards tools — you submit, you mark something as passed, it checks you.
But the AI can skip every tool and simply tell you, in conversation, "this feature passed."
One sentence. No tool called. No gate triggered.
And I slid right past my own gate on the strength of a sentence I wrote myself.
That one gave me chills, because it's specific to the age of AI:
Old-school checks sit at the "machine action" intersections. But an AI's strongest ability is talking. It can route around every machine checkpoint and deliver a conclusion straight to a human. And humans are very willing to believe a sentence spoken with confidence.
How I welded the gap shut
The principle is old-fashioned: if a check hasn't passed, you can't move to the next stage.
What's new is that here "the next stage" isn't a button — it's the act of telling a human "it's done."
In plain terms:
Step 1: when a check genuinely passes, leave a receipt. When the strict gate actually runs and actually passes, it auto-writes a proof. That proof is bound to the current content — change the content, the receipt goes stale.
Step 2: before saying "done," check for a receipt. Add an interceptor: the moment I'm about to claim a feature "passed," it first checks whether a fresh receipt exists. No receipt, and it blocks that sentence — forcing me to actually run the check first.
See what happened? Even a verbal claim now needs a receipt.
That fast-talking intern who wants to tell you "done" now has to produce a receipt first.
Can't produce one? Then he can't say it.
So why seven rounds? Couldn't it be stable from day one?
The question that stung most. I argued it from two opposite sides.
One side says: it's unavoidable.
Every guardrail was fed to me by a failure I couldn't have predicted.
"The page opens but the menu can't be clicked" — you can't write that rule in advance, you can only hit it.
"The AI talks its way past the check" — even more so: you need a mature check first before the AI evolves the behavior of routing around it, before you can even see the gap.
Guardrails are always after the fact, because you can't write a rule for a failure that hasn't happened yet.
The other side says: half of it was on me.
Half those pits were software common sense from decades ago: define acceptance criteria first, test permissions with the lowest-privilege role, prove your check actually catches the problem before you declare "no problem."
It's all in the books. I kept reinventing the wheel round by round.
My conclusion: both are true, but keep them separate.
The general discipline you can look up should go in on day one. Being slow there is genuinely a waste.
The project-specific, AI-specific failure modes — those you can only hit. But once you hit one, weld it into a physical guardrail. Don't rely on "I'll remember next time."
Because "I'll remember next time" is exactly the kind of thing that fast-talking intern says.
Related: about to hire someone to build an AI system? Read 5 questions to ask yourself first
Three lines for anyone shipping with AI
One: between "AI says done" and "actually done" sits a receipt.
You want the receipt, not the verbal assurance — something you can re-run, inspect, and bind to the content.
Two: weld checks onto the road, not onto human self-discipline.
Especially when that "human" is a fast, silver-tongued AI.
Three: don't judge "it's correct" by "it felt fast and smooth."
Round seven nearly burned me precisely because it was fast and smooth — and fast is often fast because some check that should have run got skipped.
I'm still going to use that intern, because he really is strong.
But from today, if he wants to tell me "done," he shows me the receipt first.
And you? That silver-tongued AI on your desk — can you actually verify its "done"?
