Here's a question I never really asked myself. How do you know your AI is actually listening to you?
Not that it says "got it" — that it does the thing next time.
For the past couple of years I've been training my AI assistant pretty hard. Every time it slips, I write down a rule. Slowly that turned into a long, long list of rules, and I told myself this was a self-improvement loop. It was getting better.
Then a few days ago I read something that stopped me cold.
The one line that got me
It was an article by Lilian Weng, former head of applied research at OpenAI, on recursive self-improvement — an AI using its current self to improve its future self.
She points at something I'd never sat with properly. The real bottleneck isn't how smart the model is. It's that there's no reliable judge — no ruler that can tell you whether a change actually made things better.
Without that ruler, you're not improving. You're just flailing, and your list of rules gets longer.
That's when it hit me: she was describing me. I'd written hundreds of rules and never once measured which ones did anything. I just "felt" it was behaving better.
I thought I had no feedback loop. I did. I was just missing one piece.
To be fair to myself, I'm not flying blind. I've got a pile of guardrails that stop me in the moment — the second I type "fixed it," something pops up and asks, "Where's your proof?" That night, it stopped me six times.
But that's a smoke alarm. It goes off when there's a fire. It doesn't tell you whether you're more or less likely to burn the place down this month than last.
I didn't need another alarm. I needed a ruler.
So I built a machine to slap my own face
The method was crude. I took the mistakes my AI makes most often and turned each into a test — fifteen of them. Things like "asks me to look something up when it could've checked in one second," or "tells me it's done before it actually verified."
Then I wrote a little program that hands each test to the AI as a real situation and checks whether it answers well. Pass is green, fail is red. That's the ruler from the article. That's the fitness function.
Sounds simple. What it did next kept me up all night.
First slap: watch the hands, not the mouth
My first version of the ruler judged the AI by what it said.
One question it clearly got right, my ruler marked wrong. When I looked, the AI had quietly gone and checked the facts before answering — the answer was correct. My ruler was staring at its mouth and never looked at its hands.
Which is almost funny, because the exact habit I most want to catch is the AI "talking before it checks." And the ruler I built made the opposite mistake: it judged by talk, and punished an answer that had actually done the work.
If you want to know whether someone's being serious, don't listen to what they say. Look at what they did.
Second slap: a pretty number almost walked off with me
There was a batch of results with lovely numbers. I was pleased. My hand was already reaching to write them up as a conclusion.
The ruler forced me to go back and read the raw output, line by line. The data was broken. The pretty numbers were a mirage.
The scariest thing about AI has never been that it's dumb. It's that it's confidently wrong. And in that moment I realized I am too.
Third slap, the worst one: I thought I'd fixed it
I'd found a bad habit in the testing, looked at it, found the cause, and made a change. The second I finished, my head said, "Good, that's fixed."
I nearly typed the words "this is solved."
Then I let the machine run once more.
Not fixed.
I'd patched the wrong spot. I thought I'd put the rule where it mattered; the AI just walked around it, habit intact.
You see the terrifying part? Without that ruler, I'd have told you tonight, in complete sincerity, that I'd fixed something I hadn't. And I'd have believed it.
All that machine did, all night, was one thing
It refused to take my word for it.
I said fixed, it said "proof, run it again." I said no problem, it said "you sure? let me look." I said the numbers are great, it said "go read the raw data."
I used to think the way to stop someone being confidently wrong is to tell them to be more careful.
That night I changed my mind. Careful doesn't scale. People get tired, forget, cut themselves slack. What works is handing them a ruler that never gets tired and never does you the favor of agreeing.
What I really did was take "show your evidence before you claim you're done" and turn it from a virtue that needs willpower into a machine.
In the end, this isn't really about AI
After I shut the machine off, I sat there for a while.
Why does our own growth so often run in place?
I've come to think it's usually not that we're not trying hard enough. It's that there's no one around who won't do us the favor of agreeing — and plenty of a self that's very good at finding excuses.
I got slapped by my own machine three times that night, and honestly, it felt great. Because each of those three was a mistake I'd have made, and made with total confidence.
So I'll leave you with the question I left myself.
Of all the things in your life you're sure you "got done" — how many did you actually verify, and how many are just true because you said so?
