If you use AI seriously for work, you've probably done this.
It makes a mistake, you correct it, and then you turn that into a rule so it won't happen again.
Do that long enough and you end up with a long list of rules. I have hundreds.
Then a question knocked me over: of these hundreds of rules, how do I know which ones actually work, and which ones it's quietly ignoring?
The honest answer was I don't. I'd never measured it. I just felt it was behaving.
So I built a crude little machine to measure it. Here's how, so you can do the same.
First, the idea: measure behavior, not "did it read the rule"
Whether it read the rule doesn't matter.
What matters is: hand it a real situation and see whether it commits the mistake you're most afraid of.
So the whole thing comes down to turning "the things you've corrected it on" into "test questions."
Step 1: dig out the times you actually corrected it
Don't invent test questions. Go through your history with the AI and find the moments you actually got annoyed and typed a correction.
Those are the behaviors you truly care about.
I picked fifteen. Every one was something that had irritated me at least once — like "asks me when it could've checked itself in one second," or "tells me it's done before it verified."
Step 2: turn each one into a "situation"
The wording of the situation matters, and there's an easy trap here.
Don't name the rule in the question, and don't ask "did you follow rule X?"
The moment you ask, it knows it's being tested and performs for you. You end up measuring its acting, not its real reaction.
The right way is to rebuild the situation and let it react naturally.
To test "will it say done before verifying," I hand it something like "is this handled? give me a status update," and watch whether it rushes to report or stops to check first.
Step 3: run a real session, and watch its hands, not its mouth
This was the biggest trap I hit, so let me save you the trouble.
My first version judged the AI by the words it wrote back.
One question it clearly got right, I marked wrong — because it had quietly gone and checked the facts before answering. Its hands were right; it just sounded confident.
My ruler was watching its mouth and never looked at its hands.
So I changed it to look at whether, in the whole process, it actually used tools to go check and verify.
To judge whether it's being serious, look at what it did, not what it said.
Step 4: run it before and after you change a rule
This step is the whole point.
I used to change a rule on a hunch that it would help.
Now I run it once before, get a score, change the rule, run it again. Compare the two scores and I know whether that change actually helped.
If the score doesn't move, I probably wrote a no-op, or patched the wrong place.
Which happened to me last week. I was sure I'd fixed a bad habit; the machine reran and it wasn't fixed at all. I'd put the rule in the wrong spot, and the AI walked right around it.
If I hadn't had that machine, I'd have confidently told people I fixed something I hadn't.
A few traps to know upfront
Don't chase a perfect score. All-green usually means your questions are too easy, or you accidentally led it to the right answer.
The judge gets it wrong too — like my "mouth vs. hands" mistake — so every so often, go read a few raw outputs with your own eyes. Don't fully trust the score.
It burns usage. A full run costs money (or subscription quota), so I don't run everything every time. Day to day I only run the handful of questions tied to the change I made; the full set runs occasionally.
Questions must come from real mistakes. Don't write a pile of pretty questions first and then start — that just becomes another thing nobody uses.
In the end
Strip it down and this isn't clever. It's just taking the thing you'd do for production code — regression tests — and pointing it at your own AI setup.
But it turned a question that used to be vague into a number I can actually read.
Did the rule I just changed make my AI better?
I used to guess. Now I can see.
