On a Thursday morning this month, we could not merge a one-line change.

Every gate we had built for ourselves was satisfied. The secret scanner was green. The container scan was green. The contrast guard, the claim guard, the change-ticket linker, the SBOM generator — all green. And the status check named after our AI code reviewer, the one that reads the diff and tells you what it thinks, was green too.

The merge button was still dead.

It took an embarrassing amount of time to find the number that explained it, because it was not on any of the screens we had been looking at. You have to ask the API a different question to see it:

The number nobody was looking at

reviews: 0. The reviewer’s status check had reported success on every pull request for days. In that time it had submitted exactly zero reviews.

Nothing had broken. Nothing had lied. The bot had done precisely what it was configured to do, and it had done it well — it posted a clear, genuinely useful summary of each change, every single time. The check went green because the bot successfully finished its work.

Finishing its work was never the same thing as approving the change. It had simply never occurred to us to check the difference, because for months the green check and the approval had arrived together, and we had quietly started treating one as the other.

The check was honest. Our reading of it was not.

This is worth sitting with, because the instinct is to look for the villain and there isn’t one.

The status check answered a question precisely: did the review service run to completion? Yes. It did not answer, and never claimed to answer, the question we were actually relying on it for: did a competent reviewer examine this change and accept it?

Those two questions look identical on a dashboard. They produce the same green circle in the same row. One is a statement about a process finishing. The other is a statement about a judgment being made. We had built a merge gate on the second and were reading the first.

What we thought green meant What green actually certified
A qualified reviewer examined this change A service finished its run without erroring
Defects were looked for and not found A summary comment was posted
Someone accepted the change An API call returned success
There is a record of a judgment There is a record of an execution

Read that right column again and notice that every line of it is true. The system was not malfunctioning. It was reporting, accurately, on the only thing it had ever been measuring.

Why it took days to notice

The reason this ran for as long as it did is the part worth generalizing, and it has nothing to do with code review.

Absence does not render. A failing check turns red and puts itself in front of you. A review that was never submitted produces nothing at all — no row, no icon, no warning banner. The interface for “a judgment did not happen” is a blank space that looks exactly like every other blank space on the page.

So the failure mode of automated oversight is not a red light that people ignore. Everyone notices red lights. The failure mode is a green light standing in for a question nobody asked, in a place where the missing answer has no shape.

We only found it because an unrelated rule — a branch protection setting requiring one approving review — happened to ask the right question and refused to proceed. A control we had installed for compliance reasons caught a governance failure that none of our monitoring caught. If that rule had not been there, we would have merged for weeks on the strength of a green check that certified nothing about the code, and we would have felt well governed the entire time.

This is the standard shape of AI oversight right now

Swap out the specifics and you have most of the AI safety and governance tooling running in production today.

In each case the artifact being surfaced to a human is a liveness signal about a service. The thing the human believes they are seeing is a decision about an action. These are different objects with different lifetimes, different failure modes, and radically different value in a room with an auditor in it.

Property Status signal Decision record
What it describes A service run A specific action
Names the policy version applied No Yes
States what was decided, not just that it ran No Yes
Survives the next run Usually overwritten Retained
Checkable by someone who distrusts you No Yes, if signed
Still answers questions six months later No Yes

Status signals are not worthless. They are excellent at what they do, which is telling you your infrastructure is alive. The mistake — ours, and nearly everyone’s — is mounting accountability on top of them, because they are the artifact that happens to be on the screen.

Three questions that break the illusion

You can test any AI oversight surface, ours included, with three questions. They take about five minutes and they are uncomfortable in proportion to how much the surface is worth.

1. When this goes green, what function returned success?

Not what the label says. What actually executed, and what would have had to happen for it to go red. If the honest answer is “the service responded,” you have a liveness probe wearing the costume of a control. That is fine, as long as nobody is making merge, release, or deployment decisions on it.

2. Show me the decision, not the status.

For one specific action, on one specific day: what was proposed, what was decided, and which version of which policy produced that outcome. If the system can only produce an aggregate — a pass rate, a score, a count — it has telemetry, not evidence. Aggregates cannot answer questions about individual cases, and every question that matters after an incident is about an individual case.

3. Could someone who does not trust us verify this?

This is the one that separates real evidence from a well-formatted log. A log is a claim made by the system about itself, and it is exactly as trustworthy as the system that wrote it. Evidence is something an outside party can check independently, without your cooperation, and reach the same conclusion. If verification requires logging into your dashboard, you have not produced evidence. You have produced a nicer-looking assertion.

We have the same bug. Here it is.

It would be convenient to end there, with the moral pointed outward. But we found a version of this in our own evidence format, and the honest thing is to put it next to the other one.

When our system evaluates a proposed agent action, it produces a signed record of what happened. That record contains a field called decision_status. Rehearsing a demo for a deliberately skeptical buyer, we walked through a scenario where a hostile action was stopped — the gate did its job, nothing executed — and pulled up the record.

The field read AUTHORIZED.

Correct, and badly named

decision_status records the authority decision at the moment the action was requested. Whether the action then ran is a different question, answered by different fields: executed, executed_count, blocked. The record was accurate. The top-line field was still the first thing a reader’s eye lands on, and it reads like an outcome when it is not one.

That is the same defect as the green check, produced by the same reasoning, in a product built specifically to stop this class of problem. A prominent field that is true about one thing and gets read as true about a bigger thing. We caught it by watching a hostile reader’s face, not by testing, because no test we would have thought to write asserts “a correct field is not misleading at a glance.”

Two things came out of it. The narrow fix was to say the distinction out loud before showing anyone the record. The real fix is a design rule we now apply to every field we expose: a name that can be misread as an outcome must either be an outcome or not be prominent. Accuracy is not the standard. Accuracy plus an accurate first impression is the standard, and the second half is the part that fails silently.

What a decision record has to survive

The reason we care about the distinction at all is that a decision record has to keep working under conditions a status check never faces. It has to hold up when the person reading it wants it to be wrong.

We test ours by attacking it. Take a valid signed record and flip the disposition from blocked to executed: the binding breaks and verification fails. Clear the list of violated principles: fails. Break the link to the previous record in the chain: fails. Recompute the content hash so the tampered payload hashes correctly, which defeats the obvious check: verification now fails at the signature instead. Re-sign the whole thing with your own key so hash and signature agree: it fails because the key is not the one the verifier pins.

Each of those is a different attacker with a different amount of sophistication, and every one of them has to end in a refusal rather than a plausible-looking record. That is the bar for evidence. A green check clears none of it, and was never designed to.

Concretely, a record that earns the word evidence states what was proposed, what was decided, which policy version decided it, under whose authority, and carries a signature that a stranger can check offline. Miss any one of those and you have something that is useful for debugging and useless in a dispute.

The part that should worry you

Here is why this is worth more than a war story about a merge queue.

Green checks are nearly free. A decision record costs something to produce: you have to decide what the policy is, evaluate it at the moment of action, capture the result, and sign it so it survives contact with someone who disagrees. Every incentive in a shipping organization pushes toward the cheap artifact, and the cheap artifact looks identical to the expensive one at a glance. That is not an accident of bad tooling. It is the economics.

Now scale it. Organizations are handing judgment to machines faster than they are building any way to inspect that judgment. The artifacts coming back are overwhelmingly status: uptime, pass rates, throughput, a wall of green. Underneath, the number of consequential decisions made per hour without any durable record of the reasoning is going up sharply.

For three days that cost us a merge queue. Nobody was harmed, and the story is mildly funny. Run the same misreading through an agent that moves money, denies a claim, or files something with a regulator, and the moment you need to reconstruct what happened, you will find you have a very reassuring history of services that ran successfully.

The test

Pick the most consequential thing your AI systems did last week. Not a category — one action, on one day. Ask for the record: what was proposed, what was decided, which policy governed it, and who authorized it. If what comes back is a dashboard, you do not have oversight. You have a status page with a good reputation.

We got lucky. A compliance rule we had installed for unrelated reasons asked a question our monitoring never thought to ask, and refused to move until it got an answer. That is what a control is: something that can actually say no, to you, on a Thursday, when saying no is inconvenient.

Everything else is a light that turns green.

See what a decision record looks like

EVE CoreGuard evaluates a proposed AI action against your policy pack before execution and returns a disposition with a signed, policy-bound evidence record. Inspect a real signed certificate at the verification portal, read the evidence architecture at EVE Proof, scan an agent’s declared authority at the Authority Lab, or talk to us about a governed pilot.

Frequently asked questions

What is the difference between a status check and a decision record?

A status check reports that a service ran to completion. A decision record states what was decided about a specific action, which policy version produced that outcome, under whose authority, and carries a signature that can be verified independently. A status check answers “is the machinery alive?” A decision record answers “what did we decide, and can you prove it?” Most AI oversight in production surfaces the first and is relied on as though it were the second.

How did an AI code reviewer report success without reviewing anything?

The status check reflected the completion of the bot’s run, not the submission of a formal review. The bot posted a summary of each change, which is what it was configured to do, and the check went green because that work finished successfully. No approving review was ever submitted, so a branch rule requiring one approval could never be satisfied. Nothing malfunctioned; the signal simply did not mean what the humans reading it assumed it meant.

Why did monitoring not catch a missing review?

Because absence has no user interface. A failed check renders as a red row that demands attention, while a judgment that never happened renders as nothing at all. Detecting it required querying the review count directly rather than reading the checks panel. This is the general failure mode of automated oversight: not a warning that gets ignored, but a green indicator standing in for a question nobody thought to ask.

Is an audit log the same as verifiable evidence?

No. A log is a claim a system makes about itself, and it is exactly as trustworthy as the system that wrote it. Verifiable evidence can be checked by an outside party who does not trust the issuer, without the issuer’s cooperation, and yields the same result. The practical test is whether verification requires logging into your dashboard. If it does, you have a well-formatted assertion rather than evidence.

What should a decision record contain?

Five things: what action was proposed, what was decided, which policy version produced that decision, under whose authority it was taken, and a cryptographic signature that a third party can verify offline. Omitting any one of them leaves an artifact that is useful for debugging and useless in a dispute, because every question that matters after an incident is about one specific action rather than an aggregate.

How can I test whether my AI oversight is real?

Ask three questions. When the indicator goes green, what function actually returned success and what would turn it red? For one specific consequential action on one specific day, what was decided and under which policy version? And could someone who does not trust you reach the same conclusion independently? Any surface that cannot answer all three is telemetry, which is valuable, but it is not a control and should not be gating consequential decisions.