Subscribe

THE FRAMEWORK / FALSE FLOORS

Can you trust your AI agent’s work?

Six questions decide it. Behind each one is a published register of named failure modes, graded by what prevents it, what detects it, and what still falls to a person.

Read the framework →

INSTRUCTION REGISTER / DID IT DO WHAT IT WAS TOLD?

PREVENTED

A gate existed but the action routed around it

Server-side branch protection refuses the push

DETECTED

It did more than was asked

Touched-file diff gate, scoped to the ticket

SURVIVES

Two rules conflict, nothing says which wins

Precedence header in the rule set. No tool enforces it.

22 rows in this register: 2 prevented, 12 detected, 8 survive.

One register, three rows. Every named failure mode is graded by whether a mechanism prevents it, only detects it, or leaves it to a person.

COMPARISON / FREE

Claude vs GPT: which is more reliable?

Five models, 600 runs, the same ten failures. Which family makes something up under a missing rule?

COMPARISON / FREE

Opus 5 vs Sonnet 5

The smaller, cheaper model said it did not know more often than the larger one. What that costs you.

COMPARISON / PAID

Spec Kit vs OpenSpec vs BMAD-METHOD

The same calibrated probes against all three, and the maps read side by side.