The final essay in the series asks how much autonomy an agent has earned, using the same graduated schedule that governs a new human hire. It closes the series' two-codes question: the conduct converges, but the way you verify it never does.
This series opened with a question – can a worker that is not human be brought under a workable code of conduct? The previous essay answered it with a register: every way the worker breaks the workplace's unwritten deal, written down and graded by what answers it. All failures mapped to three classes: refused outright, caught every time, or nothing – in which case it falls to a person to judge.
So you now know how this worker fails, in greater detail than for a human hire. But you still do not know how much work you can safely give it: how much autonomy a worker like that gets, on whose say-so, and when it gets more.
Answering that requires a comparison: the agentic worker versus a human worker, the one you already know how to trust.
Before we run the comparison, we need to set the rules, because a sloppy comparison flatters whichever side you already preferred.
One correction runs against the agent: it compresses time. A human employee applying a wrong rule makes perhaps a dozen mistakes before Friday. An agent executes the same bad rule 400 times before morning review. A silent failure at machine speed is a different object from the same failure at human speed.
The other correction runs in the agent's favour: agents delete whole classes of human failure. An agent does not embezzle for personal gain – it has no personal gain. It does not hold grudges, play politics, harass a colleague, or pad expenses. You can delete the entire wing of your control environment built to thwart motivated insiders.
One caution: prompt injection reintroduces an adversary with motive, that is, an attacker speaking through the agent's inputs. That is real, and it is governed by the security stack, not by this essay. But the everyday failure profile, the one you live with between attacks, is motiveless.
So the fair ledger reads: alien capabilities – recall, speed, parallelism, cost, a worker who never sleeps and never resigns – bought at the price of alien failure properties, against human capabilities bought at the price of motivated failure, fatigue and turnover. It is a ledger, not a verdict: a tally of what you gain and what you risk with each worker, not a ruling on which one is better.
The ledger sets up the real question: what do you do with a motiveless worker that fails silently?
Professional practice has already dealt with unmotivated silent failure, in various guises. Trained and well-rested airline pilots still forget to lower the landing gear. So the airline industry stopped treating memory as a control and built the checklist.
Surgeons fought their checklist for a decade, then the complication numbers came in. The bank reconciliation exists because even honest bookkeepers copy numbers wrongly – $1,523 becomes $1,253 – so accountants check the books against the bank statement rather than trusting anyone's arithmetic. The shift handover exists because the night shift cannot remember what the day shift learned: the knowledge was never in their heads to forget, so it is written down and handed over.
We have a long, successful tradition of wrapping fallible workers in an exoskeleton of mechanism. These mechanisms always have the same goal: stop asking the worker to be reliable, and make the system reliable around the worker.
Seen from that tradition, the last essay's 11 alien-shaped rows stop looking alien at all. A worker whose memory resets every session is not a broken colleague – it is shift work, and the kit is the shift log, the start-of-shift briefing and the handover note, which in register language are the session-start digest, the auto-loaded rule file and the decision log re-injected after compaction.
A fan-out of subagents with no gather stage is a site with no practical-completion certificate. The register does not ask you to forgive the agent for its amnesia, any more than aviation asked passengers to forgive pilots theirs. It asks you to install the thing that makes the amnesia not matter.
Which leaves the ladder. No workplace decides trust in one go. Trust is handed out gradually, on a schedule. A new hire does not get the corporate card on day one. They get probation: a graduated release of autonomy, gated by demonstrated performance, granted per decision domain – trusted with the card and not with hiring, trusted to draft and not to send.
Probation is the human-native answer to exactly the problem the first essay's study names: a worker whose capability boundary you cannot see from their fluency.
A genre has grown up around this analogy – KPMG publishes an agent lifecycle, Finextra runs a hire-to-retire series with an instalment on your AI's probation period. But the analogy doesn't survive contact with the reality of an agent. Nothing binds an agentic worker through consequences.
Probation works because the employee wants to pass it. An agent cannot want to pass. So its rungs can only be read off mechanism: not "has it behaved for 30 days" but "which failures in this decision domain are prevented or detected, and which still survive".
Unattended runs occur only when there are no rows open in Authority and Recovery that are graded as "survives". Send-capable only when the outbound gate is installed, because nothing inside a session can un-send an email.
The autonomy ladder is real, but it is derived – read off an inspectable register, domain by domain – or it is theatre.
The human side of the schedule is broken too. In the experiment that defined algorithm aversion, people watched a statistical model beat a human forecaster at making predictions. They were still less likely to bet on the model than people who had never seen it perform at all (Dietvorst, Simmons and Massey, 2015).
Watching it win did not make them trust it. Good performance does not buy an agent trust the way it buys a person trust – which is one more reason the ladder cannot run on impressions. It has to be read from something inspectable.
One thing the ladder must never become is a single number. A trust score out of 100 averages what must not be averaged: an agent with a perfect audit trail of wrong work is not 70% trustworthy – it is auditable and wrong. Averaging is how a well-written wrong answer gets promoted to a dashboard.
So the two-codes question this series opened has an empirical answer. (With the caveat that this framework is early release, based on a single repository.)
At row level, the class of failures that are both alien and unclosable is empty. All 11 alien-shaped failure modes are closable by a named mechanism. And the 43 rows that survive everything split close to even. Of those, 24 are the shared judgement calls no control ever closed for people either – which of two rules wins, whether the finding is real, who calls the rollback. The other 19 survive only because nobody has built or installed the fix yet: an alien mechanism still unbuilt, or a human-workplace institution not yet wired in for this worker.
One code of conduct is reachable. Strain out the alien failure modes, mechanism by mechanism, and what remains is a worker who fails the way humans fail, judged by the code you already run.
But two codes of trust are permanent. A human earns autonomy through a deterrence-backed track record – the desire to pass is real and operates as a guide for behaviour daily. An agent's autonomy can only ever be read off a mechanism. The conduct converges, but the method for verifying that conduct never does.
The irony of the fool is that, according to the archetype, they appear to have a licence to operate outside the boundaries. Unlike anyone else in court, they can impart wisdom to and insult the king, and get away with it. But in fact they observe a code of their own.
In technology terms, the fool's licence is scoped and revocable: a defined latitude, granted for a defined function, and withdrawn if the fool oversteps. Five hundred years later, that is the ladder that applies to agents – a governed duality, autonomy granted per domain to a worker who cannot be trusted the ordinary way.
In this essay's terms: aim for the mapped floor, not the convincing one. The registers do not promise a floor with no hollow boards. They promise that every hollow board is either barred, or alarmed, or marked – because a "survives" row that is written down, published and owned is no longer a false floor. It is a marked opening, and you can cross a marked floor safely.
Bruce Schneier argued in his 2023 essay "AI and Trust" that AI trust is a question about systems, not persons – argued at the scale of societies and regulation.1 This is the same argument at the scale of one worker, one production repository, one register row at a time, with the rows published so the argument can be checked.
With the right guards in place, the distance from agent to digital employee is shorter than one would expect. Because we never fully trusted our human colleagues either. We trusted the system around them. And now we just have to extend it.
The process is straightforward. Identify a failure, categorise it, and look for a control, and build the code of conduct one check at a time.
Bruce Schneier, "AI and Trust" (2023), also published in Belfer Center and Harvard versions. Its argument is that we confuse interpersonal trust with social trust, and that AI should be made trustworthy through systems and regulation rather than through anything resembling character. Recorded in this project as a neighbouring statement rather than a source, distinguished by scale – society and regulation there, one worker and one repository here. Provenance: essay-citation-and-reference-plan-2026-08-11.md and research/46-essay-prior-art-sweep-2026-08-11.md, which names it as one of two nearest prior statements. No verbatim quotation is used and none may be added without a full-text pass – the sweep worked from abstracts. ↩
Interviews, registers and essays. Free.