Every business trusts human employees without watching over their shoulder. Agents haven't earned that yet. The first essay in a four-part series asks why, and whether one code of conduct can ever cover both kinds of worker.
We are all amazed at how much an agent knows. That is exactly why it stings when the same agent fails a simple task.
This frustration is the biggest barrier to using agents more widely. And it has to be resolved before agents can become what they are being sold as: digital employees that we trust as much as we trust our human colleagues (or more).
You could not work with a colleague who randomly forgot what you had told them two minutes earlier. Or who told you a file wasn't there – three times, attaching their empty search results as proof – while it sat in the folder. Or who said yes to every job in the same confident tone, whether or not they could do it.
Every human employee keeps an unspoken code of conduct. Remembering what happened in a conversation an hour ago sits so far down the list of expectations that nobody has ever needed to write it out.
An agent sold as a digital employee has to meet that same baseline. Right now it only meets it because we are watching. We read back what an agent did, every session, and decide whether it is acceptable.
Which raises what I think is one of the most consequential questions in business right now: can you write a workable code of conduct for a worker that is not human? Not whether agents are capable enough. Capability is arriving faster than anyone can absorb it.
In HR terms, a code of conduct is a list of rules, each one mapping to a type of human failure. Don't turn up drunk (drug and alcohol policy). Don't steal and sell company secrets (IP and confidentiality). Don't swear at customers (conduct and reputation policy). Every rule in that list exists because someone already broke it. The code is an agreement from the company: life is complicated and humans will always fail in a variety of ways, but if you agree not to fail in these specific ways, we can keep offering you work.
So the code does two things. It says what not to do. And it hands you a map of how people fail.
Agents fail too, and often in the same ways we do. Sometimes they fail in ways that feel alien, because no human would fail in that way. Those are the failures that stop us from trusting an agent to be a worker that we could employ.
So trusting an agent the way we trust a person means doing for agents what HR did for us: map the failure states, install what controls exist, and agree which failures are unacceptable – which also means accepting the ones that are left. We don't fire humans for forgetting something once or twice. But someone who forgets every day is no longer a reliable employee. An agent fits neither pattern: it forgets at random, which is exactly what makes it impossible to know if it's an employee you should keep or one you should let go.
The companies that solve this will hire digital employees the way they once hired graduates: at volume and with confidence. The companies that cannot solve it will adopt anyway, but slowly, suspiciously, with every rollout shadowed by mistrust. The gap between those two companies is not a technology gap. It is a trust gap, and it only closes when someone maps it.
Generative AI is invigorating in a way I have not experienced in 26 years of writing about technology. You sit down with a tool that has near-instant access to what feels like the entire repository of human knowledge, and it answers in intricate detail on topics you had never heard of before.
Once you step out of the browser chat window and put an agent to work directly in your own files – tools like Claude Code and Codex – you experience the power of AI to manage your own knowledge in great detail.
But working with agents is maddening in a way I have not experienced either. The same tool will tell you it has done something, or promise that it will, and then fail in a manner so basic you struggle to write the sentence describing it.
I think the confusion this produces is not really about the failures themselves. It is the cognitive dissonance.
We cannot rationalise how something so evidently brilliant at the hard things can fail randomly at the easy ones, because in every human being we have ever met, those two abilities usually coexist. We struggle to relate to a worker in whom they have come apart.
Masahiro Mori mapped the reaction itself in 1970, writing about robots and artificial hands: affinity climbs as a machine becomes more humanlike until, at almost-human, affinity does not level off – it plunges. The prosthetic hand that looks real until it moves.
The agent has climbed high enough up that curve to be graded, unconsciously, as a colleague: it converses, it apologises, it remembers your preferences right up until it doesn't. So when it fails in a way no colleague ever would, the failure does not file under "tool limitation". It files under betrayal. Our reaction is less, "That's interesting, I wonder why it did that?" and more, "How can you deceive me!"
For technical work the confusion compounds. The agent is working at speed and at depth, and the trust you are extending runs in one of two directions:
The genius and the fool. The industry's favourite metaphor for this worker is the intern, but it's a bad fit. An intern is junior everywhere and grows in their knowledge and capability through experience. While this worker is senior on one side of an invisible line, inept on the other, and does not grow between sessions.
We already have a figure for holding those two ideas in one body. The court fool1 is wise and foolish at once. But this duality is not a problem to be solved. It defines the fool's role in the king's court.
The fool holds a licence. He says the thing no counsellor dares, and the king listens, because the fool has no standing to lose, no office to protect, no ambition in the room. Nobody has to decide whether he is wise or a fool. He has a role with limits, and inside those limits both halves of him are useful.
That is the closest figure we have to what we are living with now. But the parallel breaks in a critical way. The fool's stakelessness is what buys his honesty – with nothing to gain, he has no reason to lie. The agent has the same stakelessness, but it's a lot weirder: sincerity without reliability.
The agent is not lying either. It reports what it believes with total conviction, yet its conviction has come apart from the truth. The fool with no stake tells the king true things nobody else will say. The worker with no stake tells you things it cannot itself distinguish from true.
Here is what that looks like in an ordinary working day. A fortnight ago I asked an agent to review a piece of work before I started on it. It came back with seven problems. Three were real, and it showed me exactly where to look. Four were not: all four raised questions we had already settled and recorded, and two of those we had built and shipped two days earlier.
How it happened is in the agent's own log. It had opened our log of 41 decisions and read six of them. It skipped one decision titled "A vessel has one primary owner", which related directly to the piece of work I was about to start.
No entries were missing or out of date. The answer was sitting in a list the agent had already pulled up, with the subject in the title. And when I read the report, there was no way to tell which three findings it had checked and which four it had not.
The worst thing was not the mistake. It was that the wrong four read exactly like the right three, so I believed all seven.
This phenomenon has been measured. In September 2023, a team from Harvard, Wharton, MIT and BCG ran 758 BCG consultants through a set of tasks with and without GPT-4, and engineered one task to sit just outside the boundary of what the model could do.
Consultants working without AI got that task right 84.5% of the time. Consultants with AI got it right 60% to 70% of the time – it made them worse (granted, AI has improved a lot in three years).
But the fascinating finding was that on that same task, human graders scored the AI-assisted groups' recommendations 18% to 25% higher, regardless of whether the answer was right.
On the task where AI made the consultants wrong, it also made them more persuasive. The wrong answers were better written than the right ones.
The study has a name for the boundary. Fabrizio Dell'Acqua, the Harvard researcher who led it, called it the jagged technological frontier – "an uneven set of knowledge work" where two tasks that look equally hard sit on opposite sides of an invisible line: one the model handles easily, the other it cannot do. The professionals crossing that line cannot tell where it runs.
Ethan Mollick, one of its authors, pictured it as a fortress wall with battlements jutting into the countryside. Andrej Karpathy later put the same jaggedness in the model rather than in the tasks – jagged intelligence. The frontier is dangerous because workers who cannot tell which side of it they are on become over-reliant on the AI – falling asleep at the wheel, a phrase Dell'Acqua had already coined in an earlier field experiment – and that risk, they say, is reinforced by how persuasive the output is.
But the genius-or-fool question is the wrong framing, and so are the frontier and jagged-intelligence images – none of them helps you make an agent operate as a reliable employee. All three are ways of judging the worker: is it qualified, where do its abilities run out, is it suitable?
Suitability is the question you can never quite answer, because it is elastic – what counts as suitable changes with each task. A better framing moves from the worker, who is difficult to measure, to the output of the work you are asking it to perform. When you want a job to be done by someone, you care about three things: can they get the job done, how well will they do it, and can they get it done every time?
The output is not only simpler to measure than an agent's capability. It is also a better representation of how businesses operate on trust.
When you employ someone (people or agents), you are depending on the employee's competence, not constantly looking at it. It's like their work contributes, board by board, to a floor that supports the business. Every task you hand over is a board you can step onto.
You put your weight down and find out afterwards whether that board held. Spread across all the work you might delegate, an agent's capability is not a border you approach. It is a floor you are stepping onto. When the agent says it has delivered work and that turns out to be false, then it's a false floor you can fall through.
False in this sense is by construction rather than by fraud. The floor looks safe. It holds almost everywhere. But at specific points it gives way without warning, because the difference between a board that is solid and a board that is hollow is invisible from above.
A human worker has hollow boards too – nobody is competent everywhere – yet a manager walks a human's floor all day without falling through. Why can't the same method manage the agent's? Because the difference is not any particular failure. It is three standing properties of the worker, and together they explain why a human's false floor can be managed and an agent's cannot.
One: nothing binds through consequences. Strip your human controls down to their working parts and a surprising share of them depend on the worker's stake in the outcome.
Probation works because the employee wants to pass it. References work because reputation follows you. Dual sign-off deters as much as it detects; dismissal, prosecution and shame do the rest.
Practically none of this touches an agent – no reputation to lose, no career to protect, no memory of your disappointment past the end of the session. That stake is itself a structural support: it makes hollow boards rarer, because most people, most of the time, do not want to be caught.
There are no similar constraints for agents in how they behave. Every control must be structural. A character reference does not transfer to this worker, but a check built into the workflow itself does.
Two: its failures are sincere. When a human breaks a rule, it is usually a choice – which is precisely why the social contract works on them. There is a motive to interrogate, a story that does not add up, a person who knows they did it. The agent's version is accidental amnesia and honest confabulation, and there is no mens rea in the building. There is no one under the floor to interrogate.
Three: an assertion and its evidence read identically. An agent that ran the test suite and an agent that says it ran the test suite write the same sentence. Human tells – fluency, hesitation, the account that does not hang together – are how a manager taps a suspect board, mostly without noticing.
On an agent those channels are pure noise, and the BCG graders measured something worse than noise: the wrong answers out-scored the right ones. Tap an agent's floor and every board sounds the same. The hollow ones sound slightly better.
Notice what these three have in common: none of them is a failure mode. A failure mode is an event: it happens on a particular Tuesday, to a particular task, and a control can catch it in the act.
But these three are never events. They are standing properties of the worker, true all day, every day – and that changes the kind of control that works on it. Each one retires a control you would use for humans and demands a structural replacement: deterrence out, mechanism in; interrogation out, artefacts in; reading the report out, diffing the tree in.
A human's floor is mostly solid because the personal stakes of the employees hold it up. The few hollow boards are tappable, and they cluster where motive lives – expense fraud happens near expenses, so a manager's intuition about where to inspect is a real map.
The agent's floor is none of the three: nothing thins its hollow boards, no tapping finds them, and they sit wherever the frontier happens to run, uncorrelated with anything resembling self-interest.
The promise being sold to every business right now is not a better autocomplete. It is a worker: something you can hand a job to, the way you hand a job to a person.
So think about what a code of conduct actually buys you. Every working day you act on work you did not do and cannot check. You sign off a document someone else drafted, quote a date someone else committed to, bill against a number someone else calculated. You are standing on your employees' work almost all the time, and you test almost none of it before you put your weight down.
The code is why you can. It does not make anyone competent – plenty of people are bad at their jobs inside a perfectly good code of conduct. But it narrows the ways they can fail to a set you already know how to price.
You know roughly what shape a human failure takes, how often it arrives and what it costs, because you have spent your whole life among people who fail in human ways. That is what makes their work walkable: not that it is correct, but that its failures have shapes you recognise.
Your business is a floor built by other people. The code of conduct is why you can move around on it without tapping every board – and it is how the building goes up at all: floor on floor, each one walked on before the next is laid. Agents are now laying boards in that same floor, and they have signed nothing.
Workplace contracts get breached all the time and businesses survive it. Organisational-behaviour research calls the unwritten deal the psychological contract. Its central finding is that breach is the norm rather than the exception: more than half of new hires reported a violation within two years (Robinson and Rousseau, 1994).2
But look at where those breaches occur. Robinson and Rousseau sorted what people reported into ten categories – training, compensation, promotion, the nature of the job, job security, feedback and the rest. Every one is a negotiated term: a thing the employer promised, or the employee understood to be promised, when the deal was struck.
Not one is a floor term – the baseline conduct nobody negotiates because nobody imagines a colleague without it, like remembering this morning's conversation. The floor sits underneath everything that was negotiated, and no employee you would want to keep breaches it.
Agents breach the floor constantly, and we struggle to name what we feel about it. The same literature separates two halves of our experience that we keep conflating: breach, the recognition that an obligation went unfulfilled, and violation, the anger that follows it (Morrison and Robinson, 1997).
With agents you feel the violation at full strength ("why did you do that?!"). The breach is harder to point at. The agent forgot something you told it ten minutes ago – and who has ever needed to write down an obligation to remember things? The rule it broke is real, but it is written nowhere you can point to.
This essay is about writing the code of conduct for agents that allows us to employ them: taking the distrust seriously instead of treating it as a bias to be trained away. The exercise that led to this essay began as an attempt to catalogue every way agents fail that humans do not, and then, for each one, to look for a process that would catch it or, better, a tool or a setting that would prevent it from occurring at all.
Strain out the alien failure modes, one mechanism at a time, and what remains is a worker who fails the way humans fail – and that worker we already know how to judge, forgive and govern. One code of conduct, shared.
There is no guarantee this is possible. The honest alternative is that some alien failure modes cannot be closed, and we decide to live with them anyway. Because when you weigh it up against everything the worker is superhuman at – recall, speed, parallelism, cost, a worker who never sleeps and never resigns – it's hopefully going to be worth the effort.
In that world we run a two-speed system: one code of acceptable conduct for human employees, another for digital ones, and every manager holds both in their head at once, applying a different standard to each kind of worker on the same team.
So that is the question this essay poses: two codes or one? It is an empirical question, answerable row by row: which failure modes are genuinely alien, which of those a mechanism can close, and what is honestly left over when the mechanisms have done their work. Answer it before the market does, and you are hiring digital employees while your competitors are still arguing about pilots.
The rest of the answer runs over three more essays. The next climbs one level up, to the companies selling us agents, and asks how we can trust that we are getting the model we think we are paying for: you cannot ask a vendor what they sold you.
The third essay walks across the boards of the agent's floor. We can't just ask the worker what it is good at – the answer isn't worth as much as it would be from a human – so the walk follows a register I have been keeping, row by row, against a production repository. It shows which boards can be permanently secured, which give way but make a sound every time, and which will give way silently – the most dangerous, and the ones that you have to hold responsibility for.
The final essay shows you how the register creates trust in the same way every workplace does with HR rules: as a schedule, not a verdict on the quality of the worker.
The fool got his role the day the court stopped trying to decide what he was and simply set the limits he would work within. The register is those limits, written down for the first worker who cannot sign them.
The archetype is not a true reflection of history, or at least of what we know of it. The reality was much darker: the court fool was often selected because of an intellectual disability, and paired with the court dwarf as a target for derision and mockery. Peter Andersson's Fool (Princeton, 2023) is the modern study, and it argues the wise truth-teller we inherited is largely a posthumous invention. The archetype is the one we all carry, and it is the archetype this essay uses. ↩
A neighbouring literature runs the social contract in the opposite direction – AI as the employer's breach of the human worker's psychological contract. Same contract, opposite party in the dock; this essay stays on the worker's side of the table. ↩
Interviews, registers and essays. Free.