AI subscriptions can degrade under a stable label, and buyers have no way to tell whether the model they're paying for this month is the one they paid for last month. Episode two proposes a procurement discipline: seal a golden set before you sign, and index the price to what it measures.
When I first began using Claude Fable, I couldn't get enough of it. My subscription jumped from $30 a month to $150 a month, and finally $350 a month, just for a few more hours in its company.
But now… it doesn't feel the same. And I don't know whether the fading lustre is because my expectations have risen or that Fable's performance has fallen.
It's hard to quantify why, but it feels consistently worse. Sessions run longer than they used to. The agent overbuilds simple jobs. It writes tests to check its tests, which get stuck in an endless cycle of repair, test, break, repair.
Reddit is convinced that the model has been "nerfed" (reduced in effectiveness). The conspiracy theories amount to accusations of bait-and-switch: that the Fable you pay for now has been quietly swapped for something less.
Anthropic has said nothing about this latest drop in performance – though twice this year, users caught Claude's models changing in ways the company had not announced.
Ultimately I want to know what everyone else wants to know: am I still getting what I'm paying for?
I have spent three decades writing about technology vendors, and I can't recall another product where the customer has no way to answer that question.
You can check whether your broadband delivers the speed on the plan. You can weigh a bag of flour. But halfway through a $350 month, I have no idea if the model I'm paying for this month is as powerful as the one I paid for last month.
This is not just a $350 problem. Companies are now spending hundreds of thousands of dollars a year on AI subscriptions and tokens, and the business case rests on how the model performed at launch: the early experience that convinced everyone an agent this capable could deliver a leap in productivity.
If the model degrades after that, the central premise of the investment fails. Yet the token bill is still very real – and the board will still expect the productivity gains that justify the spend. But how can you be as productive without the brilliant model you thought you had paid for?
There are two popular explanations, and both are believable.
The first is acclimatisation. Nothing changed except you. The astonishment of launch week wore off, your standards rose to meet the tool, and the same output that thrilled you in March reads as ordinary in July. This happens with every technology.
The second is detuning. The vendor is serving you something cheaper under the same name – fewer thinking tokens, lower serving precision, a heavier system prompt – because your $350 is fixed and their compute bill keeps going up.
There are two more explanations that rarely pop up on Reddit. You're asking more of it than you were before. The tasks you hand the agent in July are harder than the ones you handed it in March, partly because the March ones worked so well.
And ordinary randomness: the same model, on the same task, has good days and bad days, and one bad week is nothing more than that.
So there are at least four explanations, possibly more. But they all have one thing in common: you cannot test any of them.
A $350-a-month product is defective when the customer cannot tell whether it got worse or they just got used to it. The real defect is not that the quality has declined. It is that there is no way to find out.
Software has a 40-year-old convention: v2.3.1 is v2.3.1 forever. Change one byte and the version number must change. It's a rule we all take for granted.
A model label does not carry that promise. Behind the name on your invoice sits a stack: the model weights, the numerical precision they are served at (quantisation), the routing that decides which servers answer you, the thinking budget, the system prompt, and the software harness around it all. Every layer can move without the label changing.
This is not a secret. The vendors themselves have discussed how the system works – and that the performance of a model can change within months of it launching. OpenAI engineer Ted Sanders has said publicly that model behaviour is so multi-dimensional that it is impossible to guarantee consistent performance for every use case.
Anthropic has spelled this out in even greater detail. In September 2025, the company published a postmortem admitting that three separate infrastructure bugs had degraded Claude's output quality over more than a month.
One routing fault at its peak affected 16% of Sonnet 4 requests, and nearly a third of Claude Code users had at least one request caught by it.
At first, the vendor could not see its own product getting worse. Weeks of user reports pushed it to investigate – and the investigation found the users were right. From the report: "The evaluations we ran simply didn't capture the degradation users were reporting."
In April 2026 they published a second one. Roughly seven weeks of degraded Claude Code output, caused by three changes completely separate to the model weights: the default reasoning effort was lowered, a caching bug wiped thinking history, and a system-prompt change made responses terse. There was no disclosure at the time from Anthropic, and customers kept paying for a worse-performing model.
This wasn't a case of the vendor fiddling with the system economics on the sly. Yet it happened twice in eight months.
Model degradation, it turns out, is just something that you as the buyer have to plan for. A pipeline can degrade on its own, without intent, but the label on the product remains the same.
Ultimately we can't know whether a vendor would deliberately downgrade its models, but it doesn't really matter. We know two things: that capability can change under a stable label, and that the vendor may find out about it after you do.
Ethan Mollick (of "jagged frontier" fame – see essay one) made the strongest counter-argument back in 2024. Models don't generally "get worse" – updates change behaviour, your prompts stop fitting, and there is no way to know when an update has an impact.
Getting used to the tool is real, and prompts do go stale. But Mollick's claim predates both postmortems, and neither of these incidents fits it. A routing bug is not your prompt going stale, and a lowered thinking budget is not your imagination.
Economists have a name for a product whose quality you cannot judge even after you have consumed it: a credence good.
An apple you inspect before you buy. A restaurant meal you can't inspect first, but you know by dessert whether it was any good. An AI subscription sits in a third category: you never find out. You cannot re-run last March to check whether this July is worse, and the output alone cannot tell you what produced it. Every month you pay on faith.
Economists also know what happens to markets like this. George Akerlof, the economist whose "market for lemons" analysis won a Nobel Prize, showed that when sellers know the quality and buyers cannot verify it, quality falls – because the seller who quietly cuts quality can still charge the same price, and the buyer cannot reward the seller who doesn't.
The escape from that trap is detection: a buyer who can measure quality can reward it. And AI customers can catch degradation through firsthand experience, if anecdotally. In the postmortems, the users were able to perceive the issue. But they lacked proof.
I am not claiming any vendor has cut quality on purpose. I am claiming that economic theory says the market rewards the vendor who does; the postmortems prove quality can fall without anyone intending it; and without a measurement, you cannot tell the difference.
Economists have already applied this framing to AI subscriptions.2 The missing piece is what a buyer should do about it.
First, a distinction that procurement almost never makes. "The model" is three different products in one:
None of the three comes with proof. As a mainstream customer of a closed service, you receive no independent evidence that the labelled model and the advertised configuration produced the answer to each request you make of your agent.1 Research on verifying this exists, but nothing you are paying for uses it. All you have is the label.
Read your vendor's public terms and a pattern appears. Availability gets a commitment: uptime percentages, latency tiers, service credits when the service is down.
But output quality gets a disclaimer: as-is, accuracy is your responsibility. They promise the thing they can measure, and disclaim the thing you are paying for.
Journalists have already named half of the consequence: "AI shrinkflation" – the same subscription with tightening usage limits at an unchanged price.2 That half you can at least see; a rate limit is visible. The capability half is invisible: same price, same label, less delivered, and nothing on the bill shows it.
You did not buy the vendor's costs. You bought a result. If the vendor's compute gets expensive, that is the vendor's problem. If the delivered capability falls, that is currently your problem. An uptime SLA refunds downtime because downtime is measured. Nothing refunds degraded capability, because nothing measures it.
This is really a procurement conversation, not a support issue. A support ticket asks the vendor to investigate your anecdote. Procurement asks what the contract does when delivered capability falls – and today, in the terms a buyer can actually see, the answer is nothing.
That is not because nobody has thought of it. Buyers were told to demand drift monitoring and performance guarantees for AI as far back as 2020, and procurement lawyers are drafting the clauses now.4 Capability protection is not impossible. It is just not standard, and it will stay that way until buyers arrive with a measurement in hand.
The remedy is not to ask the vendor harder questions. You cannot audit their stack, and you don't need to. Stop asking how the model is running. Measure whether the job you paid for is still being done, and let the price follow the measurement.
At contract start, build a test set from your own work: 50 to 200 real tasks, each with a way to score the answer. Don't use public benchmarks – they leak into training data, and they measure the vendor's marketing rather than your use. Record the full configuration beside the tasks: which product, which model label or pinned ID, which settings, how many runs per task, how each is scored. Then seal both, so neither you nor the vendor can change them. This is your golden set.
Re-run it monthly, and again whenever the vendor discloses a change. Hold the inputs constant. Keep the raw outputs as evidence.
Keep a second, rotating test set beside the sealed one, and update it freely as your work changes. This will show you whether capability is improving or declining, without touching the sealed baseline.
There are some well-developed resources for creating golden sets. The statistics that separate a real drop from ordinary noise are published and open-sourced – one framework detects drops as small as 0.3%. Public projects already re-run fixed tests against the same model daily. All you need is a script and a calendar reminder to run it.
Just be clear what a golden set tells you. In a failed month, it won't tell you which layer moved – weights, routing, precision, thinking budget, prompt. That is fine. Your contract is with the vendor, not with their routing layer.
The measurement says one thing: the job you paid for fell below the agreed line. Finding out why is the vendor's job. And because the remedy does not depend on proving fault, it can be automatic.
Almost every rung on the ladder below is already in buyer-side clause guidance published by law firms and procurement advisers. The one exception is flagged in the footnote.4 Agree them before signature, in escalating order:
The vendor also gets protections: they can inspect and challenge the test set before it is sealed, a bad week is treated differently from a bad quarter, and changing the baseline needs to be agreed by both parties.
Expect the vendor side to push back with the framing Gartner is already giving technology CEOs: drift is a natural property of AI, contracts shouldn't treat it as a breach, and monitoring belongs with the customer.
Take that deal. Accept the monitoring – the sealed set is exactly that – and in exchange, make the consequence of a measured drop automatic instead of renegotiated. The vendor keeps the freedom to change their pipeline. You keep a price that tracks what actually arrives.
If you are signing or renewing an AI contract this quarter, the sealed set costs you a week of one engineer's time. Build it before signature. A baseline recorded after you already suspect a decline can't prove anything about what you originally bought.
A sealed set tells you when your own service moves. It cannot tell you whether it moved for everyone else too. I am building a shared benchmark from enterprise golden sets.
Each company runs its own sealed set in its own environment, and Scale100 pools the scores (de-identified) by industry and by task type. The tasks never leave your building, so nothing leaks into training data. This will let you know whether any declines in performance are shared by your industry peers, or if it's an issue specific to your organisation.
If you are building a golden set and want to compare results, email me directly at goldenset@scale100.co – and I'll take you through the project.
Vendors of agent platforms sell availability as if the product were a pipe. But buyers purchase capability as if the product were a colleague – and in fairness, the vendors are marketing their agents as digital employees. A 200 response (the status code that says "request served") proves the pipe worked. It does not prove that you're talking to the talented digital employee you thought you hired.
The first essay in this series ended on a rule: judge the work, not the worker.3 This essay applies the same rule one level up: judge the model by its delivery, not its label.
The vendor is the easy case. It is a company, and companies can be held to contracts.
But while you can hold the vendor to a contract, the agent actually doing the work cannot make promises about its own performance. And anyway, that is not how you handle a colleague. You find out what they are good at, often because they tell you, and then give them work they will do well. An agent cannot tell you its strengths and weaknesses. So the next essay maps them from the outside: every way this worker fails, row by row against a production repository, and what catches each one – the workplace's unwritten contract, finally written down.
The engineering literature covers this already. Tian Pan's "Silent Quantization" and "The Vendor SLA Gap" (2026) describe the verification gap and propose nightly fingerprint probes – and conclude that a capability SLA cannot be obtained through procurement. This essay argues the opposite: what procurement cannot get is proof of what runs inside the stack, and a delivery measurement doesn't need it. The identity problem itself was posed formally by Semiotic Labs (2024): a provider can charge for model A and serve cheaper model B, and both return plausible prose. ↩
The economic diagnosis is prior art. Zhang, Li and Bao ("AI as a Credence Good", 2026) described the subscription case; Ward (2026) states the credence-good claim directly; Falco's "Cloud LLM Market" (April 2026) is an essay-length version of the argument; Mittal (2026) applies Akerlof's lemons to drifting agent capability; Chen, Zaharia and Zou (2023) is the canonical academic study of model drift; and "AI shrinkflation" is Stokel-Walker's phrase (Inc., May 2026). None of them looks at the buyer-side mechanism: the sealed baseline at contract start, the monthly same-label re-run, and the fault-neutral, capability-indexed price entitlement. ↩↩
"Verify the work, not the model" was popularised as a rule for agent outputs (Nate's Newsletter, July 2026); Antony Ma had earlier proposed a golden dataset as a "contractual yardstick" for agent vendors. The object here is the vendor rather than the output, which is why the measurement has to be sealed and repeated: one verified output tells you about today, while the label claims to describe next month. ↩
The World Economic Forum's AI procurement guidance asked buyers for drift-monitoring KPIs and performance guarantees in 2020, before any of this tooling existed. The ladder's rungs come from published buyer-side guidance: Mayer Brown described AI performance warranties in 2023 and, for agent deals in 2026, statistical acceptance testing against agreed evaluation sets with thresholds and change governance; Dentons set out AI-specific service levels in 2024; the US health sector's third-party AI risk guide (2026) carries sample notification language for performance falling an agreed percentage below baseline; and practitioner guidance already proposes drift clauses that hold accuracy within an agreed band of the deployment baseline, with a golden dataset as the contractual yardstick. Cure windows, rollbacks, service credits and termination rights are standard service-contract machinery. The one rung not found as standard practice anywhere: the automatic price step-down indexed to the buyer's sealed measurement. That is what this essay is proposing. ↩↩
Interviews, registers and essays. Free.