The Plausibility Subsidy
AI made looking good free — so looking good stopped meaning anything. The market for lemons is now inside your company, and the fix is evidence that buys freedom.
For the entire history of knowledge work, one signal did more management labour than any other: polish. A finished-looking document meant somebody spent hours on it. Clean code meant somebody sweated the details. A crisp analysis meant somebody thought hard. Polish was expensive, so polish was information — a rough proxy for effort, and effort a rough proxy for quality. Every manager alive was trained, mostly without knowing it, to read that signal.
AI just made polish free. A finished-looking anything now costs minutes. And a signal that costs nothing to produce carries no information — which means the oldest instrument in the management toolkit silently stopped working, everywhere, at roughly the same time.
I call what’s left behind the plausibility subsidy: the cost of producing plausible work has collapsed, while the cost of producing correct work has barely moved. The gap between those two costs is a subsidy, and it flows to whatever looks good — regardless of whether it is.
The lemons come inside
Economists have a name for markets where the seller knows the quality and the buyer can’t tell: markets for lemons. When buyers can’t distinguish good from bad, they pay average prices, sellers of genuinely good goods exit, and the bad drives out the good — not through anyone’s malice, but as the equilibrium.
Run that logic inside a company. The evaluator (a manager, a reviewer, a stakeholder) can no longer distinguish good work from plausible work at a glance, because the glance was calibrated on polish and polish is free. Plausible is now roughly 10x cheaper to produce than correct. So the rational move for a producer under deadline pressure drifts toward plausible — and the conscientious, who still pay full price for correct, get outcompeted on visible volume by colleagues who don’t.
This is worth saying plainly: adverse selection for slop is not a character failure. It’s the equilibrium. Moralizing about it changes nothing. Only re-pricing the signals does.
And it isn’t one signal that broke — it’s all four at once. Effort signals (hours, visible struggle) decoupled from output quality. Volume metrics now measure AI usage, not contribution. Review can’t scale to re-verify everything — that’s the capacity crisis I wrote about last week. And self-report (”I checked it”) is an assertion, and I’ve already written down what assertions are worth when they meet a deadline.
Quality made legible
The fix is a design principle, not a policy: move the organization from trusting assertions about work to reading evidence attached to work.
Good evidence instruments share four properties. They’re machine-checkable where possible. They travel with the artifact — the evidence is attached, not filed somewhere. They’re cheap to read — a glance, not an audit. And the property that does the real work: they’re produced by the process, not the author. A test result generated by the pipeline can’t be embellished by the person whose work it describes. That’s the same principle as a commit gate whose receipt only a passing run can produce — the receipt is trustworthy precisely because no human writes it.
Five instrument classes, roughly in ascending order of trustworthiness per dollar:
Provenance — what produced this: human, agent, which model, what context.
Verification evidence — which checks ran and passed: tests, evals, gate receipts.
Review record — who reviewed it, at what depth, what changed.
Outcome linkage — did it work downstream: acceptance rates, error rates, rework. Lagging, but it’s the ground truth that keeps every other instrument honest.
Calibration history — the track record of the author’s judgment: how often did their “done” survive review? How often was their “ship it” right?
That fifth one deserves a beat of silence, because it’s the uncomfortable one. It’s the earned-autonomy ladder — the one we built for AI agents — applied to humans. And it surfaces something genuinely new: AI didn’t make senior judgment redundant. It made senior judgment measurable. For twenty years, “good judgment” was a reputation. It’s becoming a number: calibration, tracked across decisions. Performance review drifts from how much did you ship toward how often were you right about what was ready to ship — which, if you think about it, is what we always claimed to be evaluating and never actually could.
The two traps
Every naive version of this dies in one of two ways, and both deserve respect.
Goodhart’s trap. Any legible metric gets gamed: acceptance rate, by submitting only safe work; eval scores, by teaching to the test. The mitigations are the classic ones — multiple instruments, outcome linkage as the un-gameable anchor, and unpredictable human sampling, audit-style. A quality signal you can fully predict is a quality signal you can fully manufacture.
The bossware trap. This is the fatal one. If quality legibility reads as surveillance — dashboards of keystrokes, activity scores, managers watching — trust collapses, adoption dies, and you deserve both. The design line that keeps you out of it: evidence attaches to artifacts, never to people-watching. You instrument the work, not the worker. And the framing that survives contact with an actual workforce is the same one the legal system arrived at from the other direction: the evidence protects the author. Your “done” becomes provable. Your calibration becomes your case. When something goes wrong downstream, the record shows what you checked and when — which is precisely the position you want to be in.
Evidence buys freedom
Which points at the design’s incentive-compatible core, and the reason it can work where quality bureaucracy fails: evidence buys freedom.
The deal, stated honestly: producers accept evidence requirements; in exchange, demonstrated calibration earns lighter review, faster shipping, more autonomy. The ladder pays out. Under that deal, the evidence is a career asset — your track record, portable and provable. Without the payout — legibility as pure requirement, autonomy never granted — it’s a tax, and taxed people route around systems, which is the third failure mode of the review economy arriving on schedule.
Legibility without freedom is surveillance. Freedom without legibility is the plausibility subsidy. The two together are just... trust, with receipts.
The one-question diagnostic
If you want to know where your organization stands, ask one question of any team adopting AI:
Can a manager here distinguish an AI-checked artifact from an unchecked one without redoing the work?
Most organizations, asked honestly, answer no. Which means every artifact carries the average trust of all artifacts — the lemons equilibrium, already priced in. The fix isn’t better managers or sterner policies. It’s making the difference legible — and then paying for it in the only currency producers actually value.
