The Selection Problem
In July I said AI would explode the variance between workers. The research says the opposite on most work. What it disperses is judgment — and the fear was never replacement.
Earlier this month I left a comment on someone else’s post that I’ve since had to take apart. The argument was that a superpower had become generally available, and that superpowers don’t level — they amplify. Equal access, unequal amplification. The mean rises and the variance explodes at the same time, the way it did with literacy and electricity: civilization gets richer while individuals get sorted.
It’s a satisfying line. I believed it. Then I went and read the research, and the most-cited experiments in the field say close to the opposite.
When Brynjolfsson, Li and Raymond put a generative assistant in front of 5,179 customer-support agents, resolutions per hour rose about 14% on average — but the gain was roughly 34% for novices and close to nothing for the most experienced. Their description of the mechanism is the part that should have stopped me: the tool “disseminates the best practices of more able workers.” Noy and Zhang ran writing tasks with several hundred professionals, found time down about 40% and quality up, and reported the result in one sentence I can’t argue my way around: inequality between workers decreased. The GitHub Copilot trial points the same direction — the authors note it shows promise for helping people transition into software careers.
That isn’t a caveat to the amplifier story. For that class of work, it’s a refutation of it.
The historical half didn’t survive either. I reached for literacy and electricity because they sound like they prove the point. But the largest education expansion in the modern record did the opposite of sorting: Goldin and Katz show the educational wage premium narrowing for roughly sixty-five years, from 1915 to 1980, as the high-school movement pushed skill supply up faster than technology pushed demand. And when I went looking for the study showing that electrification widened productivity gaps between firms, I couldn’t find one. The dispersion evidence I could actually verify is all from the computing era, not the dynamo. I’d been citing a vibe.
So the frame was wrong. What’s underneath it is more interesting.
Start with where the compression comes from, because it explains its own limit. On a bounded, well-specified task — a support ticket, a press release, a standard memo — the model already contains something close to the expert’s answer. Handing that to a novice closes the gap almost by construction. The expertise was never the scarce part; retrieving it was.
Now change the task. Otis, Clarke, Delecourt, Holtz and Koning randomized 640 Kenyan small-business owners into access to a GPT-4 business assistant. Across the whole sample there was no detectable effect on revenue or profit. Underneath that flat average, the sample split: the effect for the initially weakest performers came out nearly a quarter of a standard deviation below the effect for the strongest, with the low performers ending up somewhat worse off and the high performers meaningfully better. Same tool. Opposite sign.
The mechanism is the whole essay. It wasn’t that the strong performers asked better questions, and it wasn’t that the model gave them better advice. The researchers checked both. The difference was which pieces of the AI’s advice each owner selected and implemented.
Identical input. Identical output. Opposite outcome. The variable was judgment.
The BCG consultant study points at the same seam from another angle. Inside the set of tasks the model handles well, the assisted consultants were faster and better. On a task deliberately placed outside that boundary — one where the model produces a confident, wrong answer — the consultants using AI were about 19% less likely to get it right. The skill under test there wasn’t producing an answer. It was knowing when to refuse one.
Which gives a line I can actually defend: these tools compress performance on bounded work and disperse it on open-ended work. Not because they amplify existing talent in some general way, but because as work gets less specified, the binding constraint moves from producing an answer to selecting one, implementing it, and knowing when to override it. That version predicts the compression results instead of dodging them, which is the test my original framing failed.
Now the second scale, and the part I got closest to right for the wrong reason.
The same week, Noam Segal and Lenny Rachitsky published their second annual survey of how tech workers are feeling — thousands of them, across product, engineering, design, research, data and sales. The headline is a workforce splitting in two: one half describing itself as amplified by AI, the other as shaken by it. And that split, they report, predicts how people feel about their career better than role, seniority, or company size.
But the numbers underneath are less symmetric than the headline. Asked how AI had shifted how they see themselves professionally, 49% chose amplified and 27% chose redefined — my role is changing shape. Only about 19% landed in the two genuinely negative buckets. Most people are not being displaced. Most people are being expanded, or reshaped.
So what exactly is everyone anxious about? The survey answers it plainly, and it isn’t the thing the discourse assumes. Only 22% say they worry about losing their job to AI. What they actually worry about is being expected to do more for the same pay — 51%, the single largest fear on the list — followed by getting trapped in an unsustainable pace, and by the quality of their own work declining. Meanwhile significant burnout rose from 44.7% to 55.7% in a single year.
Read those together and the story stops being about capability. Almost nobody thinks the machine is going to replace them. Half of them think it already made them better. What they expect is that the surplus their amplification produces will be quietly absorbed — that they’ll be handed the multiplier and then handed more work at the same price.
That is not a fear about technology. It’s a fear about who banks the gain.
And it’s the same question I’d spent the same week arguing at a completely different altitude. In a thread about token pricing, I’d made the case that metered inputs are the best deal enterprises have been offered — because when the input is metered, all the surplus above the meter belongs to the buyer, the way the grid doesn’t take a cut of the factory’s output. With the honest catch attached: the surplus only materializes if you can actually operate the capability. Otherwise you’ve rented a meter and produced nothing to sit above it.
Individual and institution, the same sentence. Amplification creates surplus. The surplus goes to whoever can operate the capability and show what it produced. Everyone else is holding a multiplier they can’t cash.
Which is where this stops being an essay about labor economics and becomes one about how you actually run things. If the scarce input is judgment — which advice to take, which to discard, when the confident answer is wrong — then the organization that gets ahead is not the one with better model access. Everyone has that; that’s what “generally available” means. It’s the one that can find where good selection is already happening and make it institutional rather than individual.
That is what a review record is for. Not compliance. It’s the only way to see, across a hundred delegated decisions, which overrides were right and which were reflex, whose accepts hold up a month later, and where the plausible answer got waved through because nobody was measuring the difference between plausible and correct. I’ve argued before that plausible work is cheaper than correct work and drives it out, and that review is the scarce input nobody prices. Otis is the closest thing to empirical backing I’ve seen for either: a field experiment where the entire spread came from what people chose to implement.
Let me be clear about what I’m not claiming.
The adoption research — Danish register data showing women, older, and longer-tenured workers using these tools substantially less, even among colleagues doing identical work — points to widening gaps too. But that’s a claim about who picks the tool up, not about what the tool does to people who use it. Those are different arguments and I’ve stopped letting them lean on each other.
The firm-level dispersion evidence is real and large — frontier firms pulling away from laggards at several times the rate, strongest in computing-intensive services — but it’s from the ICT era, and using it to describe what AI is doing right now is a forecast, not a finding.
And I don’t have a client corpus of my own to put behind this. We’ve built the instrument — the review record, the calibration history, the accepted-versus-proposed diff — and I can tell you exactly what it would measure. I can’t yet tell you what it measured across other people’s businesses, and I’d rather say that than imply otherwise.
Which leaves the question I most want answered, and haven’t asked the one person who might have the data. If the split runs through selection and implementation, is that a trait or a practice? A trait sorts people permanently and there’s not much to do about it. A practice is coachable, transferable, and institutional — and if it’s a practice, then the whole game is building the record that lets you see it, name it, and spread it.
Otis leans toward practice. The gap wasn’t in what people were told. It was in what they did with it.
The survey’s own closing line is that everyone agrees the ground is moving, and nobody’s sure yet whether it’s an earthquake or a launch. Reading the research together, I think it’s both, and which one you get isn’t decided by the tool. It’s decided by whether the judgment that makes the tool worth anything leaves a trace — because a surplus nobody can attribute is a surplus somebody else will bank.
