133 patent lawyers. Eleven firms. Three months. On Monday Marginal Revolution surfaced a pre-registered NBER trial from David Autor and co-authors that hands practicing attorneys a custom AI drafting assistant and then does the thing most AI productivity studies skip: it takes the tool away and checks what stayed.

While the tool was on, the result matches every other white-collar study of the last two years. Quality up 0.34 standard deviations at ten days, 0.38 at ninety, scored blind by expert patent attorneys, with the larger gains going to junior lawyers. That is the number the vendors will quote. It is not the finding.

Treated lawyers outperformed controls by 0.32 SD (p = 0.04), but this advantage was concentrated entirely among senior lawyers (0.45 SD, p = 0.02). Junior lawyers showed no average gain; their scores instead bifurcated, with sharply fewer mediocre scores offset by more poor and more good ones.
NBER

Read it twice. The lawyers who gained the most while using the tool retained the least once it was gone. The seniors, who gained less on the drafting benchmark, walked away with better judgment on a redlining task the assistant never touched. The authors' own gloss is that foundational expertise may be a prerequisite for extracting durable skill from AI-assisted practice. In plainer terms: the tool compounds what you already have. If there is nothing to compound, it produces output and leaves you where it found you, or somewhere to either side at random.

The tool compounds what you already have.

This is the answer to a question every engineering manager I know has been guessing at for two years. Not whether AI makes the team faster. It does; that number was in a year ago. The question is what happens to the person who joined in 2025 and has never written a migration by hand. The bifurcation is the tell. Some juniors used the assistant as a tutor and came out sharper than the control group. Others used it as a vending machine and came out worse, having spent three months shipping work better than they could evaluate. Same tool, same firms, same ninety days. The variable was whether the person already had enough judgment to check the output, and building that judgment is what the junior years are for.

Tyler Cowen's caveat on the post is the fair objection: an RCT holds the allocation of people to tasks constant, and over ten years the allocation moves until more of the humans are more productive. Grant it. Reallocation is real. But the reallocation is happening at the hiring desk, and it is moving in one direction. Firms are cutting the junior rung because a senior plus the tool is cheaper than a senior plus a junior. That is a rational reading of the ten-day number. It ignores the ninety-day one. The seniors who extract durable skill from the tool exist because someone once paid them to be mediocre for three years. Remove that rung and the pool of people who can compound the tool stops refilling.

The tool pays out to whoever already has the skill. It does not make the skill. That is still the job of the years nobody wants to fund.