October 7, 2026
David Autor, Technology & Society Visiting Fellow, and Tanya Rodchenko, Principal, AI & Economy
In a three-month randomized controlled trial with practicing patent attorneys, we found that, while routinely using AI lifts work quality across the board, its effects on on-the-job learning depend on seniority: senior lawyers who used AI for 90 days demonstrate stronger judgment than those who didn’t, whereas junior lawyers show no average skill gain, their scores splitting into more strong scores and more low scores.
Judgment is what sets senior professionals apart from their junior peers. Building judgment requires thousands of hours learning on the job, usually through tedious but formative engagement with relatively routine work, often under guided supervision by experienced workers. But AI is already changing how on-the-job learning works. On the one hand, empirical studies in radiology, business problem-solving, job-seeker writing, and legal education show that when AI tools are thoughtfully embedded into training workflows, less-experienced workers can improve their independent performance.
On the other hand, recent experiments with software engineers, management consultants, high school students, and clinicians demonstrate that while AI serves as a temporary exoskeleton that boosts immediate output, those gains often fail to persist once the tool is removed. Understanding how AI can erode or enhance expertise requires three elements rarely found together: (1) sustained AI usage embedded in ordinary professional work over months rather than hours, (2) a credible evaluation of unassisted judgment when AI is unavailable, and (3) blinded grading by domain experts.
In “Does AI Assistance Enhance or Erode Expertise? Evidence from a Three-Month Field Experiment in Patent Drafting”, published by the National Bureau of Economic Research, we describe a field experiment used to capture changes in both short-term productivity and longer-term skill building that occur as a result of AI usage in the highly expertise-intensive profession of intellectual property law. In our experiment, we randomized access to a then-unreleased Google Labs AI patent writing assistant tool, now part of Gemini Notebook, among 133 lawyers at eleven intellectual property law firms that are engaged in regular, non-exclusive business with Google. Two-thirds of lawyers at each firm (“treatment” group) were given early access to the tool, while the remainder ("control" group) were given some training on how to use AI but were withheld from using the tool until after the three-month study period.
After 10 days of AI assistance, we sent all participating lawyers a packet of inventor materials for a hypothetical invention and asked them to draft a patent. We did the same after 90 days, with a different set of simulated inventor materials. We saw the expected AI boost in performance on the drafting tasks at both time periods. At 10 days, AI tool access raised drafting scores (graded on a rubric with Likert scales for five different aspects of quality by independent legal experts) by 0.34 standard deviations, equivalent to a 10-point climb in percentile ranking among control group scores; by 90 days, this improvement increased to 0.38 SD, implying an 11-percentile point increase. These improvements occurred with remarkable consistency across all quality dimensions. Tellingly, better performance came from lower frequency of poor scores and higher frequency of good scores, and no change in the frequency of excellent scores. Formally, scores among lawyers allocated AI access moved from the bottom to the lower-middle and middle quintiles, but there were no noticeable changes in the number of exceptional scores. This pattern was particularly clear for junior lawyers, whose improvements in quality were also matched by significant time savings (18 minutes faster than the control group’s average speed of 124 minutes on the 10-day task, for example).
But patent lawyers don’t just draft patents; they also review each other’s work for critical mistakes (sometimes called “patent profanity”) that can make patents unenforceable in court — for instance, by overclaiming the novelty or scope of an invention. We therefore added another task at 90 days, which required subjects to redline (to manually mark up and correct) an existing hypothetical patent that contained many errors, both substantive and stylistic. This test task mirrors the kind of professional judgment that more senior lawyers routinely use and for which they most often do not have the luxury of relying on AI. Critically, while treatment group lawyers were encouraged to use the unreleased AI tool or any other AI tool on the drafting tasks, no one was allowed to use AI for the redlining task, making their scores on this task a measure of how well their skills had developed. For all tasks, drafting as well as redlining, we worked with third party patent professionals to score submissions on five dimensions of quality: (1) enforceability, (2) accuracy, (3) strategic ambiguity (the tactical scoping of claims), (4) completeness, and (5) clarity.
When the AI assistant was removed for the 90-day redlining task, the leveling effect vanished. Lawyers given AI tool access outperformed controls by 0.32 SD on this unassisted task (a 9-point equivalent climb in percentile rankings), but this effect was driven entirely by senior lawyers, who outperformed control subjects by 0.45 SD (a 13-point climb), while juniors showed no discernible improvement. Instead, scores of juniors bifurcated: more very low scores, fewer mediocre ones, more good ones, and no more excellent ones.
What do these results mean for learning? Seniors who had been given AI tool access clearly harvested significant value in terms of their professional judgment from their use of the tool over the previous three months. But the junior-level picture is more nuanced. Some scores improved, hinting at AI’s ability to leapfrog learning. Other scores spread to the lower rungs of the distribution, showing that in some scenarios AI can help juniors deliver adequate work, but their higher machine-assisted performance does not translate into faster acquisition of the expertise that makes senior professionals distinctly valuable.
Examining qualitative editing patterns provides some insight into how professionals of different experience levels engage with generative AI tools. Across both treatment and control groups, junior submissions followed rigid formalism: juniors worked sequentially from top to bottom, frequently exhausting their time copy-editing low-stakes introductory sections before reaching the main patent claims. Juniors also focused on surface-level synonym swapping rather than commercial scope improvement. Even when juniors spotted serious flaws, they often left descriptive comments diagnosing the defect rather than executing the redline to fix it. Because this "diagnose-without-execute" posture was equally common among unassisted control juniors, it represents a baseline junior deficit that three months of relying on AI to execute rewrites did not remediate.
Senior lawyers, by contrast, benefited consistently from AI exposure. Treated seniors spent longer on redlining than juniors, bypassing low-stakes prose to rebuild the claims from scratch, strip out language that could inadvertently narrow legal rights or limit protective scope, and pair many edits with explicit legal doctrines. In follow-up interviews, seniors described treating AI output not as a finished product, but as a "logic auditor". It weakened attachment to existing prose and forced them to articulate the why and how of structural edits, activating and sharpening their foundational expertise.
Improving performance and building durable professional judgment are different goals. For junior professionals in particular, they may be in tension with one another. Foundational expertise may be the missing link that allows AI-assisted repetition to translate into seasoned professional judgment over the course of a career. Such judgment will likely continue to be important even as model capabilities advance: lawyers will continue to use their judgment in their interactions with judges, clients and colleagues. They will not be able to rely on an AI tool to assist them in all settings. A critical challenge of this AI moment is ensuring that tools designed to produce better work today do not prevent junior professionals from becoming the experts we need tomorrow.
Studying professionals in their working environments is a challenge for researchers. Our experiment includes a limited sample of patent lawyers, a reflection of the high hourly billing rate of top-tier professionals in the field. We only observe them over three months, a small window compared with the years they would need to develop durable expertise. Moreover, the rapid advance in both model capabilities and baseline AI familiarity (as well as AI-enabled product penetration in the legal domain) over the past year might affect our results if we re-ran the trial now. These limitations give rise to opportunities for further research, some of which we hope to pursue in the months ahead. Nevertheless, a clear takeaway from our work is that understanding the effect of AI assistance on professional skill requires tests that separate the effect of using tools from the effect of unassisted user performance.
A huge thank you to co-authors Josh Martin, Zanna Iscenko, Scott Strand, David Pearl and Melissa Ferere; to James Manyika, Ruth Porat, Kent Walker, Josh Woodward, Anu Madgavkar, Daniel Rock, and Fabien Curto Millet for their guidance and support; to Tom Shane, Mauricio Carneiro, Johnny Jacobs, Victor Lavrenko, Yinghao Sun, Samuel Yang, Daniel Nishi, Eric Wang, Kevin Li, Hao Zheng and Franz Och of Google Labs for partnering with us on the experiment; to Kiera Russell for product access and trusted tester support; to Steve Gong and Taylor Montgomery for legal guidance and recruitment support; and to our Comms, Marketing, and Research Blog partners including Kerry McHugh, Zoe Ortiz, Emily Short, Joseph Mack, Esmeralda Cardenas and Leanne Trujillo for driving the launch.