METR analyzed 5,305 Claude Code transcripts from 7 technical staff during January 2026 to estimate coding productivity gains from AI agents. Rather than running expensive controlled uplift studies, they used transcript analysis as a cheaper proxy — measuring “time with AI” (summing 10-minute windows of user activity) and estimating “time without AI” via GPT-5 as an LLM judge. The resulting time savings factors ranged from 1.5x to 13x across individuals. But the headline finding is that these numbers represent a “soft upper bound” on actual productivity, not ground truth — for reasons the authors are admirably transparent about.
Key Insights
The most striking finding is that roughly 47% of estimated task time involved work users wouldn’t have done without AI — a massive task substitution effect. This means the raw time savings numbers dramatically overstate productivity gains because they’re measuring a shifted task distribution. Staff gravitated toward AI-tractable work, meaning the transcript sample is systematically biased toward tasks where AI shines. The one outlier who achieved 13x savings was running 2.3 concurrent agents — essentially parallelizing themselves — which is a genuinely different workflow pattern rather than just “coding faster.” The validation approach (GPT-5 as judge, r-log of 0.83 against 34 human ground-truth samples) is creative but thin, and the authors know it.
Our Take
This paper is valuable precisely because it’s honest about its own limitations. The AI productivity discourse is plagued by either breathless “10x developer” claims or dismissive “it just autocompletes” takes, and METR charts a careful middle path: yes, coding agents provide real time savings; no, you can’t naively extrapolate from transcript analysis to workforce-level productivity. The 47% task substitution finding is the number that should travel — it suggests that a large fraction of “AI productivity gains” are actually “AI enabling people to do different things” rather than “AI making existing work faster.” For prediction markets on AI economic impact, this distinction matters enormously: models that assume direct labor substitution will systematically overestimate near-term effects.