Who Actually Measured AI Coding Productivity? I Classified the 12 Studies Everyone Cites. The Vendors Ran Most of Them.
Jul 8, 2026 · 6 min read · by Jordan Kwan
TL;DR: I pulled the 12 most-cited studies on AI coding productivity and classified how each one actually measured: 7 are RCTs or controlled experiments, 3 are surveys, 2 are telemetry. Five of the seven RCTs were run by the vendor selling the tool, and every headline speedup (21% to 55.8% faster) comes from that vendor-run five. The only fully independent RCT, METR's, found experienced developers 19% slower with AI while believing they had been sped up. Nobody has the real number. But the spread of claims maps suspiciously well onto who paid for the study.
Everyone has a strong opinion about whether AI coding assistants make engineers faster. Almost nobody has read where the numbers come from. So I did: the 12 studies that dominate every "AI productivity" argument, classified by methodology and by who ran them. (If your question is which tool rather than whether, that's its own comparison.)
Who actually measured what?
The census, as of August 6, 2026: 7 RCTs or controlled experiments, 3 surveys, 2 telemetry analyses. Now the part the keynote slides skip: 5 of the 7 RCTs were run or co-authored by the vendor whose tool was being tested. GitHub's Copilot lab experiment, GitHub's code-quality RCT, the GitHub/Accenture enterprise trial, Google's internal RCT on its own engineers, and the three-company field experiment with Microsoft co-authors. Widen "interested party" to anyone with a product stake (Google owns DORA, and the telemetry firms sell the measuring instruments) and it's 9 of 12. Fully independent of any product: METR, twice, and a self-selected developer survey.
That does not make the vendor numbers fake. It makes them a sales floor, not a science ceiling, and the pattern in the data bears that out.
What did the only independent RCT find?
The famous one. METR randomized real issues in mature open-source repos for 16 experienced maintainers, 246 tasks, AI allowed versus not. Result: with AI tools, developers took 19% longer. Before the study they predicted AI would make them 24% faster. Afterwards, having been measured slowing down, they still believed it had sped them up by 20%.
Read that twice, because it is the most important finding in the entire genre: self-reported AI productivity is not evidence. The people closest to the work misjudged not just the size of the effect but its direction.
The honest update: in February 2026 METR published follow-up data, now 57 developers across 143 repos and 800+ tasks. For the original cohort the estimate is a -18% speedup with a confidence interval of -38% to +9%; for newly recruited developers, -4% (-15% to +9%). Two things worth saying like an adult: the interval now spans zero, so "AI slows experienced devs down" is no longer a settled fact, and METR itself flags that developers unwilling to work without AI are dropping out, which biases the estimate down. Several blogs converted that minus sign into "METR now finds an 18% speedup." It does not. Read the primary.
What do the vendor experiments say, and how should you read them?
The vendor-run trials are real experiments with real results, and they form a gradient that tells its own story:
- Peng et al. (GitHub/Microsoft authors): 95 developers hired for one greenfield lab task, an HTTP server from scratch. Copilot group finished 55.8% faster. The biggest number in the genre comes from the least realistic task.
- Google's internal RCT: 96 Google engineers on an enterprise-grade task, roughly 21% faster, with a wide confidence interval.
- The three-firm field experiment (Microsoft, Accenture, a Fortune 100; Microsoft co-authors; peer-reviewed in Management Science): 4,867 developers, 26.08% more completed tasks (standard error 10.3%), with the gains concentrated in less experienced developers.
- GitHub/Accenture's enterprise trial: 8.69% more pull requests, a 15% higher merge rate, and around 30% of suggestions accepted.
Watch the effect shrink as the setting gets realer: 55.8% on a toy task, 26% across field deployments mostly lifting juniors, 21% inside one company, 8.69% in PR terms at an enterprise, and -18%-ish for experienced maintainers in their own gnarly repos. That gradient is the finding. AI coding help is largest where the work is most generic and the developer least familiar, and it decays toward zero, possibly through it, as expertise and codebase mess increase.
What about the quality side of the ledger?
The measurements that track what ships, rather than how fast, lean the other way:
- DORA's 2024 report (survey, ~3,000 respondents) estimated that a 25% increase in AI adoption came with a 1.5% decrease in delivery throughput and a 7.2% decrease in delivery stability, even as individuals reported feeling more productive. The 2025 edition (~5,000 respondents) saw throughput turn positive but stability still negative, with 90% of respondents using AI and 30% reporting little or no trust in its output.
- Uplevel's telemetry on ~800 enterprise developers found no significant efficiency change from Copilot access, and a 41% higher bug rate in the with-Copilot group.
- GitClear's analysis of 211 million changed lines, 2020-2024: copy-pasted lines rose from 8.3% to 12.3% of changes while refactoring-associated lines collapsed from 25% to under 10%. The codebase equivalent of eating dessert first.
- The Stack Overflow 2025 survey: 84% of developers use or plan to use AI tools, yet 46% actively distrust its accuracy versus 33% who trust it, positive sentiment is down to 60% from 70%+, and 66% name "almost right, but not quite" answers as their top frustration. Adoption up, affection down, which matches the broader workplace pattern.
So does AI make engineering teams more productive?
The defensible answer: on generic, greenfield, junior-shaped work, probably yes, and possibly by a lot; on expert work in complex codebases, unproven and possibly negative; and everyone's self-assessment, including yours, is unreliable in the optimistic direction. Anyone quoting you a single clean percentage is selling something, and in five of the seven controlled experiments, literally.
What to do with that if you run a team:
- Never accept self-reported speedup as data. METR's perception gap is the one replicated-feeling result in the pile: measured slower, felt faster. METR's other famous output gets misread in the opposite direction, because the doubling chart plots a 50% success rate and the 80% line is roughly ten times shorter.
- Point the tools at the gradient's good end. Boilerplate, scaffolding, unfamiliar-but-standard territory. That's where even the discounted numbers stay positive.
- Track the stability tax. DORA, Uplevel, and GitClear all find the cost lands downstream, in reverts, bugs, and duplicated code. If you only measure velocity, the tax compounds where you aren't looking.
- Demand the methodology line. "A study found 55% faster" and "the tool's maker timed 95 freelancers building a toy server" are the same sentence at different zoom levels.
The 10x crowd is quoting a vendor lab task as the average. The useless crowd is quoting an early METR headline whose own authors have widened it to "maybe nothing." The truth is a gradient with error bars, run mostly by interested parties. Which is a less quotable answer, and the only honest one on offer.
No hype. Just the methodology column.
Written by Jordan Kwan, founder of Reachium.
I build Reachium, the LinkedIn outreach platform behind the tactics you just read. Same brain, live product.
See what Reachium does ↗