Research review · updated 2026-07-28

Does AI actually make developers faster? What the research says.

The honest answer: it depends, and the measured effect is smaller and far more variable than vendor claims suggest. The most rigorous controlled study to date — METR's 2025 randomized trial — found experienced open-source developers were about 19% slower with AI tools even though they believed they were 20% faster, a roughly 39-point perception gap; a 2026 follow-up found effects whose ranges cross zero. Across the broader literature, AI's effect on delivery is real but modest and highly team-dependent, and it routinely carries a quality cost: Faros AI measured incidents per PR rising ~3× (+242.7%) and bugs per developer up 54% on teams that pushed throughput hardest. The defensible reading is that AI can help, but the size of the help is an empirical question for each team — which is exactly why every Spendassay productivity figure ships as a range with its the cost of going faster attached.

The evidence at a glance

Eight primary sources, what each measures, and the headline effect. Every study links to its primary source below.

StudyYearSample / scopeHeadline effectWhat it measures
METR2025Experienced open-source developers, randomized controlled trial19% slower on real tasks while believing they were 20% faster (~39-point perception gap); a 2026 follow-up found −18% / −4% with ranges crossing zeroTask completion time, with and without AI tools
Faros AI202622,000 developersIncidents per PR +242.7% (~3×); bugs per developer +54%Delivery quality against AI-accelerated throughput
GitClear2025Large multi-year code-change corpusRising code churn and copy-pasted (duplicated) code alongside assistant adoptionCode churn and duplication over time
DORA2025DORA global survey cohortAI amplifies the system it lands in: strong delivery gets faster, weak delivery breaks fasterDORA delivery and stability metrics under AI adoption
CACM · Ziegler et al.2024GitHub Copilot users (survey + telemetry)Suggestion acceptance rate tracks perceived productivity, not measured outputPerceived productivity vs. acceptance rate
LinearB · APEX2026Engineering-metrics framework (LinearB dataset)Documents the adoption J-curve and the review/verification tax; payback typically 6–18 monthsAdoption curve and payback window
Microsoft · Viva2026Team-analytics privacy guidanceMinimum-cohort floors prevent re-identifying individuals from aggregatesMinimum group size for reporting
DX (now part of Atlassian)2026400+ organizations, 14 monthsAI value is multi-dimensional and resists collapsing into a single productivity numberLongitudinal engineering velocity across dimensions

Study by study

METR: Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity

2025–26 · Experienced open-source developers, randomized controlled trial

Effect size: 19% slower on real tasks while believing they were 20% faster (~39-point perception gap); a 2026 follow-up found −18% / −4% with ranges crossing zero

The perception gap is the point: developers were about 39 percentage points off about their own AI speedup — they estimated a 20% gain on real tasks while the original measurement found a 19% slowdown. A follow-up did not reproduce that slowdown (METR, 24 Feb 2026 update): returning developers came in at -18% (range -38% to +9%) and new developers at -4% (range -15% to +9%), both ranges crossing zero, and METR now says developers are more sped up from AI tools in early 2026. Why we report ranges and never trust self-report.

Faros AI: Acceleration Whiplash

2026 · 22,000 developers

Effect size: Incidents per PR +242.7% (~3×); bugs per developer +54%

Teams accelerating epic throughput with AI saw incidents per PR rise roughly 3x. The finding that makes the cost of going faster non-negotiable.

GitClear: AI Copilot Code Quality Research

2025–26 · Large multi-year code-change corpus

Effect size: Rising code churn and copy-pasted (duplicated) code alongside assistant adoption

Large-corpus evidence of rising churn and copy-pasted code alongside assistant adoption. Grounds our churn and duplication metrics.

DORA: State of DevOps Report

2025 · DORA global survey cohort

Effect size: AI amplifies the system it lands in: strong delivery gets faster, weak delivery breaks faster

AI amplifies the system it lands in: strong delivery gets faster, weak delivery breaks faster. Why we compare against your existing DORA baseline.

LinearB · APEX: APEX Framework

2026 · Engineering-metrics framework (LinearB dataset)

Effect size: Documents the adoption J-curve and the review/verification tax; payback typically 6–18 months

Documents the adoption J-curve and the verification tax. Source of our 6–18 month payback window and the single-counting rule.

Measure it on your own team, honestly

The literature says the answer is team-specific. Spendassay reports your delivery gains next to the rework, bugs, and incidents they cost — as a range, never a single number.

More: how we measure · the complete AI spend management guide