Benchmark
What AI-authored code costs teams in review
The distribution of review cost across published Ambera Forge reports: how many more review rounds pull requests carrying agent authorship take, and how much of merged code carries attribution at all.
Attributed share of merged commits
≥1%
25th pct
≥3%
Median
≥11%
75th pct
392 teams
Time to merge, attributed ÷ rest
0.2×
25th pct
0.6×
Median
1.1×
75th pct
211 teams
Drawn from 392 reports — 392 public-repository scans and 0 published by teams from their own repositories — covering 38,457 merged pull requests.
How to read this honestly
- It is a self-selected sample. Every report here exists because somebody chose to scan a repo or publish their team’s numbers. That is a sample of curiosity, not of software.
- Attributed share is a floor. Only agents that sign their commits can be counted, so “the rest” contains AI-assisted work no repository artifact can reveal.
- It measures delegated work, not all AI-written code. Tools that stamp attribution are disproportionately the agentic ones, so read every multiple as the cost of work handed to an agent.
- Causation is not on offer. Teams hand agents the work that was already hardest to review. This distribution shows co-movement; it does not establish direction.
- Only comparable measures are pooled. Measures the two report sources compute differently are excluded, and a report needs at least 10 analyzed pull requests to count.
- Degenerate ratios are dropped. A team’s ratio joins a distribution only when its comparison side is large enough to divide by — on a repo where almost nobody requests changes, 0.0 ÷ 0.06 is a rounding error, not a multiple, and a median built from those would be too.