Methodology
How Forge measures AI authorship — and where it is wrong
The attributed share is a lower bound. The comparison group contains AI-assisted work we cannot see. The set we can detect skews toward agentic tools. All three are stated on every surface that shows the number, and all three are below in detail.
What is actually detected
Attribution comes only from marks a tool leaves on a commit. Nothing is inferred from diff shape, code style, commit size, or timing — those would be guesses dressed as measurement.
| Signal | Strength | What it is |
|---|---|---|
| Co-author trailer | Strong | A Co-Authored-By trailer whose value names a known agent. |
| Commit identity | Strong | The commit's git author or committer name/email is a known agent. |
| Bot account | Strong | The account that authored the commit is a known agent bot. |
| Branch prefix | Weak | The head branch carries an agent prefix. Never increments a count. |
| PR author | Weak | The pull request was opened by an agent bot. Never increments a count. |
Only strong signals can increment a count. A branch prefix says nothing about which commits an agent wrote, so letting it raise the number would fabricate attribution. Review bots that comment on pull requests without authoring them are deliberately excluded — counting them would fold review load into authorship and corrupt the exact comparison being made.
Limitation 1 — inline completion leaves no trace
This is the big one. A developer who accepts hundreds of inline completions commits under their own identity, with no trailer. That work is genuinely AI-written and lands in the unattributed bucket. There is no repository artifact that separates it from hand-typed code.
Two consequences: the attributed share understates AI involvement, and the comparison group is contaminated. If unattributed-but-AI-assisted pull requests also cost more review, they pull the comparison group upward and the measured gap narrows — so the reported multiple is, if anything, conservative.
Leaves a commit trail
Claude Code, GitHub Copilot coding agent, Cursor agent, OpenAI Codex, Devin, Aider, Google Jules, Gemini CLI, Amp, OpenHands
Undetectable by design
GitHub Copilot completions, Cursor tab completion, Windsurf, Sourcegraph Cody, Continue, Tabnine, Supermaven
Limitation 2 — the detectable set skews agentic
Tools that stamp attribution are disproportionately the agentic ones. Tools that do not are disproportionately inline-completion. So the attributed population is not a random sample of AI-written code — it is a sample of agent-delegated work, which tends to be larger-scoped and handed off more autonomously.
Read every result as “work delegated to a coding agent costs N× the review”, not “AI code costs N× the review”. The narrower claim is the one the data supports.
Limitation 3 — selection runs the other way
Teams hand agents the work they least want to do by hand, which is often the work that was already hardest to review. A higher review cost on attributed pull requests is consistent with agents producing harder-to-review code AND with agents being pointed at harder problems. This measurement cannot separate those, and no surface here claims to.
Sizing the blind spot
Detection cannot measure its own gap, so teams using Forge can record which tools they actually use and roughly what share of their merged code is AI-assisted. When the self-reported share lands within 15 points of the detected share, the read is marked corroborated and drops the ≥. Otherwise it stays a floor and names the tools it cannot see.
There is one more measurable tell: pull requests that came off an agent branch while none of their commits carried attribution. Those are counted and reported separately as evidence of undercounting — never added to the attributed total.
When a number is withheld
- A measure is reported only when both sides have at least 3 pull requests that can supply it. A one-sided row invites being read as a comparison, so it is omitted entirely rather than shown as a zero.
- When no commit carries attribution, the share is withheld rather than reported as 0% — absence of a signal is not evidence of absence.
- A report needs 10 analyzed pull requests to enter the benchmark, and the benchmark needs 12 qualifying reports before any percentile is published.
- Measures the two report sources compute differently are excluded from pooled statistics, even where both can produce a number.
The capacity audit — what the holiday numbers rest on
A second measurement lives here: how much of a planning window a team actually has once each member’s own country is subtracted, and whether what they have committed fits inside it. Same posture as above — the limitations are the interesting part, and this one is a calculator you can run.
- Holidays are computed from national rules, not fetched from a dataset and not shipped as a table. 14 countries are covered, each with its own rule set, so any year resolves without a runtime dependency on anyone else’s data. A country is listed only when its whole national set is derivable from rules that hold for any year. Where it is not — a lunar or state-gazetted calendar, or an Orthodox Paschalion this engine does not compute — the country is left out rather than shipped as a partial list, because a partial list is not a conservative floor here, it is a wrong number.
- Rules encode statute, and statute changes. Juneteenth arrived in 2021, Ireland’s St Brigid’s Day in 2023, Poland’s Christmas Eve in 2025 — so a rule set is a dated snapshot, and an undated one decays without anyone noticing. Each country below carries the law it encodes and when that law was last read. A monthly job re-checks every country against an independent reference and reports disagreements; it never rewrites a rule from one, because a community-maintained source is a reason to go read the statute, not an authority to adopt silently.
- Regional, state, municipal and religious holidays are excluded. That gap is largest in Germany, where Epiphany, Corpus Christi, Assumption, Reformation Day and All Saints are set per Bundesland; in Spain, where every worker also gets regional and municipal fiestas; in Canada, where provincial holidays vary widely; and in the UK, where Scotland and Northern Ireland differ from England and Wales. Each country carries its own caveat naming exactly what it leaves out, and that caveat is rendered beside every number derived from its dates.
- Weekend-substitution rules differ per country and are encoded per country rather than defaulted. The UK, Ireland and Canada move a holiday forward to the next free weekday; US federal practice observes a Saturday holiday on the Friday before; Spain transfers a Sunday holiday to the following Monday, as the Estatuto de los Trabajadores requires, and loses a Saturday one; most of continental Europe simply loses the day. Where the entitlement is real but has no computable date — Poland’s employer-scheduled compensating day — nothing is moved, which understates shrinkage rather than overstating it.
- Feasts that always fall on a weekend are omitted where the country loses such a day anyway, since it never consumes a working day. For a country that substitutes, the day is emitted so its rule can move it onto a weekday — dropping it would hide a day the team really loses. Any holiday whose observed date lands on a weekend inside the audited window is left out of the count too, so a count here can legitimately be shorter than an official calendar.
- Time off is stated, never observed. Booked vacation and part-time arrangements come from the person filling in the form; nothing infers them. The published country pages assume no time off at all, so the shrinkage they show is a floor.
- The commitment check uses the team’s own pace and nothing else. There is no dataset of how often commitments like yours land, so no such claim is computed, shown or implied anywhere. Pace is the points the team says its recent sprints completed, divided by the person-days a full calendar would have given that team — not by this window’s available days. Dividing by the available days would cancel against the correction and leave the verdict identical however many holidays the window held. Below 2 sprints there is nothing to average, and the points verdict is withheld rather than filled in from a default: story points carry no meaning across teams for a default to borrow.
What each country’s rules encode, and when that was last read
| Country | Statute | Last read |
|---|---|---|
| United States | 5 U.S.C. §6103 | 2026-08-03 |
| Canada | Canada Labour Code s.195 | 2026-08-03 |
| United Kingdom | Banking and Financial Dealings Act 1971 | 2026-08-03 |
| Ireland | Organisation of Working Time Act 1997 s.21 | 2026-08-03 |
| Germany | Feiertagsgesetze of the Länder | 2026-08-03 |
| Netherlands | Algemene termijnenwet art. 3 | 2026-08-03 |
| France | Code du travail art. L3133-1 | 2026-08-03 |
| Spain | Estatuto de los Trabajadores art. 37.2 | 2026-08-03 |
| Portugal | Código do Trabalho art. 234 | 2026-08-03 |
| Italy | Legge 260/1949 | 2026-08-03 |
| Poland | Ustawa o dniach wolnych od pracy (1951) | 2026-08-03 |
| Czechia | Zákon č. 245/2000 Sb. | 2026-08-03 |
| Sweden | Lag om allmänna helgdagar (1989:253) | 2026-08-03 |
| Brazil | Lei 662/1949 and Lei 10.607/2002 | 2026-08-03 |
What is read, and what is stored
A public scan reads only public repositories, through GitHub’s API, up to the 100 most recently merged pull requests. A private repository is refused, not read. Reports published by a team from their own repositories carry no repository name, no member, and no organization identifier — only a coarse team-size band the publisher chooses.
Scanner IP addresses are salted-hashed before storage and used only for rate limiting. Published payloads are frozen at publish time, so a report always shows the numbers it was published with.
Found a flaw in this method? That is genuinely useful — tell us.
Want to check the method against your own repository — including a private one? The same read runs locally through your own GitHub credentials, and sends nothing to us:
npx @ambera/review-tax