The useful part.
And the limits.
Thirty selected agent pull requests across nine repositories. 203 claims labelled by humans, then read by a deterministic checker. These are published source measurements, not fresh measurements from this redesign.
Report #1 records 145 citation checks with no fabricated citations. Read the measured basis ↗
Published precision for labels that require attention. A selected golden set, not a population guarantee.
0.618
The checker missed supported claims. Both figures remain visible, not only the stronger precision number.
125 of 203 claims agreed with the human labels. Deterministic tier only; model-judged pass disabled.
12 PRs had a blind second labeller; 219 shared units, kappa 0.83. The other 18 had one labeller.
The reading, in the open.
203 claims.
Five that were not true.
Three absent and two contradicted human labels. The report shows the cases, the missed detection, and results by repository and agent.
Read method, data, and findings ↗“I told you
at the beginning.”
39 agent PRs, 75 human comments, 16 corrections, one confirmed repeat. An example of why a repository-owned rule matters—not an incidence rate.
Read the correction and proposed rule ↗Nine repositories. A published rubric.
oxc, Next.js, Home Assistant, VS Code, Kubernetes, PyTorch, Transformers, curl, and Go. Report #1 carries the complete per-repository table, agent attribution method, and limitations.
Explore the repository breakdown ↗