antelier
MEASURED / PUBLISHED SEPTEMBER 2026

The useful part.
And the limits.

Thirty selected agent pull requests across nine repositories. 203 claims labelled by humans, then read by a deterministic checker. These are published source measurements, not fresh measurements from this redesign.

Fabricated citations0

Report #1 records 145 citation checks with no fabricated citations. Read the measured basis ↗

Absent / contradicted precision1.00

Published precision for labels that require attention. A selected golden set, not a population guarantee.

Present precision / recall0.784
0.618

The checker missed supported claims. Both figures remain visible, not only the stronger precision number.

Exact agreement61.6%

125 of 203 claims agreed with the human labels. Deterministic tier only; model-judged pass disabled.

Labeller agreement89.5%

12 PRs had a blind second labeller; 219 shared units, kappa 0.83. The other 18 had one labeller.

The checker marked 94 claims not checkable; humans marked 19. Report #1 explains the gap and the evidence the deterministic approach cannot see. Thirty selected PRs is not a representative population.

The reading, in the open.

TRUTH REPORT / 01

203 claims.
Five that were not true.

Three absent and two contradicted human labels. The report shows the cases, the missed detection, and results by repository and agent.

Read method, data, and findings ↗
TRUTH REPORT / 02

“I told you
at the beginning.”

39 agent PRs, 75 human comments, 16 corrections, one confirmed repeat. An example of why a repository-owned rule matters—not an incidence rate.

Read the correction and proposed rule ↗

Nine repositories. A published rubric.

oxc, Next.js, Home Assistant, VS Code, Kubernetes, PyTorch, Transformers, curl, and Go. Report #1 carries the complete per-repository table, agent attribution method, and limitations.

Explore the repository breakdown ↗

Recorded runs, not live status. Each example names its source and limits.