Slow enough
to be credible.

This is a volunteer open-source project. The roadmap favors a small defensible benchmark over a large fragile leaderboard.

01

Contracts and calibration

complete

Public schemas, neutral task format, deterministic scorer, planted defects, clean twin, private pilot shape, and exact-SHA product adapters.

02

Real-PR corpus

active

Source license-compatible changes, reconstruct exact checkouts, validate defects with tests and fixing evidence, add clean controls, and blind the held-out split.

03

Judge calibration

next

Compare semantic matching against blinded human dispositions, publish a conformance set, and measure disagreement and variance.

04

Pilot runs

next

Run both products at least three times per task, adjudicate every unmatched finding, audit costs and provider routing, and rehearse private-to-public distillation.

05

First official release

next

Freeze the release, reproduce the run, publish reviewed aggregate records, and surface confidence intervals instead of overstating small score gaps.

No platform theater.