Quality
Defect recall, precision, clean-review restraint, severity judgment, and consistency across repeated attempts.
A reproducible benchmark for the quality, accuracy, speed, and cost of automated code review—measured in that order.
A reviewer earns credit for identifying an evidence-backed defect once. Rephrasing it, guessing loudly, or padding the review does not help.
Defect recall, precision, clean-review restraint, severity judgment, and consistency across repeated attempts.
One-to-one matching, proof-backed gold findings, human adjudication, and visible uncertainty for incomplete labels.
Latency is reported after review quality, with failures and budget exhaustion kept separate from successful reviews.
Observed token and cost data where available. Unknown cost is null—never quietly treated as free.
Hold prompt, tools, budgets, context, scorer, and task release still. Change the model or provider endpoint.
Measure the whole review system: its model, prompt, tools, parsing, and orchestration. OpenRouter Review Bot and Grok Code Review Bot are first.
Current state
The schemas, calibration fixtures, scorer, and methodology are available now. Official rankings begin only after the real-PR corpus and adjudication gates are ready.