Pilot systems online · no official ranking yet

Code review claims
should survive review.

A reproducible benchmark for the quality, accuracy, speed, and cost of automated code review—measured in that order.

02
review products initially
02
public calibration tasks
06
private pilot tasks
03+
attempts per official task

Useful findings, not comment volume.

A reviewer earns credit for identifying an evidence-backed defect once. Rephrasing it, guessing loudly, or padding the review does not help.

01

Quality

Defect recall, precision, clean-review restraint, severity judgment, and consistency across repeated attempts.

02

Accuracy

One-to-one matching, proof-backed gold findings, human adjudication, and visible uncertainty for incomplete labels.

03

Speed

Latency is reported after review quality, with failures and budget exhaustion kept separate from successful reviews.

04

Affordability

Observed token and cost data where available. Unknown cost is null—never quietly treated as free.

Separate the model from the product.

A

Fixed harness / model

Hold prompt, tools, budgets, context, scorer, and task release still. Change the model or provider endpoint.

Planned
B

End-to-end product

Measure the whole review system: its model, prompt, tools, parsing, and orchestration. OpenRouter Review Bot and Grok Code Review Bot are first.

Pilot

Every public number carries its receipts.

The plumbing is public.
The leaderboard waits for evidence.

The schemas, calibration fixtures, scorer, and methodology are available now. Official rankings begin only after the real-PR corpus and adjudication gates are ready.