Publication pipeline active · no official comparison yet

Code review claims
should survive review.

A reproducible benchmark for the quality, accuracy, speed, and cost of automated code review—measured in that order.

00
immutable records
00
official records
00
published configurations
00
tracks represented

Useful findings, not comment volume.

A reviewer earns credit for identifying an evidence-backed defect once. Rephrasing it, guessing loudly, or padding the review does not help.

01

Quality

Defect recall, precision, clean-review restraint, severity judgment, and consistency across repeated attempts.

02

Accuracy

One-to-one matching, proof-backed gold findings, human adjudication, and visible uncertainty for incomplete labels.

03

Speed

Latency is reported after review quality, with failures and budget exhaustion kept separate from successful reviews.

04

Affordability

Observed token and cost data where available. Unknown cost is null—never quietly treated as free.

Separate the model from the product.

A

Fixed harness / model

Hold prompt, tools, budgets, context, scorer, and task release still. Change the model or provider endpoint.

Planned
B

End-to-end product

Measure the whole review system: its model, prompt, tools, parsing, and orchestration. OpenRouter Review Bot and Grok Code Review Bot are first.

Awaiting run

Every public number carries its receipts.

The plumbing is public.
The comparison waits for evidence.

The schemas, calibration fixtures, scorer, and methodology are available now. Primary comparisons begin only after the real-PR corpus and adjudication gates are ready.