Results status

Evidence first.
Rankings later.

The leaderboard is intentionally empty until real-PR tasks, repeated product runs, and held-out adjudication clear the release gates.

No official runs yet

Pilot

This blank table is a feature.

Calibration output can validate plumbing, but it cannot establish that one reviewer is better than another. The first row appears only after the corpus and publication gates are satisfied.

Review the gates

Two products. Same frozen work.

OR

OpenRouter Review Bot

A multi-model OpenRouter harness with read-only repository tools and structured finding output.

View repository ↗
GX

Grok Code Review Bot

A Grok CLI review product running with a strict Linux sandbox and a deliberately limited read-only tool set.

View repository ↗

Quality leads. Recall without precision rewards noise. Precision without recall rewards silence. Clean controls test restraint.

Intervals beat tiny gaps. Repeated runs and uncertainty matter more than ordinal rank when scores are close.

Conditions define a row. Model, provider, prompt, tools, budget, product SHA, task release, and scorer are part of the identity.