OpenRouter Review Bot
A multi-model OpenRouter harness with read-only repository tools and structured finding output.
View repository ↗The leaderboard is intentionally empty until real-PR tasks, repeated product runs, and held-out adjudication clear the release gates.
End-to-end product track
Calibration output can validate plumbing, but it cannot establish that one reviewer is better than another. The first row appears only after the corpus and publication gates are satisfied.
A multi-model OpenRouter harness with read-only repository tools and structured finding output.
View repository ↗A Grok CLI review product running with a strict Linux sandbox and a deliberately limited read-only tool set.
View repository ↗How to read future rows
Quality leads. Recall without precision rewards noise. Precision without recall rewards silence. Clean controls test restraint.
Intervals beat tiny gaps. Repeated runs and uncertainty matter more than ordinal rank when scores are close.
Conditions define a row. Model, provider, prompt, tools, budget, product SHA, task release, and scorer are part of the identity.