Screening v0.1 · 12 models compared

Find the reviewer
worth running.

An open, practical comparison of code-review models in one real review harness. Quality first, then completion, speed, and cost.

12
models screened
04
complete cohorts
11
held-out reviews
$43.89
total sweep spend

Three models to keep testing.

Sol is the best balanced candidate, Grok 4.6 is the complete-cohort recall leader, and Opus 5 has the strongest partial recall signal. Opus remains conditional because two reviews were unscoreable.

01

GPT-5.6 Sol

openai/gpt-5.6-sol · Azure

60% 7/11 encountered gold matched
11/11 Complete
1.6 minper attempted review
$2.59 provider reported
02

Grok 4.6

x-ai/grok-4.6 · xAI

70% 8/11 encountered gold matched
11/11 Complete
6.2 minper attempted review
$3.87 provider reported
03

Claude Opus 5

anthropic/claude-opus-5 · Amazon Bedrock

100% 10/10 encountered gold matched
9/11 Partial
4.3 minper attempted review
$13.23 provider reported
Compare all 12 models

Useful findings, not comment volume.

A reviewer earns credit for identifying an evidence-backed defect once. Rephrasing it, guessing loudly, or padding the review does not help.

01

Quality

Defect recall, precision, clean-review restraint, severity judgment, and consistency across repeated attempts.

02

Accuracy

One-to-one matching, proof-backed gold findings, human adjudication, and visible uncertainty for incomplete labels.

03

Speed

Latency is reported after review quality, with failures and budget exhaustion kept separate from successful reviews.

04

Affordability

Observed token and cost data where available. Unknown cost is null—never quietly treated as free.

Separate the model from the product.

A

Fixed harness / model

Hold prompt, tools, budgets, context, scorer, and task release still. Change the model or provider endpoint.

Screening live
B

End-to-end product

Measure the whole review system: its model, prompt, tools, parsing, and orchestration. OpenRouter Review Bot and Grok Code Review Bot are first.

Planned

Every public number carries its receipts.

Useful now.
Designed to improve.

The first sweep is a small, one-pass screening run—not a universal measure of coding ability. It is already useful for choosing models in the OpenRouter Review Bot, and future runs will add harder cases, better provider controls, and repeated finalist tests.