OpenRouter Review Bot model comparison v0.4 · 12 models compared

Find the reviewer
worth running.

An open, practical comparison of code-review models in one real review harness. Quality first, then completion, speed, and cost.

12
models screened
11
complete cohorts
33
reviews per model
$131.29
published spend

Claude Opus 5 leads the current quality-first ranking.

Claude Opus 5 currently ranks first at 80% known-issue recall and 25.4% AI-assessed provisional precision. Muse Spark 1.2 follows at 78.3% recall. Completion, speed, cost, provider behavior, and each row's limitations remain visible alongside quality.

01

Claude Opus 5

anthropic/claude-opus-5 · Amazon Bedrock

80% AI-assessed recall · baseline scorer recall 73.3%
33/33 Complete
6.1 minper attempted review
$41.42 provider reported
02

Muse Spark 1.2

meta/muse-spark-1.2 · Meta

78.3% AI-assessed recall · baseline scorer recall 75%
33/33 Complete
1.5 minper attempted review
$3.71 provider reported
03

GLM 5.3

z-ai/glm-5.3 · Fireworks

70% AI-assessed recall · baseline scorer recall 70%
33/33 Complete
17.6 minper attempted review
$14.20 provider reported
Compare all 12 models

Useful findings, not comment volume.

A reviewer earns credit for identifying an evidence-backed defect once. Rephrasing it, guessing loudly, or padding the review does not help.

01

Quality

Defect recall, precision, clean-review restraint, severity judgment, and consistency across repeated attempts.

02

Accuracy

One-to-one matching, proof-backed gold findings, clearly labeled evidence review, and visible uncertainty for incomplete labels.

03

Speed

Latency is reported after review quality, with failures and budget exhaustion kept separate from successful reviews.

04

Affordability

Observed token and cost data where available. Unknown cost is null—never quietly treated as free.

Separate the model from the product.

A

Fixed harness / model

Hold prompt, tools, budgets, context, scorer, and task release still. Change the model or provider endpoint.

12 models live
B

End-to-end product

Measure the whole review system: its model, prompt, tools, parsing, and orchestration. OpenRouter Review Bot and Grok Code Review Bot are first.

Planned

Every public number carries its receipts.

Useful now.
Designed to improve.

The current board is a small, 3-attempt model comparison—not a universal measure of coding ability. It is useful for choosing models in the OpenRouter Review Bot, and future runs can add harder cases and newly released models. When a row needs targeted repair, successful original reviews are preserved and only missing cells are rerun; that history and its all-in cost stay visible.