Claude Opus 5
anthropic/claude-opus-5 · Amazon Bedrock
An open, practical comparison of code-review models in one real review harness. Quality first, then completion, speed, and cost.
Current recommendation
After model-blind review of every unmatched finding, Claude Opus 5 led assessed recall at 80.0% but reached only 12.1% assessed precision. GLM 5.3 followed at 70.0%. Grok 4.6 offered the strongest high-recall balance at 60.0% recall and 42.8% precision; GPT-5.6 Sol traded a little quality for much better speed and lower cost.
anthropic/claude-opus-5 · Amazon Bedrock
z-ai/glm-5.3 · Fireworks
x-ai/grok-4.6 · xAI
A reviewer earns credit for identifying an evidence-backed defect once. Rephrasing it, guessing loudly, or padding the review does not help.
Defect recall, precision, clean-review restraint, severity judgment, and consistency across repeated attempts.
One-to-one matching, proof-backed gold findings, clearly labeled evidence review, and visible uncertainty for incomplete labels.
Latency is reported after review quality, with failures and budget exhaustion kept separate from successful reviews.
Observed token and cost data where available. Unknown cost is null—never quietly treated as free.
Hold prompt, tools, budgets, context, scorer, and task release still. Change the model or provider endpoint.
Measure the whole review system: its model, prompt, tools, parsing, and orchestration. OpenRouter Review Bot and Grok Code Review Bot are first.
Current state
The current board is a small, three-attempt model comparison—not a universal measure of coding ability. It is useful for choosing models in the OpenRouter Review Bot, and future runs can add harder cases and newly released models.