Does the review
help the next fix?

Astra Flex caught useful follow-up problems in this small live trial. That is a reason to keep testing it, not a verdict on every model.

Three PRs, not fifteen independent experiments

This snapshot contains 15 posted review rounds with all three models configured, across 3 merged PRs: one feature PR and two configuration PRs. It was selected from a private capture of 19 merged PRs and 63 archived rounds. The merge window is September 5, midnight–4:43 p.m. Eastern; all archived rounds of selected PRs are included. Unposted or still-open work is outside this capture.

The configuration work involved repeated migrations and recovery attempts. Its ten trial rounds are kept separate from the feature PR's five rounds. Initial and follow-up reviews also stay separate because their scope and budgets differ. Same-round models share the change being reviewed, but this was not randomized and the PR evolved between rounds.

What we actually checked

Four deliberately selected finding appearances received an additional AI audit with executable code witnesses. Three Astra follow-up findings were supported while both peer lanes completed with no new findings. They concerned an incomplete fix or a regression introduced by a fix. One Grok claim was rebutted by a native-runtime control. This is a selected case study: the other appearances remain unassessed, and these counts do not establish precision, recall, or three independent new bugs.

Six new private before/fixed packets now cover those three mechanisms, alongside a separate negative control. The evaluator fails the target invariant before each fix and passes it afterward. Packets retain exact source and evaluator hashes, keep answers outside reviewer input, and test whether an author's rebuttal agrees with the code. They have not been run against models and are development regressions, not additions to the held-out leaderboard.

Same-round observations

Paired medians use only rounds where all three lanes completed with recorded durations. Completion and known cost cover every attempt in that group, including unpaired rounds. A failed lane is never assigned a zero-second duration or a free request.

Feature work · follow-up verification

1 PR(s) · 3 rounds · 3 fully timed paired round(s)

ModelCompletedPaired medianKnown cost, all attempts
Grok 4.6 3/3 24.2s $0.2402
GLM 5.3 Flash 3/3 47.6s $0.0156
GPT-6 Astra · Flex 3/3 18.6s $0.6900

Feature work · initial review

1 PR(s) · 2 rounds · 1 fully timed paired round(s)

ModelCompletedPaired medianKnown cost, all attempts
Grok 4.6 1/2 134.6s $0.7300 + unknown (1 attempt)
GLM 5.3 Flash 1/2 575.9s $0.0275 + unknown (1 attempt)
GPT-6 Astra · Flex 2/2 39.8s $2.4000

Configuration · follow-up verification

2 PR(s) · 7 rounds · 7 fully timed paired round(s)

ModelCompletedPaired medianKnown cost, all attempts
Grok 4.6 7/7 12.3s $0.1695
GLM 5.3 Flash 7/7 57.1s $0.0191
GPT-6 Astra · Flex 7/7 14.0s $0.8361

Configuration · initial review

2 PR(s) · 3 rounds · 3 fully timed paired round(s)

ModelCompletedPaired medianKnown cost, all attempts
Grok 4.6 3/3 259.1s $1.1208
GLM 5.3 Flash 3/3 326.5s $0.0470
GPT-6 Astra · Flex 3/3 40.4s $1.1800

Useful evidence, with a narrow reach

Astra returned structured results in all 15 captured trial attempts, each with explicit Flex confirmation. The most useful comparison here is feature follow-up verification: three matched rounds, with median lane times of 18.6 seconds for Astra, 24.2 for Grok, and 47.6 for GLM. Astra's reported cost for those three attempts was $0.69, versus $0.2402 and $0.0156 respectively. Flex was fast in this sample; it was not the cheapest lane.

GLM was the slowest completed lane in 13 of the 14 fully timed matched rounds here. That suggests testing a shorter or selective GLM pass, but this sample cannot show what defects would be lost by removing it. The feature initial-review comparison has only one fully timed round; another initial round had two failed lanes and unknown charges. No statistical significance or general latency promise is claimed.

Times are reported lane durations, not whole workflow or time-to-merge measurements. Costs are rounded amounts from posted reviews and exclude the judge, fixing agents, runner minutes and unposted work. Resolved threads and clean verification are not correctness labels. This analysis spent no new model tokens; its evidence came from reviews already run and local executable checks.

The next decision should use marginal useful findings and total loop delay across more ordinary feature work. The current evidence supports keeping Astra in a bounded trial; it does not yet justify automatically replacing Grok or dropping GLM.

View the separate benchmark leaderboard → · Read the retrospective method ↗