Resources
Model bench.
19 scored models and 1 unreachable model row use the same fixture corpus, evaluator, and binary. The raw report records the inputs needed to compare another run with this one.
Quality and total run cost
Every point uses the same 70 fixtures:57 seeded-defect pull requests and 13 clean pull requests. Cost is the USD total for that one screening run, not a provider rate card. Gate-verdict correctness excludes cases without a valid envelope, which the table records under No envelope. The chart describes this fixture corpus and screening route only.
Raw results
The table retains the denominators, validity failures, total cost, and p95 latency for every attempted model.
| Model | Detected | Gate correct | Silent on clean | False findings | No envelope | Total cost | p95 |
|---|---|---|---|---|---|---|---|
openai/gpt-5.6-luna | 96.5% | 90.0% | 13/13 | 0 | 0 | US$0.063 | 21.4s |
openai/gpt-5.6-terra | 94.7% | 85.7% | 13/13 | 0 | 0 | US$0.640 | 17.6s |
moonshotai/kimi-k2.7-code | 94.7% | 82.6% | 13/13 | 0 | 1 | US$0.511 | 18.1s |
z-ai/glm-5.2 | 94.7% | 76.8% | 13/13 | 3 | 1 | US$0.633 | 62.1s |
openai/gpt-5.6-luna-pro | 93.0% | 85.7% | 13/13 | 0 | 0 | US$0.303 | 38.8s |
moonshotai/kimi-k2.6 | 91.2% | 84.3% | 13/13 | 1 | 0 | US$1.61 | 57.8s |
google/gemma-4-31b-it | 91.2% | 77.9% | 10/13 | 3 | 2 | US$0.285 | 11.7s |
qwen/qwen3.8-27b | 91.2% | 73.9% | 13/13 | 1 | 1 | US$0.982 | 67.5s |
mistralai/mistral-small-3.2-24b-instruct | 91.2% | 73.1% | 6/13 | 7 | 3 | US$0.034 | 27.9s |
thinkingmachines/inkling-small | 89.5% | 82.1% | 13/13 | 0 | 3 | US$0.247 | 63.2s |
google/gemma-3-27b-it | 87.7% | 79.4% | 4/13 | 8 | 7 | US$0.043 | 41.8s |
deepseek/deepseek-v4-pro-0813 | 86.0% | 78.6% | 12/13 | 2 | 0 | US$0.705 | 48.4s |
qwen/qwen3-32b | 86.0% | 77.8% | 8/13 | 12 | 7 | US$0.132 | 19.6s |
meta/muse-glimmer-30b | 84.2% | 82.9% | 13/13 | 1 | 0 | US$0.281 | 54.5s |
nvidia/nemotron-3.5-lightning | 84.2% | 79.4% | 13/13 | 0 | 2 | US$0.173 | 87.1s |
google/gemini-3.7-flash | 75.4% | 80.3% | 13/13 | 0 | 4 | US$0.232 | 22.0s |
bytedance-seed/seed-2.0-code | 73.7% | 83.3% | 10/13 | 2 | 16 | US$0.660 | 131.6s |
deepseek/deepseek-v4-flash-0731 | 68.4% | 63.2% | 13/13 | 0 | 2 | US$0.065 | 74.4s |
z-ai/glm-5.3 | 63.2% | 75.0% | 12/13 | 0 | 18 | US$1.24 | 134.6s |
qwen/qwen3.8-max | No zero-retention endpoint; every request refused before inference. | ||||||
Interpreting the results
Detected is the share of defect fixtures with a finding overlapping the seeded path and line range. The evaluator does not verify the diagnosis. The report's scoring rules define the location match and unmatched-finding count. Gate correct is the share of valid envelopes whose gate result matches the fixture outcome. Silent on clean counts the clean pull requests the model left alone. False findings counts unmatched envelope findings across both clean and defect fixtures. No envelope counts fixture runs without a valid structured output.
Repeated screening runs
The report contains four screening runs for each of three models. The chart shows each observed detection rate and marks the degraded run separately. These repeated runs use screening routing and do not establish a ranking for a different provider route.
Cost measurement
Total cost is the observed USD charge for one full screening run. It includes the model output required for these fixtures and does not represent a provider price list or the cost of another corpus.
Provider route measurements
The screening results allow routed endpoints. This table records the screening p95 latency and the p95 latency observed on the named pinned provider routes for two models. A deployment with provider, retention, or price constraints needs measurements under the same contract.
| Model and pinned provider | Detected | Gate correct | p95 screening | p95 pinned |
|---|---|---|---|---|
openai/gpt-5.6-luna on Azure | 89.5% | 85.5% | 21.4s | 17.6s |
moonshotai/kimi-k2.7-code on CoreWeave | 94.7% | 85.5% | 18.1s | 58.2s |
Scope
The fixtures are synthetic pull requests with seeded defects and clean changes. The report measures this review pipeline on those fixtures; it does not establish performance on a separate codebase, corpus, or provider route.
The fixture corpus is maintained with the reviewed system. Read the guide to comparing benchmarks when evaluating this report or another benchmark.
Run it yourself
Start with the public CLI benchmark README. It describes mock mode and live inference. Live mode spends real inference tokens.
cargo build --release
cd bench && bun install --frozen-lockfile
export MODEL_API_KEY=...
REVIEW_MODEL=openai/gpt-5.6-luna \
bun run bench:live -- --json-out .runs/luna.jsonRun one model per report. Compare only runs whose fixtureCorpusSha256 and evaluatorSha256 match yours; a different corpus is a different test.
Report identity and reproduction data
Report timestamp 2026-08-19; CLI version postil 0.8.26; fixture corpus 8e4c2cb9ad5a7efd; evaluator ebc72a8d7d60c08b; binary 37d0c7ef4b49ab27.
Download the raw report (JSON) or read the model catalogue and local inference guide.