Postil

Resources

Model bench.

19 scored models and 1 unreachable model row use the same fixture corpus, evaluator, and binary. The raw report records the inputs needed to compare another run with this one.

Quality and total run cost

Every point uses the same 70 fixtures:57 seeded-defect pull requests and 13 clean pull requests. Cost is the USD total for that one screening run, not a provider rate card. Gate-verdict correctness excludes cases without a valid envelope, which the table records under No envelope. The chart describes this fixture corpus and screening route only.

Gate correctness (%)65%70%75%80%85%90%$0.05$0.10$0.25$0.50$1.00Total cost for 70 fixtures (USD, log scale), cheaper to the rightgpt-5.6-lunakimi-k2.7-codeglm-5.2
Gate-verdict correctness and total run cost for 19 scored models. Higher is more correct; right is less expensive. Each point is one screening run across 70 attempted fixtures. Gate correctness excludes cases without a valid envelope; the results table on the benchmark page lists those failures. Focus or hover a point to read its row.

Raw results

The table retains the denominators, validity failures, total cost, and p95 latency for every attempted model.

ModelDetectedGate correctSilent on cleanFalse findingsNo envelopeTotal costp95
openai/gpt-5.6-luna96.5%90.0%13/1300US$0.06321.4s
openai/gpt-5.6-terra94.7%85.7%13/1300US$0.64017.6s
moonshotai/kimi-k2.7-code94.7%82.6%13/1301US$0.51118.1s
z-ai/glm-5.294.7%76.8%13/1331US$0.63362.1s
openai/gpt-5.6-luna-pro93.0%85.7%13/1300US$0.30338.8s
moonshotai/kimi-k2.691.2%84.3%13/1310US$1.6157.8s
google/gemma-4-31b-it91.2%77.9%10/1332US$0.28511.7s
qwen/qwen3.8-27b91.2%73.9%13/1311US$0.98267.5s
mistralai/mistral-small-3.2-24b-instruct91.2%73.1%6/1373US$0.03427.9s
thinkingmachines/inkling-small89.5%82.1%13/1303US$0.24763.2s
google/gemma-3-27b-it87.7%79.4%4/1387US$0.04341.8s
deepseek/deepseek-v4-pro-081386.0%78.6%12/1320US$0.70548.4s
qwen/qwen3-32b86.0%77.8%8/13127US$0.13219.6s
meta/muse-glimmer-30b84.2%82.9%13/1310US$0.28154.5s
nvidia/nemotron-3.5-lightning84.2%79.4%13/1302US$0.17387.1s
google/gemini-3.7-flash75.4%80.3%13/1304US$0.23222.0s
bytedance-seed/seed-2.0-code73.7%83.3%10/13216US$0.660131.6s
deepseek/deepseek-v4-flash-073168.4%63.2%13/1302US$0.06574.4s
z-ai/glm-5.363.2%75.0%12/13018US$1.24134.6s
qwen/qwen3.8-maxNo zero-retention endpoint; every request refused before inference.

Interpreting the results

Detected is the share of defect fixtures with a finding overlapping the seeded path and line range. The evaluator does not verify the diagnosis. The report's scoring rules define the location match and unmatched-finding count. Gate correct is the share of valid envelopes whose gate result matches the fixture outcome. Silent on clean counts the clean pull requests the model left alone. False findings counts unmatched envelope findings across both clean and defect fixtures. No envelope counts fixture runs without a valid structured output.

Repeated screening runs

The report contains four screening runs for each of three models. The chart shows each observed detection rate and marks the degraded run separately. These repeated runs use screening routing and do not establish a ranking for a different provider route.

70%75%80%85%90%95%100%gpt-5.6-lunaopenai/gpt-5.6-luna run 1: 96.5% seeded-region hits; 0/70 unavailable resultsopenai/gpt-5.6-luna run 2: 91.2% seeded-region hits; 0/70 unavailable resultsopenai/gpt-5.6-luna run 3: 93.0% seeded-region hits; 0/70 unavailable resultsopenai/gpt-5.6-luna run 4: 75.4% seeded-region hits; 16/70 unavailable results21.1 pt spreadkimi-k2.7-codemoonshotai/kimi-k2.7-code run 1: 94.7% seeded-region hits; 1/70 unavailable resultsmoonshotai/kimi-k2.7-code run 2: 89.5% seeded-region hits; 1/70 unavailable resultsmoonshotai/kimi-k2.7-code run 3: 98.2% seeded-region hits; 0/70 unavailable resultsmoonshotai/kimi-k2.7-code run 4: 94.7% seeded-region hits; 1/70 unavailable results8.8 pt spreadglm-5.2z-ai/glm-5.2 run 1: 94.7% seeded-region hits; 1/70 unavailable resultsz-ai/glm-5.2 run 2: 89.5% seeded-region hits; 1/70 unavailable resultsz-ai/glm-5.2 run 3: 87.7% seeded-region hits; 2/70 unavailable resultsz-ai/glm-5.2 run 4: 89.5% seeded-region hits; 1/70 unavailable results7.0 pt spread
Each dot is one screening run against the same 70 fixtures. Runs with equal percentages stack vertically at the same horizontal position. Every dot contributes to the displayed range, including the hollow dot for Luna's run with 16 unavailable results out of 70 attempts. Its full range is 21.1 percentage points; no failed run is excluded. The case evidence lists output availability for every repeated run. Percentages use all 57 defect cases.

Cost measurement

Total cost is the observed USD charge for one full screening run. It includes the model output required for these fixtures and does not represent a provider price list or the cost of another corpus.

Provider route measurements

The screening results allow routed endpoints. This table records the screening p95 latency and the p95 latency observed on the named pinned provider routes for two models. A deployment with provider, retention, or price constraints needs measurements under the same contract.

Model and pinned providerDetectedGate correctp95 screeningp95 pinned
openai/gpt-5.6-luna on Azure89.5%85.5%21.4s17.6s
moonshotai/kimi-k2.7-code on CoreWeave94.7%85.5%18.1s58.2s

Scope

The fixtures are synthetic pull requests with seeded defects and clean changes. The report measures this review pipeline on those fixtures; it does not establish performance on a separate codebase, corpus, or provider route.

The fixture corpus is maintained with the reviewed system. Read the guide to comparing benchmarks when evaluating this report or another benchmark.

Run it yourself

Start with the public CLI benchmark README. It describes mock mode and live inference. Live mode spends real inference tokens.

cargo build --release
cd bench && bun install --frozen-lockfile
export MODEL_API_KEY=...
REVIEW_MODEL=openai/gpt-5.6-luna \
  bun run bench:live -- --json-out .runs/luna.json

Run one model per report. Compare only runs whose fixtureCorpusSha256 and evaluatorSha256 match yours; a different corpus is a different test.

Report identity and reproduction data

Report timestamp 2026-08-19; CLI version postil 0.8.26; fixture corpus 8e4c2cb9ad5a7efd; evaluator ebc72a8d7d60c08b; binary 37d0c7ef4b49ab27.

Download the raw report (JSON) or read the model catalogue and local inference guide.