Postil

Blog

Choosing a code review model

· Postil team

Detection credit goes to a defect case when at least one finding overlaps the planted bug's location, and each case earns at most one. In the Postil screening benchmark, GPT-5.6 Luna receives it on 55 of 57 defect cases; GLM-5.2 receives it on 54.

Of the numbers this benchmark reports, this single-run location-detection count is the least useful standalone number for choosing a reviewer, because it measures line overlap alone. A comment on the right lines can receive credit while explaining the wrong problem, and a location match does not guarantee that a model assigns the severity that produces the correct merge decision.

A fixture is a prepared change with known files, contracts, and an expected outcome. The historical screening run submits 70 fixtures: 57 defect cases built around a known faulty file and line range, and 13 clean cases that need no finding under their supplied code and contracts.

Headline screening results

This run uses unconstrained provider routing, meaning the upstream provider serving each request is not fixed. An attempted case is any fixture submitted to the review pipeline; it is either scored, yielding a usable review to grade, or unavailable, yielding none. A gate verdict is the pass or fail decision the review's findings produce under the benchmark's severity policy; the benchmark's gate is error-only, so an error finding blocks a change and warnings alone let it pass.

One screening run per model, using unconstrained provider routing
MeasurementGPT-5.6 LunaGLM-5.2Denominator
Defect cases with a location match555457 defect cases
Correct gate verdicts / scored cases6353Scored cases (70 for Luna, 69 for GLM)
Cases without a scored result0170 attempted cases
Recorded run cost$0.0628$0.6329one 70-case benchmark run
95th-percentile case latency21.4s62.1sScored cases' final invocation

These totals cover one benchmark run, not a hosted subscription or a typical pull request. The 95th-percentile latency times each scored case's final invocation, from diff preparation to construction of the result, including model calls and internal repairs; it excludes unavailable cases, earlier attempts, retry backoff, process startup, and the surrounding CI or developer workflow.

GLM's one unscored case is a cache-expiry change that returns milliseconds where seconds are required. Its output is an operational error labeled review/invalidOutput, so the pipeline has no usable review to score; the label does not identify the underlying cause. The case stays in the 70 attempted cases, earns no detection credit, and drops out of gate scoring, leaving GLM with 69 scored cases.

Repeated runs and provider routing

The headline run is one of four on the same corpus, and the four do not agree.

Location-detection rates across repeated runs (denominator: 57 defect cases per run)
ModelRun 1Run 2Run 3Run 4
GPT-5.6 Luna96.5%91.2%93.0%75.4% (output failures)
GLM-5.294.7%89.5%87.7%89.5%

The headline run's rates match Run 1: 55 of 57 matches is 96.5%, and 54 of 57 is 94.7%. In Run 4, Luna degrades because of invalid output: 16 of its 70 attempts are unavailable, including five clean cases; the remaining eleven are defect cases that earn no detection credit. Luna's observed detection range is 21.1 percentage points across its four runs.

Luna's and GLM's detection rates overlap across these runs, so the first run does not establish a stable ranking between the two models. These runs all use unconstrained provider routing; separate provider-pinned results, where the upstream provider is fixed, describe different conditions and must not be substituted into this comparison.

Gate outcomes and severity policy

A gate is an automated rule that turns review findings into a pass or fail merge decision. An authored verdict is the pass or fail decision that a fixture's author expects under this severity policy. Gate correctness asks whether the review result agrees with that authored verdict. A model can find a defect's exact location and still assign a severity, or a finding kind, that produces the wrong gate action; the gate reads both.

Gate outcomes across all 70 attempted cases relative to authored verdicts
OutcomeGPT-5.6 LunaGLM-5.2Denominator
Correct pass181970 attempted cases
Correct block453470 attempted cases
Pass when the fixture requires a block21270 attempted cases
Block when the fixture expects a pass5470 attempted cases
Unavailable review result0170 attempted cases

The largest operational difference between the models appears in cases that should block under benchmark policy: Luna passes 2 of these changes, GLM passes 12.

In the archive-deletion case, both models identify the replacement of archiveFile with storage.delete but label the finding a warning. The fixture requires preserving the recovery copy and assigns error severity; because the gate only blocks on errors, both models let a change pass when the fixture expects a block.

Disagreement also runs the other way. In the disabled-timeout fixture, the authored verdict is a warning, which should pass an error-only gate. Both models label the finding an error, so both gates block the change.

These outcomes record agreement with this benchmark's severity policy, not a universal rule about which risks every repository must accept. The case evidence records the authored targets, source excerpts, model findings, and gate decisions for each case.

Clean cases across two experiments

A silent final review is a final review result that contains no findings. A suppressed candidate is a concern a model raises internally that Postil's pipeline discards before producing the final review. Because suppression happens inside the pipeline, a silent final review does not by itself show whether a model generated any candidate concern at all. The benchmark tests clean behavior in two separate experiments: the historical 13-case sample and an expanded 25-case bank; neither separates the two models on final review silence.

The 13 clean cases, drawn from the historical 70-case screening report, include comment and documentation edits, variable and type renames, test maintenance, formatting updates, and safe changes to authorization, cache expiry, and concurrent fetching; a variable-rename fixture, for example, alters a name while preserving the return value. The fixture suite contains the changes and expected results. In this run, both models produce 13 silent final reviews out of 13 clean cases, with no findings and no unavailable results. This sample is small and curated.

Clean cases with a silent final Postil review
ResultGPT-5.6 LunaGLM-5.2
Silent final reviews13 / 1313 / 13
Final reviews with findings00
Unavailable reviews00
The 13 clean cases from the historical 70-case screening report. Any retained finding counts against silence, even a nonblocking warning. Suppressed candidates are separate from final findings. Unavailable means no usable review result and does not count as silence.

The expanded clean bank adds 12 executable examples to those same 13 cases, forming a separate 25-case corpus on a different binary. These additions include extracting a tenant check without altering read permissions, preserving a boundary cache-expiry limit, and keeping a zero-valued retry setting distinct from an absent one. In one example, replacing an explicit null-or-undefined check with config.retries ?? 3 preserves zero as a request for no retries, where a truthiness check would instead turn zero into three; the fixture tests whether a model accepts that simplification under a stated configuration contract.

Clean cases with a silent final Postil review
ResultGPT-5.6 LunaGLM-5.2
Silent final reviews25 / 2525 / 25
Final reviews with findings00
Unavailable reviews00
Suppressed findings01
Provider routeazure/euz-ai/fp8
One run per model on the separate 25-case clean bank, using the same inputs for both models. Any retained finding counts against silence, even a nonblocking warning. Suppressed candidates are separate from final findings. Unavailable means no usable review result and does not count as silence.

Both models see identical fixture inputs in this experiment, with matching evaluator and binary hashes, three concurrent cases, no outer retries, and disabled provider fallbacks. Luna uses the Azure EU provider route; GLM uses the Z.AI route.

The recorded charges are $0.003744202 for Luna and $0.0387586 for GLM; these figures describe only this clean-only experiment and must not be substituted for the 70-case screening-run costs reported above. This 25-case run does not change the 57-defect and 13-clean denominators of the historical report; it measures the complete review pipeline, including suppression, rather than whether either model generates any candidate concern at all.

Supplied runtime contracts shape these results. In an optional-field fixture without an explicit runtime guarantee, Luna retains a compatibility warning about Object.hasOwn in its final review; that warning is not an established false alarm, because the fixture leaves runtime support for the method unspecified. The 25-case comparison states that runtime guarantee directly in the diff, and the final review stays silent. The evidence file keeps the result from the variant without the guarantee so the two inputs are not mistaken for repeated runs of the same fixture. One observation per case and model, in the 25-case bank, cannot establish a stable false-positive rate.

The extra-findings counter

The benchmark report records falsePositives: 3 for GLM. That name is misleading. The evaluator awards at most one detection credit per defect case, and this counter increments on every additional finding on that case, whether the finding duplicates the credited match or raises an unrelated concern; it is an extra-findings counter, not a tally of verified false alarms. The versioned case evidence shows that GLM's 3 extra findings consist of one duplicate diagnosis and two unverified compatibility questions.

In the configuration-loader case, the code replaces return defaultConfig with throw err beside a comment that still promises fallback to defaults. GLM reports both the misleading comment and the lost fallback behavior at the same line, src/config/load.ts:26. Both reports describe the same seeded fallback-contract defect; the second is a duplicate, not a separate bug.

In the provider-client case, the seeded defect disables a request timeout. GLM also flags a metric rename from provider.request to provider.requests and an Accept header change from application/json to application/vnd.api+json. Its metric warning asserts that existing consumers lose data, though the supplied context identifies neither such consumers nor a stable-name contract. Its header warning holds only if endpoints reject the requested type; the supplied endpoint contract does not establish that rejection. Both are compatibility questions, not demonstrated failures in the supplied code.

Choosing a reviewer

The location-detection score shows only that both models produce findings overlapping the seeded defect locations. It does not show whether a model explains a defect correctly, whether it assigns a severity that matches a repository's merge rules, or how consistently it produces usable output across repeated runs, and it says nothing about recorded cost or response time.

The aggregate report identifies the binary, fixture corpus, and evaluator behind these figures, and the suite instructions explain how to run the harness.

Choosing a code review model for a repository also requires evidence that its findings explain real defects and that its gate decisions fit that repository's own policy. A high location-detection score supplies neither of those judgments.