Judging and Verdicts
The composite score, its five dimensions, the session floor, the tie band, and exactly what each verdict entitles you to do.
Judging and Verdicts
The paired judge
A trial, a replay, an experiment and a watch are all scored by the paired judge: one comparative call that sees the served answer and the candidate's answer together, at a randomized presentation order.
The order matters enough to be reported. Every candidate's verdict carries the observed split of how often it was shown first, which is what makes the randomization verifiable rather than asserted.
Trial comparisons are paired -- both arms answered the same request -- so the composition imbalances that force other measurements into stratification cannot arise between the arms.
Compare uses a different rubric, the absolute judge, which scores one answer on its own merits. Its scores are never read as a delta and are never written to the pair-score store at all. The two rubrics carry different judge versions, and that difference is not cosmetic.
The composite and its dimensions
Every answer is scored on five dimensions, each on a one-to-five scale. The composite is their mean, and it is the single number every delta, floor, band and verdict is expressed in.
| Dimension | |
|---|---|
quality | Overall answer quality. |
faithful | Faithfulness to the input. |
concise | Concision. |
instruct | Instruction following. |
clarity | Clarity. |
A verdict reports the composite delta and the per-dimension deltas, so you can see which dimension moved rather than only that something did.
This composite is not the efficiency score. They measure different things and share no scale.
The unit is the session
A session's judged pairs count as one observation, however many turns it has. Rows within a session are not independent, so counting them as rows would overstate how much evidence a verdict rests on.
Every verdict therefore reports sessions beside pairs, and a session count is what the floor is measured against.
Do not read a pair count as a measure of how much traffic was observed. Pairs scale with the slate; sessions do not.
The three thresholds
| Threshold | Value | Effect |
|---|---|---|
| Session floor | 20 judged sessions, per candidate | Below it the verdict is insufficient_data and the judge refuses to speak. |
| Tie band | 0.10 on the composite | A delta inside the band is a tie. The band is the measured composition noise floor of this composite, not a comfort margin. |
| Parity floor | -0.10 | A candidate at or below it is degraded and never adoptable, whatever its cost delta. |
The session floor is per candidate and is not divided by the slate. Adding candidates does not lower the bar for any of them.
The parity floor is not a target to aim at, and a client-side copy of it is never authoritative -- the verdict reports the floor it actually applied.
The verdicts
| Verdict | Condition | Adoptable |
|---|---|---|
insufficient_data | Fewer than 20 judged sessions for this candidate. | No |
degraded | Composite delta at or below the parity floor. | No |
parity | Composite delta inside the tie band. | Yes |
improved | Composite delta above the tie band. | Yes |
A cost advantage is not a verdict. Cost alone must never read as a recommendation, which is why a degraded candidate is unadoptable however much cheaper it is.
And adoptable is not a recommendation, nor evidence that a promotion happened. It means the evidence does not stand in the way.
Provisional verdicts
A verdict is provisional while the trial sits paused at its spend cap. It names a winner, which is exactly why the label matters: a provisional verdict is not final, and the remedy is to raise the cap and resume.
A verdict paused by a person is not provisional. It simply has not finished yet.
The frontier and the recommendation
Each candidate is marked for whether it is on the cost/quality frontier, whether it is dominated, and whether it is the recommendation.
The recommendation is the verdict-of-verdicts: the cheapest adoptable frontier candidate, with prose reasoning a human can read -- a headline of the form "model -- 71% cheaper, quality within the noise band". It is absent when nothing is adoptable.
Frontier membership is not the same as adoptable. A candidate can sit on the frontier and still be degraded.
The population a verdict speaks for
Every verdict reports the population it was measured on: the workload mix of the pairs that were actually judged.
That is there to be read, not skipped. A verdict speaks for the traffic in its population and nobody should generalise it beyond that.
Readable evidence
Behind the numbers are the turns themselves: the prompt, both responses in full, both arms' scores, the judge's reasoning, and which position the candidate occupied. You can filter to losing pairs only, to winning pairs only, to one candidate, or to a sampling window. Asking for both losing-only and winning-only is a contradiction and is refused rather than silently intersected into nothing.
| Property | Detail |
|---|---|
| Retention | The trial's lifetime plus 30 days, enforced by the store. Rows past it are gone from the response, not deprecated. |
| Content | Everything retained passes through a versioned credential and PII deny-list first. Hits are replaced with typed placeholders, and a sample with too many redactions, or any high-severity hit, is dropped entirely rather than stored as placeholder soup. |
| Provenance | Every evidence row carries the version of the scrub rules that produced it. A rule change is a new version and means re-collecting, never patching in place. |
| Availability | Evidence exists only where content uplink is on. On a content-free deployment you get scores without readable evidence -- by design, not a bug. |
An empty evidence list is not a verdict with no basis. The rows may simply have aged out of their retention window while the verdict, which is a decision record, remains.
Trial-evidence retention is additionally subject to enablement for your team -- check with DevZero whether the readable-turns view is available on your account.
Unpriced is not free
Wherever a cost figure is involved, a model whose rate could not be resolved is named rather than priced at zero. A quote or an estimate carrying any unpriced model is a floor rather than a price, and the response says so.
That distinction runs all the way through: unpriceable and free are different answers, and the product does not let one impersonate the other.