Judging and Verdicts

The composite score, its five dimensions, the session floor, the tie band, and exactly what each verdict entitles you to do.

Judging and Verdicts

The paired judge

A trial, a replay, an experiment and a watch are all scored by the paired judge: one comparative call that sees the served answer and the candidate's answer together, at a randomized presentation order.

The order matters enough to be reported. Every candidate's verdict carries the observed split of how often it was shown first, which is what makes the randomization verifiable rather than asserted.

Trial comparisons are paired -- both arms answered the same request -- so the composition imbalances that force other measurements into stratification cannot arise between the arms.

Compare uses a different rubric, the absolute judge, which scores one answer on its own merits. Its scores are never read as a delta and are never written to the pair-score store at all. The two rubrics carry different judge versions, and that difference is not cosmetic.

The composite and its dimensions

Every answer is scored on five dimensions, each on a one-to-five scale. The composite is their mean, and it is the single number every delta, floor, band and verdict is expressed in.

Dimension
qualityOverall answer quality.
faithfulFaithfulness to the input.
conciseConcision.
instructInstruction following.
clarityClarity.

A verdict reports the composite delta and the per-dimension deltas, so you can see which dimension moved rather than only that something did.

This composite is not the efficiency score. They measure different things and share no scale.

The unit is the session

A session's judged pairs count as one observation, however many turns it has. Rows within a session are not independent, so counting them as rows would overstate how much evidence a verdict rests on.

Every verdict therefore reports sessions beside pairs, and a session count is what the floor is measured against.

Do not read a pair count as a measure of how much traffic was observed. Pairs scale with the slate; sessions do not.

The three thresholds

ThresholdValueEffect
Session floor20 judged sessions, per candidateBelow it the verdict is insufficient_data and the judge refuses to speak.
Tie band0.10 on the compositeA delta inside the band is a tie. The band is the measured composition noise floor of this composite, not a comfort margin.
Parity floor-0.10A candidate at or below it is degraded and never adoptable, whatever its cost delta.

The session floor is per candidate and is not divided by the slate. Adding candidates does not lower the bar for any of them.

The parity floor is not a target to aim at, and a client-side copy of it is never authoritative -- the verdict reports the floor it actually applied.

The verdicts

VerdictConditionAdoptable
insufficient_dataFewer than 20 judged sessions for this candidate.No
degradedComposite delta at or below the parity floor.No
parityComposite delta inside the tie band.Yes
improvedComposite delta above the tie band.Yes

A cost advantage is not a verdict. Cost alone must never read as a recommendation, which is why a degraded candidate is unadoptable however much cheaper it is.

And adoptable is not a recommendation, nor evidence that a promotion happened. It means the evidence does not stand in the way.

Provisional verdicts

A verdict is provisional while the trial sits paused at its spend cap. It names a winner, which is exactly why the label matters: a provisional verdict is not final, and the remedy is to raise the cap and resume.

A verdict paused by a person is not provisional. It simply has not finished yet.

The frontier and the recommendation

Each candidate is marked for whether it is on the cost/quality frontier, whether it is dominated, and whether it is the recommendation.

The recommendation is the verdict-of-verdicts: the cheapest adoptable frontier candidate, with prose reasoning a human can read -- a headline of the form "model -- 71% cheaper, quality within the noise band". It is absent when nothing is adoptable.

Frontier membership is not the same as adoptable. A candidate can sit on the frontier and still be degraded.

The population a verdict speaks for

Every verdict reports the population it was measured on: the workload mix of the pairs that were actually judged.

That is there to be read, not skipped. A verdict speaks for the traffic in its population and nobody should generalise it beyond that.

Readable evidence

Behind the numbers are the turns themselves: the prompt, both responses in full, both arms' scores, the judge's reasoning, and which position the candidate occupied. You can filter to losing pairs only, to winning pairs only, to one candidate, or to a sampling window. Asking for both losing-only and winning-only is a contradiction and is refused rather than silently intersected into nothing.

PropertyDetail
RetentionThe trial's lifetime plus 30 days, enforced by the store. Rows past it are gone from the response, not deprecated.
ContentEverything retained passes through a versioned credential and PII deny-list first. Hits are replaced with typed placeholders, and a sample with too many redactions, or any high-severity hit, is dropped entirely rather than stored as placeholder soup.
ProvenanceEvery evidence row carries the version of the scrub rules that produced it. A rule change is a new version and means re-collecting, never patching in place.
AvailabilityEvidence exists only where content uplink is on. On a content-free deployment you get scores without readable evidence -- by design, not a bug.

An empty evidence list is not a verdict with no basis. The rows may simply have aged out of their retention window while the verdict, which is a decision record, remains.

Trial-evidence retention is additionally subject to enablement for your team -- check with DevZero whether the readable-turns view is available on your account.

Unpriced is not free

Wherever a cost figure is involved, a model whose rate could not be resolved is named rather than priced at zero. A quote or an estimate carrying any unpriced model is a floor rather than a price, and the response says so.

That distinction runs all the way through: unpriceable and free are different answers, and the product does not let one impersonate the other.

On this page