Scoring and quality gates¶
How per-grader scores become a question score, how question scores become a target verdict, how the quality gate is applied, and what ends up stored. For operators reading a result and for anyone wiring a gate into CI.
The three levels¶
flowchart TD
G[Grader scores, each on its own declared scale] --> C[Mapped onto the composite 0 to 5 scale]
C --> Q[Question avg_score, the mean of the mapped scores]
Q --> P{avg_score >= question pass threshold?}
P -->|yes| QP[Question passed]
P -->|no| QF[Question failed]
Q --> T[Target avg_score, the mean over measured questions]
QP --> R[Target pass_rate, the share that passed]
QF --> R
T --> GATE{Gate metric}
R --> GATE
GATE --> V[passed_gate]
Question score¶
A question's avg_score is the plain mean of its grader scores, rounded to two
decimal places. Three rules decide which scores enter that mean.
Scores are mapped onto a shared scale first. Graders declare their own range. Averaging a binary judge's 1 with a gradient judge's 4 as raw numbers would be meaningless, so every score is converted onto the composite 0 to 5 scale using that grader's declared minimum and maximum. A binary yes becomes 5, a binary no becomes 0. A grader that already declares 0 to 5 goes through an explicit identity path and its number is not touched at all, so results from before per-grader scales existed reproduce exactly. See graders for which grader uses which scale.
Null scores are excluded, not counted as zero. A grader that does not apply
to the question (no expected_facts, no expected_tools, no executable rubric
checks) contributes nothing to the mean.
Errored graders are excluded and flag the question. A grader that could not
run is a different matter from one that did not apply. It is excluded from the
mean and it sets an error key on the question reading
grader error: <grader ids>. That flag is what keeps the question out of the
gate later, so a PASS built only on the deterministic graders while every judge
was unreachable cannot happen.
A question with no valid scores is marked unscored. It remains distinct from both a genuine zero and a grader error, is excluded from the target average and pass-rate denominator, and is reported in the run's unscored count.
This is reachable with an ordinary configuration. Attach only
fact_match_grader and tool_usage_grader, then include a question that
carries neither expected_facts nor expected_tools, and it is marked
unscored because neither grader applies. Check your grader list against the
shape of your dataset.
The question passes when it is scored, has no errored grader, and its avg_score meets
the question pass threshold. The threshold resolves in this order: the target's
own question_pass_threshold, then the site-wide question_pass_threshold
setting, then 3.5.
There are no per-grader weights¶
Every grader that produced a usable score counts equally. Attaching four judges and one deterministic check gives the judges four fifths of the question score, and the only way to change the balance is to change the grader list.
Weights do exist one level down, inside a rubric: a rubric's scoring block
can combine its checks with weighted_avg and a per-check weight map, and the
combined verdict then arrives as the single rubric_checks grader score. See
rubrics.
Target score¶
Questions carrying an error key or an unscored marker are dropped before aggregation. Legitimate
zeros are kept. A question that genuinely scored 0.0 is a measurement and
belongs in the average; a question whose graders could not run is an absence of
measurement and does not.
From the surviving questions:
| Figure | How it is computed |
|---|---|
avg_score |
Mean of the surviving question avg_score values, rounded to 2 dp |
pass_rate |
Share of surviving questions that passed, rounded to 2 dp |
pass_rate_ci_low / pass_rate_ci_high |
Wilson score interval at 95% (z = 1.959964) over the same counts |
valid_count |
How many questions survived |
The Wilson interval is stored and displayed alongside every pass rate, so a rate from twelve questions visibly carries more uncertainty than one from four hundred.
The coverage guard¶
If fewer than half of a target's questions could be scored, the surviving
sample is too small to trust and the gate fails regardless of what the
survivors scored. A total outage, where nothing survived, is the degenerate
case of the same rule and produces a row of zeros with passed_gate false.
This is the one place where a verdict is forced rather than computed. A transient judge outage must not read as a PASS on a lucky subset.
The quality gate¶
A gate is a metric plus a threshold, and the two live on different scales. The
comparison is always measured figure >= threshold.
| Metric | Scale | Default threshold | Displayed as |
|---|---|---|---|
avg_score |
0 to 5 | 4.0 | 4.0 |
pass_rate |
0 to 1 | 0.8 | 80% |
The same number means opposite things on the two scales. 0.8 is a sensible
pass_rate gate and a near-guaranteed pass as an avg_score gate; 4.0 is a
sensible avg_score gate and unreachable as a pass_rate gate.
Validation is metric-aware. The widget is not. The threshold field uses a
single step of 0.05 whatever the metric, because it does not re-render when
you change the metric and 0.05 divides both scales. What does look at the
metric is validation on save: it blocks a gate that could never pass, a
pass_rate above 1.0 or an avg_score above 5, and it warns without blocking when
you set an avg_score threshold of 1.0 or less, because that passes almost any
run and usually means you wanted pass_rate. There is no equivalent warning for
an implausibly small pass_rate.
Any stored metric value that is not exactly pass_rate resolves to
avg_score, including an empty or unrecognized one. That keeps an old or
hand-edited config readable rather than fatal.
Since every grader is mapped onto the composite scale before averaging, the avg_score ceiling really is 5. A grader declaring a wider range no longer lifts it, so an avg_score threshold above 5 is unreachable by any grader selection.
Hard and soft gates¶
A target's gate field is hard or soft, defaulting to hard.
drush ai-eval:run exits 1 when any target with a hard gate failed its gate,
and prints One or more hard gates failed.. A soft gate is reported the same
way in the output but does not affect the exit code. Nothing else in the module
changes behavior based on hard versus soft. See running for the
CI wiring.
What does not move the verdict¶
Several things qualify a gate verdict without changing it. Judge trust states, the judge and target sharing a model family, and a dataset below the sample-size floor all produce warnings on the run and on the result screens, and none of them flips a PASS to a FAIL or changes the exit code. See judges for trust states.
Runs are gated against the recorded threshold¶
A stepped run started from the UI is scored against the metric and threshold captured in its run-identity snapshot at start time, not against the live target config. Editing a target's gate halfway through a run does not decide the verdict of the run that began under the old one. A run started before snapshots existed falls back to the live config.
What is persisted¶
One row per completed target run in ai_eval_result:
| Column | Contents |
|---|---|
target_id, timestamp |
Which target, and when the row was written |
avg_score, pass_rate |
The target-level figures above |
pass_rate_ci_low, pass_rate_ci_high |
The Wilson bounds |
passed_gate |
The gate verdict, as 0 or 1 |
question_count |
Every question attempted, including errored ones |
error_count |
How many questions carry an error key |
scores_json |
The full per-question detail, described below |
source |
How the run was triggered (manual, cron, ui) |
duration |
Seconds the run spent working. A run from the UI that was interrupted and resumed leaves the pause out. UI runs recorded before active-time tracking stored the wall-clock span instead, which the result page labels as wall clock |
run_id |
The stepped-run row that produced it, when run from the UI |
config_json |
The run-identity snapshot: provider, model, prompt, dataset and rubric hashes, judge config, retries |
scores_json holds one entry per question with:
question_id,input,avg_score,passed, anderrorwhen the question was flagged;grader_scores, keyed by grader ID, each carryingscore,reason,grader_id,error,min_scoreandmax_score. The score is the grader's raw answer, not the mapped one, and the recordedmin_scoreandmax_scoreare what make a stored 1 from a binary judge readable as a full pass without looking up a plugin definition that may have changed since;elapsed, andraw_responsecapped at the target's response character limit or 8000 characters, whichever is smaller;token_usage,model_idandmodel_id_resolvedin chat mode, where the provider reports them;identity, carrying the retry count, the rubric reference and its SHA-256 when one was used, and per-judge prompt versions, decoding-pin outcome and whether the response was truncated for the judge.
What is not persisted¶
The coverage guard result. insufficient_coverage is computed at scoring
time and forces passed_gate to false, but it is not stored as its own column.
The result screen reconstructs it from scores_json: both errored and unscored
questions are missing measurements, so error_count alone cannot identify a
coverage failure. The CLI has the in-memory scoring result and reports the same
forced failure directly.
Tool-call arguments. They reach the graders and judges for the run and are never written to the result row, so they are not an audit trail. See graders.
The full response. Only the capped raw_response survives.
The judge prompts themselves. identity records which judge prompt
version scored the question, not the text that was sent.
Reading a result¶
The summary band starts with N of M examples met your expectations. Here, M includes only scored examples. Unscored and errored examples are reported separately and are never folded into that headline as failures or successes. The question table puts Needs attention first so failed, unscored, and errored cases are available before cases that met expectations.
Comparing two runs¶
Choose Compare with... on a result to inspect two runs of the same test.
- Direct comparison means the recorded examples, graders, question and gate thresholds, and judge configuration match. The evaluated provider, model, agent, the evaluation mode (agent or chat), system prompt, token limit, temperature, and resolved model IDs may differ. Those are the inputs a direct comparison is intended to examine. A run that recorded identity drift is never a direct comparison.
- Side-by-side inspection means one or more recorded evaluation conditions differ; the verdict names how many.
- Inspection only means at least one run has no valid configuration snapshot, typically a result recorded before snapshots existed.
Paired examples are classified as improved, regressed, or unchanged from their composite score. If either side is unscored or errored, that pair is unscored and never counts as an improvement or regression. The comparison is descriptive only. It makes no statistical significance claim.