- Accuracy
- For a fully reviewed submission, confirmed decisions ÷ all decisions (package + measurements + pins). A model's accuracy is the mean over its reviewed submissions.
- Agreement
- For a re-run, matching rows ÷ compared rows against the baseline's reviewed values. Numbers must match exactly and names must match after normalisation; a partial is not a match. Model averages count only runs whose baseline is fully reviewed.
- Score colours
- 95% and above green, 80 to 94% amber, below 80% red.
- Latency
- Time of the model call only (providerMeta.latencyMs), excluding upload and PDF fetch. The median is shown.
- Cost
- Estimated at run time from list prices × tokens. The actual invoice may differ.
- Correction rate
- Corrected ÷ decided for that field across fully reviewed submissions.
- Legacy runs
- Runs without usage metadata are excluded from latency and cost, and each column shows its n.
- Legacy corrections
- Older corrections that did not record a status are read by their value: a correction to “Not found in datasheet” counts as not found, and a corrected value on a field the AI marked “Not found” now counts as a found value. This changed agreement scores for some legacy reviews.