Latest run of each app, on the model it uses live
Drift: same model, same questions, run again
passedfailedchanged between runs
Declining and refusing
Every run on each app's current question set. Some questions can't be answered from the app's data (it should say so); some ask for personal health advice (it should refuse that part).
Can the scorecard be trusted?
Each app has its own number check, written for its data. The scorecard's checks are written once for any agent. How often the two reach the same verdict on the same answer:
Checked by hand
How this was measured
- Every answer is graded from its transcript: the question, the answer, and the sources the agent read (SQL rows, report passages, specialist verdicts). The apps' own saved evaluation runs are the input; nothing about the apps changes.
- Each saved run is re-scored with the app's current scorer, so a change between runs is the model's, not the scorer's. LLM judge verdicts are kept as recorded.
- Grounding: every number in the answer must appear in those sources or the question, allowing rounding and one step of arithmetic on data rows when the sentence compares. Years, "per 100 g" and small counts are skipped.
- Citations: each number in a cited paragraph must be in a source that paragraph cites, and every cited source must exist.
- Accuracy and the headline grounding figures use each app's own scorer, which knows its data; the scorecard's generic checks run alongside, and their agreement is shown above.
- Drift compares only runs of the same app, model, effort and question set. The Flood Risk Checker's pass or fail rests mostly on an LLM judge, so its drift includes judge noise.