AI Quality Scorecard.

How good are my AI apps' answers, and do they stay good? Every saved evaluation run is graded the same way: is the answer right, does every number come from what the agent read, do its citations hold up, does it decline what it can't answer, and what changed since the last run.

Latest run of each app, on the model it uses live

Drift: same model, same questions, run again

passedfailedchanged between runs

Declining and refusing

Every run on each app's current question set. Some questions can't be answered from the app's data (it should say so); some ask for personal health advice (it should refuse that part).

Can the scorecard be trusted?

Each app has its own number check, written for its data. The scorecard's checks are written once for any agent. How often the two reach the same verdict on the same answer:

Checked by hand

    How this was measured

    • Every answer is graded from its transcript: the question, the answer, and the sources the agent read (SQL rows, report passages, specialist verdicts). The apps' own saved evaluation runs are the input; nothing about the apps changes.
    • Each saved run is re-scored with the app's current scorer, so a change between runs is the model's, not the scorer's. LLM judge verdicts are kept as recorded.
    • Grounding: every number in the answer must appear in those sources or the question, allowing rounding and one step of arithmetic on data rows when the sentence compares. Years, "per 100 g" and small counts are skipped.
    • Citations: each number in a cited paragraph must be in a source that paragraph cites, and every cited source must exist.
    • Accuracy and the headline grounding figures use each app's own scorer, which knows its data; the scorecard's generic checks run alongside, and their agreement is shown above.
    • Drift compares only runs of the same app, model, effort and question set. The Flood Risk Checker's pass or fail rests mostly on an LLM judge, so its drift includes judge noise.