Technical · 13 min · reviewed August 2026
Evaluating a model before a pilot
A test suite built from labelled history, run before anyone sees a demo — and an honest account of what a passing score does not prove.
Why before
A pilot that starts without an evaluation suite has no way to distinguish a good system from a persuasive one, and the demo will be persuasive. Building the suite first also settles an argument that is otherwise deferred indefinitely: what counts as a correct answer, decided by the business rather than by whoever is holding the model.
Building the set from labelled history
Take decisions the institution has already made, with their outcomes, across a window long enough to include a bad quarter. Hold out a slice nobody may look at. Where the labels are contested — and in exception handling they usually are — the process of resolving them is itself the most valuable output of the engagement, because it produces a written decision policy where there was a convention.
The three numbers
- The error rate the business will acceptAgreed in advance, in writing, by the person accountable for the process. Not chosen after seeing the results.
- The cost asymmetryWhat a false positive costs against a false negative. These are rarely equal and the threshold should reflect it.
- The coverageThe share of cases the system is allowed to handle at all. Narrow and reliable beats broad and supervised.
What a passing score does not prove
It does not prove the system will behave the same on next quarter’s distribution, that the corpus will still be current, or that the humans around it will keep reviewing what they are meant to review. Those are operating properties, and they are measured after deployment or not at all. The suite is the entry condition for a pilot, not evidence that the pilot will succeed.
Terms used here
Author
Name pending · practice lead. Reviewed by the editorial owner.
Cite
MLG Blockchain, “Evaluating a model before a pilot,” 2026. TechArticle, machine-readable. https://mlgblockchain.com/insights/ai-transformation/evaluating-a-model-before-a-pilot
Prints cleanly, with URL and date in the running head.