Find out that quality moved before your users do
Test sets, regression checks and monitoring, so a change in output quality is something you detect rather than something your customers and your regulator report to you. Built in at the start, not bolted on after an incident.
How we measure an AI system
The same approach on every engagement, published so you can check it before you buy anything.
Golden datasets
A maintained set of at least 100 labelled examples per system, built from real cases rather than invented ones, and versioned alongside the code that is graded against it.
Pass or fail, not one to five
Binary grading against an agreed definition of correct. A five-point scale hides disagreement inside the middle three points and makes a regression impossible to detect.
Judge calibration
Where a model does the grading, it is calibrated against human labels before it is relied on, and re-calibrated on a schedule. An uncalibrated judge is an opinion with a number attached.
Regression checks on every release
The suite runs before a change ships. Below the acceptance threshold means it does not ship, and that threshold is agreed before the build rather than negotiated after it.
Drift detection, including judge drift
Input distribution drift, silent model version changes by a supplier, and movement in the evaluator’s own behaviour. The third is the one most teams never watch.
Open standards throughout
Instrumented on open standards from day one, so you are never locked into whoever happens to be monitoring you, and can take the measurement with you.
Everything else depends on this one
An acceptance threshold you cannot measure is a wish. An outcome-linked commercial model you cannot measure is uninsurable. A monthly run contract with nothing to report against is a retainer. Evaluation is what makes the other services real, which is why it is instrumented into every engagement rather than sold as an afterthought.
Need governance and assurance around what you are building?
We engineer it. Certified helps you govern, evidence and prepare it for assurance.
The group’s specialist compliance, cyber-assurance and AI-governance business. Where a technology programme needs formal governance, certification readiness, privacy or security-assurance support, Certified can help scope the requirement, coordinate appropriately credentialed specialists and support the route to independent assessment where required.
Questions worth answering
Why grade AI output pass or fail rather than on a one-to-five scale?
Because a five-point scale hides disagreement inside the middle three points, which makes regressions impossible to detect reliably. Binary grading against an agreed definition of correct forces the definition to be written down, and makes a change in quality visible the moment it happens.
What is judge drift and why does it matter?
Judge drift is movement in the behaviour of the model doing the grading, as opposed to drift in the input data or a silent version change by a supplier. It matters because it silently invalidates your quality measurements: the system looks stable while the instrument measuring it has moved.
When should evaluation be built into an AI system?
At the start. Evaluation designed in from day one is what makes an acceptance threshold enforceable and a support-and-run contract possible. Evaluation bolted on after an incident can only tell you what is happening now, not what changed.
Want to know whether your system still works?
The baseline measures a system that already exists just as readily as one that does not. We tell you what it is doing now, and what it would cost to keep it honest.