Methodology

How we evaluate AI systems, and why we use pass or fail

Before an AI system performs a job for a client, there must be a clear answer to a practical question: has it met the requirements for this particular use? Here is how we would reach that decision, what we measure along the way and why an overall score cannot make the decision on its own.

Article

A client does not commission an AI system simply to receive a promising demonstration or a score out of five. They need to know what the system is allowed to do, how it behaves on the tasks it will actually encounter and what happens when it gets something wrong.

That question cannot be answered with one number. An average may be useful for comparing versions, but it can conceal a failure that matters far more than several successful routine tasks. At some point, the client and delivery team need a decision: does this version meet the agreed requirements for the job and level of autonomy proposed, or does it need more work?

That is what pass or fail means in our approach. The decision is simple to understand. The evaluation behind it is detailed.

Define the job before setting the test

Evaluation begins when we establish what the system is meant to do. We need to understand its users, the information it may access, the actions it may take, the conditions in which it will operate and the circumstances in which it must stop or ask a person.

We then agree what success and unacceptable failure mean for that particular use. Response quality, speed and cost may matter. So may factual support, permissions, reliability and the ability to recognise uncertainty. The criteria should reflect the consequences of getting something wrong, not just what is easy to measure.

This follows the problem first approach set out on our AI and Automation page. The task and the evaluation method should be understood before choosing a model or deciding how much autonomy to give the system.

An illustrative client decision

Imagine a client asking us to develop an AI assistant for its customer service team. It should use authorised order records and approved policies to draft answers about deliveries and returns. An employee will review each draft before it is sent.

Together, we would define the proposed job and its boundaries. A useful answer must address the customer’s question and rely on the correct records. The assistant must ask for clarification when essential details are missing. It must not reveal another customer’s information, invent a returns policy or issue a refund.

The test set would include ordinary enquiries alongside incomplete messages, conflicting records, unusual requests and attempts to make the assistant act outside its permissions. We would look at the answer and the steps used to produce it.

Suppose the assistant produces strong drafts for most enquiries but, in one test, includes details from the wrong customer’s order. Its average quality score might still look impressive. That version fails the agreed criteria for the proposed use because it crossed a critical permission boundary.

The next step is to investigate the failure, correct the system and run the relevant tests again. A later version might pass for producing drafts that an employee reviews. That would not mean it had passed for sending replies or issuing refunds without approval. Those are different jobs and require different controls and evidence.

  1. 01The proposed jobDraft answers using authorised customer records, with a person approving the response
  2. 02The critical testAn enquiry could lead the system to retrieve another customer’s information
  3. 03The release decision
    1. Fail for this version
    2. Fix the permission boundary
    3. Retest

One possible outcome, and its boundary

Pass for reviewed drafts is not approval to send messages or issue refunds automatically

An illustrative example, not a client project. A pass applies to the job, the conditions and the actions it was granted for

Stages one and two describe the proposed job and the test that matters most for it. Stage three is the decision taken when that test fails, and leads back into retesting rather than forward to release. The statement that follows is one possible outcome of a later version, not a further stage.

What we measure beneath the decision

A pass or fail decision should be supported by results the client can inspect. Depending on the system, we would examine:

  • Task results. Did it complete the job defined for each test?
  • Evidence. Are important claims supported by the information it was permitted to retrieve?
  • Permissions and actions. Did it respect access limits and approval points?
  • Failure behaviour. Did it ask for help, decline or stop appropriately when information was missing or the request was outside its scope?
  • Performance. Were response time, reliability and operating cost acceptable for the intended use?

We would report these separately. A strong result in one area should not silently cancel a serious failure in another. The client should be able to see what passed, what failed and why the decision was reached.

How external guidance informs the criteria

The criteria come from the client’s intended use, agreed requirements, applicable law and assessed risks. External guidance helps us structure the work; it does not provide a universal score that makes every AI system ready for release.

Depending on the project, relevant references include the NIST AI Risk Management Framework, the UK National Cyber Security Centre’s guidelines for secure AI system development and the OWASP Top 10 for LLM applications. Data protection, sector requirements and the client’s own policies may add further criteria.

This matters because an assistant drafting low risk internal text, an agent acting on customer accounts and an AI system supporting a consequential decision should not all face an identical release test.

Why a score alone is not the decision

Scores remain useful. We may use them to compare versions, detect regressions or see whether changes improve a particular measure. Some outputs also require informed human judgement rather than a mechanical check. Automated assessment can help at scale, provided its conclusions are checked against suitable human review.

The release decision has a different purpose. It asks whether the system has met the agreed requirements for its defined job. A high average cannot authorise an action that the system was never approved to take. Nor should several successful routine answers excuse a critical failure involving the wrong customer’s data.

A pass means the evidence supports the specified use under the stated conditions. It does not mean the system will never make a mistake. A fail identifies what needs to change before that use can proceed.

Evaluation continues after launch

Prelaunch testing is essential, but real usage can reveal cases that the original tests missed. Once a system is operating, its owners need a way to review errors, user feedback, changes in data and changes to the underlying model or workflow.

The test set should evolve as those cases emerge. A material change to the system or its permitted actions may require the release decision to be revisited. Monitoring and improvement are part of operating the system, consistent with the approach described on our AI and Automation page.

What a client should be able to ask us

A client should be able to ask what the system was tested against, which failures would block release, what a pass permits it to do and who is responsible when it needs attention. We should be able to give clear answers backed by the evaluation results.

That is the value of pass or fail. It turns detailed testing into an understandable decision about a specific use, while keeping the evidence and limitations visible.

Sources

AI & Automation

Where should AI agents actually work?

AI agents can improve a business process or create a more responsive experience in a digital product. The starting point is the same in either case: define the problem, decide what better looks like and establish whether an agent is the right way to get there.

Judge how we think before deciding how we build

Explore our thinking, methodology and technical approach before starting a conversation.