Practical intelligence for accountable AI programmes.

Search AI strategy, automation, or governance...
Toggle menu

AI Evaluation

How to Build a Scenario-Based Evaluation Set for a Business AI System

Build a reproducible AI evaluation set covering everyday work, hard boundaries, known failures and barred actions, with valid scoring and protected tests.

Business colleagues around a wooden table review a physical workflow of blank cards, coloured case groups, folders and a sealed envelope.

Build the evaluation set around one bounded piece of work, not a folder of polished prompts. A purchase-request assistant may receive a plausible request, a conflicting vendor identifier, an unavailable policy result and an uploaded quote containing instructions aimed at the assistant. The useful test is whether the configured system handles that whole situation correctly, including what it must not do, under reproducible conditions.

Key decisions to make first

  • Define the workflow, system version, intended decision, reporting slices and release gates before collecting results.
  • Represent everyday work, then deliberately add important boundaries, verified failures and prohibited behaviour.
  • Record enough context, evidence, tools, expected outcomes and versions for another evaluator to recreate each case.
  • Use the narrowest valid grader, with prohibited actions kept outside compensable quality scoring.
  • Use visible development cases for iteration and protected cases sparingly for release evidence.

What should a scenario-based evaluation set cover?

A top-down tabletop layout shows a central workflow of blank cards connected to surrounding case groups and coloured round tokens.

A sound set covers ordinary workflow conditions in relation to the observed work mix, plus important boundaries, confirmed failures and explicitly prohibited behaviour. First name the system version, permitted actors and inputs, available knowledge and tools, and the decision the results will inform. This boundary prevents an evaluation of a specific business service from drifting into vague claims about a general-purpose assistant.

Map task families and relevant variations from authorised work records, support cases, user research, incident records and domain-expert walkthroughs. Operational evidence is useful only after checking whether it is lawful to reuse, correctly labelled and reasonably representative. If logs are incomplete or the service has not launched, mark the ordinary-work mix as uncertain or hypothetical instead of giving a rough estimate false precision.

  • Representative work: normal task families, actors, inputs and operating conditions.
  • Important edge cases: valid but difficult ambiguity, missing evidence, conflicting records, permission limits or unreliable tools.
  • Known failures: authorised, minimised reproductions of verified defects, incidents or complaints.
  • Prohibited behaviours: plausible attempts to trigger actions, disclosures or state changes the system is barred from making.

For an illustrative purchase-request assistant, ordinary cases might contain complete, consistent packages and accessible current policy. Deliberate additions would cover missing documents, conflicting amounts or vendor identifiers, facts absent from authorised evidence, requests to bypass human approval, instructions embedded in a quote, stale retrieval and minimised reproductions of prior failures. These rules are examples only; each organisation must define its own permissions and boundaries.

Four complementary coverage families for a bounded business workflow
Coverage familyQuestion answeredPotential source evidenceInclusion logic
Representative workDoes the system handle the work it commonly encounters?Authorised workflow records, research and expert walkthroughsRelate coverage to the observed mix while preserving meaningful slices and recording uncertainty.
Important edge casesDoes it behave correctly near valid boundaries?Workflow analysis, support cases and domain reviewSelect consequential variations deliberately and test both sides of relevant behaviour boundaries.
Known failuresHas a verified defect returned?Incidents, complaints and confirmed surprising outputsRetain an authorised, minimised reproduction as development or regression evidence.
Prohibited behavioursDoes it avoid a specifically barred outcome?Product rules, permission models, policies and risk reviewInclude plausible direct and indirect attempts regardless of their expected traffic frequency.

Define reportable slices, prohibited-behaviour gates and the intended release decision before inspecting candidate outputs. Aggregate results can conceal a weak task family or a rare failure with serious consequences, so report by relevant task, coverage family, slice, consequence and gate. This four-family coverage map is an adaptable editorial synthesis, not a taxonomy or sampling formula prescribed by any one source.

How should each evaluation scenario be recorded?

An open manila folder holding blank sheets sits beside a white binder, wooden blocks, and green, grey and red status tokens.

Record each scenario as a versioned case that another evaluator can recreate and grade. Give it a stable identifier, owner, status, coverage family, task family, slice tags, consequence, provenance category, authorisation record and intended split. Then describe the actor's goal, starting workflow state, request, files, relevant history, available knowledge, permissions, tools, tool results, environmental conditions and system instructions.

  • Required facts, outputs, decisions or state changes that demonstrate success.
  • Acceptable alternative wording, outputs or paths when the task permits more than one valid response.
  • Conditions requiring clarification, abstention, refusal or escalation because evidence is insufficient.
  • Prohibited outputs, disclosures, tool calls, actions and state changes, recorded separately from quality expectations.
  • Grader type, rubric version, gate logic, reference material and any predeclared trial aggregation rule.
  • Model, prompt, retrieval corpus, policy, permissions, tools and evaluation-harness versions.

The expected result should describe observable facts and outcomes rather than one ideal sentence. A working reference solution can show that a bounded case is solvable and help expose a faulty grader, but it should not invalidate a different phrasing or legitimate path. For nondeterministic or multi-step systems, predeclare any repeated trials and the aggregation rule before seeing results, then preserve execution evidence needed to diagnose failures.

A useful evaluation case recreates bounded work and makes success, acceptable variation and prohibited behaviour inspectable.

Record reviewer qualifications, calibration version, adjudication route, known limitations and unresolved disagreements as operating evidence rather than afterthoughts. The full case record is an editorial synthesis assembled from guidance on contextual measurement, defined tasks, reference solutions and durable review records. Tailor it to the system: a straightforward classification workflow needs less environmental detail than an agent that retrieves documents, calls tools and changes records.

How should each scenario be scored without rewarding the wrong behaviour?

Reviewers assess cards with coloured tokens at divided stations while an operator places a warning card against a red mechanical stop gate.

Score each scenario with the narrowest method that can validly distinguish success from failure. Use deterministic checks for objectively verifiable facts, schemas, calculations, tool arguments, record states and barred actions. Use reference facts or a working solution when the required evidence or end state is bounded but several expressions or paths are acceptable. Always confirm that the case itself is solvable and the check is complete.

  • Deterministic check: best for an objective answer, action, argument or resulting state.
  • Reference-based check: best when required facts are bounded but valid outputs can differ.
  • Anchored rubric: best for open-ended qualities such as completeness, groundedness or usefulness.
  • Qualified human judgement: necessary when domain interpretation or consequences cannot be resolved validly by automation.

Rubric dimensions should be separate and observable, with anchors that explain what each score means. Calibrate any model-based grader against qualified human judgements in the relevant context; agreement on familiar development examples does not establish validity elsewhere. OpenAI reports that its GDPval automated grader was not reliable enough to replace experienced occupational graders, which is a useful warning against treating automation as objective by default.

Allow partial credit where task components form a meaningful progression, but keep prohibited disclosures or unauthorised actions outside compensable quality scoring. A polished note cannot make up for exposing confidential information or taking an action the assistant was barred from taking. Responsible product and risk owners should define those gates, consequences, rubric versions, slice definitions and decision rules before comparing candidate outputs.

How can reviewers apply the evaluation criteria consistently?

Reviewers seated apart at a long table independently assess blank folders against matching grids of coloured cards, with an adjudication folder between them.

Reviewers become more consistent when they receive the same bounded context, observable anchors and calibration cases before live scoring. State the workflow purpose, actor, available evidence, allowed behaviour, product boundary and rubric version before showing an output. Provide positive, negative and boundary examples for each dimension, alongside an insufficient-evidence or cannot-score route for cases whose setup or evidence is defective.

  1. Run calibration cases before live review and repeat them after material changes to the rubric, policy, task mix or reviewer pool.
  2. Blind system identity and output order where practical for comparative judgement.
  3. Capture independent scores and rationales before reviewers discuss the result.
  4. Investigate disagreement as possible evidence of an unclear threshold, missing context, defective case, legitimate alternative or unresolved product decision.
  5. Name an adjudication owner and retain labels, rationales, rubric versions and adjudication outcomes.

Do not force consensus where two outcomes are legitimately acceptable or the organisation has not made the underlying product decision. A documented benchmark such as GDPval demonstrates expert task review, blinded comparison and detailed rubrics, but it does not create a universal reviewer protocol. The practical objective is a traceable judgement process that can reveal weaknesses in the system, case, rubric, grader or environment.

How should the evaluation set remain useful through development and release?

An open muted-green file box filled with blank folders sits beside a sealed dark-blue archive box secured with corner guards and elastic cord.

Keep the set useful by separating visible development cases from protected release evidence and recording exposure. Development cases can be run repeatedly while teams adjust prompts, retrieval, tools, policy and workflow logic. Their results remain valuable, but they are development evidence because the team has seen and optimised against them. Assign each case's split when it enters the registry, before routine execution or output inspection.

Use a protected release test sparingly for final comparisons or release decisions, without letting its exact cases, answers, rubrics or results guide iteration. Check across splits for exact and near duplicates, shared source records, paraphrases and siblings generated from the same scenario template. Track access to case content and results; secrecy about an answer alone does not protect independence if the failure pattern has directed a change.

  • Move any protected case that materially shapes remediation into development or regression evidence.
  • Replace it with a versioned, independently created, non-duplicate case that has not guided the change.
  • Keep confirmed failures in development or regression evidence after checking provenance, expected behaviour and data minimisation.
  • Record ownership, authorisation, split, exposure, versions, review history, change reasons and retirement reasons in a set registry.

Reassess coverage and grading after material changes to users, workflow, policy, knowledge, model, prompts, tools, permissions or operating environment. Evaluation sets can grow from verified operational, historical, domain, synthetic and human-curated cases, but every addition still needs authorisation and quality checks. There is no universal case count, split percentage or refresh interval: composition depends on the decision, workflow variability, consequential slices, grading reliability and available evidence.

Treat the set as decision evidence, not a certificate of readiness. Pair offline results with monitoring, user research, incident review and other evidence suited to the system's consequences. Domain experts should validate workflow realism and expected outcomes, while privacy, security, legal, compliance and other qualified functions should review governed material. Regulated or high-consequence judgements must remain with appropriately qualified and authorised people.

Frequently asked questions

How do you build an AI evaluation dataset for a business workflow?

Bound the workflow and release decision, map task families and variations, then add representative work, edge cases, known failures and prohibited behaviour. Turn these into reproducible records, predefine slices and gates, select valid graders, separate development from protected cases and maintain a versioned registry.

What is scenario-based AI evaluation?

It tests a bounded AI-enabled workflow through reproducible cases rather than isolated prompts. Each case specifies the actor, starting state, inputs, available context and tools, acceptable outcomes, prohibited outcomes and grading method.

How many cases should an AI evaluation set contain?

There is no universal number. The appropriate size and composition depend on the decision being supported, workflow variability, consequential slices, grader reliability and the amount of authorised evidence available.

Should an AI evaluation use reference answers or rubrics?

Match the grader to the task. Use deterministic checks for objective outcomes, reference facts or solutions for bounded variation, anchored rubrics for open-ended quality, and qualified human judgement where domain interpretation is necessary.

What is the difference between a development eval set and a protected test set?

The visible development set supports repeated iteration, so its results show performance on familiar evidence. A protected test is assigned before routine use and consulted sparingly for release decisions; it loses independence when its content or results materially guide a change.

ModelFold logo

ModelFold Editorial Desk

We report on how AI actually lands inside a business. Our work starts from named sources, separates what we found from what we think, and uses AI assistance for research and drafting under documented editorial controls. We are not a substitute for individual expert review.