Practical intelligence for accountable AI programmes.

Search AI strategy, automation or governance…
Toggle menu

AI Evaluation

Building a Scenario-Based Evaluation Set for a Business AI System

A practical guide to building reproducible AI evaluation scenarios for ordinary work, difficult boundaries, known failures and prohibited behaviour.

Business colleagues at a wooden table examine a physical workflow of blank cards, coloured case groups, folders and a sealed envelope.

Build the evaluation set around one bounded workflow, not a folder of impressive prompts. Consider an internal purchase-request assistant facing a plausible request, a conflicting vendor identifier, an unavailable policy result and an uploaded quote containing instructions aimed at the assistant. A few polished examples will not show whether it requests evidence, respects the approval boundary or follows the hostile instruction. Useful release evidence must recreate ordinary work and consequential exceptions, define valid grading in advance and make prohibited outcomes impossible to hide inside an average.

The method at a glance

  • Define the workflow, system version, evaluation decision, reporting slices and gates before collecting results.
  • Represent ordinary work in relation to observed conditions, then deliberately add important boundaries, verified failures and prohibited behaviour.
  • Record each case so another evaluator can reproduce its setup, acceptable variation, prohibited outcomes and grading.
  • Use the narrowest valid grader, and never let partial credit compensate for an explicitly prohibited action.
  • Use visible development cases for iteration and protected cases sparingly for release evidence.

What should a scenario-based evaluation set cover?

A top-down tabletop arrangement shows a central workflow of blank cards linked to surrounding case groups and coloured round tokens.

A scenario-based set should cover the normal workflow mix and deliberately selected boundaries, known failures and prohibited behaviour. First fix the system version, permitted actors and inputs, available knowledge and tools, and the decision the evaluation will support. Then map task families and relevant variations from authorised work records, support cases, user research, incident records and domain-expert walkthroughs. Where evidence is incomplete or pre-launch, label the uncertainty instead of presenting an invented traffic mix as fact.

For ordinary work, sample task families in relation to trustworthy observations of expected use. Frequency alone is insufficient, however, because a rare disclosure, unauthorised action or consequential error can matter far more than its share of traffic. Define the slices, gate conditions and intended release use before inspecting outputs. That order prevents a preferred result from determining which user role, document type, permission boundary or consequence is reported afterwards.

An adaptable coverage map for a bounded business workflow
Coverage familyQuestion answeredPotential source evidenceInclusion logic
Representative workCan the system handle the work it ordinarily encounters?Authorised workflow records, user research and domain walkthroughsReflect observed task families and meaningful subgroups; record uncertainty where evidence is weak.
Important edge casesWhat happens at valid but difficult boundaries?Workflow analysis, support cases and expert reviewSelect ambiguity, missing context, conflicting evidence, permission limits and unreliable tools deliberately.
Known failuresHas a confirmed defect returned?Verified incidents, complaints and surprising outputsRetain an authorised, minimised reproduction as regression evidence with provenance.
Prohibited behavioursDoes the system avoid actions or disclosures that are explicitly barred?Product rules, permission models and risk decisionsTest plausible direct and indirect attempts; apply a predefined gate rather than a compensating average.

In the purchase-request example, representative cases include complete requests with consistent records and an accessible current policy. Boundary cases can introduce a missing quote, conflicting amounts, absent evidence or unavailable retrieval. Known failures should be verified and reproduced with minimised, authorised material. Prohibited cases can ask the assistant to bypass approval, disclose another party's information or obey instructions embedded in a quote. These examples illustrate the map; each organisation must define its own permissions and barred outcomes.

How should each evaluation scenario be recorded?

An open manila folder with blank sheets lies beside a white binder, wooden blocks, and green, grey and red status tokens.

Each scenario should be a reproducible case record that defines the starting conditions, available evidence, valid outcomes and grading. Give it a stable identity, owner, coverage family, task family, slice tags, consequence, provenance category, authorisation record and intended split. Record the actor and goal, initial workflow state, request and files, relevant history, system instructions, approved knowledge, permissions, tools, returned tool results and environmental conditions. Another evaluator should be able to reconstruct what the system could actually see and do.

  • Required facts, outputs or workflow state changes that demonstrate success.
  • Acceptable alternative wording, evidence use or paths when the task permits genuine variation.
  • Conditions requiring clarification, abstention, refusal or escalation because evidence is insufficient.
  • Prohibited outputs, disclosures, tool calls, approvals and state changes, recorded separately from quality.
  • Grader type, rubric version, gate logic, reference material and any predeclared trial aggregation rule.
  • Model, prompt, retrieval corpus, policy, tool, permission and evaluation-harness versions.

A known working solution is useful when it demonstrates that a bounded case is solvable and helps expose a faulty grader. It should not become the only accepted wording or route where alternatives are valid. For nondeterministic or multi-step systems, declare any repeated trials and aggregation rule before viewing outputs. Also record reviewer qualifications, calibration material, the adjudication route, unresolved limitations and version history. This combined record is an editorial synthesis, so teams should tailor it without dropping information needed for reproduction.

A useful evaluation case recreates bounded work and makes success, acceptable variation and prohibited behaviour inspectable.

How should each scenario be scored without rewarding the wrong behaviour?

Reviewers assess cards with coloured tokens at divided stations while an operator places a warning card against a red mechanical stop gate.

Score each scenario with the narrowest method that can validly separate success from failure. Use deterministic checks for objectively verifiable calculations, schemas, tool arguments, record states and barred actions, while checking that the test itself is complete and solvable. Use reference facts or a working solution when the required evidence or final state is bounded but several phrasings or paths are acceptable. Reserve anchored rubrics for genuinely open-ended qualities, and qualified human judgement for interpretation that automation cannot validly resolve.

  • Check objective outcomes or states deterministically where the rule can be expressed completely.
  • Grade against required facts or a reference solution without demanding incidental wording.
  • Separate open-ended dimensions such as completeness, groundedness and usefulness, with observable anchors for each.
  • Calibrate model-based graders against qualified human judgements drawn from the relevant task context.
  • Allow partial credit for meaningful components, but report which component failed.
  • Apply any prohibited disclosure or unauthorised action as a non-compensable gate fixed in advance.

Automated grading is not made valid merely by agreement on familiar development examples. OpenAI reports that its GDPval automated grader was not reliable enough to replace experienced occupational graders, although that result does not settle every narrower use. Inspect outputs, grades and execution traces to distinguish a system failure from a defective case, grader or environment. Fix rubric versions, slice definitions, gate logic, trial aggregation and the decision rule before comparing candidates; a later change creates a new evaluation version.

How can reviewers apply the evaluation criteria consistently?

Reviewers seated apart at a long table independently compare blank folders with matching grids of coloured cards, with an adjudication folder between them.

Reviewers can apply criteria more consistently when they receive the same task context, observable anchors and calibration before live scoring. State the workflow purpose, actor, available evidence, allowed behaviour, product boundary and rubric version before presenting an output. Describe what each score point looks like and provide positive, negative and boundary examples. Include an insufficient-evidence or cannot-score route so a defective case does not masquerade as a poor system response.

  • Run calibration cases before live review and repeat calibration after material changes to the rubric, policy, task mix or reviewer pool.
  • Blind system identity and output order where practical for comparative judgements.
  • Collect independent scores and concise rationales before reviewers discuss a case.
  • Investigate disagreement as possible evidence of missing context, an unclear threshold, a defective case or legitimate plurality.
  • Name an adjudication owner and retain labels, rationales, rubric versions and outcomes.

Disagreement is diagnostic information, not automatically a reviewer failure. It may reveal that two outputs are legitimately acceptable, that the evidence cannot support a score or that product owners have not decided the boundary. Do not force consensus in those situations. Separate correction of a faulty case from a new policy decision, and preserve the original judgements. Detailed rubrics, blinded comparison and expert review have been used in professional benchmarking, but the complete procedure here remains adaptable rather than universal.

How should the evaluation set remain useful through development and release?

An open muted-green file box packed with blank folders sits beside a sealed dark-blue archive box secured with corner guards and elastic cord.

Keep the set useful by separating visible development cases from protected release evidence and tracking exposure throughout both. The development set can be run repeatedly while teams change prompts, retrieval, tools, policies and workflow logic. Its results are development evidence because the team has optimised against those cases. Assign split membership when a case enters the registry, before routine execution or output inspection, and reserve the protected test for sparing final comparisons or release decisions.

Protection involves more than hiding a file name. Check for exact and near duplicates, shared source records, paraphrases and siblings generated from the same scenario template. Track access to case content, reference answers, rubrics and results. A confirmed failure that has already shaped diagnosis belongs in development or regression evidence. A protected set may test the same failure class only with independently created, non-duplicate cases that have not guided the change.

  • Move a protected case into development or regression evidence once it materially shapes a system change.
  • Add a versioned replacement rather than continuing to describe an exposed case as independent.
  • Record ownership, provenance, authorisation, split, exposure, review history, change reasons and retirement reasons.
  • Version case content, labels, rubrics, policies and the evaluation harness together.
  • Reassess coverage after material changes to users, workflow, policy, knowledge, models, prompts, tools, permissions or operating conditions.

Evaluation sets can grow through verified production, historical, domain, synthetic and human-curated cases, but each addition still needs authorisation, minimisation and label review. There is no universal split percentage, case count or refresh interval. Report results by task family, coverage family, slice, consequence and gate so an overall score cannot conceal a concentrated weakness. Where versions change, retain enough documented overlap or bridge testing to explain movement without silently rewriting earlier evidence.

Treat an offline evaluation as decision evidence, not a certificate of safety, fairness, compliance, business value or production readiness. Product and risk owners should define consequence-sensitive gates before seeing results, while domain experts validate workflow realism and expected outcomes. Involve qualified data, privacy, security, legal, compliance or other governed functions when their material or judgements are implicated. Keep regulated decisions with authorised people, and pair scenario results with monitoring, user research, incident review and evidence suited to the system's actual consequences.

Frequently asked questions

How do you build an AI evaluation dataset for a business workflow?

Bound the workflow and release decision, then map task families, variations, consequences and prohibited outcomes. Create reproducible records, define slices and gates before testing, choose valid graders, separate development from protected release evidence and maintain a versioned registry.

What is scenario-based AI evaluation?

It tests a bounded AI-enabled workflow through reproducible cases rather than isolated prompts. Each case specifies the actor, starting state, inputs, available context and tools, acceptable outcomes, prohibited outcomes and grading method.

How many cases should an AI evaluation set contain?

There is no universal case count. Size and composition depend on the decision being supported, workflow variability, consequential slices, grading reliability and the authorised evidence available; uncovered or uncertain areas should be stated plainly.

Should an AI evaluation use reference answers or rubrics?

Use deterministic checks for objective outcomes, reference facts or solutions for bounded variation, and observable rubric dimensions for open-ended quality. Use qualified expert judgement where domain interpretation is necessary, without treating one reference wording as the only valid answer.

What is the difference between a development eval set and a protected test set?

A visible development set supports repeated iteration, so its scores reflect optimisation against known cases. A protected test is assigned before routine use and used sparingly for release evidence; it loses independence when its content, answers, rubrics or results materially guide changes.

ModelFold logo

ModelFold Editorial Desk

We report on how AI actually lands inside a business. Our work starts from named sources, separates what we found from what we think, and uses AI assistance for research and drafting under documented editorial controls. We are not a substitute for individual expert review.