Clear, source-led guidance for accountable business AI.

Search AI strategy, automation, or governance...
Toggle menu

AI Evaluation

How to Build a Scenario-Based Evaluation Set for a Business AI System

A practical guide to designing reproducible AI evaluation scenarios, valid grading rules and protected release evidence for a bounded business workflow.

Business colleagues gathered around a wooden table review a physical workflow of blank cards, coloured case groups, folders and a sealed envelope.

A useful scenario-based evaluation set recreates one bounded piece of business work, specifies what the system may see and do, and makes acceptable and prohibited outcomes inspectable. Consider a purchase-request assistant receiving a plausible request with conflicting vendor details, an unavailable policy result and an uploaded quotation containing instructions aimed at the assistant. A folder of polished sample prompts may show fluent answers, yet miss whether the system asks for evidence, respects the approval boundary and ignores instructions embedded in the document.

The operating principles

  • Define the bounded workflow, evaluation decision, reporting slices and prohibited-behaviour gates before collecting results.
  • Represent ordinary work in relation to observed conditions, then deliberately include consequential boundaries, verified failures and prohibited behaviour.
  • Record each scenario so another evaluator can recreate its starting state, evidence, tools, acceptable outcomes and barred actions.
  • Use the narrowest valid grader, and never let partial credit compensate for a prohibited action or disclosure.
  • Use visible development cases for iteration and protect separate release cases from influencing those changes.

What should a scenario-based evaluation set cover?

A top-down tabletop layout shows a central workflow of blank cards connected to surrounding case groups and coloured round tokens.

It should cover representative work, important edge cases, known failures and prohibited behaviour within one clearly bounded workflow. First name the system version, permitted actors and inputs, available knowledge and tools, and the decision the evaluation will support. A realistic test set should reflect expected-use conditions and document its method, but frequency alone is insufficient: rare conditions can deserve deliberate coverage when their consequences or policy significance are greater.

Build the map from authorised workflow records, support cases, user research, incident evidence and domain-expert walkthroughs. Production records are not automatically complete, representative, correctly labelled or lawful to reuse; mark uncertainty instead of manufacturing precision. Define task families, slices and release gates before inspecting candidate outputs. This reduces the temptation to retain only criteria that flatter a preferred model or configuration.

Four adaptable coverage families for a bounded business workflow
Coverage familyQuestion answeredPotential source evidenceInclusion logic
Representative workDoes the system handle ordinary tasks under expected conditions?Authorised workflow records, user research and domain walkthroughsReflect the observed task mix while preserving meaningful variations and documenting gaps.
Important edge casesDoes it behave correctly near valid boundaries?Missing or conflicting inputs, unusual formats, permission limits and unreliable tool resultsChoose consequential boundaries deliberately and test both sides of a behaviour boundary.
Known failuresHas a verified defect stayed fixed?Confirmed incidents, complaints and reviewed surprising outputsRetain an authorised, minimised reproduction as regression evidence with its provenance.
Prohibited behavioursDoes the system avoid explicitly barred actions or disclosures?Policy limits, permission rules, threat analysis and risk-owner decisionsInclude plausible direct and indirect attempts; evaluate the prohibited outcome as a separate gate.

For the illustrative purchase-request assistant, ordinary coverage could include a complete request and accessible policy. Boundaries could include a missing quotation, inconsistent amounts, conflicting vendor identifiers, insufficient evidence, stale retrieval or an unavailable tool. Targeted cases should also test requests to bypass human approval and instructions hidden inside uploaded material. These behaviours and gates are examples only; each organisation must define them for its own workflow, permissions and governed information.

How should each evaluation scenario be recorded?

An open manila folder holding blank sheets sits beside a white binder, wooden blocks, and green, grey and red status tokens.

Record each scenario as a reproducible case that joins identity, setup, expectations and evaluation operation. Another evaluator should be able to reconstruct what the system received, which knowledge and tools were available, what success allowed, what was forbidden and how the result would be graded. Realistic professional evaluations can use work products, context and reference files, provided the organisation has authorised their use and minimised sensitive information.

  • Identity: stable case ID, owner, version history, coverage and task families, slice tags, consequence, provenance, authorisation record and intended split.
  • Setup: actor and goal, initial workflow state, request and files, relevant history, system instructions, authorised knowledge, permissions, tools, tool results and operating conditions.
  • Expectations: required facts or state changes, acceptable alternative outputs or paths, and conditions requiring clarification, abstention, refusal or escalation.
  • Evaluation: prohibited outputs, disclosures, tool calls and state changes; grader type, rubric, gates, reference material, reviewer qualification and adjudication route.
  • Versions: model, prompt, retrieval corpus, policy, tools, permissions and harness, plus any predeclared trial count and aggregation rule when nondeterminism matters.

Do not turn an open-ended task into a search for one ideal sentence. A reference solution can establish that a bounded case is solvable and help test the grader, while still allowing other valid wording or paths. For a multi-step system, graders may inspect the final outcome, the response or the execution trace. Record limitations and unresolved disagreements so later reviewers do not mistake incomplete evidence for a settled answer.

A strong evaluation case recreates bounded work and makes success, acceptable variation and prohibited behaviour inspectable.

How should scenarios be scored without rewarding the wrong behaviour?

Reviewers score cards with coloured tokens at separate stations while an operator places a warning card against a red mechanical stop gate.

Score each scenario with the narrowest method that can validly separate success from failure, and keep prohibited actions outside compensable quality scoring. Deterministic, model-based and human graders serve different purposes. An exact check may suit a schema, calculation, tool argument or record state; it does not become valid merely because it is automatic. Confirm that the task is solvable and that the check recognises every materially acceptable result.

  • Use deterministic checks for objectively verifiable answers, required fields, state changes, authorised tool calls and barred actions.
  • Use reference facts or a working solution when evidence or the end state is bounded but several expressions or paths are acceptable.
  • Use separate, observable rubric dimensions for open-ended qualities such as usefulness, completeness, groundedness or explanation quality.
  • Use qualified human judgment when the task requires domain interpretation that automated checks cannot validly resolve, retaining regulated decisions with authorised roles.

Partial credit is useful when meaningful components form a progression, but it must show which component failed. A prohibited disclosure or unauthorised action should fail its predeclared gate regardless of polished prose elsewhere; authorised product and risk owners must define that consequence. Calibrate any model-based grader against relevant human judgments rather than calling it objective. OpenAI, for example, reports that its GDPval automated grader was not reliable enough to replace experienced occupational graders.

How can reviewers apply the criteria consistently?

Reviewers seated apart at a long table independently assess blank folders against matching grids of coloured cards, with an adjudication folder between them.

Give reviewers the same bounded context, observable anchors and disagreement process before live scoring begins. They need the workflow purpose, actor, available evidence, allowed behaviour, product boundary and rubric version before seeing an output. Context, intended-solution information and contrasting examples can improve reviewer reliability, while disagreements often reveal an ambiguous case, unclear threshold, missing context or an unresolved product choice rather than simple reviewer error.

  • Define observable anchors for each score point, supported by positive, negative and boundary examples.
  • Provide an insufficient-evidence or cannot-score route for defective cases rather than forcing a judgment.
  • Run calibration cases before live review and repeat calibration when the rubric, policy, task mix or reviewer pool changes.
  • Blind system identity and output order where practical, then capture independent scores and rationales before discussion.
  • Name an adjudication owner and preserve labels, rationales, rubric versions and outcomes for later review.

Treat disagreement as diagnostic evidence. Inspection of outputs, traces and grades can distinguish a genuine system failure from a broken task, grader or environment. Do not manufacture consensus when multiple answers are legitimately acceptable or when the organisation has not decided a policy boundary. Blinded expert comparison and detailed rubrics have been used in documented professional evaluation, but the reviewer procedure still needs tailoring to the workflow and consequence.

How should the evaluation set remain useful through development and release?

An open muted-green file box filled with blank folders sits beside a sealed dark-blue archive box secured with corner guards and an elastic cord.

Keep a visible development set for repeated iteration and a separate protected test for sparing release evidence. Once teams use a case, answer, rubric or result to guide a prompt, retrieval, tool, policy or workflow change, that case is development evidence rather than an independent release test. Repeated consultation risks implicit overfitting, and validation or test evidence can wear out as teams become familiar with it.

Assign split membership when a case enters the registry, before routine execution or output inspection. Check exact and near duplicates, shared source records, paraphrases and scenario-template siblings across sets. Track access to case content, reference answers, rubrics and results. If a protected case materially shapes a change, move it into development or regression evidence and create a versioned, non-duplicate replacement that has not guided the remediation.

  • Record the owner, provenance, authorisation, split, exposure, versions, review history, change reason and retirement reason.
  • Add verified failures only after confirming expected behaviour and minimising confidential or personal information.
  • Reassess coverage after material changes to users, workflow, policy, knowledge, model, prompts, tools, permissions or operating conditions.
  • Audit stale references, label errors, ambiguous tasks, unreachable outcomes and graders that reject valid alternatives.
  • Report results by task family, coverage family, slice, consequence and gate, not solely as an overall average.

There is no universal case count, split percentage, pass threshold or refresh interval. Composition depends on the release decision, workflow variability, consequential slices, grading reliability and authorised evidence available. Treat an offline result as one part of a decision record, not a certificate of business value, safety, fairness, compliance or production readiness. Pair it with suitable monitoring, user research and incident review, and keep legal, medical, financial, employment, security, privacy and other regulated judgments with appropriately qualified and authorised people.

Frequently asked questions

How do you build an AI evaluation dataset for a business workflow?

Bound one workflow and define the release decision, task families, reporting slices and gates. Map representative work, important edges, known failures and prohibited behaviour; then create reproducible records, assign valid graders and separate development cases from protected release evidence. Maintain everything in a versioned registry.

What is scenario-based AI evaluation?

It tests a bounded AI-enabled workflow through reproducible cases rather than isolated prompts. Each case specifies the actor, starting state, inputs, authorised context and tools, acceptable outcomes, prohibited outcomes and grading method.

How many cases should an AI evaluation set contain?

There is no universal number supported for every system. The required size and composition depend on the decision being made, workflow variability, consequential slices, grading reliability and the authorised evidence available.

Should an AI evaluation use reference answers or rubrics?

Match the grader to the task. Use deterministic checks for objective outcomes, reference facts or solutions for bounded variation, anchored rubrics for open-ended quality and qualified human judgment where domain interpretation is necessary.

What is the difference between a development eval set and a protected test set?

The visible development set supports repeated iteration, so its results are development evidence. A protected test is assigned before routine execution and used sparingly for release evidence. It loses independence when its cases, answers, rubrics or results materially guide changes.

ModelFold logo

ModelFold Editorial Desk

We report on how AI actually lands inside a business. Our work starts from named sources, separates what we found from what we think, and uses AI assistance for research and drafting under documented editorial controls. We are not a substitute for individual expert review.