Practical intelligence for accountable AI programmes.

Search AI strategy, automation, or governance...
Toggle menu

AI Use-Case Discovery

How to Observe Real Work Before Proposing an AI Use Case

A practical field-research method for testing workflow complaints against real cases before choosing an AI experiment or a simpler operational change.

Colleagues exchange a buff document folder across an office table spread with case files and an open dark blue binder.

A complaint is a research lead, not an AI use case. Imagine an operations team reporting that client-onboarding briefs take too long and asking for an AI summariser. A workflow map drawn from memory makes drafting look like the problem, yet completed cases may show that difficult briefs wait for missing commercial decisions and are rewritten when contradictory fields are resolved. Automating the visible writing step would leave that constraint untouched. The sounder starting point is a bounded sample of real work that connects reported frustration to recurring cases, operational consequences, counterevidence and a choice between AI and simpler changes.

What to carry into the field

  • Treat a complaint as a lead to investigate, not a ready-made AI use case.
  • Bound the workflow and select only the evidence methods needed to answer the research question.
  • Use interviews, observation, artefacts, diaries and records as complementary partial views.
  • Call a bottleneck credible only when recurrence, location, consequence, mechanism, counterevidence and actionability align.
  • Compare AI with rules, clearer ownership, process redesign, training, better information and no intervention.

What must be defined before observing a workflow?

Colleagues arrange printed sheets and coloured folders into a left-to-right process marked by small black arrows.

Define the work by its trigger, end condition, output, downstream user, participating roles and relevant case types before discussing AI capabilities. Start with recurring work and real outputs, then write the questions the field round must answer. GOV.UK interview guidance likewise starts planning from the research questions and the processes or technologies to understand. Keep the boundary provisional: observing an apparently self-contained task may expose an upstream decision or downstream acceptance step that materially changes the unit of study.

  • Name the event that starts the workflow and the condition that ends it.
  • Describe the accepted output and who depends on it next.
  • Include routine, incomplete and exception cases rather than only the expected path.
  • Record the roles, handoffs, variations and unresolved questions inside the boundary.

For the onboarding example, the useful boundary runs from receipt of an approved sales handoff to acceptance of the brief by a delivery lead. It includes routine, incomplete and changed-scope cases, not merely the coordinator's drafting session. The evidence protocol here is a practical editorial synthesis, not a standardised package validated by one source. Choose the smallest proportionate mix of interviews, observation, artefacts, diaries and operational records that can answer the question and expose important blind spots.

How do recent-case interviews reveal what happened?

A man points to a page in an open ring binder while a woman takes notes beside a laptop and loose office papers.

Recent-case interviews reveal a sequence of actions, decisions and handoffs when they stay anchored to one completed instance. Ask what triggered the case, what arrived, what happened next, which decisions were made, who received each handoff, what output was accepted and where events departed from the expected route. GOV.UK guidance favours stories and concrete examples over accounts of how work should happen, using open, neutral follow-up questions. One reconstructed case remains illustrative, so contrast routine and difficult cases.

  • What happened next?
  • What told you to act?
  • Who did you contact, and what did they need?
  • What was uncertain or unusually difficult?
  • Can you show the item you used?

Probe the information cues, alternatives, uncertainty and difficult judgements behind visible steps, while presenting this as compact cognitive-demand probing rather than formal ACTA. In the running example, interview coordinators, sales operations and delivery leads so the handoffs are represented. People who deliver a service together may also be interviewed in pairs or small groups, and an interview can accompany an observed task. One original workflow study found that demonstrations and artefact inspection supplied details that verbal descriptions had not produced.

What should you watch in the normal work setting?

A woman files a white card into a thick client folder while her colleague observes and writes in a spiral notebook.

Watch the task with its usual equipment, documents, information, interruptions and dependencies. Note source switching, waiting, verification, rework, handoffs, workarounds and moments when someone interprets incomplete or conflicting material. Contextual research is valuable precisely because normal surroundings can expose barriers and support activity that retrospective descriptions miss. It is still a small, interpretive sample: the observer's presence may change behaviour, and a few observed cases cannot represent an entire workforce.

Choose the observation mode deliberately and record it. Silent observation preserves more of the natural flow but can leave motives unclear. Occasional questions provide context with limited interruption. Continuous explanation yields depth while changing the activity more substantially. Whichever mode is used, separate what was seen from what the researcher inferred. At the end, ask the participant to confirm, correct or dispute the reconstructed flow; shared interpretation improves the account without converting it into population-level proof.

  • Case context and observation mode
  • Observed action and referenced artefact
  • Decision, uncertainty or information cue
  • Handoff, interruption or workaround
  • Immediate consequence
  • Researcher interpretation, clearly labelled

Frame the exercise as research into the workflow, never as an assessment of individual productivity. Obtain informed participation, minimise researcher influence, reconfirm consent before introducing any new recording and store personal data securely. Collect only what the research question requires, restrict access and involve the appropriate privacy, security, legal, workforce-relations, accessibility or domain owners where relevant. These are safeguards for responsible research, not a legal or compliance determination, and sensitive or regulated work may require qualified professional guidance before observation begins.

Which methods reveal what interviews or observation miss?

A researcher sorts printouts and pastel sticky notes into groups around large divider folders on a broad worktable.

Artefacts, brief task diaries and operational records reveal parts of work that a scheduled interview or observation may not capture. Ask participants to explain templates, checklists, spreadsheets, messages, drafts, paper notes and queue views used in a real case. Their format and markings may expose state, priorities, memory aids and coordination practices absent from an official process map. Confirm meaning with the participant: possession of an artefact does not show how, why or how often it is used.

Use a short diary when relevant task instances are intermittent or distributed across time, then follow entries with clarification. An applied GOV.UK project captured feedback during or just after real use and combined the diary with interviews and observed tasks; it does not establish a universal study duration or prompt frequency. Diary entries remain self-reported. Where usable case, activity and timestamp records exist, inspect recurrence, waiting, rework, duplicated handling, deviations, missed deadlines and recurring quality problems.

What each evidence method contributes—and where it needs support
Evidence methodWhat it can revealWhat it can missHow to corroborate
Recent-case interviewSequence, rationale, uncertainty and handoffsForgotten detail and recurrence across casesAsk for artefacts, demonstrations and contrasting cases
Contextual observationTools, interruptions, workarounds and tacit supportInfrequent events and unaffected behaviourConfirm the reconstruction and compare other cases
Workflow artefactsState, memory aids, revisions and coordinationMeaning, frequency and unrecorded contextAsk the user to explain each item in a case
Brief task diaryReal instances spread across timeUnreported activity and retrospective interpretationClarify entries in follow-up interviews
Operational recordsTiming, variants, queues, returns and reworkOffline activity, motives and inconsistent case definitionsReconcile logs with people, observation and artefacts

Logs are examples of recorded behaviour, not the whole workflow. They may omit negative cases and offline clarification, sit across disconnected systems or lack a stable case definition. A timestamp can show when a case waited but cannot establish why it waited. Less common exceptions also deserve examination before they are discarded: their effort or consequence may matter even when their frequency is modest. For onboarding briefs, compare intake material, messages, checklists, draft trails and correction comments with available queue age, returns and missing-field reasons.

When does a pain point become a credible bottleneck hypothesis?

A research team compares rows of coloured case folders while moving round wooden markers among them, with wooden arrows at the table edge.

A pain point becomes a credible bottleneck hypothesis when multiple relevant cases or evidence sources place a recurring constraint at the same workflow point and connect it to an observable consequence. Operational records can reveal waiting, deviations and rework, but incomplete logs must be reconciled with observation, artefacts and the interpretations of people doing the work. Keep cases that contradict the emerging account. Contextual inquiry supports checking researcher interpretations with participants while cautioning against broad inference from small, subjective samples.

  1. Recurrence: Does the constraint appear across more than one relevant case, role or source?
  2. Location: Where does the queue, wait, information gap, rework loop or judgement overload enter?
  3. Consequence: What observable delay, repeated handling, correction, inconsistency, avoidable effort or exposure follows?
  4. Mechanism: How might the constraint produce that consequence, and do practitioners confirm or correct the account?
  5. Counterevidence: Which cases avoid the problem, what differs and which alternative explanations remain?
  6. Actionability: Could a change alter the outcome, and is AI preferable to a rule, clearer ownership, training, better information, process change or no intervention?

This six-question test is an editorial decision rule, not a formal statistical test and not proof of causation. In the onboarding example, the hypothesis shifts from slow summarisation to missing decisions and contradictory fields in difficult cases. Routine briefs are completed quickly, longer briefs are not consistently slower, and system timestamps omit offline clarification. That counterevidence prevents an inflated conclusion. If recurrence, location or consequence remains unclear, retain the complaint as an open lead and gather more evidence rather than manufacturing a confidence score.

Do not ask where AI fits until you can show the recurring constraint, its consequence and the uncertainty that remains.

What belongs on an evidence-backed AI opportunity card?

Colleagues lean over a blank opportunity sheet centred between separate piles of supporting and contradictory case papers.

An opportunity card should be a bounded decision record, not a polished sales pitch for AI. Record the workflow trigger, end condition, output, downstream user and roles, followed by the observed period, case types, represented roles and method mix. Separate the stated pain from the observed pattern. Link substantive statements to field-note identifiers, redacted artefacts, diary entries or record queries; label anything else as an assumption. Preserve disagreement, negative cases, gaps and alternative explanations rather than compressing them into one score.

  • Operational consequence and the evidence connecting it to the pattern
  • A bounded opportunity statement naming who needs help, with which task and towards what outcome
  • Information availability, data quality, permissions, security, privacy, workforce, accessibility and human-oversight constraints
  • AI assistance compared against deterministic rules, ownership or workflow changes, training, better information and no intervention
  • The next smallest test, its assumption, relevant cases, success and failure evidence, owner and review date

The NIST AI RMF Playbook's Map function calls for documentation of purpose, users, operational setting, business value, limitations, impacts, application scope and human oversight. It also asks teams to consider non-AI and non-technology alternatives, without prescribing a winner. The guidance is voluntary and context-dependent, not a compliance certification. Similarly, official use-case discovery guidance retains an evidence requirement, an owner, a next milestone and a later test rather than treating agreement in a workshop as evidence of value in operational use.

Finish with one decision: test AI, test a non-AI change, gather more evidence or stop. For the onboarding example, the next sensible test is required intake fields, an explicit handoff owner and a visible exception queue. Drafting assistance can be considered later for fact-complete cases, but it should not be presented as the remedy for missing upstream decisions. Propose AI only when the evidence supports a bounded task, a recurring constraint, an observable consequence, workable information and oversight, and a testable advantage over simpler alternatives.

Frequently asked questions

How do you identify a good AI use case in a business workflow?

Define a recurring workflow, inspect real cases and locate a constraint with an observable operational consequence. Preserve uncertainty and counterevidence, then compare AI with non-AI alternatives. Propose only the smallest test that can distinguish between the competing explanations or interventions.

What is the difference between a pain point and an AI opportunity?

A pain point is a valid report of frustration or difficulty. An AI opportunity is narrower: evidence places a recurring constraint within a bounded task, connects it to a consequence, identifies workable information and oversight, and gives AI a plausible advantage worth testing.

How should teams observe employees without creating workplace surveillance?

Frame observation as workflow research rather than individual performance assessment, obtain informed participation and state the observation mode. Minimise collection, restrict access and let participants correct the reconstruction. Involve the appropriate organisational owners before handling sensitive, confidential or regulated information.

Do you need interviews, job shadowing, diaries, artefacts and process logs for every workflow?

No. This method is a practical synthesis, not a compulsory checklist. Choose the smallest combination that answers the research question and compensates for known blind spots; add another method only when it resolves a material gap.

Can workflow observation prove that a bottleneck causes delay or rework?

No. Observation, interviews, artefacts and records can support a coherent bottleneck hypothesis, but they do not establish causation by themselves. Preserve contrary cases and alternative explanations, then run the smallest discriminating test instead of claiming proof.

ModelFold logo

ModelFold Editorial Desk

We report on how AI actually lands inside a business. Our work starts from named sources, separates what we found from what we think, and uses AI assistance for research and drafting under documented editorial controls. We are not a substitute for individual expert review.