How to Observe Real Work Before Proposing an AI Use Case
A practical field-research method for verifying workflow pain, locating recurring constraints, and deciding whether AI or a simpler change merits a test.
A complaint is a research lead, not an AI use case. Imagine an operations team reporting that client onboarding briefs take too long and requesting an AI summarizer. A workflow map drawn from memory makes drafting look like the obvious target. Completed cases may tell a different story: difficult briefs wait for missing commercial decisions, coordinators reconcile contradictory fields, and drafts are rewritten when the facts finally arrive. The reported frustration remains valid, but it does not by itself establish where the recurring constraint sits or what consequence it creates. The responsible next move is to observe a bounded sample of work, preserve conflicting evidence, and compare AI with simpler changes before proposing a test.
Key takeaways
A complaint is a research lead, not yet an AI use case.
Observe a bounded sample and use only the evidence methods needed to answer the workflow question.
Treat interviews, observation, artifacts, diaries, and records as complementary partial views.
A credible bottleneck hypothesis preserves recurrence, location, consequence, mechanism, counterevidence, and actionability.
Compare AI with rules, ownership changes, process redesign, training, better information, and no intervention.
What must you define before observing a workflow?
Define the work by its trigger, end condition, output, downstream user, participating roles, and relevant case types before collecting evidence. Official discovery guidance begins with recurring workflows and real outputs, including their frequency, difficulty, and dependencies. Interview guidance likewise starts with explicit research questions and the processes or technologies that must be understood. Keep the boundary provisional: observation may reveal that an upstream decision or downstream acceptance step belongs inside the inquiry.
Trigger: the event that starts the case
End condition: the point at which the work is accepted
Output and downstream user
Roles, handoffs, and decision owners
Routine, incomplete, and exception cases
Questions the field round must answer
For the onboarding example, set the boundary from an approved sales handoff to a brief accepted by a delivery lead. Include routine, incomplete, and changed-scope cases, then record what the work produces, how it varies, and where handoffs occur. The method described here is a practical editorial synthesis, not a standardized protocol: choose the smallest proportionate mix of interviews, observation, artifacts, diaries, and records that can answer the question.
How do recent-case interviews reveal what actually happened?
Recent-case interviews reveal the sequence, decisions, and exceptions in a completed instance more reliably than a broad request to describe the usual process. Ask for one specific case, use open and neutral prompts, and follow the evidence from trigger to accepted output. Interview people across the handoffs—such as coordinators, sales operations, and delivery leads—either separately or, where appropriate, in pairs or small groups. One case illustrates work; contrasting cases and roles are needed before inferring a pattern.
What triggered this case?
What information arrived first?
What happened next?
Which decisions were made, and from what cues?
Who received each handoff?
Where did the case depart from the expected path?
Can you show the item you used?
Probe beyond visible actions. Ask what was uncertain, which alternatives were considered, what indicated that a decision was safe to advance, and where experience mattered. These are compact cognitive-demand prompts, not a claim to perform the formal ACTA protocol. Demonstrations and artifacts can be especially useful: an original workflow study found that some activities became sufficiently detailed only when participants showed what they did. Avoid asking participants to design an AI feature while reconstructing the case.
What should you watch when work happens in its normal setting?
Watch the task with its usual equipment, documents, data, interruptions, and dependencies. Contextual observation can expose source switching, waiting, verification, rework, workarounds, support activity, and moments when someone interprets incomplete or conflicting information. Select and record the observation mode. Silent observation preserves more natural flow but may leave motives unclear; occasional questions add context with some interruption; continuous explanation produces richer detail while changing how the activity is performed.
Case context
Observed action
Artifact reference
Decision or uncertainty
Handoff
Interruption
Consequence
Researcher interpretation
Keep observed action separate from interpretation so participants and later reviewers can challenge an inference. Frame the work as research into the workflow, never an assessment of individual productivity. Obtain informed participation, minimize researcher influence, reconfirm consent before any new recording, and handle personal data securely. Collect only what the question requires, restrict access, and involve appropriate privacy, security, legal, labor, accessibility, or domain owners where applicable. At the end, ask the participant to confirm or correct the reconstructed flow. The result remains a small, interpretive sample—not an objective account of the workforce.
Which evidence methods reveal what interviews or observation miss?
Artifacts, short task diaries, and operational records reveal parts of work that a scheduled interview or observation may miss. Ask participants to explain the templates, messages, drafts, notes, spreadsheets, checklists, and queue views used in a real case. An original field study found that artifact markings and formats prompted explanations of personal organizing and tracking strategies. Confirm each artifact's meaning with its user; possession alone does not show why or how often it is used.
Use brief diary entries when relevant events are intermittent or distributed across time, then follow up for clarification. One applied GOV.UK project captured feedback during or just after real use and paired the diary with interviews and observed tasks, but that example establishes no universal study duration or prompt schedule. When usable case, activity, and timestamp records exist, inspect recurrence, waiting, rework, duplicated handling, deviations, deadlines, and quality problems while documenting activity outside the recorded systems.
What each workflow evidence method can reveal—and where it needs corroboration
Evidence method
What it can reveal
What it can miss
How to corroborate
Recent-case interview
Sequence, experience, decisions, uncertainty, and handoffs
Implicit steps, recall gaps, and recurrence across cases
Ask for artifacts, observe a task, and compare contrasting cases
Contextual observation
Tools, interruptions, workarounds, verification, and support activity
Infrequent events, private reasoning, and work outside the session
Ask neutral questions, review notes with the participant, and use diaries
Workflow artifacts
State, priority, memory cues, revisions, and coordination practices
Meaning, frequency of use, and activity that leaves no artifact
Have participants explain markings and connect artifacts to a specific case
Short task diary
Task instances and experience near the point of use across time
Unreported events, incomplete entries, and details needing clarification
Follow entries with interviews, artifacts, or targeted observation
Operational records
Timing, recurrence, queues, deviations, duplicated handling, and rework
Offline work, motives, negative cases, and inconsistent case definitions
Reconcile logs with practitioners, observation, artifacts, and exceptions
For onboarding briefs, compare the intake form, source records, messages, checklist, draft trail, and correction comments. Diary entries can capture the missing inputs or revisions tied to real briefs, while records may show queue age, returns, and missing-field reasons. Logs remain incomplete examples of behavior: they may span disconnected systems, omit offline clarification, or lack a shared case definition. Preserve uncommon exceptions until their effort and consequence are understood; do not invent a universal cutoff for importance.
When does a reported pain point become a credible bottleneck hypothesis?
A pain point becomes a credible bottleneck hypothesis when evidence places a recurring constraint at a specific workflow point, connects it to an observable consequence, and survives a deliberate search for counterevidence. Operational records can reveal waiting, deviations, and rework, but incomplete logs must be reconciled with observation, artifacts, and practitioner interpretation. Use the following six questions as an editorial decision rule, not a causal or statistical test.
Recurrence: Does the constraint appear across more than one relevant case, role, or evidence source?
Location: Where does the queue, wait, rework loop, information gap, or judgment overload enter?
Consequence: Is the pattern connected to delay, repeated handling, correction, dropped work, inconsistency, avoidable effort, or risk?
Mechanism: How could the constraint produce that consequence, and do people doing the work confirm or correct the explanation?
Counterevidence: Which cases avoid the problem, what differs, and which alternate explanations remain?
Actionability: Could changing this point alter the outcome, and is AI more suitable than a rule, clearer ownership, process change, training, better information, or no intervention?
In the hypothetical onboarding sample, difficult cases repeatedly wait for missing decisions and require contradictory fields to be reconciled. Routine briefs are completed quickly, longer briefs are not consistently slower, and system timestamps omit offline clarification. The hypothesis therefore shifts from slow summarization to incomplete and conflicting inputs, but it remains a hypothesis. If recurrence, location, or consequence is missing, keep the complaint open as a research lead. If the pattern is coherent, design the smallest test that could distinguish the leading explanation from its alternatives.
Do not ask where AI fits until you can show where a recurring constraint appears, what consequence follows, and what remains uncertain.
What belongs on an evidence-backed AI opportunity card?
An evidence-backed opportunity card records the bounded work, observed pattern, counterevidence, constraints, alternatives, and next test without presuming that AI is the answer. NIST's voluntary AI RMF Playbook calls for documenting intended purpose, users, operational setting, business value, limitations, impacts, application scope, and human oversight. It also asks teams to consider non-AI and non-technology alternatives. The card translates those concerns into a compact decision record tied directly to field evidence.
Workflow boundary: trigger, end condition, output, downstream user, and roles
Observed sample: period, case types, roles represented, and method mix
Stated pain: whose account it is and the recent cases behind it
Observed pattern: steps, decisions, handoffs, workarounds, and variants
Operational consequence
Evidence references: field notes, redacted artifacts, diary entries, and record queries
Counterevidence, gaps, disagreements, and alternate explanations
Opportunity statement: who needs help with which bounded task and toward what outcome
Alternatives considered, including no intervention
Next smallest test, relevant cases, and success and failure evidence
Owner and review date
Link each substantive statement to an evidence identifier or mark it as an assumption. Preserve disagreement instead of collapsing the record into a composite confidence score. Official discovery guidance similarly retains an evidence requirement, owner, next milestone, and later test rather than treating workshop agreement as proof. For onboarding, first test required intake fields, an explicit handoff owner, and a visible exception queue. Consider drafting assistance later, only for fact-complete cases, if it offers a testable advantage over the revised non-AI workflow.
End with one decision: test AI, test a non-AI change, gather more evidence, or stop. Propose AI only when the evidence supports a bounded task, recurring constraint, observable consequence, workable information and oversight, and a testable advantage over simpler alternatives. Before observing sensitive work or handling employee, customer, confidential, or regulated information, involve the appropriate organizational owners and seek qualified guidance for applicable obligations or high-stakes decisions. This research method is not a compliance determination.
Frequently asked questions
How do you identify a good AI use case in a business workflow?
Start with bounded, recurring work rather than an AI capability. Reconstruct real cases, observe the workflow where appropriate, and connect a recurring constraint to an operational consequence while recording counterevidence. Compare AI with rules, ownership changes, process redesign, training, better information, and no intervention before selecting a controlled test.
What is the difference between a pain point and an AI opportunity?
A pain point is a valid report of frustration or difficulty. An AI opportunity is a bounded task supported by evidence about recurrence, workflow location, consequence, available information, constraints, and a plausible intervention. The gap matters because the most visible frustration may be downstream of the actual constraint.
How should teams observe employees without creating workplace surveillance?
Frame observation as workflow research, not individual performance scoring, and require informed participation. Record the observation mode, collect only evidence needed for the research question, restrict access, and let participants confirm or correct the reconstruction. Involve appropriate privacy, security, legal, labor, accessibility, and domain owners where applicable.
Do you need interviews, job shadowing, diaries, artifacts, and process logs for every workflow?
No. The combined method is a practical synthesis, not a universal checklist. Choose the smallest proportionate mix that answers the research question, then add another method only when a known blind spot could materially change the decision.
Can workflow observation prove that a bottleneck causes delay or rework?
No. Interviews, contextual observation, artifacts, diaries, and operational records can support a bottleneck hypothesis, but observational evidence does not by itself establish causation. Preserve negative cases and alternate explanations, then run the smallest discriminating test instead of presenting the hypothesis as proof.
References & Sources
This article was researched using the following sources:
We report on how AI actually lands inside a business. Our work starts from named sources, separates what we found from what we think, and uses AI assistance for research and drafting under documented editorial controls. We are not a substitute for individual expert review.
Practical guide to mapping real workflows, redesigning exceptions, testing approvals, defining handoffs, and deciding when automation is ready to deploy.
Practical framework for mapping tasks, testing AI assistance, tracking hidden effort, and redesigning roles only after workload evidence is stable enough.
Build an operating model by assigning decision rights, evidence, escalation, and review across standards, funding, delivery, risk, operations, and reuse.