How to Observe Real Work Before Proposing an AI Use Case
A practical field method for observing real workflows, testing reported bottlenecks and framing evidence-backed AI opportunities without mistaking complaints for proof.
A complaint is a research lead, not a ready-made AI use case. Suppose an operations team says client onboarding briefs take too long and asks for an AI summariser. A workflow map drawn from memory makes drafting look like the problem. Completed cases may show something different: difficult briefs wait for missing commercial decisions, then require rewriting when contradictory source fields are resolved. The complaint remains valid evidence of frustration, but the proposed remedy has arrived before the recurring constraint, its consequence and the uncertainty have been established.
What to carry into the field
Treat a complaint as a lead to investigate, not proof that AI is the answer.
Bound the workflow and choose only the evidence methods needed for the research question.
Use interviews, observation, artefacts, diaries and records as complementary partial views.
Preserve counterevidence and alternate explanations when forming a bottleneck hypothesis.
Compare AI with simpler interventions before selecting the next test.
What must you define before observing a workflow?
Define the workflow by its trigger, end condition, output, downstream user, participating roles and relevant case types. Start with recurring work and actual outputs rather than asking broadly where a team could use AI. Write down the research questions, expected handoffs and known variations, while treating the boundary as provisional. Observation may reveal that a supposedly local delay begins upstream or that rework only becomes visible to a downstream team.
Trigger and end condition
Output and downstream user
Roles and handoffs
Routine, incomplete and exception cases
Questions the field round must answer
For onboarding, a workable boundary runs from an approved sales handoff to a brief accepted by a delivery lead. Include routine, incomplete and changed-scope cases so the field round does not study only the happy path. The method is a practical synthesis, not a standardised package: choose the smallest proportionate mix of interviews, observation, artefacts, diaries and records that can answer the questions and compensate for known blind spots.
How do recent-case interviews reveal what actually happened?
Reconstruct one specific, recently completed case from trigger to accepted output. Ask what arrived, what happened next, which decisions were made, who received each handoff and where the case departed from the expected path. Concrete stories preserve the participant's experience while giving the researcher something that can later be compared with artefacts, observation or records. A generic account of how work should happen cannot provide the same case-level sequence.
What triggered this case, and what arrived?
What happened next, and what told you to act?
What was uncertain, and which alternatives did you consider?
Who did you contact or hand the work to?
Can you show the item you used?
What output was finally accepted?
Keep prompts open and neutral; do not ask the participant to design an AI feature during reconstruction. Probe information cues, difficult judgements and uncertainty as well as visible actions, without calling the sequence a formal ACTA protocol. For onboarding, contrast a routine brief with a difficult one and include coordinators, sales operations and delivery leads. Interviews can cover people across a service and be paired with task demonstrations, but one case remains illustrative rather than representative.
What should you watch when work happens in its normal setting?
Watch the work with its usual equipment, documents, data, interruptions and dependencies. Note source switching, waiting, verification, rework, handoffs, workarounds and moments when incomplete or conflicting information requires judgement. Context matters because a neat demonstration away from the work setting can omit the interruptions and support activity that shape a real case. Observation still affects behaviour, so record how it was conducted rather than presenting it as an untouched view.
Case context and observed action
Artefact reference and information cue
Decision, uncertainty or workaround
Handoff, interruption and consequence
Researcher interpretation kept separate from observation
Choose an observation mode explicitly. Silent observation preserves more natural flow but may leave motives unclear; occasional questions add context with limited interruption; continuous explanation provides depth while changing the activity. Separate what was seen from what the researcher inferred, then ask the participant to confirm or correct the reconstructed flow. That checking improves the shared account, but a small, interpretive sample still cannot represent an entire workforce or establish objective truth.
Frame the exercise as research into the workflow, not an assessment of individual performance. Obtain informed participation, minimise collection, restrict access, reconfirm consent before any new recording and handle personal data securely. Involve the appropriate privacy, security, legal, labour, accessibility or domain owners when the work or evidence warrants it. These are safeguards for responsible research, not a substitute for qualified advice about Australian obligations, organisational policy or high-stakes decisions.
Which methods reveal work that interviews or observation miss?
Use artefacts, short task diaries and operational records to extend the field view across forms of work and periods that scheduled observation may miss. Ask participants to explain templates, checklists, spreadsheets, messages, drafts, paper notes and queue views from the case. Their markings or arrangement may expose tracking, memory and coordination practices absent from an official process map. Confirm their meaning with the user; possession of an artefact does not establish why or how often it is used.
What each evidence method contributes, misses and needs for corroboration
Evidence method
What it can reveal
What it can miss
How to corroborate
Recent-case interview
Sequence, decisions, experience and handoffs
Forgotten steps or broader recurrence
Compare contrasting cases, artefacts and records
Contextual observation
Tools, interruptions, workarounds and judgement
Infrequent events and unobserved periods
Ask for clarification and confirm the reconstruction
Workflow artefacts
State, priority, memory cues and coordination
Meaning, frequency and surrounding activity
Have participants explain use in a real case
Short task diary
Intermittent work near the point of occurrence
Unrecorded detail and self-reporting gaps
Follow entries with interviews or observation
Operational records
Recurrence, waiting, rework and variants
Offline activity, motives and inconsistent case definitions
Reconcile with practitioners, artefacts and field notes
Use a brief diary when the question concerns intermittent or distributed events, and follow entries with clarification rather than prescribing a universal duration. Where usable case, activity and timestamp records exist, inspect waiting, repeated handling, deviations, returns and recurring quality problems. Document activity outside the recorded systems and do not assume a timestamp explains why a delay occurred. Less common exceptions also deserve examination because their effort or consequence may matter even when their frequency is modest.
When does a reported pain point become a credible bottleneck hypothesis?
A pain point becomes a credible bottleneck hypothesis when evidence places a recurring constraint at a particular workflow location, connects it to an observable consequence and leaves counterevidence visible. This is triangulation, not a vote among methods and not causal proof. Interview accounts, observations, artefacts, diaries and records answer different questions. The test below is an editorial decision rule for organising those partial views, not a formal statistical test or a confidence score.
Recurrence: does the same constraint appear across relevant cases, roles or evidence sources?
Location: where does the queue, wait, rework loop, information gap or judgement overload enter?
Consequence: is it connected to observable delay, repeated handling, correction, inconsistency, avoidable effort or risk?
Mechanism: how might the constraint produce that consequence, and do people doing the work confirm or correct the account?
Counterevidence: which cases avoid the problem, what differs and which alternate explanations remain?
Actionability: could changing this point alter the outcome, and is AI preferable to a rule, clearer ownership, training, better information or no intervention?
In the onboarding example, the hypothesis shifts from slow summarisation to missing decisions and contradictory fields in difficult cases. Routine briefs are completed quickly, longer briefs are not consistently slower, and system timestamps omit offline clarification. Those negative cases and gaps matter. The team can reasonably test whether a clearer handoff changes waiting and rework, but it cannot yet claim that missing information causes every delay or that drafting assistance will fix the upstream problem.
Do not ask where AI fits until you can show where a recurring constraint appears, what consequence follows and what remains uncertain.
What belongs on an evidence-backed AI opportunity card?
An evidence-backed opportunity card should be a bounded decision record, not a sales pitch for a capability. Record the workflow boundary, observed period, represented roles, case types and method mix. Separate the stated pain from the observed pattern, and link substantive assertions to field-note identifiers, redacted artefacts, diary entries or record queries. Mark assumptions plainly, preserve disagreement, and describe confidence through evidence, gaps and counterevidence rather than a composite score.
Workflow trigger, end condition, output, downstream user and roles
Observed cases, period, participants and evidence methods
Stated pain and concrete recent examples
Observed pattern, operational consequence and evidence references
Counterevidence, missing coverage and alternate explanations
Bounded task, intended users and desired outcome
Information, quality, permission, security, privacy, labour, accessibility and oversight constraints
AI, deterministic, workflow, ownership, training, information and no-intervention alternatives
Next test, addressed assumption, relevant cases, success and failure evidence
Owner, review date and current decision
The NIST AI RMF Playbook supports documenting purpose, users, operational setting, benefits, limitations, impacts, scope and human oversight, while also considering non-AI and non-technology alternatives. Its guidance is voluntary and does not decide which option should win. Apply the same work evidence to each candidate intervention. A broad promise to ‘use AI for onboarding’ is not comparable with a bounded test of drafting assistance on fact-complete cases.
For the onboarding team, first test required intake fields, an explicit handoff owner and a visible exception queue. Define which assumption the test addresses, which cases it covers, what success and failure would look like, who owns it and when the evidence will be reviewed. Finish with one decision: test AI, test a non-AI change, gather more evidence or stop. Propose AI only when the bounded task and evidence show a testable advantage over simpler alternatives.
Frequently asked questions
How do you identify a good AI use case in a business workflow?
Start with a bounded, recurring workflow and examine real cases rather than beginning with an AI capability. Look for a constraint that recurs at a known point, produces an observable consequence and can be addressed through a controlled test. Compare AI with rules, ownership changes, training, better information and no intervention before choosing.
What is the difference between a pain point and an AI opportunity?
A pain point is a valid report of frustration or difficulty. An opportunity adds evidence about recurrence, workflow location, operational consequence, constraints and a plausible intervention. It becomes an AI opportunity only when a bounded AI assist deserves testing against credible non-AI alternatives.
How should teams observe employees without creating workplace surveillance?
Frame the work as non-evaluative workflow research, obtain informed participation and collect only evidence required by the research question. State the observation mode, control access, confirm interpretations with participants and involve appropriate organisational owners. Do not use covert monitoring or individual productivity scoring.
Do you need interviews, job shadowing, diaries, artefacts and process logs for every workflow?
No. The protocol is a practical synthesis, and each method has different strengths and blind spots. Choose the smallest proportionate mix that answers the workflow question, then add another method only when it can test an important gap, contradiction or inference.
Can workflow observation prove that a bottleneck causes delay or rework?
No. Observation and operational records can support a bottleneck hypothesis by showing recurrence, location and associated consequences, but they do not prove causation. Preserve counterevidence and alternate explanations, then run the smallest test capable of distinguishing between plausible accounts.
References and sources
This article was researched using the following sources:
We report on how AI actually lands inside a business. Our work starts from named sources, separates what we found from what we think, and uses AI assistance for research and drafting under documented editorial controls. We are not a substitute for individual expert review.
A practical guide to assigning AI standards, funding, delivery, risk and operations, then turning pilot evidence into portfolio and strategy decisions.