Practical intelligence for accountable AI programmes.

Search AI strategy, automation, or governance...
Toggle menu

AI Adoption and Work Redesign

Redesigning Roles by Tasks When Introducing AI Assistance

A practical method for mapping tasks, testing AI assistance, tracing redistributed effort and redesigning roles only when ownership and workload are clear.

A cross-functional team rearranges coloured task cards, work samples and wooden markers across a large office table.

Redesign a role by examining its tasks, not by declaring the whole job automated or augmented. An assistant might help a service agent produce a reply faster while a team leader acquires a larger review queue, a specialist receives harder escalations and a knowledge owner corrects faulty material. Counting only the quicker draft mistakes local production savings for usable capacity. A defensible redesign traces demand, effort, authority, exceptions and learning across everyone affected before changing duties or expectations.

What to remember

  • Redesign the task bundle, not the job title.
  • Production time saved is not capacity until review, judgement, exceptions, coordination, learning and monitoring are counted.
  • Test one recorded AI configuration on representative cases before assigning human and AI responsibilities.
  • Every transferred or newly created duty needs an owner, authority, queue and capacity.
  • Replace expertise-building practice deliberately when AI removes the work that once provided it.

Why should role redesign begin with tasks rather than job titles?

The same service employee sorts a packet, checks a binder and files a folder at separate stations while a manager watches.

Tasks are the useful unit of evidence because one job contains work that may change in different ways. AI might support retrieval, leave a delegated decision with a person, automate a tightly bounded routing step and create new monitoring work. A single label for the role conceals those differences. The eventual decision still concerns a complete job and its relationships, but it should follow evidence about the component work rather than precede it.

The ILO's refined exposure index uses task-level occupational information and says transformation is generally more likely than replacement. The OECD likewise describes jobs in which AI can automate, complement or introduce tasks, with many workers encountering change through their tasks and working environment. Neither source turns occupational exposure into a local automation rate, productivity estimate or staffing conclusion. Treat exposure as a prompt to inspect actual work, operating conditions and consequences.

What belongs in a current-state task ledger?

A seated warehouse coordinator compares paperwork with a metal component while a standing observer records notes at the workbench.

A current-state ledger should describe what people actually do, why the work arrives and what makes it complete. Write each task as an action and object, then identify its trigger, inputs, output, current owner, receiver and exception route. Split a step when its evidence, judgement, consequence or owner changes. Combine small motions only when they will change together and remain under the same control. This format is an adaptable editorial method, not an O*NET standard.

  • Demand volume and pattern over a stated period
  • Human touch time, elapsed waiting time and queue size
  • Variation, boundary cases and exception families
  • Rework, returns and known failure modes
  • Dependencies, handovers, consultations and approvals
  • Consequences of error or delay
  • Judgement, authority, independence and domain knowledge required
  • How proficiency is learned, practised and refreshed
  • Evidence source and confidence for each baseline

Build the baseline from several forms of evidence: observation reveals side work, interviews explain judgement, work samples expose variation and system records show volumes or queues. None substitutes for the others. Keep human touch time separate from elapsed waiting time; waiting matters to service flow but is not automatically labour effort. O*NET can supply useful occupational task language and completeness prompts, yet its US occupational data cannot establish what happens, how often or at what cost in a particular UK organisation.

How should each task be tested before work is assigned to people or AI?

A quality specialist reviews a tabbed case folder while a colleague sorts wooden person markers into trays at an office table.

Test each task against one fixed, recorded AI configuration before choosing its future owner. Record the model, tools, prompts, permitted data, controls, cases and test date so the result describes an identifiable system rather than AI in general. A preregistered experiment with 758 consultants found better measured performance on selected tasks inside the tested capability frontier but lower correctness on a selected complex task outside it. Those results belong to that study, not to another organisation.

  • Routine, high-frequency cases
  • Legitimate variants and ambiguous boundaries
  • Rare but consequential cases
  • Missing, conflicting or stale inputs
  • Known historical failures
  • Relevant differences in experience, product, channel, customer group or region

Judge quality, effort and failure patterns by case family rather than relying on one average. The consulting study describes the capability frontier as uneven, changing and difficult to locate in advance; a permanent task label is therefore weak operating evidence. For the chosen configuration, define completion authority, allowed inputs and actions, review or monitoring duties, exception routes, stop conditions and fallback arrangements. NIST's voluntary AI RMF supports differentiated responsibilities, proficiency and oversight without prescribing a staffing model.

Four adaptable configurations for tested tasks
Task configurationHuman responsibilityAI responsibilityRequired operating design
Human-performedPerforms and owns the task.Absent or limited to unrelated support.Record current ownership and why evidence or benefit does not justify a change.
AI-assisted humanFrames the work, checks required evidence and remains the decision-maker.Provides bounded retrieval, drafting, analysis or transformation.Define allowed inputs, review criteria, completion authority and non-delegable decisions.
AI-first with human decision or reviewReviews defined cases or makes the consequential decision.Produces or routes the initial output.Name queue ownership, evidence, triggers, priority, exception routes, authority and peak capacity.
Bounded automationMonitors performance and resolves exceptions.Completes a narrowly defined task within approved conditions.Define boundaries, observability, stop conditions, change control, exception ownership and fallback.

Which effort changes must be measured beyond production time?

Office workers assemble paperwork, inspect a metal sample, discuss a binder and hand over a folder in a shared operations room.

Measure five separate effort accounts: production, verification, judgement, exception handling and coordination. They are an editorial accounting device, not a research standard. Production covers creating or executing the initial output. The other four expose work retained by the original owner, transferred elsewhere or created by the new process. Record effort for observed demand over a stated period, name its destination owner and examine queues and peaks as well as averages.

  • Verification: checking sources, calculations, completeness, policy fit and downstream usability
  • Judgement: framing, interpreting context, choosing outcomes, applying authority and accepting responsibility
  • Exception handling: resolving ambiguity, missing inputs, tool failures, overrides, corrections and escalations
  • Coordination: transferring context, routing work, answering questions, reconciling outputs, coaching and maintaining shared knowledge
  • Production: finding, composing, transforming, entering or executing the initial work product

Time is necessary but insufficient. The OECD evidence review says AI may reduce tedious work while also increasing work pace, reducing autonomy or narrowing the task set, with effects varying by worker and setting. Review workload intensity, task variety and retained responsibility alongside minutes saved. NIST also calls for clear responsibilities, communication, feedback and monitoring, which reinforces the need to assign oversight and coordination work explicitly rather than leaving it as invisible overhead.

AI does not remove work in one clean block; it changes where effort, authority, exceptions and learning live.

How can changed tasks become workable roles without losing ownership or expertise?

An older technician guides a younger colleague through a mechanical housing inspection while another technician works at a rear bench.

Recompose changed tasks by turning every necessary duty into named work with criteria, authority, communication routes and capacity. A generic human-in-the-loop label is not an operating control. The design must say what evidence is reviewed, which cases trigger attention, who can decide or intervene and how the queue performs under peak demand. NIST calls for documented roles, differentiated responsibilities, operator proficiency, oversight, feedback and monitoring, but it does not mandate one review or staffing pattern.

  • Output reviewer with defined evidence, criteria and authority
  • Exception owner with priorities, recovery options and escalation routes
  • Knowledge maintainer responsible for sources, procedures and expired guidance
  • System steward responsible for the approved configuration and change record
  • Performance monitor tracking quality, workload, overrides, defects and feedback
  • Learning owner responsible for practice, coaching and proficiency assessment

Protect the work that develops expertise. Identify where people currently gain observation, repetition, feedback, causal understanding and exposure to difficult cases. If assistance removes that practice, replace it deliberately through sampled original work, shadowing, simulations, case review, rotation, coaching and progressively consequential decisions. These are design recommendations, not universally tested interventions. The aim is not to preserve low-value production for its own sake, but to ensure developing staff can diagnose and recover when the assistant is unavailable or wrong.

Consider a hypothetical customer-support team. Tested retrieval and drafting assistance may reduce routine production while agents retain diagnosis, delegated decisions and the customer-facing action. Team leaders or specialists may inherit review, difficult exceptions, knowledge correction and coaching. Evidence from 5,179 support agents found materially different measured productivity effects by experience, while a separate experiment with 776 product-development professionals changed performance and expertise integration across tested team arrangements. Neither result supplies a transferable staffing ratio; both argue for segmented local evidence.

When is there enough evidence to change capacity, roles or service commitments?

A work team reviews trays of sorted components as the woman in the centre points to samples in a light-industrial workspace.

There is enough evidence only after a measured pilot reconciles the complete workload and shows reasonably stable demand. For a stated period, estimate gross production effort reduced for observed demand, then subtract added verification, judgement, exceptions, coordination, learning, monitoring and knowledge maintenance. Adjust for changed demand, service levels, quality, rework, peak queues and effort transferred downstream. Treat any remainder as provisional capacity, not an immediate headcount or service commitment.

  • Quality, defects, overrides and recovery outcomes
  • Demand changes, rework and peak queue performance
  • Effort and workload distribution across destination owners
  • Worker feedback, autonomy and task variety
  • Ability to diagnose and recover without reliable AI
  • Changes to the model, prompt, data, controls, workflow or task mix

Run the redesigned process long enough for side work and exception demand to surface. NIST calls for feedback, continuing measurement and monitoring, while experimental evidence shows that capability can differ across nearby tasks and change over time. Reassess the design whenever its configuration or operating context changes. Do not alter staffing, job descriptions, performance measures or service commitments while representative evidence, workload ownership, effective review, exception capacity or the expertise plan remains incomplete.

Use the ledger as a pilot decision tool, not a shortcut to a predetermined workforce answer. Keeping the present role or staffing model is a valid outcome when the evidence is unresolved. Employment, labour-relations, accessibility, discrimination, privacy, safety, legal, regulatory and professional determinations belong with qualified and authorised organisational owners. Task analysis can make operational consequences visible, but it does not make those judgements or guarantee that a human review step will be effective.

Frequently asked questions

How do you redesign a job for generative AI?

Break locally observed work into tasks and baseline demand, effort, judgement, dependencies and consequences. Test a fixed AI configuration on representative cases, measure all five effort accounts and then recompose duties, authority and queues. Change the role only after the complete workload and expertise plan are credible.

What is task-based AI workforce planning?

It is workforce planning that starts with task demand, effort, judgement, consequences, dependencies and tested AI effects rather than an occupation-wide automation estimate. It traces where work is reduced, retained, transferred or newly created before drawing conclusions about roles or capacity.

How should tasks be divided between people and AI?

Choose among human-performed, AI-assisted human, AI-first with human decision or review, and bounded automation after testing the task. Each choice needs explicit authority, allowed actions, review or monitoring, exception handling, stop conditions and fallback arrangements.

How do you measure AI workload redistribution?

Record production, verification, judgement, exception handling and coordination effort for each task over a stated period. Attribute transferred or new effort to its destination owner, then examine demand, queues, peaks and rework rather than relying on an average time saving.

When can AI time savings be treated as workforce capacity?

Only provisionally, after a measured pilot reconciles complete workload, changed demand, quality, queues, transferred work, monitoring and expertise-development needs. If ownership, representative evidence or stable demand is missing, keep current staffing and commitments unchanged.

ModelFold logo

ModelFold Editorial Desk

We report on how AI actually lands inside a business. Our work starts from named sources, separates what we found from what we think, and uses AI assistance for research and drafting under documented editorial controls. We are not a substitute for individual expert review.