Practical intelligence for accountable AI programmes.

Search AI strategy, automation, or governance...
Toggle menu

AI Adoption and Work Redesign

Redesigning Roles by Task When Introducing AI Assistance

A practical South African guide to mapping tasks, testing AI assistance, tracing hidden effort and redesigning roles after a measured pilot.

A cross-functional team rearranges coloured task cards, work samples and wooden markers across a large office table.

Redesigning a role for AI assistance starts by examining the work it contains, not by deciding whether its job title is automatable. A service agent may draft a reply faster while a team leader inherits more reviews, a specialist receives harder escalations and a knowledge owner must correct the material feeding the assistant. Counting the agent's faster production alone makes that transferred work disappear on paper. A defensible redesign therefore maps local tasks, tests a recorded AI configuration, follows effort to its destination and changes the role only after the whole operating picture is visible.

The operating principles

  • Redesign the task bundle, not the job title.
  • Production time saved is not team capacity until verification, judgement, exceptions, coordination, learning and monitoring are counted.
  • Test one recorded AI configuration on representative cases before choosing a human-AI task arrangement.
  • Every transferred or newly created duty needs a named owner, authority, queue and capacity.
  • If AI removes work that developed expertise, replace that practice deliberately before changing the role.

Why must role redesign begin with tasks rather than job titles?

The same service employee sorts a packet, checks a binder and files a folder at separate stations while a manager watches.

Tasks are the useful starting point because one job combines work that may remain human-performed, gain bounded assistance, move to an AI-first flow or become narrowly automated. The ILO's refined exposure index uses task-level occupational information and describes transformation as the more likely general outcome than replacement. The OECD likewise reports that AI can automate, complement or introduce tasks within a job. Neither source tells a South African employer how much local capacity will change, however; exposure is a prompt to investigate, not a staffing result.

The design question is therefore local: what does a person actually do, under which conditions, and what changes when a particular system is introduced? After answering that, the organisation can decide how tasks should be bundled into complete roles and how those roles should relate to one another. This avoids reducing judgement, customer accountability, coaching or exception work to miscellaneous leftovers. It also keeps the eventual role decision connected to observable demand, consequences and authority rather than to a broad occupational label.

What belongs in a credible current-state task ledger?

A seated warehouse coordinator compares paperwork with a metal component while a standing observer records notes at the workbench.

A credible ledger describes each task as an action and object, then records what starts it, what information it uses, what it produces, when it is complete, who owns it, who receives it and where exceptions go. Split a process where evidence, judgement, consequence or ownership changes. Combine small motions only when they will change together and share the same control. That level of detail is enough to test a meaningful unit without turning the ledger into a catalogue of every click.

  • Measure demand and its pattern over a stated period, including peaks and queues.
  • Record human touch time separately from elapsed waiting time; waiting may affect service flow without being labour effort.
  • Capture variation, rework, returns, dependencies, handovers, consultations and approvals.
  • Describe the consequence of error or delay, required authority and domain knowledge.
  • Record how proficiency is learned and refreshed, plus the evidence source and confidence for every baseline.

Use interviews with workers and managers, direct observation, work samples and system records together. Each exposes something the others can miss: records may show volume but not informal recovery work, while interviews may recall difficult cases without establishing their frequency. O*NET can provide task language, ratings, work activities and context as a completeness check, but it is a United States occupational database rather than a description of one local operation. The ledger format and baseline fields here are an adaptable editorial method, not an O*NET standard.

How should each task be tested before work is assigned?

A quality specialist reviews a tabbed case folder while a colleague sorts wooden person markers into trays at an office table.

Test each candidate task against one identifiable AI configuration before assigning responsibility. Record the model, tools, prompts, data, controls, test cases and date, then keep them fixed for that test round. Include routine high-volume work, legitimate variants, ambiguous boundaries, rare consequential cases, missing or conflicting inputs and known failures. Where it matters, separate results by experience, product, channel, customer group or region. A team average can conceal a weak boundary or an uneven effect on the people expected to operate it.

  • Define what counts as a usable output and who has completion authority.
  • State which inputs and actions are allowed and what may not be delegated.
  • Set review or monitoring duties, exception routes and stop conditions.
  • Specify fallback and recovery when the assistant is unavailable or wrong.

The caution is empirical. In a preregistered experiment with 758 consultants, AI improved measured performance on selected tasks inside the tested capability frontier but reduced correctness on a selected complex task outside it. The study also describes that frontier as uneven, changing and difficult to locate in advance. Those results belong to the study's people, model and tasks; they do not predict another workplace. NIST's voluntary AI Risk Management Framework nevertheless reinforces the need to differentiate responsibilities, operator proficiency and oversight processes.

Four adaptable task configurations and the operating design each requires
Task configurationHuman responsibilityAI responsibilityRequired operating design
Human-performedPerforms and owns the taskAbsent or limited to unrelated supportKeep ownership explicit and record why the arrangement remains appropriate
AI-assisted humanFrames the work, checks required evidence and makes the decisionProvides bounded retrieval, drafting or analysisDefine allowed inputs, review criteria, completion authority and non-delegable work
AI-first with human decision or reviewReviews defined cases or makes the consequential decisionProduces or routes the first outputName the queue owner, evidence, triggers, priority, authority, exception route and peak capacity
Bounded automationMonitors performance and resolves exceptionsCompletes a narrowly defined task within approved conditionsDefine boundaries, observability, stop conditions, change control, exception ownership and fallback

Which effort changes matter beyond production time?

Office workers assemble paperwork, inspect a metal sample, discuss a binder and hand over a folder in a shared operations room.

Measure five separate effort accounts: production, verification, judgement, exception handling and coordination. This ledger is an editorial accounting device, not a research standard. Production covers creating or executing the first work product. Verification covers checking sources, calculations, completeness, policy fit and downstream usability. Judgement covers framing, interpretation, choice, authority and responsibility. Keeping these accounts distinct prevents a short drafting time from obscuring the effort needed to decide whether the draft is appropriate.

  • Exception handling includes ambiguous inputs, tool failures, overrides, correction, recovery and escalation.
  • Coordination includes routing context, answering questions, reconciling outputs, coaching and maintaining shared knowledge.
  • Every moved activity should name its destination role or team, together with queue demand and peak workload.
  • Review workload intensity, autonomy, task variety and retained responsibility alongside time.

These wider measures matter because task change can alter the quality as well as the quantity of work. The OECD review notes that AI may reduce tedious activity but may also increase pace, reduce autonomy or narrow the task set, with effects varying by worker and setting. NIST calls for clear responsibilities, communication, feedback and monitoring, although it does not prescribe these five accounts. The practical point is to make all work visible, including duties that land with leaders, specialists or support functions.

AI does not remove work in one clean block; it changes where effort, authority, exceptions and learning live.

How can changed tasks become workable roles without eroding expertise?

An older technician guides a younger colleague through a mechanical housing inspection while another technician works at a rear bench.

Recompose changed tasks by naming the duties, authority and capacity that make the new operation work. Depending on the design, that may require an output reviewer, exception owner, knowledge maintainer, system steward, performance monitor, learning owner and escalation authority. A generic “human in the loop” label is not enough: specify the evidence reviewed, triggers, competence, decision rights, queue owner and performance under peak demand. NIST supports differentiated responsibilities and oversight, but it does not mandate one staffing or review model.

  • Identify work that currently supplies observation, repetition, feedback and difficult-case exposure.
  • Replace lost practice, where needed, through sampled original work, shadowing, simulations, case review, rotation and coaching.
  • Give developing staff progressively consequential decisions under suitable supervision.
  • Keep experts involved in evaluation, exception analysis, knowledge maintenance and updates to the task boundary.

A hypothetical support team illustrates the redistribution. AI-assisted retrieval and drafting may reduce production on tested routine cases, while agents still diagnose ambiguity and make decisions within delegated authority. Team leaders or specialists may receive more reviews, escalations, knowledge corrections and coaching. Evidence should also be segmented by experience: a study of 5,179 support agents found markedly different productivity effects across experience groups, while a field experiment with 776 product-development professionals raised unresolved questions about expertise development. Neither result should be transferred as a staffing ratio.

When is the evidence strong enough to change roles or capacity?

A work team reviews trays of sorted components as the woman in the centre points to samples in a light-industrial workspace.

There is enough evidence only after a measured pilot has reconciled the complete workload and produced a reasonably stable operating picture. For a stated period, estimate gross production effort reduced at observed demand, then subtract added verification, judgement, exception handling, coordination, learning, monitoring and knowledge-maintenance effort. Adjust for changes in demand, service levels, quality, rework, peak queues and downstream transfers. Treat any remainder as provisional capacity, not an immediate promise of fewer posts or permanently higher service volumes.

  • Track quality, defects, overrides, recovery, workload distribution and worker feedback.
  • Check autonomy, task variety and whether people can diagnose failures without the assistant.
  • Reassess after changes to the model, prompt, data, controls, workflow, task mix or operating context.
  • Keep the existing arrangement when ownership, representative evidence, workload or the expertise plan remains incomplete.

Run the redesigned process long enough for side work and less common exceptions to appear; a brief improvement in average production does not establish stable team capacity. NIST supports ongoing monitoring and feedback, while the consulting experiment shows why changed configurations require reassessment. The pilot gate and reconciliation are nevertheless operational synthesis, not formulae prescribed by NIST, the OECD or the experiments. Their purpose is to delay structural commitments until the organisation can see who is doing the work and under what conditions.

Do not change staffing, job descriptions, performance measures or service commitments while the task ledger, review design, exception capacity, ownership or learning path remains unresolved. Employment, labour-relations, accessibility, discrimination, privacy, safety, legal, regulatory and professional determinations belong with qualified and authorised organisational owners. A task ledger can expose the evidence those owners need, but it cannot make their judgements. Used this way, role redesign becomes a disciplined operating decision rather than a shortcut from an AI demonstration to a workforce conclusion.

Frequently asked questions

How do you redesign a job for generative AI?

Break locally observed work into testable tasks and baseline their demand, effort, judgement, dependencies and consequences. Test a fixed AI configuration, record changes across the five effort accounts and then rebuild the role with named duties, authority and capacity. Confirm the design through a measured pilot before changing staffing or expectations.

What is task-based AI workforce planning?

It is workforce planning based on actual task demand, effort, conditions and tested AI effects rather than an occupation-wide automation estimate. It follows retained, reduced, transferred and newly created work across roles. The result is a local workload picture, not a universal forecast.

How should tasks be divided between people and AI?

Choose among human-performed work, an AI-assisted human, AI-first work with human decision or review, and bounded automation. Base the choice on representative tests of a recorded configuration. Define authority, inputs, review, exceptions, monitoring, stop conditions and fallback for the selected arrangement.

How do you measure AI workload redistribution?

Measure production, verification, judgement, exception handling and coordination separately for each task. Record where every reduced, retained or added activity lands, including the destination owner and queue. Compare demand, peaks, quality and rework over the same stated period.

When can AI time savings count as workforce capacity?

Only provisionally, after a pilot reconciles reduced production effort with review, judgement, exceptions, coordination, learning, monitoring and knowledge maintenance. Changed demand, service quality, rework, peak queues and transferred work must also be included. If the evidence or ownership remains incomplete, retain the current capacity and role design.

ModelFold logo

ModelFold Editorial Desk

We report on how AI actually lands inside a business. Our work starts from named sources, separates what we found from what we think, and uses AI assistance for research and drafting under documented editorial controls. We are not a substitute for individual expert review.