Redesign a role by examining its tasks and the work around them, not by declaring the whole job automated or retained. An AI assistant may help a service agent draft a reply sooner while creating extra review for a team lead, harder escalations for a specialist and maintenance work for the person responsible for shared knowledge. Counting only the faster draft mistakes displaced effort for genuine capacity and can change expectations before the new operation is dependable.
What leaders should carry into the pilot
Redesign the task bundle, not the job title.
Production time saved is not capacity until retained, transferred and newly created effort is counted.
Test one recorded AI configuration against representative local cases before assigning responsibility.
Every new review, exception and maintenance duty needs an owner, authority, queue and capacity.
Replace expertise-building practice deliberately when AI removes the work that previously supplied it.
Why start role redesign with tasks rather than job titles?
Tasks are the right unit of evidence because one job contains activities that AI may automate, complement, leave substantially unchanged or introduce for the first time. The OECD reports that many workers are more likely to experience AI through changed tasks and working conditions than through employment loss. The ILO's refined exposure index likewise uses task-level occupational information and identifies transformation, rather than replacement, as the more likely broad outcome.
Neither finding establishes what will happen in a New Zealand workplace. Exposure is not proof of local automation, productivity, job loss or spare staffing capacity. Begin with observable work: what triggers it, who performs it, which evidence and judgement it requires, and where its output goes. Only then should leaders ask how the changed tasks fit together as a complete role and how that role interacts with managers, specialists and adjacent teams.
What belongs in a useful current-state task ledger?
A useful ledger describes actual local work, its demand and its operating conditions in enough detail to test change. Write each task as an action and object, then record the trigger, inputs, output, completion condition, current owner, receiver and exception route. Split steps when their evidence, judgement, consequences or owners differ. Combine small motions only when they would change together and remain under the same control.
Record demand volume and its pattern over a stated period, including peaks rather than only averages.
Separate human touch time from elapsed waiting time and identify the queue in between.
Capture variation, rework, known failures, dependencies, consultations, handoffs and approvals.
Describe the consequence of error or delay, required authority and relevant domain knowledge.
Note how proficiency is learned, practised, observed and refreshed in the current role.
Attach the evidence source and confidence: observation, worker and manager interviews, work samples and system records should inform one another.
Use O*NET task statements and work-context data as prompts for completeness, not as proof of local demand, effort, judgement or consequence.
How should a task be tested before people and AI divide the work?
Test each task against one identifiable AI configuration before choosing its future arrangement. Freeze and record the model, tools, prompts, data, controls, cases and test date. In a preregistered experiment with 758 consultants, AI improved measured performance on selected tasks inside the tested capability frontier but reduced correctness on a selected complex task outside it. Those results belong to that study's people, model, tasks and period, not to every workplace.
Include routine, high-frequency cases and legitimate variants.
Include ambiguous boundaries and rare but consequential cases.
Include missing, conflicting or stale inputs and known historical failures.
Segment results by experience, product, channel, customer group or region when those differences matter.
Record quality, completion, effort, overrides, failure patterns and the destination of escalated work.
Retest when the model, prompts, data, controls or operating context changes.
The study describes the capability frontier as uneven, changing and hard for workers to locate in advance. Permanent labels such as ‘safe to automate’ therefore carry more confidence than the evidence allows. For the selected configuration, define completion authority, permitted inputs and actions, review or monitoring duties, exception routes, stop conditions and fallback arrangements. NIST's voluntary AI RMF Core supports differentiated human-AI responsibilities, documented operator proficiency and defined oversight, without prescribing one staffing model.
Four task configurations for an explicit future operating design
Task configuration
Human responsibility
AI responsibility
Required operating design
Human-performed
A person completes the task and retains its authority.
Absent or limited to incidental support.
Record why evidence, suitability or net benefit supports keeping the current arrangement.
AI-assisted human
A person frames the task, checks required evidence and remains the decision maker.
Provides bounded retrieval, drafting, transformation or analysis.
Define allowed inputs and actions, review criteria, completion authority and non-delegable elements.
AI-first with human decision or review
A named person reviews defined cases or makes the consequential decision.
People monitor performance and resolve exceptions.
Completes a narrowly defined task within approved conditions.
Define boundaries, observability, stop conditions, exception ownership, change control and fallback.
Which effort changes matter beyond production time?
Measure five separate effort accounts: production, verification, judgement, exception handling and coordination. This ledger is an editorial accounting device, not a research standard. It prevents a quicker first output from being mistaken for a lighter team workload. The OECD evidence review notes that AI may reduce tedious work but may also increase work pace, reduce autonomy or narrow the task set, with effects differing by worker and setting.
Production covers finding, composing, transforming, entering or executing the initial work product.
Verification covers checking sources, calculations, completeness, policy fit and downstream usability.
Judgement covers framing, interpretation, choice, authority and responsibility for the outcome.
For every change, record whether effort fell, remained, moved or appeared anew, and name its destination owner.
Measure queue length and peak demand wherever work moves, particularly when a faster frontline task feeds review or specialist escalation. Do not count elapsed waiting time automatically as labour effort, but retain it as a flow measure because it can reveal congestion. NIST calls for clear responsibilities, communication paths, feedback and monitoring; the five accounts turn those duties into visible workload without suggesting that NIST prescribes this particular method.
AI rarely removes work in one clean block; it changes where effort, authority, exceptions and learning live.
How can changed tasks become workable roles without losing expertise?
Recompose changed tasks by assigning named duties, real authority and enough capacity to perform them. Where required, identify the output reviewer, exception owner, knowledge maintainer, system steward, performance monitor, learning owner and escalation authority. A vague ‘human in the loop’ label is not an operating control: specify what evidence is reviewed, which cases enter the queue, who may decide or intervene, and how the queue behaves during peak demand.
Protect expertise by identifying which current tasks provide repetition, observation, feedback, causal understanding, supervised practice and difficult-case exposure. One study of 5,179 customer-support agents found that measured productivity effects differed materially by experience and offered suggestive evidence that the assistant spread practices associated with more able workers. A separate field experiment involving 776 product-development professionals found changed performance and expertise integration across tested individual and team arrangements, while leaving longer-term expertise development unresolved.
Use sampled original work, shadowing, simulations and case review when routine production no longer supplies enough practice.
Rotate developing staff through diagnosis, knowledge maintenance and difficult cases with appropriate supervision.
Progress responsibility deliberately rather than measuring prompt fluency as a substitute for domain competence.
Keep experts involved in evaluation, recurring exception analysis and updates to the tested task boundary.
Check whether staff can diagnose and recover when the assistant is unavailable or wrong.
Treat these learning arrangements as adaptable recommendations, not universally tested interventions.
Consider a hypothetical support team. AI-assisted retrieval and drafting may reduce production on tested routine enquiries, while agents retain diagnosis, contextual checks and decisions within their delegated authority. Team leads or specialists may inherit quality review, ambiguous escalations, knowledge correction and coaching. Assign those queues before raising service expectations, and segment pilot results by experience rather than transferring the customer-support study's effects or proposed mechanism to the local team.
When is the evidence strong enough to change capacity or roles?
There is enough evidence only after a measured pilot reconciles the complete workload and shows a reasonably stable operating pattern. For a stated period, estimate gross human production effort reduced for observed demand. Then subtract added verification, judgement, exception handling, coordination, learning, monitoring and knowledge-maintenance effort. Adjust for changed demand, service levels, quality, peak queues, rework and work transferred to downstream owners.
Treat any remainder as provisional capacity, not an immediate staffing conclusion. Run the redesigned process long enough to expose side work and credible demand patterns, while tracking quality, overrides, defects, workload distribution, worker feedback, autonomy and task variety. Also test whether people can diagnose and recover when AI is unavailable or wrong. OECD evidence that task changes can affect job quality and workers differently makes an average production measure an incomplete basis for redesign.
Confirm representative cases have been tested against the recorded configuration.
Confirm review and exception queues have named owners, decision rights and workable peak capacity.
Confirm monitoring, feedback and knowledge-maintenance duties appear in the workload ledger.
Confirm the expertise plan preserves difficult-case capability and meaningful supervised practice.
Confirm quality, service and workload results are stable enough to support the proposed decision.
Reassess after changes to the model, prompt, data, controls, workflow, task mix or operating context.
Make no structural change while essential evidence, ownership or recovery capability remains unresolved.
NIST's Core calls for documented responsibilities, feedback mechanisms and ongoing measurement and monitoring of AI risks and impacts. It does not provide a universal pilot duration or capacity threshold. Keep the current staffing model, job descriptions, performance measures and service commitments unchanged when the ledger or expertise plan is incomplete. Employment, privacy, accessibility, discrimination, safety, legal, regulatory and professional determinations belong with qualified and authorised organisational owners.
Questions about task-based role redesign
How do you redesign a job for generative AI?
Map the locally observed job into tasks, baseline demand and effort, and test a fixed AI configuration on representative cases. Measure changes across production, verification, judgement, exceptions and coordination before recomposing duties. Assign each retained, transferred or new duty to an owner with suitable authority and capacity.
What is task-based AI workforce planning?
It is workforce planning grounded in task demand, effort, judgement, consequences, dependencies and tested local effects. It does not apply an occupation-wide exposure or automation estimate directly to staffing. The task evidence is eventually recomposed into complete roles and team relationships.
How should work be divided between people and AI?
Choose among human-performed, AI-assisted human, AI-first with human decision or review, and bounded automation only after testing. For the selected arrangement, define permitted actions, completion authority, review evidence, exception routes, monitoring, stop conditions and fallback. The boundary should change when new evidence warrants it.
How do you measure AI workload redistribution?
Track production, verification, judgement, exception handling and coordination separately for each task. Record whether effort fell, remained, moved to another owner or appeared as new work. Include queue demand and peaks so faster frontline production does not conceal congestion elsewhere.
When do AI time savings become workforce capacity?
Time savings remain provisional until a measured pilot reconciles the complete workload, changed demand, quality, rework, queues, monitoring and expertise needs. The redesigned process must also operate long enough to reveal side work and a credible demand pattern. If that evidence remains incomplete, do not change staffing or service commitments.
References and Sources
This article was researched using the following sources:
We report on how AI actually lands inside a business. Our work starts from named sources, separates what we found from what we think, and uses AI assistance for research and drafting under documented editorial controls. We are not a substitute for individual expert review.