Redesign a role by tracing how each task changes, then rebuilding the complete bundle of work around measured demand, authority and expertise. An AI-assisted service agent might draft faster while a team leader inherits more reviews, a specialist receives harder escalations and a knowledge owner corrects the material behind the assistant. Counting only the agent's saved minutes mistakes local speed for team capacity. A defensible redesign records work that disappears, remains, moves or is newly created before changing duties, targets, staffing or service commitments.
The operating rules
Redesign the task bundle, not the job title.
Production time saved is not capacity until verification, judgement, exceptions, coordination, learning and monitoring are counted.
Test one recorded AI configuration on representative cases before allocating work.
Every transferred or new duty needs a named owner, authority, queue and capacity.
Replace expertise-building practice deliberately when AI removes it.
Why start role redesign with tasks rather than job titles?
Start with tasks because a job title conceals work that AI may automate, support, leave unchanged or create. A role such as claims officer, service adviser or operations coordinator combines retrieval, interpretation, drafting, decisions, hand-offs and exception handling under one label. Those activities differ in evidence, consequence and authority. The design decision eventually concerns a whole role and its relationships, but the evidence must first show what happens to each observable unit of work.
Occupational exposure is a prompt for investigation, not a local capacity estimate. The ILO's task-level index describes potential exposure and says transformation is generally more likely than replacement; it does not measure staffing effects for an individual employer. The OECD likewise reports that AI can automate, complement or introduce tasks within a job. Use that research to reject a single automate-or-retain label, then examine actual Australian workplace conditions, volumes, controls and delegated authority.
What belongs in a useful current-state task ledger?
A useful ledger describes each task as an action and object, then records the conditions needed to understand its workload and control. Write “validate a returned item against the order”, for example, rather than “manage returns”. Split steps when evidence, judgement, consequences or owners differ. Combine small motions only when they would change together and sit under the same control. This keeps the inventory detailed enough to test without turning every click into a separate task.
Trigger, inputs, systems, output, receiver and completion condition
Current owner, consultations, approvals and exception route
Demand volume and pattern over a stated period
Human touch time, elapsed waiting time, queues and peak load
Variation, rework, returns and known failure modes
Consequence of error or delay, required authority and domain knowledge
How proficiency is learned, practised, observed and refreshed
Evidence source and confidence for every baseline
Build the baseline from observation, worker and manager interviews, work samples and system records; no single source substitutes for the others. Keep touch time separate from elapsed waiting time because a day in a queue is not automatically a day of labour. O*NET can supply occupation-linked task statements, ratings and work-context vocabulary, but it describes occupations in the United States. It cannot establish what happens locally, how often, with what effort or consequence.
How should each task be tested before allocating it to people or AI?
Test each task against one frozen, identifiable AI configuration before deciding its future owner. Record the model, tools, prompts, approved data, controls, cases and test date. Include routine high-volume work, valid variants, ambiguous boundaries, rare consequential cases, missing or conflicting inputs and known failures. Where relevant, segment results by experience, product, channel, customer group or region. A team average can conceal a configuration that assists one group while shifting burdens or errors to another.
This discipline matters because apparently similar tasks can behave differently. In a preregistered experiment with 758 consultants, AI assistance improved measured performance on selected tasks inside the tested capability frontier but reduced correctness on a selected complex task outside it. The result belongs to that study's people, model, tasks and period, not to every knowledge-work role. Its uneven, changing frontier supports versioned evidence and reassessment rather than permanent labels.
Four adaptable configurations for a tested task
Task configuration
Human responsibility
AI responsibility
Required operating design
Human-performed
Performs and owns the task
Absent or limited to unrelated support
Record why evidence, suitability or burden favours current ownership
AI-assisted human
Frames the task, inspects required evidence and makes the decision
Provides bounded retrieval, analysis or drafting support
Define allowed inputs and actions, review criteria, authority and prohibited delegation
AI-first with human decision or review
Reviews defined cases or makes the consequential decision
Produces or routes the first output
Name queue ownership, evidence, priority, review scope, exceptions and peak capacity
Bounded automation
Monitors performance and resolves exceptions
Completes a narrow task inside approved conditions
Define boundaries, stop conditions, observability, change control and fallback
Choose among these configurations only after the test exposes an operating boundary. For each choice, name completion authority, permitted inputs and actions, review or monitoring duties, exception routes, stop conditions and fallback arrangements. NIST's voluntary AI Risk Management Framework supports differentiated human-AI responsibilities, operator proficiency and defined oversight, but it does not select a staffing model. “Human in the loop” is not a design until the evidence, authority, queue and available time are explicit.
Which effort changes must be measured beyond production time?
Measure five separate effort accounts: production, verification, judgement, exception handling and coordination. They are an editorial accounting device, not a research standard. Their value is practical: a shorter first draft cannot hide the time spent checking it, resolving failures or moving context across teams. Record demand and effort for each account against the same period, then name the destination owner whenever work moves. Also inspect queues and peak demand, because transferred work can be small on average yet operationally decisive.
Production: finding, composing, transforming, entering or executing the initial work product
Verification: checking evidence, calculations, completeness, policy fit and downstream usability
Judgement: framing, interpreting, choosing, exercising authority and accepting responsibility
Time is not the only design measure. Review workload intensity, autonomy, task variety and retained responsibility alongside throughput. The OECD evidence review notes that AI may reduce tedious work but can also increase pace, reduce autonomy or narrow a worker's task set, with effects varying by person and setting. NIST also calls for clear responsibilities, communication paths, feedback and monitoring. Neither source prescribes these five accounts, but both reinforce the need to look beyond gross production.
AI does not remove work in one clean block; it changes where effort, authority, exceptions and learning live.
How can changed tasks become workable roles without losing expertise?
Recompose changed tasks by assigning named duties, decision rights, communication paths and realistic capacity. Where the operating design requires them, identify an output reviewer, exception owner, knowledge maintainer, system steward, performance monitor, learning owner and escalation authority. Each duty needs criteria and permission to act, not merely a name in a process diagram. NIST calls for documented risk-management roles, differentiated responsibilities, proficiency, oversight, feedback and monitoring, while leaving the actual staffing and review model to the organisation.
Consider a hypothetical customer-support team using AI for retrieval and first drafts. Agents may spend less time producing routine responses while retaining diagnosis, context checking and decisions within their delegated authority. Team leaders or specialists may inherit larger review and escalation queues; other staff may need to correct knowledge, analyse recurring failures and update the workflow. Before increasing service expectations, measure those destination queues and confirm that their owners have authority, competence and capacity.
Segment the results by experience. A study of 5,179 customer-support agents found sharply different measured productivity effects between experience groups and suggestive evidence that its assistant disseminated practices associated with more able workers. Those findings are confined to that organisation, assistant and performance measure. A separate experiment involving 776 product-development professionals found changes in performance and expertise integration across tested individual and team configurations, while leaving longer-term expertise development unresolved. Neither study supplies a universal staffing ratio.
Map which current tasks provide observation, repetition, feedback, causal understanding, supervised practice and exposure to difficult cases. If AI removes that practice, replace it deliberately through sampled original work, shadowing, simulations, case review, rotation, coaching and progressively consequential decisions. These are adaptable design options, not proven universal interventions. Keep experts involved in evaluation, exception analysis and knowledge maintenance, and test whether developing staff can diagnose and recover when the assistant is unavailable or wrong.
When is the evidence strong enough to change capacity or roles?
Change capacity, roles or service commitments only after a measured pilot reconciles the complete workload and the redesigned operation is reasonably stable. For a stated period, estimate production effort reduced at observed demand, then subtract added verification, judgement, exceptions, coordination, learning, monitoring and knowledge maintenance. Adjust for changed demand, service levels, quality, rework, peak queues and work transferred downstream. Treat any remainder as provisional capacity, not an immediate staffing conclusion.
Quality, defects, overrides and recovery outcomes
Demand, rework, waiting queues and peak workload by owner
Workload distribution, autonomy, task variety and worker feedback
Performance when inputs are ambiguous or the AI is unavailable
Developing workers' ability to diagnose difficult cases independently
Changes to the model, prompt, data, controls, workflow or task mix
Operate the pilot long enough for side work and legitimate variation to emerge; a brief quiet period may not reveal exception or coordination demand. NIST calls for ongoing measurement, feedback and monitoring, while OECD evidence shows that task and job-quality effects can differ between workers. Reassess the role whenever the tested configuration or operating context changes, because an earlier result no longer describes the same system. The reconciliation and pilot gate remain an operational synthesis, not a formula prescribed by those sources.
Keep the current staffing or role model when representative evidence, stable workload, named ownership, effective review, exception capacity or expertise development remains unresolved. The ledger is a decision tool, not a shortcut to restructuring. Employment, labour-relations, accessibility, discrimination, privacy, safety, legal, regulatory and professional determinations belong with qualified and authorised organisational owners. Task analysis can show where work and responsibility move; it cannot make those determinations or guarantee that a redesigned process is fair, lawful or safe.
Frequently asked questions
How do you redesign a job for generative AI?
Map the locally observed job into tasks, establish demand and effort, and test one recorded AI configuration on representative cases. Measure production, verification, judgement, exception and coordination changes before recomposing duties, authority and ownership. Use a measured pilot before changing capacity or role expectations.
What is task-based AI workforce planning?
It is workforce planning based on actual task demand, effort, judgement, consequences, dependencies and tested AI effects. It avoids applying an occupation-wide exposure or automation estimate directly to a local role. The final decision still considers the complete role, team and operating context.
How should work be divided between people and AI?
Choose among human-performed, AI-assisted human, AI-first with human decision or review, and bounded automation after representative testing. For every configuration, define permitted actions, completion authority, evidence, review, exceptions, monitoring, stop conditions and fallback. Do not rely on a generic human-in-the-loop label.
How do you measure AI workload redistribution?
Record production, verification, judgement, exception-handling and coordination effort by task over the same stated period. Attribute moved or newly created effort to its destination owner and measure that owner's queue and peak demand. Keep human touch time separate from elapsed waiting time.
When can AI time savings count as workforce capacity?
Only provisionally, after a pilot reconciles complete workload, changed demand, quality, rework, queues, monitoring, transferred duties and expertise development. The process should operate long enough to expose side work and reasonably stable demand. If ownership or representative evidence remains incomplete, do not change staffing or service commitments.
References and sources
This article was researched using the following sources:
We report on how AI actually lands inside a business. Our work starts from named sources, separates what we found from what we think, and uses AI assistance for research and drafting under documented editorial controls. We are not a substitute for individual expert review.
A practical field method for observing real workflows, testing reported bottlenecks and framing evidence-backed AI opportunities without mistaking complaints for proof.
A practical guide to assigning AI standards, funding, delivery, risk and operations, then turning pilot evidence into portfolio and strategy decisions.