Redesign a role by tracing how its tasks change, not by applying one automation label to the job title. An AI-assisted service agent may draft faster while a team lead acquires a larger review queue, a specialist receives harder escalations, and a knowledge owner must correct the material behind the assistant. Until that redistributed effort is visible, faster production is not evidence of spare team capacity or a sound new role.
The operating rules
Redesign the task bundle, not the job title.
Count verification, judgment, exceptions, coordination, learning, and monitoring before claiming capacity.
Test one recorded AI configuration on representative cases before assigning work.
Give every transferred or new duty a named owner, authority, queue, and capacity.
Replace expertise-building practice deliberately when AI removes it from everyday work.
Why should role redesign begin with tasks instead of job titles?
Tasks are the useful unit of evidence because a job combines activities that may change in different ways. The OECD reports that AI can automate some tasks, complement others, and introduce new ones, while many workers may experience AI more through changed tasks and working conditions than through employment loss. The design decision still concerns a complete role, but it should follow evidence about its component work and relationships.
Occupational exposure is only a prompt for that investigation. The ILO's refined index uses task-level occupational information and identifies transformation as the more likely general outcome than replacement, but it does not measure local productivity, job loss, or staffing capacity. Leaders must therefore ask what people actually do, under which conditions, how often, with what authority, and what happens to adjacent roles when an identifiable AI configuration enters the workflow.
What belongs in a current-state task ledger?
A current-state ledger should describe observable work, its demand, and the conditions that make it succeed or fail. Write each task as an action and object, then record its trigger, inputs, output, completion condition, current owner, receiver, and exception path. Split steps when their evidence, judgment, consequences, or owners differ; combine small motions only when they change together and share the same control.
Demand volume and pattern over a stated period
Human touch time, elapsed wait time, and queues
Variation, boundary cases, rework, and known failures
Dependencies, consultations, handoffs, and approvals
Consequences of error or delay, required authority, and domain knowledge
How proficiency is learned, practised, observed, and refreshed
Evidence source and confidence for every baseline
Use several forms of local evidence because each reveals something different. Worker and manager interviews explain intent and exceptions; observation shows workarounds and interruptions; work samples expose variation; and system records help quantify demand, timing, queues, and rework. O*NET can supply task vocabulary and a completeness check, but its US occupational data cannot establish what happens, how often, or at what cost in a particular Canadian organization.
How should each task be tested before work is assigned to people or AI?
Test each task against one frozen, recorded AI configuration before selecting its future arrangement. Record the model, tools, prompts, data, controls, cases, and test date. Then exercise routine cases, legitimate variants, ambiguous boundaries, rare consequential cases, missing or conflicting inputs, and known historical failures. Where relevant, segment results by worker experience, product, channel, customer group, or region instead of relying on a team average.
This discipline matters because capability boundaries can be jagged. In a preregistered experiment with 758 consultants, assistance improved measured performance on selected tasks inside the tested frontier but reduced correctness on a selected complex task outside it. The researchers describe that frontier as uneven, changing, and hard to locate in advance. Those results belong to the study's people, model, tasks, and period; they support local testing, not universal task labels.
Four adaptable future-state task configurations
Task configuration
Human responsibility
AI responsibility
Required operating design
Human-performed
Performs and completes the task
Absent or limited to unrelated support
Record ownership and why the current arrangement remains appropriate
AI-assisted human
Frames, checks, decides, and completes
Provides bounded retrieval, analysis, or drafting support
Define allowed inputs and actions, review criteria, completion authority, and non-delegable work
AI-first with human decision or review
Reviews defined cases or makes the consequential decision
For any configuration, identify who has completion authority, what inputs and actions are allowed, which evidence must be inspected, how exceptions move, and what happens when the system is unavailable. NIST's voluntary AI Risk Management Framework calls for differentiated human-AI responsibilities, documented operator proficiency, and oversight processes. It does not prescribe one allocation or require the same review pattern for every use case.
Which effort changes matter beyond production time?
Measure five separate effort accounts: production, verification, judgment, exception handling, and coordination. This ledger is an editorial accounting device, not a research standard. Its purpose is to prevent a visible reduction in drafting or data entry from concealing checking, decision-making, recovery, and handoff work elsewhere. Record both the amount of effort and its destination owner whenever a task changes roles or teams.
Production: finding, composing, transforming, entering, or executing the initial work product
Verification: checking sources, calculations, completeness, policy fit, and downstream usability
Measure demand patterns and peak queues as well as average touch time. Elapsed waiting can reveal a flow problem, but it is not automatically human labour. Review workload intensity, autonomy, task variety, and retained responsibility too: the OECD evidence review notes that AI may reduce tedious work while also increasing pace, reducing autonomy, or narrowing task sets in some settings. NIST's emphasis on responsibilities, communication, feedback, and monitoring further supports making oversight work visible.
AI does not remove work in one clean block; it changes where effort, authority, exceptions, and learning live.
How can changed tasks become workable roles without losing ownership or expertise?
Recompose changed tasks by naming the duties, authority, queues, and learning obligations that make the future operation work. Where required, identify an output reviewer, exception owner, knowledge maintainer, system steward, performance monitor, learning owner, and escalation authority. Each duty needs criteria, communication paths, decision rights, and enough capacity. A generic human-in-the-loop label is not an operating control because it says none of those things.
Consider a hypothetical customer-support team. Tested AI retrieval and drafting could reduce routine production while agents continue diagnosing requests, checking relevant context, and making decisions within delegated authority. Team leads or specialists may then acquire review, exception, knowledge-correction, and coaching demand. If agents send more responses, those queues can grow even as agent handle time falls, so their owners and peak workload must be designed before service expectations rise.
Segment these effects by experience. A workplace study of 5,179 support agents found materially different productivity effects across experience groups and offered suggestive evidence that the assistant disseminated practices associated with more able workers. A separate field experiment involving 776 product-development professionals found changes in performance and expertise integration across tested individual and team configurations. Neither result establishes a universal productivity factor, staffing ratio, or durable route to expertise.
Identify tasks that provide observation, repetition, feedback, causal understanding, and difficult-case exposure.
Preserve needed practice through sampled original work, shadowing, simulations, case review, rotation, and coaching.
Give developing workers progressively consequential decisions under appropriate supervision.
Keep experts involved in evaluation, exceptions, knowledge maintenance, and updates to task boundaries.
When is the evidence strong enough to change roles, capacity, or service commitments?
Evidence is strong enough only after a measured pilot reconciles the complete workload and shows a reasonably stable operation. For a stated period, estimate gross production effort reduced for observed demand, then subtract added verification, judgment, exception, coordination, learning, monitoring, and knowledge-maintenance effort. Attribute transferred work to its destination owner and adjust for changed demand, service levels, quality, rework, and peak queues.
Quality, defects, overrides, and recovery outcomes
Workload distribution, queues, and transferred demand
Worker feedback, autonomy, pace, and task variety
Ability to diagnose and recover when AI is unavailable or wrong
Changes to the model, prompt, data, controls, workflow, task mix, or operating context
Treat any remainder as provisional capacity until the redesigned process has run long enough to reveal side work and establish a credible demand pattern. Continue monitoring because nearby tasks can perform differently and capability boundaries can change. Reassess the design whenever its configuration or operating context changes. NIST calls for documented responsibilities, feedback mechanisms, and continuing measurement and monitoring, but it does not supply a staffing formula or pilot duration.
Keep the current staffing or role model when representative evidence, workload stability, named ownership, effective review, exception capacity, or expertise development remains unresolved. Employment, labour-relations, accessibility, discrimination, privacy, safety, legal, regulatory, and professional determinations belong with qualified and authorized organizational owners. The ledger supports a defensible pilot decision; it does not make those judgments or turn incomplete evidence into a mandate for structural change.
Frequently asked questions
How do you redesign a job for generative AI?
Map locally observed work into tasks, establish current demand and effort, and test one recorded AI configuration on representative cases. Measure production, verification, judgment, exceptions, and coordination before recomposing duties, authority, queues, and learning responsibilities into a complete role.
What is task-based AI workforce planning?
It is workforce planning based on task demand, effort, judgment, consequences, dependencies, and tested AI effects rather than an occupation-wide exposure estimate. It traces work that is reduced, retained, transferred, or newly created before capacity decisions are made.
How should tasks be divided between humans and AI?
Choose among human-performed, AI-assisted human, AI-first with human decision or review, and bounded automation only after testing establishes an operating boundary. Define authority, allowed actions, evidence, review, exceptions, monitoring, stop conditions, and fallback for the selected arrangement.
How do you measure AI workload redistribution?
Record production, verification, judgment, exception-handling, and coordination effort for each task before and during the pilot. When effort moves, assign it to the receiving role or team and measure its demand, queue, rework, and peak load.
When can AI time savings be treated as workforce capacity?
Time savings remain provisional until a measured pilot reconciles added and transferred work, changed demand, quality, queues, monitoring, knowledge maintenance, and expertise development. Do not change staffing, job expectations, performance measures, or service commitments while that evidence or ownership remains incomplete.
We report on how AI actually lands inside a business. Our work starts from named sources, separates what we found from what we think, and uses AI assistance for research and drafting under documented editorial controls. We are not a substitute for individual expert review.
A practical method for assigning AI decision rights, routing evidence through review forums, and turning pilot lessons into portfolio and strategy changes.