Redesign a role by tracing how AI changes its individual tasks, not by applying one automation label to the job title. A service agent may draft faster while a team lead acquires a review queue, a specialist receives harder exceptions and a knowledge owner must correct the material behind the assistant. Until those movements are measured, apparent time savings say little about usable capacity. The sound decision unit is therefore the local task; the eventual design decision is the complete role and its relationships with the wider team.
What to carry into the pilot
Redesign the task bundle, not the job title.
Production time saved is not team capacity until verification, judgement, exceptions, coordination, learning and monitoring are counted.
Test one recorded AI configuration on representative cases before choosing a human-AI task arrangement.
Every transferred or newly created duty needs a named owner, authority, queue and capacity.
If AI removes work that developed expertise, replace that practice deliberately before changing the role.
Why begin role redesign with tasks rather than job titles?
Begin with tasks because one job contains work that may be automated, complemented, retained for people or created by the introduction of AI. A title such as service agent or operations analyst conceals differences in evidence, judgement, consequence and authority. The ILO's refined exposure index uses task-level occupational information and says transformation is generally more likely than replacement. That exposure measure is a prompt to investigate, however; it is not evidence of local productivity, job loss or staffing capacity.
The OECD similarly reports that many workers are more likely to encounter AI through changed tasks and working conditions than through employment loss. Its review describes AI as capable of automating some tasks, complementing others and introducing new ones. Translate that broad finding into a local question: what happens to each observable action under the actual conditions of work? Only after answering it should leaders reassemble duties, authority and working relationships into a coherent role.
What belongs in a current-state task ledger?
A current-state ledger should describe what people actually do, what prompts the work and what proves completion. Write each task as an action and object, then identify its input, output, current owner, receiver and exception path. Split a step when its evidence, judgement, consequence or owner differs from the next step. Combine tiny motions only when they will change together and remain under the same control.
Record demand and its pattern over a stated period, not merely an average day.
Separate human touch time from elapsed waiting time and show queues and peaks.
Capture variation, rework, returns, dependencies, consultations and approvals.
Describe error consequences, delegated authority, required knowledge and how proficiency is developed.
Name the evidence source and confidence for every baseline estimate.
Use observation, worker and manager interviews, work samples and system records together; each reveals omissions in the others. O*NET can supply useful task statements, ratings, work activities and work-context vocabulary, but it describes occupations in the United States rather than work inside an Irish organisation. It cannot establish local frequency, effort or consequence. The ledger fields are an adaptable editorial method, not an O*NET standard, and assumptions should remain visible wherever evidence is thin.
How should a task be tested before work is assigned?
Test each task against one fixed, recorded AI configuration before assigning responsibilities. Record the model, tools, prompts, data, controls, test cases and date so the result remains attributable to something identifiable. Include routine cases, legitimate variants, ambiguous boundaries, consequential cases, missing or conflicting inputs and known historical failures. Where experience level, product, customer group, channel or region could matter, examine the segments instead of accepting one team average.
Define who has completion authority and what the system may receive or do.
Set review evidence, exception routes, stop conditions and fallback arrangements.
Measure quality, effort, failures and downstream usability within the tested boundary.
Record what remains unknown rather than converting uncertainty into a permanent task label.
This caution is supported by a preregistered experiment with 758 consultants. AI assistance improved measured performance on selected tasks inside the tested capability frontier but reduced correctness on a selected complex task outside it. The finding belongs to that study's people, model, tasks and period; it supplies no universal automation rule. The researchers also describe the frontier as uneven, changing and difficult to locate, which makes versioned testing and reassessment more defensible than intuition based on surface similarity.
Four adaptable task configurations and the operating decisions each requires
Task configuration
Human responsibility
AI responsibility
Required operating design
Human-performed
A person performs and completes the task.
None, or support outside the task boundary.
Keep ownership explicit and record why the current arrangement remains appropriate.
AI-assisted human
A person frames the work, checks required evidence and remains the decision maker.
Provides bounded retrieval, drafting, analysis or transformation.
Specify allowed inputs and actions, review criteria, completion authority and prohibited delegation.
AI-first with human decision or review
A named person reviews defined cases or makes the consequential decision.
People monitor performance and resolve exceptions.
Completes a narrowly defined task within approved conditions.
Define boundaries, observability, stop conditions, change control, recovery and fallback.
Which effort changes matter beyond production time?
Measure five separate effort accounts: production, verification, judgement, exception handling and coordination. They are an editorial accounting device rather than a research standard, but separation prevents faster drafting from being mistaken for a lighter operation. Record each account by task, period and destination owner. Include queue size and peak demand when work moves to another role, because a small average can still conceal a review or escalation bottleneck.
Production covers finding, composing, transforming, entering or executing the initial work.
Verification covers sources, calculations, completeness, policy fit and downstream usability.
Judgement covers framing, interpretation, choice, delegated authority and responsibility.
Exception handling covers ambiguity, failures, overrides, corrections, recovery and escalation.
Review the quality of the resulting job as well as its minutes. The OECD evidence review notes that AI may reduce tedious work but can also increase pace, reduce autonomy or concentrate people into a narrower task set, with effects varying by setting and worker. NIST's voluntary AI Risk Management Framework calls for clear responsibilities, communication paths, feedback and monitoring. It does not prescribe these five accounts, but it reinforces the need to make oversight and coordination visible.
AI does not remove work in one clean block; it changes where effort, authority, exceptions and learning live.
How can changed tasks become coherent roles without losing expertise?
Recompose changed tasks by naming every required duty, its authority and the capacity available to perform it. A generic human-in-the-loop label is not an operating design. State what evidence is reviewed, which cases enter the queue, who may decide or override, how work is prioritised and what happens at peak demand. NIST calls for documented roles, differentiated human-AI responsibilities, operator proficiency, oversight, feedback and monitoring, while leaving the staffing arrangement to the organisation.
Protect expertise by identifying which current tasks provide observation, repetition, feedback, causal understanding, supervised practice and exposure to difficult cases. A study of 5,179 customer-support agents found that measured productivity effects differed materially by experience and offered suggestive evidence that the assistant spread practices associated with more able workers. Those findings apply only to the studied organisation and assistant; they do not show that access to plausible answers creates durable expertise.
Name the output reviewer, exception owner and escalation authority.
Assign a knowledge maintainer, system steward and performance monitor.
Give a learning owner responsibility for practice, coaching and proficiency evidence.
Retain sampled original work, difficult-case review, shadowing and simulations where needed.
Use rotation and progressively consequential decisions to develop independent diagnosis.
Consider a hypothetical support team. AI-assisted retrieval and drafting may reduce routine production, while agents retain diagnosis, verification and decisions within their delegated authority. Team leads or specialists may inherit more review, exception, knowledge-correction and coaching demand. A separate field experiment with 776 product-development professionals found changed performance and expertise integration across tested individual and team arrangements, while leaving longer-term expertise questions unresolved. The practical lesson is to test team boundaries, not assume AI removes coordination.
When is the evidence strong enough to change capacity or roles?
Change capacity, roles or service commitments only after a measured pilot reconciles the complete workload. For a stated period, estimate gross production effort reduced for observed demand, then subtract added verification, judgement, exceptions, coordination, learning, monitoring and knowledge maintenance. Adjust for changed demand, service levels, quality, rework, downstream transfers and peak queues. Waiting time may reveal poor flow, but it is not automatically human labour effort.
Run the redesigned process long enough for side work and demand patterns to become visible before calling any remainder provisional capacity. Track quality, overrides, defects, workload distribution, worker feedback, autonomy and task variety. Test whether people can diagnose and recover when the AI is unavailable or wrong. NIST supports documented responsibility, feedback and continuing monitoring; the precise reconciliation and pilot gate remain an operational synthesis rather than a formula prescribed by NIST or the OECD.
Reassess after changes to the model, prompt, data or tools.
Reassess after changes to controls, workflow, task mix or operating context.
Check that review and exception queues work during peaks.
Confirm that transferred duties have capable owners and real capacity.
Keep the expertise plan active and test independent recovery.
Leave the current arrangement in place while material evidence is missing.
Do not change staffing, job descriptions, performance measures or service commitments when representative task evidence, stable workload, named ownership, effective review, exception capacity or expertise development remains unresolved. Use the ledger as a pilot decision tool, not a shortcut to a workforce conclusion. Employment, industrial-relations, accessibility, equality, privacy, safety, legal, regulatory and professional determinations belong with appropriately qualified and authorised organisational owners; this task-redesign method does not decide them.
Questions about task-based role redesign
How do you redesign a job for generative AI?
Map the locally observed job into tasks, baseline demand and effort, and test one recorded AI configuration on representative cases. Measure production, verification, judgement, exception and coordination effort before recomposing duties, authority and queues into a workable role. Treat any capacity estimate as provisional until the redesigned process has been piloted.
What is task-based AI workforce planning?
It is workforce planning grounded in task demand, effort, judgement, consequence, dependencies and tested AI effects. It does not apply an occupation-wide exposure or automation estimate directly to local staffing. The task supplies the evidence, while the complete role and team remain the design outcome.
How should tasks be divided between people and AI?
Choose among human-performed, AI-assisted human, AI-first with human decision or review, and bounded automation arrangements after testing. For each task, define allowed inputs and actions, completion authority, review, monitoring, exception routes, stop conditions and fallback. A human-in-the-loop label alone is not sufficient.
How do you measure AI workload redistribution?
Measure production, verification, judgement, exception handling and coordination separately for each task and stated period. Attribute transferred or new effort to its destination owner, including review, escalation, coaching, monitoring and knowledge maintenance. Examine queues and peaks as well as average time.
When can AI time savings be treated as workforce capacity?
Only after a measured pilot reconciles reduced production with added and transferred work, changed demand, quality, rework, service performance, monitoring and expertise development. The remainder is still provisional until demand and side work are reasonably stable. If material evidence or ownership remains incomplete, keep the current staffing and role design.
References & Sources
This article was researched using the following sources:
We report on how AI actually lands inside a business. Our work starts from named sources, separates what we found from what we think, and uses AI assistance for research and drafting under documented editorial controls. We are not a substitute for individual expert review.