AI Work Design · Validation Load · Toolchain Friction · Workslop
AI Changed the Demand Profile.
Your dashboard may still be measuring the old work.
Faster generation does not guarantee lower total demand. Work can move from producing to prompting, reviewing, correcting, coordinating, and deciding what to trust.
Emergent Skills examines both sides of the resulting execution drag: friction in the work path and reduced access to skill in the people carrying it. We find where AI-enabled work waits, loops, returns for correction, or concentrates in review queues, then test whether those conditions are also affecting judgment, focus, communication, and creativity.
The adoption dashboard says the deployment is working.
Usage is up. Drafts arrive faster. More tasks are attempted. Selected cycle times may fall. Those gains can be real.
The execution system may be recording a different cost.
Review queues grow. Weak output returns for another pass. Senior people spend more time validating. Tool switching increases. Faster production creates more work for the people who must decide what is safe to accept.
Both readings can be accurate. AI can reduce the time required for one task while increasing total work in the surrounding system. The operating question is not whether AI is good or bad. It is where the gain appears, where new demand lands, and whether the full path performs better.
A familiar measurement problem
Years before generative AI, I worked on a system serving twelve million riders. Platform dashboards could show healthy uptime and response time while the support queue showed a poor customer experience. Neither reading was false. Each measured a different part of the system.
AI creates the same risk. A dashboard built to measure adoption, output volume, or task speed may not capture review labor, correction loops, queue age, defect escape, trust, or strategic work displaced by validation. The missing answer is not another AI usage metric. It is a view of the complete work path.
What the evidence supports
AI can change workload, coordination, oversight, and cognitive demand.
The evidence below supports the logic for testing AI work design. It does not prove that every AI deployment creates capacity loss, that the ES Zones are validated categories, or that a specific organization is losing a fixed percentage of payroll.
Workload creep
An eight-month UC Berkeley Haas ethnography at a roughly 200-person U.S. technology company found that employees using generative AI worked faster, took on a broader scope of tasks, extended work into more parts of the day, and kept more work threads active at once. The researchers describe this pattern as workload creep and argue for an explicit organizational AI practice that sets boundaries around use. Evidence boundary: in-progress research from one company. It is useful operating evidence, not a representative estimate for all employers. UC Berkeley Haas summary.
Toolchain friction
GitLab's 2026 Global DevSecOps research surveyed 3,266 professionals. Respondents reported losing about seven hours per week to inefficient processes and collaboration barriers. The same survey found that 49% used more than five AI tools, but the seven-hour figure should not be presented as time caused solely by AI. Evidence boundary: vendor-sponsored survey data and self-reported time. It supports examining tool sprawl, handoffs, and fragmented workflows. GitLab report summary.
AI brain fry
A BCG study of 1,488 full-time U.S. workers identified a self-reported pattern the researchers call AI brain fry: mental fatigue associated with excessive AI use or oversight beyond perceived cognitive capacity. Workers reporting the condition also reported more decision fatigue, more errors, and higher intent to quit. The study reported that perceived productivity rose through three tools and fell among people using four or more. Evidence boundary: these are survey associations, not proof that a fourth tool causes performance decline or that three tools form a universal ceiling. BCG summary.
Workslop
BetterUp Labs and the Stanford Social Media Lab surveyed 1,150 full-time U.S. desk workers about AI-generated work that looks complete but does not contain enough substance to advance the task. Forty percent reported receiving this work in the previous month. Respondents estimated roughly two hours to resolve an incident, and the researchers modeled a monthly labor cost from those reports. Evidence boundary: the dollar estimate is a survey-based scenario, not a measured P&L loss. The useful operating signal is downstream clarification, correction, and duplicated thinking. BetterUp research summary.
Cognitive offloading
A 2025 MIT Media Lab preprint studied 54 participants completing essay-writing tasks with an LLM, search, or no external tool. The LLM group showed weaker measured brain-network connectivity, lower ownership of the essays, and poorer ability to quote their own work. The authors use the term cognitive debt to describe the possible longer-term consequence of repeated offloading. Evidence boundary: this was a small educational writing study, not an enterprise performance study. It does not establish workforce skill atrophy, and a later scholarly comment raised design, reproducibility, and reporting concerns. Treat it as a question worth testing, not a settled business outcome. Original preprint and methodological comment.
The ES interpretation
AI can create friction in the path and increase demand on the people carrying it.
These are related mechanisms, not one universal AI tax. Some can be established directly from workflow evidence. The additional capacity claim must be tested.
Work-path pattern
Validation queues
Generation speeds up while acceptance waits.
AI can create more drafts, analyses, code, and options than qualified reviewers can evaluate. The measurable drag is queue age, review time, return rate, defect escape, and the concentration of approval in a small number of senior people.
Work-path pattern
Correction and rework loops
Output moves forward before it is ready.
Weak context, unclear acceptance criteria, and pressure to show AI usage can produce work that looks finished but returns for clarification or correction. The direct evidence sits in repeated review, duplicated labor, reopened work, and downstream cleanup.
Capacity hypothesis
Oversight and switching load
More parallel work can consume the judgment needed to govern it.
Monitoring several AI tools, reconstructing context, and evaluating confident output can increase cognitive demand. ES tests whether errors, reversals, or delays cluster under these conditions before making the capacity-mediated claim.
Capacity hypothesis
Strategic displacement
Faster production can still consume the margin required for higher-value work.
If validation, exception handling, and tool coordination fill the available space, strategy, customer signals, preparation, and original thinking may keep moving to the next week. This is tested through planned-versus-completed strategic work, not inferred from AI usage.
These mechanisms can increase exposure to Meeting Tax, Decision Density Tax, Manager Load Tax, Recovery Debt Tax, and Forfeited Upside Tax. The Taxes are overlapping cost lenses, not five AI causes and not five buckets to add automatically.
The Four Tests still apply
Queue age, rework, approval concentration, and broken handoffs can establish structural execution drag directly. Baseline shift, load signature, shared conditions, and reversibility govern the additional claim that AI-related demand is materially reducing access to existing skill. A failed test withholds that capacity claim. It does not erase the work-path evidence or identify another cause by itself.
What this looks like at the executive layer
The strongest business case starts with measured operating drag, not a universal AI cost curve.
Direct extra labor
Review time, correction time, duplicated analysis, manual verification, reopened work, and exception handling that can be traced to the actual path.
Elapsed delay
Queue age, approval wait, decision cycle time, and milestone slippage. Delay is reported separately from payroll unless extra labor can be traced.
Quality and consequence
Reversals, defect escape, compliance problems, customer corrections, missed forecasts, and other downstream outcomes tied to accepted AI-supported work.
Client-supplied opportunity value
Revenue, margin, customer, or strategic value attached to delayed work. It remains a separate assumption rather than an automatic part of the visible drag floor.
This discipline prevents two common errors: treating every waiting hour as lost payroll and stacking several modeled losses that may describe the same underlying work. The Audit controls overlap, labels assumptions, and reports Forfeited Upside separately.
AI may make an individual task faster while making the surrounding system harder to govern. Measure the whole path, not just the generated output.
What an operations leader does next
Find the drag. Test the capacity effect. Change the conditions. Measure the result.
01Trace one AI-enabled path
Pick one live priority or recurring deliverable. Reconstruct where AI enters, who reviews the output, where work waits, what returns, and which decisions remain concentrated in senior people.
02Measure the review system
Track draft volume, review time, acceptance rate, return rate, queue age, defect escape, reopened work, and after-hours validation. Do not treat AI usage or token volume as performance.
03Redesign intake and acceptance
Define where AI is appropriate, the context required before generation, who owns verification, what evidence is needed, and which outputs require human judgment before they move.
04Limit work in progress
Faster starts can create more open work than the review system can absorb. Limit parallel AI threads, assign queue ownership, age the queue, and stop starting work that cannot be evaluated.
05Protect consequential judgment
Route high-consequence, difficult-to-reverse work to protected periods. Managers use declared capacity, observable load, and task consequence. They do not infer private Zones or receive individual app data.
06Run a bounded experiment
Change one or two conditions for a defined period, such as cleaner intake, fewer tools, a protected review block, or a tighter acceptance rule. Measure whether the agreed operating metric improves.
Start with one priority, not an enterprise-wide AI diagnosis.
The Stalled Priority Snapshot is a 90-minute working session on one priority that should be moving. It reconstructs the actual path, identifies the strongest observable source of drag, estimates directly traceable extra labor, reports elapsed delay separately, and designs one 14-day routing experiment.
The Snapshot creates a credible local routing hypothesis. It does not prove that AI caused the problem, classify the broader issue as capacity-driven, or establish a measured outcome. The Work Demand Diagnostic examines whether the pattern is shared. The Pilot tests reversibility.
The Capacity Cost Calculator can help frame an assumption-labeled scenario. It is not a measured loss, a universal AI tax, or proof that separate cost models can be added together.
Find out whether AI removed work, moved it, or multiplied it.
Bring one AI-enabled priority that should be moving faster. Trace the path before the organization blames the people, buys another tool, or treats faster generation as completed value.
Research supports the logic for testing AI work design. ES hypotheses remain hypotheses until client data and a measured operating experiment support them.