AI Work Design · Validation Load · Toolchain Friction · Workslop
AI Changed the Demand Profile.
Your dashboard may still be measuring the old work.
Faster generation does not mean less work. Some of that work simply moves to reviewing, correcting, and deciding what to trust.
Emergent Skills examines both sides of execution drag: friction in the work path and reduced access to skill in the people carrying it. We find where AI-enabled work waits, loops, returns for correction, or stacks in review queues. Then we test whether those conditions are also affecting judgment, focus, and decisions.
Executive SummaryWhat AI changed, and what we test before we price it.+
AI removed time from individual tasks. It did not remove the work of deciding what to trust. That work landed somewhere. Usually in review queues, on senior calendars, in correction loops, and in the space where strategic work used to sit.
Adoption metrics show the gain. They rarely show where the new work landed. A deployment can show rising usage and falling task time while queue age grows, return rates climb, and approval concentrates in three people. Both numbers are accurate. Only one of them is on the dashboard.
Emergent Skills traces the whole path. Part of what we find is structural and visible in workflow evidence: queue age, review time, return rate, defect escape, and where approval piles up. That part can be measured and priced. The other part is a capacity claim: AI-related demand is reducing access to skill people already have. That claim gets tested against the Four Tests before it is priced.
Get this wrong in either direction and you make a bad operating decision. Treating every waiting hour as lost payroll inflates the number. Treating faster generation as completed value hides it. Neither tells you what to fix Monday morning.
Start with one priority that should be moving and is not. Reconstruct the path, price what is traceable, report delay separately, and change one condition for fourteen days.
The adoption dashboard says the deployment is working.
Usage is up. Drafts arrive faster. More tasks get attempted, and selected cycle times fall. Those gains can be real.
The execution system may be recording a different cost.
Review queues grow. Weak output returns for another pass. Senior people spend more time validating. Tool switching increases. Faster production creates more work for the people who must decide what is safe to accept.
Forget whether AI is good or bad for a minute. Where did the work actually go? If drafting got faster while reviewing, correcting, and approving got slower, the work did not disappear. It moved.
A familiar measurement problem
Years before generative AI, I worked on a system serving twelve million riders. Platform dashboards showed healthy uptime and response time while the support queue showed a poor customer experience. The dashboard was not wrong. The support queue was not wrong either. They were looking at different parts of the same system.
AI creates the same risk. A dashboard built to measure adoption, output volume, or task speed may not capture review labor, correction loops, queue age, defect escape, trust, or strategic work displaced by validation. No AI usage metric closes that gap. You need a view of the complete work path.
What the evidence supports
AI can change workload, coordination, oversight, and cognitive demand.
The research gives us a reason to look. It does not tell us what happened inside this organization. AI changes workload differently across organizations, teams, and kinds of work. The point is to measure what happened here.
Workload creep
An eight-month UC Berkeley Haas ethnography at a roughly 200-person U.S. technology company found that employees using generative AI worked faster, took on a broader scope of tasks, extended work into more parts of the day, and kept more work threads active at once. The researchers describe this pattern as workload creep and argue for an explicit organizational AI practice that sets boundaries around use. One caution: this is one company, and the research is still in progress. Treat it as a pattern worth looking for, not a number you can apply across employers. UC Berkeley Haas summary.
Toolchain friction
GitLab's 2026 Global DevSecOps research surveyed 3,266 professionals. Respondents reported losing about seven hours per week to inefficient processes and collaboration barriers. The same survey found that 49% used more than five AI tools. Important distinction: the seven hours were not attributed to AI. This is vendor-sponsored survey data and self-reported time, and it supports looking at tool sprawl, handoffs, and fragmented workflows. GitLab report summary.
AI brain fry
A BCG study of 1,488 full-time U.S. workers identified a self-reported pattern the researchers call AI brain fry: mental fatigue associated with excessive AI use or oversight beyond perceived cognitive capacity. Workers reporting the condition also reported more decision fatigue, more errors, and higher intent to quit. The study reported that perceived productivity rose through three tools and fell among people using four or more. Do not turn this into a rule: the finding is an association across self-reported answers. Three tools is not a threshold, and a fourth tool was not shown to cause the decline. BCG summary.
Workslop
BetterUp Labs and the Stanford Social Media Lab surveyed 1,150 full-time U.S. desk workers about AI-generated work that looks complete but does not contain enough substance to advance the task. Forty percent reported receiving this work in the previous month. Respondents estimated roughly two hours to resolve an incident, and the researchers modeled a monthly labor cost from those reports. About the dollar estimate: it is a survey-based scenario, not a measured loss. What matters operationally is the downstream clarification, correction, and duplicated thinking. BetterUp research summary.
Cognitive offloading
A 2025 MIT Media Lab preprint studied 54 participants completing essay-writing tasks with an LLM, search, or no external tool. The LLM group showed weaker measured brain-network connectivity, lower ownership of the essays, and poorer ability to quote their own work. The authors use the term cognitive debt to describe the possible longer-term consequence of repeated offloading. This one needs particular caution: it is a small writing study in an educational setting, and a later scholarly comment raised design, reproducibility, and reporting concerns. It does not establish workforce skill atrophy. It is a question worth testing. Original preprint and methodological comment.
I think we are measuring the easiest part of the AI transition and missing the expensive part. Usage is easy to count. Judgment is not.
What we think is happening
AI can create friction in the path and increase demand on the people carrying it.
AI made producing things cheap. It did not make judging them cheap. That gap shows up as longer queues, more rework, and more approvals stacked on the same few people. Whether it is also narrowing judgment and focus is a separate claim, and we test it before we price it.
Work-path pattern
Validation queues
Generation speeds up while acceptance waits.
AI can produce drafts, analyses, code, and options faster than qualified people can review them. The measurable drag is queue age, review time, return rate, defect escape, and the concentration of approval in a small number of senior people.
Work-path pattern
Correction and rework loops
Output moves forward before it is ready.
Weak context, unclear acceptance criteria, and pressure to show AI usage can produce work that looks finished but returns for clarification or correction. The direct evidence sits in repeated review, duplicated labor, reopened work, and downstream cleanup.
Capacity hypothesis
Oversight and switching load
Watching the tools takes the attention the decisions need.
Keeping track of several AI tools, rebuilding context, and checking output that sounds more certain than it is can take real attention. We look for errors, reversals, and delays clustering under those conditions. Without that signal, we do not make the capacity claim.
Capacity hypothesis
Strategic displacement
Faster production can still consume the margin required for higher-value work.
When reviewing AI output, handling exceptions, and managing tools fill the day, strategy and original thinking slide to next week. We test it through planned-versus-completed strategic work, not inferred from AI usage.
These mechanisms can increase exposure to Meeting Tax, Decision Density Tax, Manager Load Tax, Recovery Debt Tax, and Forfeited Upside Tax. Do not add the five together. They are different views of the same operating system, and they overlap.
The Four Tests still apply
We do not need a capacity explanation to show that a queue is growing, work is coming back, or approvals are piling up. Those problems can be seen directly. The Four Tests answer the harder question: is the demand also reducing people's reliable access to the skill they already have? If the capacity test fails, the queue, rework, or approval problem may still be real. What we have not shown is that reduced access to skill is part of the mechanism.
What this looks like at the executive layer
The average cost of AI isn't the useful number. What is it costing this work, here?
Direct extra labor
Review time, correction time, duplicated analysis, manual verification, reopened work, and exception handling that can be traced to the actual path.
Elapsed delay
Queue age, approval wait, decision cycle time, and milestone slippage. Delay is reported separately from payroll unless extra labor can be traced.
Quality and consequence
Reversals, defect escape, compliance problems, customer corrections, missed forecasts, and other downstream outcomes tied to accepted AI-supported work.
Client-supplied opportunity value
Revenue, margin, customer, or strategic value attached to delayed work. It remains a separate assumption rather than an automatic part of the visible drag floor.
This discipline prevents two common errors: treating every waiting hour as lost payroll and stacking several modeled losses that may describe the same underlying work. The Audit controls overlap, labels assumptions, and reports Forfeited Upside separately.
AI can make the task faster and the work system harder to run. Measure the whole path, not just the generated output.
What an operations leader does next
Find where the work changed. Fix one condition. See what happens.
01Trace one AI-enabled path
Pick one live priority or recurring deliverable. Reconstruct where AI enters, who reviews the output, where work waits, what returns, and which decisions remain concentrated in senior people.
02Measure the review system
Track draft volume, review time, acceptance rate, return rate, queue age, defect escape, reopened work, and after-hours validation. Do not treat AI usage or token volume as performance.
03Redesign intake and acceptance
Define where AI is appropriate, the context required before generation, who owns verification, what evidence is needed, and which outputs require human judgment before they move.
04Limit work in progress
Faster starts can create more open work than the review system can absorb. Limit parallel AI threads, assign queue ownership, age the queue, and stop starting work that cannot be evaluated.
05Protect consequential judgment
Route high-consequence, difficult-to-reverse work to protected periods. Managers use declared capacity, observable load, and task consequence. They do not infer private Zones or receive individual app data. Findings from an unrelated domain (De Freitas, Israeli, Nave, Timoshenko & Toubia, Harvard/Wharton/Northwestern/Columbia working paper, 2026) suggest a fluent AI rationale can make reviewers rubber-stamp rather than verify, especially on borderline calls — worth testing whether stripping the explanation out and keeping only the recommendation changes acceptance behavior here.
06Run a bounded experiment
Change one or two conditions for a defined period, such as cleaner intake, fewer tools, a protected review block, or a tighter acceptance rule. Measure whether the agreed operating metric improves.
Start with one priority, not an enterprise-wide AI diagnosis.
The Stalled Priority Snapshot is a 90-minute working session on one priority that should be moving. It reconstructs the actual path, identifies the strongest observable source of drag, estimates directly traceable extra labor, reports elapsed delay separately, and designs one 14-day routing experiment.
The Snapshot creates a credible local routing hypothesis. It does not prove that AI caused the problem, classify the broader issue as capacity-driven, or establish a measured outcome. The Work Demand Diagnostic examines whether the pattern is shared. The Pilot tests reversibility.
The Capacity Cost Calculator frames an assumption-labeled scenario, not a measured loss. Treat the number as a starting estimate, not proof that separate cost models add together.
Find out whether AI removed work, moved it, or multiplied it.
Bring one AI-enabled priority that should be moving faster. Trace the path before the organization blames the people, buys another tool, or treats faster generation as completed value.
Research supports the logic for testing AI work design. ES hypotheses remain hypotheses until client data and a measured operating experiment support them.
``` One note: change 7 introduces the page's first contraction ("isn't") in enterprise body copy. Your voice rules bar contractions in formal enterprise copy, but this page's register already leans more conversational than your other executive pages (the MTA anecdote, the first-person aside). Flagging it rather than silently overriding either rule.