The judgment gate* assumed a fixed line: automate below it, escalate above it, and the humans above it stay as good as they were. They don't. Judgment at the escalation tier is built by volume at the routine tier, so every case routed to AI is a case your fallback never practises. Endoscopists with a mean of twenty-eight years' experience lost a fifth of their unassisted detection rate within three months of routine AI adoption β degradation in precisely the skill the human was retained to provide. This is not a training problem, and the reskilling industry that dominates the conversation is selling the wrong remedy. It is a routing problem, and routing is a lever you already own.*
The most useful number in the deskilling literature is not about knowledge workers. It is about endoscopists, and the thing to notice is how quickly it happened.
Across four Polish endoscopy centres, researchers compared the three months before routine AI polyp-detection was introduced with the three months after β looking only at procedures performed with the AI switched off. Unassisted adenoma-detection fell from 28.4% to 22.4%, an absolute drop of six percentage points and a relative decline of twenty percent, across 1,443 colonoscopies.
Three details make this worth more than its headline. The nineteen clinicians involved had each performed over two thousand colonoscopies, with a mean of twenty-eight years' experience β these were not trainees leaning on a crutch. The exposure period was three months. And the researchers were not looking for this: it emerged as a sub-study nested inside a trial designed to test whether AI improved detection, which is to say the finding is not the product of anyone setting out to prove that AI degrades clinicians.
The AI did not fail. The humans behind it quietly got worse at the exact thing they were retained to do, in a single quarter, and nobody would have known until a case arrived without the AI.
The study earns its caveats and should carry them. It is observational rather than randomised, so it establishes association, not proof of cause. Total colonoscopy volume rose across the period, which means clinician fatigue is a live alternative explanation β a point made by external commentators at the time. The before and after patient groups differed measurably in sex mix and sedation rates. And the cohort was uniformly senior, which limits what it says about everyone else. The authors flag the observational design themselves and call for further work.
What survives those caveats is still substantial. Multivariable regression identified AI exposure as an independent factor associated with lower detection, controlling for patient sex and age. The authors' own reaction is the part worth carrying forward: they expected experienced clinicians to be insulated, and found the opposite β the decline was consistent across most endoscopists and, counterintuitively, strongest at centres that had started with the highest detection rates. The better the baseline, the more there was to lose.
That is the finding to hold while re-reading any escalation architecture built in the last two years, including the one this publication described when the customer-service reversals began. The judgment gate was framed as a line: automation handles what sits below it, humans handle what sits above, and the design task is placing the line correctly. The framing was right about the line and silent about something underneath it.
It assumed the humans above the gate stay as good as they were on the day you drew it.
Routing is training, whether you intend it or not
Expertise in almost any operational discipline is built in a sequence. Handling volume at the routine tier builds familiarity. Familiarity accumulates into intuition. Intuition, applied to the unfamiliar, is what we call judgment. The escalation tier does not source its competence from somewhere else β it sources it from the routine tier, over time, case by case.
Automate the routine tier and you have not merely removed cost. You have removed the input to the tier above it. The people who staff your escalation queue in three years are being made, or not made, by the cases you route today.
This is the classic paradox of automation, named by Lisanne Bainbridge in 1983: the more reliable the automation, the less practised the operator, and the operator is summoned precisely when the automation fails. What is new is not the mechanism but the scope. Aviation and process control ran this experiment on narrow, well-instrumented tasks. It is now running across customer operations, engineering, clinical work, security response, and finance simultaneously, in organisations that have no equivalent of a flight-hours log.
The evidence is no longer anecdotal. Programmers who leaned on AI to learn an unfamiliar library gained little speed on average and ended up worse at reading and debugging code unaided. Students given an unrestricted AI tutor solved more problems while they had it, then scored below classmates who had never used it once it was withdrawn β not neutral, net-negative. Roughly half of surveyed business leaders already observe deskilling in their organisations, and more than sixty percent expect it to become a material threat within three to five years.
The exposure is largest exactly where automation has gone furthest. When around three-quarters of new code at a major platform is AI-generated and human-approved, the reviewing engineer's ability to catch what the generator got wrong is the entire safety margin β and that ability is maintained by writing code, which is the activity being displaced.
The part that isn't obvious: concentration doesn't help
Here the argument usually gets waved away, and the wave-away deserves a direct answer.
The intuitive rebuttal runs: if AI takes the easy cases, humans spend all their time on hard ones, so they should get better at hard cases. More reps where it counts. Concentration of practice on the demanding tail.
It doesn't work, for two reasons that are worth separating.
The first is feedback density. Routine cases deliver fast, frequent, unambiguous outcomes β you resolve it, you find out immediately whether you were right, you adjust. Escalations deliver slow, sparse, noisy, often unresolvable feedback: the outcome arrives weeks later, tangled with other causes, and frequently nobody ever establishes what the right answer was. Skill is built by the dense signal, not the sparse one. Routing away the routine removes the learning channel and leaves the one that teaches slowly and badly.
The second is calibration, and it is the more consequential. Recognising an unusual case requires knowing what usual looks like. That knowledge is a felt distribution, built by exposure to the whole range β it is what lets an experienced operator say "something about this one is off" before they can articulate why. Strip the routine cases out of a person's working life and you do not sharpen their instinct for the exceptional; you remove the reference distribution against which exceptional is defined. The escalation specialist who only ever sees escalations loses the base rate. They still see hard cases. They lose the ability to see which hard cases are hard in an unfamiliar way.
So concentrating humans on the tail does not produce tail experts. It produces people working without a baseline, on sparse feedback, at the exact point in the system where being wrong costs the most.
The trap: the symptom argues for more of the cause
The failure mode compounds, and the compounding is what makes it dangerous rather than merely unfortunate.
As escalation-tier performance degrades, escalated cases resolve worse and slower. That shows up in the metrics that are measured β handling time, resolution quality, cost per escalated case. The tier looks expensive and underperforming. The natural management response to an expensive, underperforming human tier is to automate more of it, or to staff it more cheaply.
Both responses accelerate the degradation. The remedy for the symptom is a larger dose of the cause, and because the mechanism operates on a multi-year lag while the metrics operate quarterly, the causal link is nearly impossible to see from inside the reporting.
Which points at why this stays invisible. Every metric in a standard operations dashboard describes throughput at the moment of contact β deflection rate, containment, handle time, first-contact resolution. Not one of them measures whether your fallback is still capable of being a fallback. That capability is only tested at the moment it is needed, which is the worst possible moment to discover the answer.
Why the usual remedy is the wrong one
Almost all the writing on this subject arrives from one constituency: reskilling consultancies, training platforms, and L&D vendors, for whom deskilling is a demand signal. Their prescription is predictable and mostly beside the point β more training, more courses, more upskilling programmes.
It is beside the point because the mechanism is not a knowledge deficit. Those endoscopists did not forget how adenoma detection works. They could pass any test you gave them on the theory. What degraded was practised discrimination, and practised discrimination is restored by practice, not by instruction. A course cannot give a support specialist the felt sense of what a routine billing dispute sounds like. Only routine billing disputes can do that.
Which is good news, in a way that the training framing obscures: you do not need to buy anything. The variable that controls this is your routing policy, and you already own it.
The practice budget
One industry has already run this to its conclusion. Aviation watched long-haul pilots lose manual flying proficiency to autopilot, understood that the loss landed precisely where the fallback mattered, and responded by mandating manual flying time. Not more ground school β more hands on the controls. A regulator concluded that practice at the routine tier had to be deliberately allocated, and priced, because it would not survive being left to fall out of the routing.
That is the instrument: a practice budget. A deliberate, costed allocation of automatable work to humans, justified not by the value of the work but by the capability it maintains. Four things make it real rather than a slogan.
Set the allocation as a policy number, not a leftover. Some measurable share of automatable cases routes to humans on purpose. It will look inefficient in every quarterly review, because it is β the efficiency was never the point. Name the line item honestly: this is what maintaining a fallback costs.
Sample across the distribution, not just the interesting part. The whole argument is about the base rate, so the human-routed sample has to include ordinary cases, drawn broadly. Routing humans only the "interesting" automatable cases recreates the problem you are solving.
Measure escalation-tier performance longitudinally, unassisted. Track how your escalation staff perform over time, periodically without AI support, on comparable case difficulty. This is the only metric that detects the degradation before the degradation is discovered by an incident. Nobody currently instruments this, which is why nobody currently knows their exposure.
Protect the pathway, not just the incumbents. The people who staff your escalation tier in five years are junior now, and if the routine tier is fully automated they will never build the intuition the role requires β deskilling's harder-to-reverse sibling, since you cannot restore a skill that was never formed. The practice budget is a hiring-pipeline instrument as much as a competence-maintenance one, which is uncomfortable to defend and cheaper than the alternative.
Bottom line
The judgment gate was described as a line to be placed. It is better understood as a line that moves, and the direction it moves is determined by how much practice you leave on the human side of it. Automate everything below the gate and the gate rises, because competence above it decays without the volume beneath it. Efficiency and fallback capability trade against each other continuously, the trade is invisible in every dashboard currently in use, and the cost lands in a quarter far enough away that nobody connects it to the decision that caused it.
The forward call: within two years, expect the first regulated industry outside aviation β most plausibly clinical or financial β to mandate a minimum unassisted-practice floor for professionals who serve as the human fallback to an automated system. The mechanism is documented, the aviation precedent exists, and the first serious incident traced to a degraded human fallback will make the argument for regulators far more efficiently than any of this will. Organisations that instrumented their escalation tier beforehand will be able to show a number. Everyone else will be explaining why they never measured whether their last line of defence still worked.
The uncomfortable version, for anyone building an escalation architecture right now: you are not choosing where to draw the gate. You are choosing how fast it climbs.
Sources: The endoscopy study is BudzyΕ K, RomaΕczyk M, Kitala D, et al., "Endoscopist deskilling risk after exposure to artificial intelligence in colonoscopy: a multicentre, observational study," The Lancet Gastroenterology & Hepatology 2025;10(10):896β903 (published online 12 August 2025; DOI 10.1016/S2468-1253(25)00133-5; PMID 40816301), verified at primary. Adenoma detection rate in standard non-AI-assisted colonoscopy fell from 28.4% (226 of 795) before to 22.4% (145 of 648) after AI exposure, an absolute difference of β6.0% (95% CI β10.5 to β1.6; p=0.0089); in multivariable logistic regression, AI exposure was an independent factor (OR 0.69, 95% CI 0.53β0.89). Conducted across four Polish centres between 8 September 2021 and 9 March 2022, comparing three months before and after implementation; 19 endoscopists (16 gastroenterologists, 3 general surgeons), each with more than 2,000 prior colonoscopies and a mean of 28 years' experience. The study is a sub-study nested within the ACCEPT trial, itself part of the University of Oslo-led EU OperA project; funded by the European Commission and the Japan Society for the Promotion of Science. Author commentary on the unexpected magnitude and the stronger effect at higher-baseline centres per Medscape (August 2025); the fatigue and patient-mix confounds per Science Media Centre commentary reported by TIME (2026); accompanying Lancet commentary by Omer F. Ahmad (UCL). The finding that programmers who relied on AI to learn a new library gained little speed and performed worse at unaided code reading and debugging, and the AI-tutor study in which students outperformed while assisted then scored below never-assisted classmates after withdrawal, per "Managing the Human Fallback: Skill Investment Under Improving AI and Worker Mobility" (arXiv 2606.29111, 2026) and the studies it reviews. The same paper supplies the survey finding that roughly half of business leaders already observe deskilling and over 60% expect it to become a real threat within three to five years, and the observation that around three-quarters of new code at Google is AI-generated and human-approved, up from about half in late 2025. That paper also frames the gap this memo builds on: existing literature "treats that erosion as a byproduct of automation rather than as something a firm controls." The paradox of automation is Lisanne Bainbridge's ("Ironies of Automation," 1983), applied to AI delegation in arXiv 2602.11865. The L1-familiarity / L2-intuition / L3-judgment pathway framing per Victor Fang, "AI Deskilling In Cybersecurity And Beyond" (Forbes, May 2026), which also cites Anthropic's 2026 labour-market research showing a 16% employment decline among younger workers in AI-exposed roles. The FAA's response to autopilot-driven proficiency loss β mandating additional manual flying time β per aviation-deskilling summaries including SigNoz's "AI Isn't Replacing SREs. It's Deskilling Them." (March 2026), which also names "never-skilling." Cross-references to prior Signal Memo coverage: the judgment gate, the deflection illusion. What is original to this memo: the endogeneity framing (routing policy as skill policy), the base-rate and feedback-density argument for why concentrating humans on hard cases does not produce hard-case expertise, the compounding trap, the practice budget as an operational instrument, and the observation that the dominant remedy is sold by the constituency that benefits from the diagnosis.