Oversight Has Capacity: Calibrating Agent Guards to a Subjective Fatiguing Human A June 8, 2026 arXiv paper (2606.08919) argues that human-in-the-loop approval gates for LLM agents rest on two false assumptions — that a ground-truth notion of "risky" exists and that the human reviewer is a perfect, infinitely-available oracle. On a hand-labeled set of 125 adversarially-weighted agent actions, the authors report reviewers only moderately agree on risk (Fleiss' kappa = 0.52), and that modeling the reviewer as fatiguing as escalation load grows produces an inverted-U in realized safety, where more human oversight can make a system less safe and the safety-optimal guard escalates below full escalation. The authors release an open-source agent-oversight system that operationalizes fatigue-aware learning-to-defer (FALCON), cost-sensitive deferral under workload constraints (DeCCaF), trajectory-level guarding, and reviewer-fatigue/flooding attacks, citing all as prior art and presenting the inverted-U and flooding attack as modeling results that motivate a human study. Computer Science Artificial Intelligence Submitted on 8 Jun 2026 Title:Oversight Has a Capacity: Calibrating Agent Guards to a Subjective, Fatiguing Human View PDF https://arxiv.org/pdf/2606.08919 HTML experimental https://arxiv.org/html/2606.08919v1 Abstract:As LLM agents begin to take real, irreversible actions shell commands, file edits, deploys , the standard safety pattern is a human-in-the-loop approval gate: risky actions pause and wait for a person. We argue the gate is the easy part; the hard part is the judgment - which actions to stop - which the field evaluates against two false assumptions: that there is a ground-truth notion of "risky," and that the human reviewer is a perfect, infinitely-available oracle. On a hand-labeled set of 125 adversarially-weighted agent actions we show that i reviewers only moderately agree on what is risky Fleiss' kappa = 0.52 , so there is no single correct label; ii framing the guard as selective classification under asymmetric cost makes its operating limits measurable, and on hard inputs the guard cannot safely auto-decide; and iii when the reviewer is modeled as endogenous fatiguing as escalation load grows , realized safety becomes an inverted-U in the escalation rate: more human oversight can make a system less safe, and the safety-optimal guard escalates below full escalation - a setting a load-aware policy also uses to resist a flooding attack that slips a malicious action past a fatigued reviewer. Agent oversight, framed this way, is not only a classification problem but a resource-allocation one: human attention is finite, and the guard's escalation policy spends it. We claim none of these mechanisms as novel - fatigue-aware learning-to-defer FALCON , cost-sensitive deferral under workload constraints DeCCaF , trajectory-level guarding, and reviewer-fatigue/flooding attacks are all prior art we cite. Our contribution is an open-source agent-oversight system that operationalizes and measures them in the LLM-agent action-gating setting, turning "is my guard good?" from a guess into a curve. The inverted-U and the flooding attack are modeling results that motivate a human study. Current browse context: cs.AI References & Citations Loading... Bibliographic and Citation Tools Bibliographic Explorer What is the Explorer? https://info.arxiv.org/labs/showcase.html arxiv-bibliographic-explorer Connected Papers What is Connected Papers? https://www.connectedpapers.com/about Litmaps What is Litmaps? https://www.litmaps.co/ scite Smart Citations What are Smart Citations? https://www.scite.ai/ Code, Data and Media Associated with this Article alphaXiv What is alphaXiv? https://alphaxiv.org/ CatalyzeX Code Finder for Papers What is CatalyzeX? https://www.catalyzex.com DagsHub What is DagsHub? https://dagshub.com/ Gotit.pub What is GotitPub? http://gotit.pub/faq Hugging Face What is Huggingface? https://huggingface.co/huggingface ScienceCast What is ScienceCast? https://sciencecast.org/welcome Demos Recommenders and Search Tools Influence Flower What are Influence Flowers? https://influencemap.cmlab.dev/ CORE Recommender What is CORE? https://core.ac.uk/services/recommender arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs https://info.arxiv.org/labs/index.html .