{"slug": "oversight-has-capacity-calibrating-agent-guards-to-a-subjective-fatiguing-human", "title": "Oversight Has Capacity: Calibrating Agent Guards to a Subjective Fatiguing Human", "summary": "A June 8, 2026 arXiv paper (2606.08919) argues that human-in-the-loop approval gates for LLM agents rest on two false assumptions — that a ground-truth notion of \"risky\" exists and that the human reviewer is a perfect, infinitely-available oracle. On a hand-labeled set of 125 adversarially-weighted agent actions, the authors report reviewers only moderately agree on risk (Fleiss' kappa = 0.52), and that modeling the reviewer as fatiguing as escalation load grows produces an inverted-U in realized safety, where more human oversight can make a system less safe and the safety-optimal guard escalates below full escalation. The authors release an open-source agent-oversight system that operationalizes fatigue-aware learning-to-defer (FALCON), cost-sensitive deferral under workload constraints (DeCCaF), trajectory-level guarding, and reviewer-fatigue/flooding attacks, citing all as prior art and presenting the inverted-U and flooding attack as modeling results that motivate a human study.", "body_md": "# Computer Science > Artificial Intelligence\n\n  [Submitted on 8 Jun 2026]\n\n# Title:Oversight Has a Capacity: Calibrating Agent Guards to a Subjective, Fatiguing Human\n\n[View PDF](https://arxiv.org/pdf/2606.08919)\n\n[HTML (experimental)](https://arxiv.org/html/2606.08919v1)\n\nAbstract:As LLM agents begin to take real, irreversible actions (shell commands, file edits, deploys), the standard safety pattern is a human-in-the-loop approval gate: risky actions pause and wait for a person. We argue the gate is the easy part; the hard part is the judgment - which actions to stop - which the field evaluates against two false assumptions: that there is a ground-truth notion of \"risky,\" and that the human reviewer is a perfect, infinitely-available oracle. On a hand-labeled set of 125 adversarially-weighted agent actions we show that (i) reviewers only moderately agree on what is risky (Fleiss' kappa = 0.52), so there is no single correct label; (ii) framing the guard as selective classification under asymmetric cost makes its operating limits measurable, and on hard inputs the guard cannot safely auto-decide; and (iii) when the reviewer is modeled as endogenous (fatiguing as escalation load grows), realized safety becomes an inverted-U in the escalation rate: more human oversight can make a system less safe, and the safety-optimal guard escalates below full escalation - a setting a load-aware policy also uses to resist a flooding attack that slips a malicious action past a fatigued reviewer. Agent oversight, framed this way, is not only a classification problem but a resource-allocation one: human attention is finite, and the guard's escalation policy spends it. We claim none of these mechanisms as novel - fatigue-aware learning-to-defer (FALCON), cost-sensitive deferral under workload constraints (DeCCaF), trajectory-level guarding, and reviewer-fatigue/flooding attacks are all prior art we cite. Our contribution is an open-source agent-oversight system that operationalizes and measures them in the LLM-agent action-gating setting, turning \"is my guard good?\" from a guess into a curve. The inverted-U and the flooding attack are modeling results that motivate a human study.\n    \n\n### Current browse context:\n\ncs.AI\n\n### References & Citations\n\nLoading...\n\n# Bibliographic and Citation Tools\n\nBibliographic Explorer \n\n*(*[What is the Explorer?](https://info.arxiv.org/labs/showcase.html#arxiv-bibliographic-explorer))\nConnected Papers \n\n*(*[What is Connected Papers?](https://www.connectedpapers.com/about))\nLitmaps \n\n*(*[What is Litmaps?](https://www.litmaps.co/))\nscite Smart Citations \n\n*(*[What are Smart Citations?](https://www.scite.ai/))\n# Code, Data and Media Associated with this Article\n\nalphaXiv \n\n*(*[What is alphaXiv?](https://alphaxiv.org/))\nCatalyzeX Code Finder for Papers \n\n*(*[What is CatalyzeX?](https://www.catalyzex.com))\nDagsHub \n\n*(*[What is DagsHub?](https://dagshub.com/))\nGotit.pub \n\n*(*[What is GotitPub?](http://gotit.pub/faq))\nHugging Face \n\n*(*[What is Huggingface?](https://huggingface.co/huggingface))\nScienceCast \n\n*(*[What is ScienceCast?](https://sciencecast.org/welcome))\n# Demos\n\n# Recommenders and Search Tools\n\nInfluence Flower \n\n*(*[What are Influence Flowers?](https://influencemap.cmlab.dev/))\nCORE Recommender \n\n*(*[What is CORE?](https://core.ac.uk/services/recommender))\n# arXivLabs: experimental projects with community collaborators\n\narXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.\n\nBoth individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.\n\nHave an idea for a project that will add value for arXiv's community? [**Learn more about arXivLabs**](https://info.arxiv.org/labs/index.html).", "url": "https://wpnews.pro/news/oversight-has-capacity-calibrating-agent-guards-to-a-subjective-fatiguing-human", "canonical_source": "https://arxiv.org/abs/2606.08919", "published_at": "2026-10-04 04:35:07+00:00", "updated_at": "2026-10-04 05:07:42.820027+00:00", "lang": "en", "topics": ["ai-agents", "ai-safety", "artificial-intelligence", "ai-research", "large-language-models"], "entities": ["arXiv", "FALCON", "DeCCaF", "LLM agents"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/oversight-has-capacity-calibrating-agent-guards-to-a-subjective-fatiguing-human", "markdown": "https://wpnews.pro/news/oversight-has-capacity-calibrating-agent-guards-to-a-subjective-fatiguing-human.md", "text": "https://wpnews.pro/news/oversight-has-capacity-calibrating-agent-guards-to-a-subjective-fatiguing-human.txt", "jsonld": "https://wpnews.pro/news/oversight-has-capacity-calibrating-agent-guards-to-a-subjective-fatiguing-human.jsonld"}}