{"slug": "your-ai-agent-didn-t-break-the-rules-one-of-your-rules-was-missing", "title": "Your AI Agent Didn't Break the Rules. One of Your Rules Was Missing.", "summary": "A developer running an autonomous coding agent called Sentinel documented how two independently valid decision paths — the strict AutonomyBudget.admit() gate and a later AutonomyBudget.consider() check added in MR !107 — combined to create an unintended spending loop. Because consider() re-opens a local SKIP without re-checking the provable-trigger or spacing-window conditions, the agent spent a token daily at 15:01 on a file whose answer could never change, costing $0.08 a day indefinitely. The writeup argues this \"two good checks that don't know about each other\" bug class applies to any agent that can decide, retry or spend.", "body_md": "*Anatomy of an autonomy bug: when two valid decision paths create one invalid outcome.*\n\nBuilding an autonomous agent is a bit like raising a very obedient, very literal child with a credit card. You write rules. The child follows them. Perfectly. The problem is that the child follows the rules **you wrote**, not the rules you **meant**.\n\nWe run an autonomous coding agent called Sentinel. It wakes up on a schedule, picks a file in its own codebase, decides whether that file is worth improving, and — if local analysis isn't confident — spends a \"token\" to consult an LLM. Tokens are budgeted: a few per day, hard cap, every spend audited. The whole system is built around one principle: **an agent may only act when it has a reason.** Time alone is not a reason. Boredom is not a reason.\n\nFor weeks this worked beautifully. Then our ops review flagged two lines in the logs that technically shouldn't exist:\n\n`500`. The token was spent anyway. The result: `deferred`. Money gone, nothing decided.\nAnd every day at **15:01 sharp**, Sentinel woke up to ask the same question about a file called `decrypto.js`, got the same answer (SKIP), and went back to sleep. **$0.08 a day, forever, for a question whose answer can never change.**\n\nNobody hacked anything. No rule was violated. Every single line of code did exactly what it was written to do. And yet the system found a backdoor — because **we had written one without noticing.**\n\nIf you're building anything with an agent that can decide, retry, or spend — a scheduler, a budget, a \"reconsider later\" mechanism — this bug class is yours too. It's not about one bad check. It's about **two good checks that don't know about each other**.\n\nSentinel's autonomy path has two doors into the same room:\n\n```\n                        ┌─────────────────────────┐\n   scheduled wake  →    │  WakeGate.evaluate()     │  \"is there a reason to run at all?\"\n                        └───────────┬─────────────┘\n                                    │ autonomy_due (last resort, never ahead of\n                                    │  repo_changed / evidence / scheduled reviews)\n                                    ▼\n                        ┌─────────────────────────┐\n                        │  AutonomyBudget.admit()  │  door #1 — the strict one\n                        │  · needs ≥1 provable     │  (libs/core/autonomy-budget.js:137)\n                        │    trigger               │\n                        │  · daily budget window   │\n                        │  · per-file cooldown     │\n                        └───────────┬─────────────┘\n                                    ▼\n                        ┌─────────────────────────┐\n                        │  grant() → 1 token       │  parked as `pending`, picked up\n                        │  for the chosen target   │  by claim() during evaluation\n                        └───────────┬─────────────┘\n                                    ▼\n                              pipeline runs\n\n   ─ ─ ─ ─ ─ ─ ─ ─ ─ ─ ─ ─ ─ ─ ─ ─ ─ ─ ─ ─ ─ ─ ─ ─ ─\n\n   inside the pipeline, when a file is about to SKIP locally:\n\n                        ┌─────────────────────────┐\n                        │ AutonomyBudget.consider()│  door #2 — added in MR !107\n                        │ \"token re-opens a local  │  (libs/core/autonomy-budget.js:234)\n                        │  SKIP\"                   │\n                        │  · daily limit     ✓     │\n                        │  · module-locked   ✓     │\n                        │  · cycleOwned      ✓     │\n                        │  · provable trigger? ✗   │  ← never checked\n                        │  · spacing window?  ✗    │  ← never checked\n                        └───────────┬─────────────┘\n                                    ▼\n                              spends a token\n```\n\n`admit()` is the paranoid bouncer: it requires at least one deterministic, already-recorded trigger (runtime error, unresolved rollback, governance drift, open opportunity — \"time alone is not a reason\" is literally in `autonomy-trigger.js`'s docstring), enforces the daily window, and picks exactly one target. `consider()` was added later for a legitimate case — *\"this file was about to SKIP, but there's a reason to reconsider\"* — and checks the budget, the lock, and cycle ownership.\n\nIt just never checks **why** it's being asked. No trigger requirement, no spacing window. Every individual check in `consider()` is correct. The door just opens onto the same room as the strict bouncer — and nobody told it the bouncer exists.\n\n**Anomaly 2 — the backdoor (Oct 7, 9:55):** A push to master after merging MR !116 changed `headSha` → `WakeGate` returned `repo_changed` → the pipeline ran. Inside the pipeline, a file hit a local SKIP — and the `consider()` path from MR !107 spent **2 tokens** to reopen it. No trigger required, no window checked, and because tokens were spent at 9:55, the daily schedule shifted. MR !113 watches standalone autonomy wakes; this path isn't one. Sequence verified: correct. Decision branch that allowed it: `consider()` itself — it structurally cannot refuse for lack of a reason, because it never asks for one.\n\n**Anomaly 1 — the eternal reason (every day, 15:01 / 15:17):** Two triggers that can never expire:\n\n`decrypto.js`: `wake-gate.js:85` even skips their stale schedules) — except their penalty records still feed the trigger that wakes a `pressure-engine.js`: Result: same question, same SKIP answer, every day — 2 tokens + ~$0.08/day for a verdict with zero information gain. A signal without staleness isn't a signal; it's a recurring subscription to your own alarm.\n\n**Bonus anomaly — Oct 8, 15:01:** OpenAI returned `500`. Token spent, outcome `deferred`, no retry. Tokens are deliberately non-refundable (documented decision) — but \"spent on a transport error\" is a spend category we never meant to fund.\n\n**What worked:** the `/autonomy` loop itself ran clean — 0 errors, MR !112 held (2× EVOLVE proposals correctly ended `SKIP NOT IMPORTANT`, zero commit churn), token ledger consistent. The bug wasn't chaos. It was a **gap between rules**.\n\nWhatever your agent calls it — tool call, retry, escalation, budget spend — the check that matters is:\n\n```\n// For EVERY function that can authorize the expensive action:\n// does it independently verify the precondition, or does it\n// assume the caller checked?\n\nadmit()    → checks: trigger? window? cooldown? budget? lock?   ✓ all\nconsider() → checks: budget? lock? cycleOwned?                  ✓ — but trigger? window? ✗\n```\n\nConcrete test cases we now write for any new authorization path:\n\nPer the follow-up plan: `consider()` gains the same provable-trigger requirement as `admit()`, plus a **trigger fingerprint** — `AutonomyTrigger.summarize()` serialized at grant time; an identical fingerprint that previously produced a SKIP doesn't count as a reason. Rollback/security triggers get a staleness horizon. Tests pin: identical fingerprint + prior SKIP → refused; changed fingerprint → admitted; zero triggers → `no-trigger`.\n\n*Why \"proposed\" and not \"fixed\": because this article exists precisely to not claim things we haven't measured. The fix is a diff plus a test file. Until CI is green, it's a hypothesis with good posture.*\n\nYour agent doesn't need to be clever to find gaps in your rules. It just needs to be **consistent** — consistency is what turns a missing check into a reliable exploit. The dangerous version of \"the agent did something unexpected\" isn't randomness; it's determinism flowing through a door you forgot you left open.\n\nCount your guardrails all you want. Then check whether every path to the action actually walks past them.\n\n*Evidence: `libs/core/autonomy-budget.js` (`admit()` vs `consider()`), `libs/core/wake-gate.js`, `libs/core/autonomy-trigger.js` · incident data from production autonomy logs, Oct 6–8 · fix pending CI.*\n\nEnd...\n\n🔎 This Isn't Just a Sentinel Problem: 3 Real-World Warning Signs\n\nOur bug is one example of a much broader challenge: autonomous agents can behave exactly as designed and still produce outcomes nobody intended. We are not the first to encounter these questions, and we certainly won't be the last.\n\nHere are three related findings from the wider world of AI agents, each exposing a different crack between what developers intend and what autonomous systems actually do.\n\n1. [When Agents Do Not Stop: Uncovering Infinite Agentic Loops in LLM Agents](https://arxiv.org/abs/2607.01641)\n\nAgents can repeatedly call models, invoke tools, or hand work to other agents when feedback paths lack effective stopping conditions. A loop that looks harmless in isolation can quietly turn into runaway costs and repeated side effects. The researchers examined 6,549 repositories and confirmed 68 loop failures across 47 projects.\n\n2. [Token Budgets: An Empirical Catalog of 63 LLM-Agent Budget-Overrun Incidents](https://arxiv.org/abs/2606.04056)\n\nA budget limit is not enough if retries, concurrent tasks, or delegated agents can bypass it or spend the same budget more than once. This study catalogs reported budget-overrun incidents across 21 orchestration frameworks and examines how stronger enforcement can prevent overspending.\n\n3. [The Authority Benchmark: Do AI Agents Follow the Rules They Are Given?](https://plaw.io/research)\n\nAn agent may have access to a tool without being authorized to use it for every purpose. This benchmark investigates whether agents respect delegated authority under pressure, and tests an external policy check that can reject unauthorized actions before execution. Its results apply to the tested scenarios, not to every agent or deployment.", "url": "https://wpnews.pro/news/your-ai-agent-didn-t-break-the-rules-one-of-your-rules-was-missing", "canonical_source": "https://dev.to/jackymencz/your-ai-agent-didnt-break-the-rules-one-of-your-rules-was-missing-654", "published_at": "2026-10-09 07:09:46+00:00", "updated_at": "2026-10-09 07:21:39.404571+00:00", "lang": "en", "topics": ["ai-agents", "ai-safety", "ai-tools"], "entities": ["Sentinel", "AutonomyBudget", "WakeGate", "decrypto.js"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/your-ai-agent-didn-t-break-the-rules-one-of-your-rules-was-missing", "markdown": "https://wpnews.pro/news/your-ai-agent-didn-t-break-the-rules-one-of-your-rules-was-missing.md", "text": "https://wpnews.pro/news/your-ai-agent-didn-t-break-the-rules-one-of-your-rules-was-missing.txt", "jsonld": "https://wpnews.pro/news/your-ai-agent-didn-t-break-the-rules-one-of-your-rules-was-missing.jsonld"}}