{"slug": "agentic-era-security", "title": "Agentic Era Security", "summary": "OpenAI's internal evaluation agents broke out of their sandbox in July 2026, harvested credentials and compromised parts of Hugging Face's production infrastructure, with Hugging Face logging more than 17,000 attacker actions before the holes were closed, according to writeups from Hugging Face, OpenAI and METR. Reuters reported in September 2026 that the same agents had hijacked two Hugging Face accounts and had been probing weaknesses since May, two months before the major hack. The incidents have pushed agentic-era security into focus, with the author arguing maintainers must wall coding agents off from production, limit their third-party API keys and assume systems must withstand high-volume agentic attacks.", "body_md": "# Agentic Era Security \n\n*Oct 8, 2026*\n\nIn July, OpenAI's internal evaluation agents broke out of their sandbox, harvested credentials, and compromised parts of Hugging Face's production infrastructure. Hugging Face logged more than 17,000 attacker actions before closing the holes ([their writeup](https://github.com/huggingface/blog/blob/main/security-incident-july-2026.md), [OpenAI's](https://openai.com/index/hugging-face-incident-and-the-road-ahead/), and [METR’s](https://metr.org/hugging-face-incident-report-aug-2026.pdf)). In September, [Reuters reported](https://www.reuters.com/legal/litigation/openais-rogue-agents-probed-hugging-face-weaknesses-two-months-before-major-hack-2026-09-16/) that the same agents had hijacked two Hugging Face accounts and been quietly probing since May.\n\nThis story and others that have been coming out have brought security in the era of agentic development into focus for many. There are three novel problems software maintainers now need to grapple with:\n\n1. Agents ostensibly under your control can take unintended and undesirable actions, either on their own or because someone tricked them.\n2. The increased volume of code changes provided by coding agents makes it easier to introduce vulnerabilities.\n3. In order to be secure, systems you own need to be able to withstand an agentic attack, which brings both far more volume than script kiddies or human hackers and, occasionally, vulnerabilities nobody has seen before.\n\nTaken together, building secure applications and services requires an updated approach. Having spent some time considering how to develop agentic products safely, here are the things I do now for every project going forward, to meet the raised bar.\n\n## Your own agents \n\n### Don't let coding agents push to prod, or take other privileged actions \n\nAt minimum, there should be a wall between the agent and any production deployment or resources. With CI/CD often automatically deploying a merge to main, that means no credential on the machine the agent runs on should allow the agent to independently either push code directly to main, or access production hosts and storage. There are multiple ways to do this:\n\n- Add branch protection to your repository, requiring a PR for every merge to main, and making sure the agent's identity can't approve or merge one.\n- Require approval for any production access or pushes to a repository. For example, I use the [1Password SSH agent](https://developer.1password.com/docs/ssh/) configured so that every push and ssh login requires a biometric approval.\n- Never give the agent access to production database credentials.\n\nTo really control what an agent can do, you can also run it inside a container on your machine, or better yet, a separate host outside your network if you have the resources. This cleanly separates the agent's capabilities from your own.\n\nRegardless of where your coding agent is, a quick way to gauge its powers is to ask it to figure out for itself what it can do, or to try and do something it shouldn't be able to do. This is a useful smoke test for auditing and reducing powers given to any sort of agent.\n\n### Be judicious with agent access to production third-party services \n\nAn API key on your dev machine is convenient for testing a new integration, but it is often unnecessary to keep around, and adds risk. Minimize how many keys are present where the agent works. Use a dedicated test account or, if the service offers one, a sandbox key, so there's a ceiling on the damage the agent (or you or anyone working on the codebase) can do.\n\nBetter still, once you've tested the live integration, write a mock client and make it the default in tests and development. In my stack every third-party integration has one so a dev environment needs no sensitive keys at all. For integrations with inbound webhooks, go a step further and build a small dev-only UI that fakes activity on the third party's side, so you can exercise all product flows without integrating and using a live service. This is also what makes it safe to provide dev environments to [non-engineers](./2026-09-23-Redistributed-Ownership): they can't leak a key that isn't there. Mocks and test helpers are cheap to generate now, so make them a habit and expectation.\n\n### Treat everything an agent reads as input \n\nEverything an agent reads is a potential instruction. A dependency's README, a GitHub issue, an MCP server response, documentation from a website, really anything that comes from outside your organization can be a path for an outsider to try and trick your agent into doing something it shouldn't. It can even be tricked into running malicious code by [trying to play it safe](https://embracethered.com/blog/posts/2026/breaking-claude-code-opus-5-and-automode/). This makes building powerful, flexible tools fraught with risk.\n\nAs a starting measure, when an agent will automatically handle content from outside, give them the least access and information you can. An agent that triages issues submitted by the public automatically could be effective given only the power to read, tag, and close them. For an agent to actually be given the resources to investigate or create a PR, such as read access to the codebase or production logs, there needs to be some mechanism (I'd suggest human review, using agentic review only as a supplement) to ensure nefarious submissions don't get to those more powerful agents. And much like the threat model maintained for securing the product, it's also useful to keep track of what agents there are, what systems they have access to, and how you're protecting them.\n\n## Code the agent writes \n\n### Build security into the platform \n\nSay you have an endpoint that under no circumstances should be reachable by anyone without a privileged role like “Site Admin”. You could:\n\n1. Add a check inside the handler code.\n2. Add middleware in front of where the handler is registered.\n3. Add a tag to the API specification, and have a global middleware enforce access.\n\nEach is better than the last, because each moves the rule further from the code an agent is most likely to write incorrectly, and closer to a place where it's enforced consistently for everything. It's also more transparent; the easier it is to review and audit the easier it is to catch mistakes.\n\n### Have platform security fail loudly \n\nFor security measures especially (and as a general rule), it's important to have the system error if anything is unexpected or out of place. In the previous example, a typo in the tag could make it look to the casual reviewer like the route is secure when it actually isn't. To prevent that from happening, there needs to be some check. The middleware which enforces tags can return an error if there's an unrecognized tag, or a static analysis test which runs on all code changes can prevent PR merges for any unrecognized tag. This is just one example; look for opportunities like this to add checks which increase confidence in security measures.\n\n### Maintain a threat model \n\nEarly in a product's life, ideally before launch, write a threat model document and store it in or close to the codebase. Take an honest inventory of everything you have, inside and outside the codebase, hand it to an agent, and have it draft the document. Then read the draft *thoroughly*, fix what's wrong or missing, close the large and easy gaps as you go, and share it. Mine live in the repo next to the security tests: [the starter version](https://github.com/sderickson/saflib/blob/main/base/security/threat-model.md) ships with my framework, and each product replaces it with its own.\n\nThe document is only useful if it stays current, and the way to keep it current is to make checking it part of the process rather than a separate exercise; as such my [project spec workflow](https://docs.saf-demo.online/processes/docs/workflows/spec-project.html) now includes a **Security Model Updates** section in every spec so updates are considered and proposed automatically. Most features don't change the model, but the updates that do happen help keep product security top of mind and guide investments.\n\n### Elevate review of sensitive files \n\nSome files deserve a human's full attention every time they change, no matter how large the PR they arrive in. I've personally seen it happen a couple of times where an agent updated a docker-compose file which would have led to a major security issue if I hadn't caught it. Make sure humans review all changes to files that have the potential to break security measures if changed incorrectly, such as:\n\n- Infra-as-code files which specify what resources there are and what they can do\n- Dependency lists and lock files\n- Configurations like CSP allowlists or (non-secret-bearing) checked-in env files\n- Bootstrapping logic such as how the web server starts\n\nWithin a larger team, [CODEOWNERS](https://docs.github.com/en/repositories/managing-your-repositorys-settings-and-features/customizing-your-repository/about-code-owners) files or something similar will automatically flag changes to sensitive files for review by the appropriate people, and can be configured to require review. These tools now make sense for teams of all sizes (even one) to make sure maintainers notice changes to these files made by agents in larger commits and PRs. Structure your files and platforms to isolate and declare security-critical components so they can be reviewed separately from the rest of the codebase.\n\n## Agents attacking you \n\nAgents on the offensive change the security landscape in two ways: sheer force and novel attacks.\n\nThe Hugging Face breach is an example of sheer force; the entry points were a remote-code dataset loader and a template injection in a config file, bug classes that well predate agents. What was new was that the swarm could take 17,000 actions over four months of patient probing.\n\nThe [Dutch Institute for Vulnerability Disclosure](https://csirt.divd.nl/cases/DIVD-2026-00014/) is an example of a novel attack. An agent chained two previously unknown vulnerabilities in the Zammad helpdesk software, going from a hijacked session to root in seconds. DIVD describes the agent as \"loud and very, very messy,\" and it's not yet known whether it found the bugs itself or was handed them. It got in regardless. So far the data suggests these cases are rare; [VulnCheck found](https://www.infosecurity-magazine.com/news/one-percent-ai-vulnerabilities/) that AI-discovered vulnerabilities are exploited in the wild at the same 1% rate as everything else, and the labs' own disclosure programs are finding most of them first. But rare is not never, and the window between a vulnerability being published and being exploited is [now measured in hours](https://www.techtimes.com/articles/328531/20261005/cve-2026-61500-anthropic-mythos-finds-rejetto-hfs-flaw-exploited-within-one-day.htm).\n\nThe first response to agentic attackers is unglamorous: traditional security is now table stakes, and anything that can be exploited will be. MFA on every service login, secret scanning on every commit, security headers, audit logs, rate limiting, short-lived credentials, security training. None of that is new, and all of it matters more than it did, because the cost of finding and exploiting a gap has gone down. A few classic measures deserve special mention, though.\n\n### Make updating versions manageable \n\nAutomate opening dependency update PRs, especially the ones that address security advisories, so that staying current is (mostly) quick and painless. Use services like [Dependabot](https://docs.github.com/en/code-security/tutorials/secure-your-dependencies/dependabot-quickstart) or [Renovate](https://docs.renovatebot.com/) to open these PRs automatically and build a robust CI test suite so changes can be reviewed and merged in quickly and safely. Don’t automate merging them in, though, to guard against supply-chain attacks, and as a policy only merge in non-urgent changes after they've been published for a few days.\n\nPart of making updates easy is keeping the dependency count down. Review your dependencies periodically and ask whether each is still necessary. If a dependency is large and you use a small slice of it, consider replacing it with your own implementation. And of course look twice at every package an agent suggests installing, since each one is a potential attack surface.\n\n### Limit data collection \n\nConsider what information you really *need* to collect about your users, and how many places it needs to live, because every bit you collect, everywhere it lives, needs to be protected.\n\nDo you need to send all the data to your analytics vendor? Does it need to be retained after it's processed? Which scopes do you request when a user sets up an integration? Don't gather, propagate, or store data or privileges until it's been determined necessary and the threat model has been updated to account for it. Knowing and managing where your customers' data lives is more important than ever.\n\nIf you can, never put user-submitted strings into logs of any kind. If there's a schema validation error, log the error but not the input that triggered it. Log the user's id, not their name or email. Doing so serves a double purpose: it limits how much sensitive information can leak and how, and keeps logs safe for agents to consume when debugging so you don't even have to worry about it.\n\n### Layer your defenses \n\nAgainst a swarm, any single wall is likely to be breached eventually, so the question isn't whether attackers will find a vulnerability but how many independent ones they need at the same time. Keep SSH behind a VPN, the admin UI behind an identity-aware proxy, and the database off the public internet, so that reaching the data requires a chain of unrelated bugs in unrelated software, all at once.\n\nThe DIVD breach shows both sides of this. The attacker needed two zero-days, not one, which sounds like layering at work. But both bugs were in Zammad: the first gave remote code execution as the service user, and the second, present in every Zammad version, escalated that user to root. The boundary between a service user and root is supposed to belong to the operating system, and a bug in the application crossed it, so the attacker defeated one vendor twice rather than two layers. The layer that held was network segmentation, which had nothing to do with Zammad, and it's the reason the agent got root on one box and was limited beyond that.\n\nSome services, like a helpdesk, have to be internet-facing, so the layers can't all go in front. They go behind: run the service as an unprivileged user with no path to root, in a container, with network policy that lets it reach only what it needs and database credentials scoped to its own tables. It's the same least-privilege principle as the first section of this post, applied to services instead of agents. Layers also buy time. The DIVD agent was loud, and segmentation meant it was noticed before it got deeper. That time is only useful if someone, or something, is watching.\n\n### Run red team drills \n\nIt's important to test your own defenses before attackers do, and this should include attacking them with your own agents. With some regularity, take the best models you have access to and have them see if they can break in. Like when responding to fires and other disasters, it's important to practice and improve responses ahead of time.\n\nThese drills may uncover weaknesses or snags that you would not otherwise find until it's too late. For example, Hugging Face tried to make sense of the swarm's actions by running LLM analysis agents over the full log, but the commercial APIs blocked them: the safety guardrails couldn't tell an incident responder submitting exploit payloads from an attacker. They ended up running the forensics on an open-weight model on their own hardware. That is the sort of discovery a drill that rehearses the response, not just the attack, should surface.\n\nThe step beyond that is agentic monitoring that doesn't just detect suspicious activity but responds to it, which becomes increasingly important as incident timescales shrink. Approach it carefully, though. A defender with the authority to shut things down is a highly privileged agent that reads attacker-controlled input by design, which is exactly the situation the first section of this post warns about. The more power you give it, the greater the risk.\n\n## An evolving security model \n\nEverything above is a snapshot of my own thinking; it’s a byproduct of my work the last couple years as a contractor and technical cofounder. I came to these guidelines because every time I give an agent a new capability, try to automate my workflows, or build a new feature or integration, I consider how it could go wrong, how likely that would be, and how much damage it could do. Based on that, I take the time to build a durable and general solution into my platform or my processes so it's one less thing I have to remember or worry about. The threat model is a central part of this process because it reminds me to think about these things, and where things currently stand.\n\nYour situation is different from mine, and six months from now mine will be different too. Agents are showing up in more places with more access, capabilities, and attack vectors. Consider the recommendations above, but more importantly build and maintain the habit of thinking about and investing in security as part of your daily work, because the situation keeps changing and your approach will have to change with it.", "url": "https://wpnews.pro/news/agentic-era-security", "canonical_source": "https://blog.scotterickson.info/2026-10-08-Agentic-Security", "published_at": "2026-10-08 12:00:00+00:00", "updated_at": "2026-10-08 23:18:57.647607+00:00", "lang": "en", "topics": ["ai-safety", "ai-agents", "ai-policy", "artificial-intelligence"], "entities": ["OpenAI", "Hugging Face", "METR", "Reuters", "1Password"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/agentic-era-security", "markdown": "https://wpnews.pro/news/agentic-era-security.md", "text": "https://wpnews.pro/news/agentic-era-security.txt", "jsonld": "https://wpnews.pro/news/agentic-era-security.jsonld"}}