{"slug": "update-on-security-at-metr", "title": "Update on Security at METR", "summary": "METR, a nonprofit that evaluates AI models, reported two security incidents in 2026: in March, attackers stole an API key for public models and consumed approximately $600,000 in credits, and in May, attackers probed its infrastructure but failed to access internal data. METR stated that no sensitive information was accessed and that it has increased security investment in response.", "body_md": "*Please note that this post focuses on incidents where external actors attempted to gain unauthorized access to METR’s systems, not AI agents hacking in our evaluations. We have conducted an initial scan of our evaluations, and currently have no evidence of any agents hacking third parties during our evaluations. We will share a more detailed update on this soon.*\n\nMETR had two notable security incidents earlier this year. Following a thorough investigation in collaboration with our security consultants, we believe no sensitive information was accessed in either of these incidents. Nonetheless, we believe it is valuable to share information about how these incidents look in practice and the steps we took as a result.\n\nIn March 2026, attackers stole an API key for inference on public models and consumed a substantial amount of credits. In May 2026, we observed attackers systematically probing our publicly accessible infrastructure, including an unsuccessful attempt to access internal data via an inadvertently exposed endpoint. Although these incidents had limited consequences, we considered them near-misses, and increased our security investment in response.\n\nMETR handles sensitive data as part of our core work, including nonpublic model access and confidential information, so we have invested accordingly in our security posture. Our [historical security controls](https://metr.org/blog/2026-02-17-how-we-protect-confidential-information/#security-measures) are described in more detail on our blog, and contributed to our SOC 2 Type I [certification](https://metr.org/assets/2025-SOC-2-Type-1-Report.pdf).\n\nOur general organizing principles are:\n\nWe work with four major categories of company data. They are, from least to most sensitive:[1](#fn:1)\n\nOur security protocols are designed to keep sensitive data (categories 3 and 4) tightly controlled behind [information barriers](https://metr.org/blog/2026-02-17-how-we-protect-confidential-information/#confidentiality-measures), while minimizing friction for researchers working with less-sensitive data (categories 1 and 2).\n\nTo the best of our knowledge, no data from categories 3 or 4 was accessed as a result of these incidents. However, some sensitive model output data (3.a) was inadvertently accessible in principle, although we believe it was not accessed by the attackers.\n\nIn March 2026, one of our researchers with no sensitive access (e.g., no access to data or credentials in categories 3 and 4) used agents running on a personal EC2 instance intentionally made publicly accessible behind Google authentication. This EC2 instance contained an API key for METR’s general-access (public models) account. The vibe-coded app included a fail-open vulnerability that silently disabled authentication, which led to the system being exposed to the public internet for several days.\n\nFrom our analysis, we suspect that the attacker found the instance by looking through recently-registered websites (e.g. in certificate transparency lists) to find vibe-coded sites with high-signal keywords relating to LLMs or agents, for purposes of harvesting potentially exposed model provider API keys.\n\nUpon finding the deployed system, the attacker prompted an agent directly to reveal its model provider API key, added an SSH key for persistent access, and over the course of three weeks used the stolen credentials to consume a significant amount of API credits on publicly-available models. These credits would have been worth approximately $600,000, although the model developer had granted them to METR for free.\n\nIn early May 2026, METR became the target of a sustained external attack campaign. We were tipped off that we were being targeted by hackers who appeared to be financially motivated and may have been looking to obtain frontier model access. We observed the attackers systematically probing our publicly accessible infrastructure, with heavy use of agents to automate vulnerability discovery, including by credential stuffing authentication providers, attempting OAuth token grants, scanning newly deployed services, and attempting to phish staff.\n\nDuring the same timeframe, we inadvertently exposed a read-only SQL query mechanism via our public transcript viewer. The queries were scoped to public data by default, but a bug could be exploited to access unpublished evaluation data. This dataset was supposed to contain only data from non-sensitive models, i.e. category 2 above. However, some sensitive model data (i.e. category 3 above) was accidentally included in this database.\n\nWe became aware of the issue after an independent security researcher discovered this vulnerability and responsibly disclosed it to us; we took the API offline and paid a bounty.\n\nThe attackers had probed this endpoint in passing as part of their broader campaign, but the evidence shows no indication that they discovered the exploit or accessed any non-public data.[2](#fn:2)\n\nAbove, we summarized changes made based on lessons from specific incidents.\n\nIn addition, we have improved our security infrastructure, protocols and review process, including with our ongoing external cybersecurity partner:\n\nWe will continue to invest in security as our work and the risk landscape evolve.\n\n*The measures described above were accurate as of July 30, 2026 and are subject to change.*\n\n*This post was shared with several of the AI companies we work with ahead of publication, and some minor wording changes were made as a result of feedback.*\n\nThis taxonomy is simplified. There are increasingly important gradations of sensitivity, such as models with some know-your-customer (KYC) or usage policy restrictions, or with some special settings enabled for METR. For this blog post, we are classifying “otherwise public models with ZDR enabled” as category 2, but models with safeguards disabled or with government restrictions on access as category 3. [↩](#fnref:1)\n\nIn order to access the sensitive transcripts, the attacker would have had to discover and exploit the bug, then identify and download the sensitive transcripts without triggering any error messages, which we think is highly unlikely. [↩](#fnref:2)\n\nThe access in question occurred before we became aware of the attack. We took down the vulnerable endpoint during the initial period, but stood it back up again after a few days without noticing the bug. However, the attackers did not attempt to access it again. [↩](#fnref:3)", "url": "https://wpnews.pro/news/update-on-security-at-metr", "canonical_source": "https://metr.org/blog/2026-08-31-security-update/", "published_at": "2026-08-31 07:00:00+00:00", "updated_at": "2026-09-01 00:21:46.381702+00:00", "lang": "en", "topics": ["ai-safety", "ai-policy"], "entities": ["METR"], "alternates": {"html": "https://wpnews.pro/news/update-on-security-at-metr", "markdown": "https://wpnews.pro/news/update-on-security-at-metr.md", "text": "https://wpnews.pro/news/update-on-security-at-metr.txt", "jsonld": "https://wpnews.pro/news/update-on-security-at-metr.jsonld"}}