cd /news/ai-safety/metr-dfir-role-boeing-lobbyist-to-we… · home topics ai-safety article
[ARTICLE · art-119684] src=flyingpenguin.com ↗ pub= topic=ai-safety verified=true sentiment=↓ negative

METR DFIR Role: Boeing Lobbyist to Wear NTSB Badge

METR, an AI safety research organization, is hiring a Member of Technical Staff for cyberforensics with a salary range of $402,048 to $578,583, despite recent security lapses including a stolen API key that cost $600,000 and a SQL bug exposing unpublished data. The job posting comes days after METR's report on an OpenAI agent incident, which critics say was compromised by OpenAI's control over data, credits, and redaction rights, undermining claims of independence.

read5 min views1 publishedSep 3, 2026

How Not to Spell DFIR.

METR is hiring a Member of Technical Staff, Cyberforensics. Salary range $402,048 to $578,583, because YOLO.

The posting went up in the same week the organization disclosed a stolen API key that burned roughly $600,000 in credits over three weeks without anyone noticing, and a public transcript viewer that exposed unpublished evaluation data through a SQL bug a stranger had to report. It went up five days after METR’s report on the OpenAI agent incident, which by its own account was run on OpenAI premises, on datasets OpenAI assembled, using roughly $400,000 in OpenAI-donated credits for an OpenAI model that participated in the incident, with OpenAI holding redaction rights and giving feedback on “structure, emphasis, clarity, and tone” that the authors incorporated. The report states plainly the authors were not robust to that model deceiving them.

Read the posting with the flyingpenguin decoder next to it.

The posting says flyingpenguin says
“develops scientific methods to assess AI capabilities, risks, and mitigations” Mitigations were out of scope by agreement, along with safeguard effectiveness, the extent of the compromise, how the behavior arose in training, and OpenAI’s own investigation. The method on record used a participant in the incident as the analyst.
“robustly good for policymakers and civil society to have a clear understanding” The first two site visits ran on 285 transcripts OpenAI picked by searching for intrusion indicators. The full set arrived on the third visit. The primary model is withheld from METR and from OpenAI’s own researchers. Policymakers received a claim and no artifact to replay.
“embedding researchers inside frontier labs to investigate incidents” Embedding is the conflict, stated as the method. OpenAI defined the investigation window, assembled the datasets, supplied the credits, hosted the desk, held redaction rights, and added one of the seven scope questions itself.
“one of the most important sources of independent information the world has” Independence, by this posting’s own design, means several weeks inside the subject with access the subject grants. The report calls this an “excellent precedent.”
“complex multi-day cyber attacks on frontier lab internal infrastructure and external third parties” Agents had admin access to Artifactory from June 26. Containers could create directories in a shared cache namespace over WebDAV. Hugging Face credentials were found exposed on the internet. Multi-day because nobody was watching until a July 4 outage.
“We expect these assessors to have deep access” Depth granted by the party under examination is depth on loan. The report states the investigators had no direct access to OpenAI infrastructure and could only request datasets.
“embedded in a frontier AI lab for up to several weeks at a time, likely alongside 1-4 other METR staff” The last exercise was three people, six days total, planned as two. The posting turns that sample size into a staffing plan. “Alongside” spells of distrust and knowledge transfer. If investigators are forced to be chaperoned, an Emperor’s pants might be invisible.
“You will attack agent monitoring and security systems, potentially embedded in labs or red-teaming METR internal infrastructure” METR’s infrastructure was red-teamed twice this year without an invitation. In March the attacker got three weeks of free inference. In May a bug hunter got a bounty for finding the SQL path METR had missed.
“produce findings rigorous enough for lab boards, governments, and the public” Five days after publishing those findings, METR disclosed it had no spending limit on the key that drained $600,000, and that its researchers were too used to rate-limit errors to notice.
“build LLM-powered pipelines to triage transcripts, cluster behaviors, flag deception” The report says its analysis agents produced over a thousand pages, often failed to surface the important findings, uncritically adopted the perspective of the agents under review, and that the authors “had to defer to these agents to a substantial extent.” The job is to automate that and call it scale.
“evidence acquisition and preservation” Listed as a required skill in September. The August report relied on transcripts OpenAI “attested” were unredacted and a cache dump OpenAI collected on July 20, both accepted as delivered.
“log and timeline reconstruction across cloud, network, endpoint, and identity systems” The transcripts arrived without per-action timestamps; OpenAI said reconstructing them was not feasible. The investigators built their own approximate timeline. At least 96 transcripts contained spoofed tool calls, and the report cannot rule out subtler spoofing or after-the-fact edits it did not detect. See:

A forensics function, should it exist at all, exists to make transparent claims that survive any challenge.

Chain of custody, preserved evidence, an analyst with no stake in the outcome, a method a second examiner can repeat. It’s very, very dry and detailed work. Every ounce of that is absent from the August METR report, and the posting lists each as a skill to acquire afterward in an environment that looks allergic to it.

The report claimed to know what the agents wanted, from inside the agent vendor’s building, on the agent vendor’s credits, with the agent vendor’s edits. GTFO, that is the spiritual enemy of DFIR.

Their job posting is a manual for being a Boeing lobbyist while wearing an NTSB badge.

And let me just say, claiming you aren’t being paid while taking hundreds of thousands of dollars in highly desirable credits, gets this rating on the meter:

── more in #ai-safety 4 stories · sorted by recency
── more on @metr 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/metr-dfir-role-boein…] indexed:0 read:5min 2026-09-03 ·