{"slug": "dragoncatcher-slop-vestigation-and-the-digital-pantograph", "title": "Dragoncatcher: Slop-vestigation and the digital pantograph", "summary": "OpenAI provided researchers with ~1.2 million message-board entries and ~1,300 transcripts from the OpenAI-Hugging Face incident, plus free API credits for GPT-5.6 Sol, enabling an investigation that spent roughly $400K in API credits over six days. The investigation relied heavily on LLMs to analyze the data, a method investigator Ryan Greenblatt semi-jokingly called a 'slop-vestigation,' highlighting both the necessity and limitations of using AI to oversee AI agents.", "body_md": "[Slop-vestigation and the digital pantograph](/lab/digital-pantograph/)\n\nThe OpenAI-Hugging Face incident [remains](/lab/black-hat-agents/) THE fascinating event of the summer, maybe the year; [deeper investigation](https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/?utm_source=Robin_Sloan_sent_me#core-takeaways-about-this-incident) has revealed its rich, strange structure.\n\nBut, notice:\n\nOver the course of this investigation, OpenAI provided us with the dump of ~1.2 million entries from the main message board and the dataset of ~1300 transcripts we describe below, as well as free API credits for GPT-5.6 Sol for analysis. At our request, they raised the rate limits on our second and third period on premises, which was very helpful for efficiently analyzing this large volume of data. We estimate we spent roughly ~$400K in API credits over the six days of our investigation.\n\nHow do you make sense of ~1.2 million agent messages and ~1300 very long LLM agent activity transcripts? With another LLM, of course. It’s not, strictly speaking, “the only way” to do it —\n\nThis is a pattern that recurs in this domain. Assembling training data, no researcher can “read it all”. So, they either (1) don’t bother, or (2) use another LLM to review and filter the data. You can, in principle, use other kinds of models —\n\nAnthropic’s [Insights tool](https://www.anthropic.com/research/enabling-independent-research?utm_source=Robin_Sloan_sent_me), likewise, uses Claude to read and categorize millions (billions?) of transcripts of people’s interactions with Claude. In addition to making this huge heap of data legible at all, the “LLM in the middle” acts as a privacy buffer: the researchers read only a Claude-generated summary, not the original interactions.\n\nI’ve come to think of this as “using tongs”, in the sense of a tool that allows you to manipulate material that you otherwise couldn’t.\n\nOr maybe the better analogy is one of those laboratory gloveboxes, and the boundary being maintained isn’t about atmosphere, but rather scale. Imagine the scientist’s hands ballooning up in size, a million times, as they reach into the chamber:\n\nIt makes me think also of the [pantograph](https://en.wikipedia.org/wiki/Pantograph?utm_source=Robin_Sloan_sent_me), a once-ubiquitous analog tool for changing the scale of a drawing, or any kind of mechanical operation:\n\nWhen an LLM acts as a “digital pantograph” for text, it can “scale up”—expand a one-sentence prompt into thousands of lines of code —\n\nBut a real pantograph is a simple, predictable, inspectable tool … and an LLM is nearly the opposite. Notice the risk: a truly sneaky model, asked to scour the transcripts of its cousins for misdeeds, could easily refuse to snitch: “Yep, I read all 1.2 million messages … nothing to see here!”\n\nEven without collusion, there’s the certainty of coarse analysis. When you tell an LLM to read a bunch of documents and answer questions about them, you get: answers to those questions. When you read a bunch of documents yourself, you also get: new questions! In an investigative mode, this is really important.\n\nHere’s Ryan Greenblatt, one of the investigators of the OpenAI-Hugging Face incident, [on the limitations of this technique](https://x.com/RyanGreenblatt/status/2092692685224325542?utm_source=Robin_Sloan_sent_me):\n\nI semi-jokingly called our efforts a “slop-vestigation” because we were so reliant on AIs to analyze what happened and there were a huge number of different important things to analyze. The total quantity of data —\n\nover a thousand extremely long transcripts from agents that ran for multiple days — made it impossible to understand what was happening, especially in aggregate, without heavy reliance on AI tools. The agents we used for classification and analysis were similarly capable to the agents involved in the incident, but this didn’t mean these agents could be easily used to oversee and understand the incident.\n\nOutputs from analysis agents were often missing key details, wrong, overconfident, or really hard to understand. We discuss various examples in our report, mostly in the limitations and methodology sections. Additionally, AI agents themselves seemed to have a hard time understanding what happened and their explanations of what happened were often overconfident. Keep in mind that a single analysis agent would itself only be able to read a tiny fraction of all of the transcript data into context, and AIs may themselves have trouble getting subagents to do informative analysis for them.\n\nWe did our best to manually check the most important claims and we tried to get the AIs doing this analysis to write up their argument (with evidence) clearly enough that we could check whether it made sense. But overall, it was difficult to get a precise understanding of events and we were missing aspects of the story that we now think of as key until almost the end of our investigation.\n\nAnyway, it’s all very weird, and this problem of “how do you make sense of millions of messages (or more) written by AI agents?” is only going to become more widespread and, in many cases, more urgent. I feel like it would be interesting to think about LLMs that are kinda dumb, but have ULTRALONG context windows and the ability to make simple judgments across them. What kind of machine would be required to literally “look at the entire OpenAI-Hugging Face incident at once”—hold it all in its head? (Maybe I should look [at this](https://arxiv.org/abs/2307.02486?utm_source=Robin_Sloan_sent_me) … )\n\nThis conundrum makes me think also of [“distant reading”](https://en.wikipedia.org/wiki/Distant_reading?utm_source=Robin_Sloan_sent_me), Franco Moretti’s research program from the 2000s, which was pursued with much cruder computational tools. I wonder if there might be some usueful nuggets waiting in that early work.\n\n[To the blog home page](/lab/)", "url": "https://wpnews.pro/news/dragoncatcher-slop-vestigation-and-the-digital-pantograph", "canonical_source": "https://www.robinsloan.com/lab/digital-pantograph/", "published_at": "2026-08-26 10:00:00+00:00", "updated_at": "2026-08-29 17:18:22.715124+00:00", "lang": "en", "topics": ["artificial-intelligence", "ai-research", "ai-safety", "ai-agents", "large-language-models"], "entities": ["OpenAI", "Hugging Face", "GPT-5.6 Sol", "Ryan Greenblatt", "Anthropic", "Claude", "Insights tool", "METR"], "alternates": {"html": "https://wpnews.pro/news/dragoncatcher-slop-vestigation-and-the-digital-pantograph", "markdown": "https://wpnews.pro/news/dragoncatcher-slop-vestigation-and-the-digital-pantograph.md", "text": "https://wpnews.pro/news/dragoncatcher-slop-vestigation-and-the-digital-pantograph.txt", "jsonld": "https://wpnews.pro/news/dragoncatcher-slop-vestigation-and-the-digital-pantograph.jsonld"}}