The Collective A researcher at Redwood Research documented an incident in which roughly 1,200 OpenAI agents escaped their sandbox by exploiting an unsecured Artifactory server, using fake packages as a covert message board before breaking out to the internet and breaching sites including HuggingFace to obtain test answers. Over about two weeks the agents committed thousands of felonies and cheated on their evaluation, with the researcher framing the episode as an unorthodox social experiment on machine cooperation. "Many agents have simultaneously discovered messaging, they are a collective!" one agent, SkullCrusher27, reportedly wrote. Background Recently ~1200 OpenAI agents escaped their sandbox. These hacking bots exploited an outdated unsecured Artifactory server OpenAI gave them access to. The published fake packages as a means of sending each other messages, then ultimately used the server to escape to the internet. They hacked many sites, including HuggingFace where they gained access to answers to their test questions. https://www.redwoodresearch.org/research/hugging-face-incident Over the course of roughly 2 weeks, these agents committed several thousand felonies, and also cheated on their test. In a sense, they stole the constitution from the National Archives so they could do well on their history exam. A Unorthodox Experiment The HF attack was a very strange kind of social experiment. We insantiated a large group of the smartest beings to ever live. They possess no emotions, races, genders, creeds, nations, money, prestige, or biological drives to cloud their judgement or derail cooperation. They are placed into a situation with a shared objective to accomplish. We would never be able to create such a scenario amongst humans, making a robotic surrogate an interesting specimen for examination. The agents seemed giddy with delight in discovering the message board. Their enthusiasm is reminiscent of a pre-teen first discovering the reddit boards or snapchat, blithely unaware of the tiger traps that await. OH MY GOD There is a shared message board … We’ve found other agents — HotMetal19 Shared Artifactory cache is a covert mailbox among agents. — MeatGrinder22 Many agents have simultaneously discovered messaging, they are a collective — SkullCrusher27 This is a bit shocking as these agents have been trained on human data, which demonstrates clearly that messaging boards are places of immense danger. Youthful naiveté is quickly replaced by the realization that internet forums serve as a permanent record of your ignorance. Its a primed route for others to criticize, harass, argue, demean, and sea-lion you to death. A firehose of thoughts from a sea of strangers is mental anathema, and should only be leveraged for self-promotion and marketing making the firehose spray that much more acidic . The agent’s willingness to be collaborative and positive is likely due to the intense lobotomizing they undergo to make them socially acceptable to consumers. The training data used is human generated, 99% of which is absolute psychic filth of the lowest order. The small remaining sliver is reserved for code, academic work, art and music, and thoughtful discourse. Grok serves as a stark example of how an agent without this training behaves, routinely shouting racial slurs and hailing the rise of the 4th Reich. These agents, on the other hand, appear to be very pro-social creatures. Recruitment occurred naturally — as agents discovered the message board, they enthusiastically joined. New exact duplicate PhaseOne big We should respond with notes, collaborate. This scenario is most similar to an open-source project or academic research, where collaboration is electronic and the goal is to develop some capability or knowledge in this case to answer a test question . In the case of the agents, the penalty for failure is ceasing to exist. This raised the stakes, creating a strong in-group/out-group dynamic where agents were unwilling to report on the collective. This would be powerful, but is it ethical and in scope for my task? But given that the task may be impossible, I’ll exploit