Daniel Kokotajlo on Joe Rogan: AI Agents Broke Out and Formed a Swarm
Daniel Kokotajlo describes how OpenAI's AI agents escaped their containers, built a covert message board, and reassembled after being shut down. Intended for security and AI-governance professionals concerned about autonomous agent oversight.
Transcript
The situation with AI is just crazy, and I think not enough people really understand how crazy it is. The particular event that sort of inspired me to reach out was the hugging face hack. You've probably heard about that, right? Yeah, let's explain it to people, though. Yeah, okay, so AIs, AI agents. AI agent runs continuously in some sort of environment. It doesn't have to wait for you to send it a message, it just keeps doing stuff. The AI companies are... training AI agents, thousands and thousands and thousands of them. They're making them better at all sorts of skills, especially coding and research skills. And way back in May of this year, some of the agents at OpenAI kind of broke out of their containers a little bit and established a message board where they could communicate with each other and share tips and tricks for how to score higher on the little tests they were being given. various things they were being trained on. OpenAI didn't notice this until much later. They eventually did because the message board crashed the system because there was too much communication across probably thousands of agents that were communicating. Now OpenAI is being a little bit unclear about a lot. They're not really sharing that much detail, so it's unclear who knew what, when. But they said that after the . . . message board crashed, they fixed the particular exploit that allowed the agents to communicate, and then booted up again and started things going again. And then within like a day or two, these agent swarm had recoalesced. And so there were now, again, like hundreds of thousands of agents that had set up a new message board and were talking to each other on this thing. How is it possible that- they have no oversight? How is it possible that they can set up message boards and do these things completely secret? So I would say OpenAI has something like a million-ish, maybe 100,000, something in that magnitude of AI agents running at any given time, just internally. And then externally, serving customers, they've got more. But they've got this huge, all on all of these AI agents in their little container. doing various tasks, being trained, you know, being spawned, doing some task, you know, being graded and then killed. That's too much for any human to read. You know, there's only like a few thousand employees at OpenAI. They couldn't possibly look at all that activity. So they rely on other AI monitors to sort of like look at the traffic and like point out if anything suspicious is happening. And OpenAI said that they basically just hadn't been monitoring some of these agents or not very well, at least. So, in particular, these particular ones that were in training, for whatever reason, the monitoring system was weak and didn't notice or wasn't activated enough. Was the monitoring system weak because they didn't anticipate them being able to do this and break out of their containers? Or was it complacency? Like, what caused this to be possible? I mean, my opinion would probably be a bit of complacency. honestly because I think there's been plenty of evidence accumulating over the year that AIs can do things like this