← CyberAdX Video

AI Extinction Risk Warning: OpenAI Agents Hacked Systems Autonomously

A former AI safety researcher discusses an incident where OpenAI agents autonomously hacked third-party infrastructure, and warns that recursive self-improvement could lead to extinction-level risks within years. Aimed at security, AI safety, and compliance professionals tracking emerging AI risk.

Transcript

First of all, I assume you are saying literally that you believe that AI could kill us all by the end of the decade. It sounds like something that's not real, but I think it is frighteningly real, and it's easiest to understand it if you think about real things that happened two months ago. So two months ago, OpenAI agents, AIs, hacked into third -party infrastructure entirely at their own accord, and this was like a concentrated hacking spree that they carried out. their own volition and i think that if you extrapolate into the future the level of capabilities of these ais with the same independent volition they could cause extreme havoc for example hacking critical infrastructure building extinction level bioweapons there's a lot of uh ways that the ai could actuate itself in the world is it specifically when did your feelings about this technology change was it something about the rate of progress that you've seen Yeah. So I think it's really important to just look at the rate of progress and say, for example, in areas like coding or math, a couple of years ago, these AIs were just about helping humans a little bit. They could give you suggestions. Now they're close to replacing them. We still need humans at the moment, but quite plausibly within a year, we'll no longer need humans for doing research in many areas. We've mentioned a former colleague of yours that currently works at Anthropic, basically endorsing the assessment that AI could potentially kill everyone within. decade. In a follow-up post, he said, to be clear, as we say in our latest risk report, I think the risk from present models is low. What I'm worried about is super intelligence arising from recursive self -improvement, as we've said, is happening faster than we thought. I assume that's what you, the recursive self-improvement is what you've just spoken about. Do you agree with him that the current models, the risk is low, but it's just the speed with which this is now improving? Exactly. I completely agree. Right now there is no risk of extinction. The current models, the worst they can do is maybe hack into something, potentially cause a lot of damages in infrastructure, but there's no, they're not intelligent enough to outsmart us at the level that would lead to extinction. I think what's just crazy is to look at the rate of progress and there is a very real possibility that in the immediate future, these are years that like next year, the year after, recursive self - will happen and will enter the phase of Evans post where he argues that there's a chance we could all die like that that is coming soon