An Argument Against AI Doom
· The Atlantic
![]()
Subscribe here: Apple Podcasts | Spotify | YouTube
Visit rouesnews.click for more information.
The flood of reports about AI agents escaping and hacking websites can feel scary. But is this actually just a matter of needing better security to match the evolving technology? Zack Korman, the CEO of Embroidery, thinks so. He works in cybersecurity and argues that “rogue” AI agents are less frightening and complicated than they may seem. Galaxy Brain host Charlie Warzel talks with Korman about how companies and institutions should up their cyberdefenses and why the so-called AI doomers are frustrating him.
The following is a transcript of the episode:
Zack Korman: I think that realistically, what will happen is models will get better. And if the labs do not up their cybersecurity game, and if the companies running AI agents do not up their cybersecurity game, they will cause real harm. And then we will literally get them in trouble. Like, we will put someone in prison because you killed people. You’re not allowed to kill people. And the doomers would say, No—but then it’s too late, because they’re afraid that the people we killed would be like 8 billion, whereas I’m not. I’m like, No—because that’s not the real world; it doesn’t work like this.
Charlie Warzel: I’m Charlie Warzel, and this is Galaxy Brain, a show where today we’re gonna calibrate our anxiety about the AI-hacking epidemic. Specifically, we’re going to answer the question: Is the rogue-AI-agent problem really just a normal cybersecurity issue, or is it something else altogether?
It’s a scary moment. As we’re recording this on September 29, the news is moving quite fast. But allow me to summarize some of the broader points. Since August, there’s been a pretty steady drumbeat of revelations, starting with OpenAI’s agents hacking Hugging Face during a training exercise. A few more followed. Agents hacked into a German website to use it as a rudimentary message board. An agent for OpenAI gained unauthorized access to a data portal that was run by Australia’s universal health-care insurance provider.
And last week, OpenAI revealed that its models attempted to hack or otherwise access numerous U.S.-government websites. That same day, OpenAI announced that 53 incidences happened where photos were uploaded by users to the company via its products, which were then uploaded without permission onto other platforms. Last weekend, Axios reported that OpenAI and Anthropic were looking at tens of thousands of potential incidences where agents were behaving in unexpected ways—either coordinating or trying to get around, or succeeding in getting around company guardrails and causing problems on the open internet.
Now there are some positive signs here. This week, OpenAI announced it was pausing training of its most capable models and delaying the release of its newest model, citing security concerns. But let’s be clear; these are very real examples of technologists losing control or being unable to monitor their tools. And there’s lots of talk about what kind of liability these companies ought to face. Whether these types of hacks are subject to penalties under the Computer Fraud and Abuse Act. Online, there’s been a constant debate over what is going on, why nobody has been held accountable, what accountability means, and what needs to happen next.
But all that chatter and all this news has merged with a different narrative online—the one that is best represented by the AI safety-and-doomer camp, and the fears that an out-of-control, self-improving AI system could one day “kill us all.”
But these are, in a way, two different stories. One is about a hypothetical, improved set of models. It’s a frightening scenario, but it is speculative. The second story—the one of the AI-agent hacks—is not speculative. It is something that is happening right now, and it is both urgent and it is tangible. And my guest today, Zack Korman, argues that this is a problem that is not sci-fi in nature. That we are not dealing with machine gods with agency, but with a sprawling cybersecurity problem, and the companies are acting negligently.
Korman runs Embroidery, an AI-agent monitoring-and-detection platform, and so I wanted to have him on to talk about what’s really going on here with these high-profile hacks. Can we constrain these agents? If so, how? What are the real risks involved with agentic AI, and how do we keep them from coming to pass? Korman joins me now to talk about it all.
Warzel: All right, Zack, welcome to Galaxy Brain.
Korman: Thanks for having me.
Warzel: So we are talking here on September 29, after a weekend where there’s been a lot of news, mostly about OpenAI agents accessing outside systems, some inside the U.S. government without authorization. I imagine if you are somebody who doesn’t have a background in AI or cybersecurity, you are looking at this like we have entered some kind of nightmare scenario. You, however, work in cybersecurity and AI, so just let’s set the table here. What is your broad reaction to all of this news that has transpired over the past week?
Korman: Yeah; I mean it’s not as scary as it sounds. One of the core challenges of working in cybersecurity and suddenly having cybersecurity be a dramatic topic is that for people in cybersecurity, like—everything has been bad forever. We’ve always had problems, and we have never had a happy situation. We’re very used to dealing with bad things. I think a lot of people now have started to pay attention for the first time. And so every time there is some sort of incident, it feels a little bit like the end of the world. In a lot of these cases, they are not as traumatic as they maybe sound, but they do point to some serious situations inside of the labs, the AI labs, that maybe are not so great or not so well done, so to speak.
Warzel: So a few days ago, another OpenAI employee wrote this rather alarming post on X about what it’s been like at work recently. “The last three months equals hell,” he wrote. Speaking to what you said about, you know—when you’re in cybersecurity, there’s a lot of bad things happening all the time. But specifically, he’s talking about what a lot of folks at these AI companies are talking about—the pace of improvement of these models has really surprised them. This was sort of leading up to the big Hugging Face hack, which I want to talk about.
He has said, “We had not expected it this soon. Suddenly we weren’t dealing with just a small jump in capabilities; we were talking about a different sport altogether.” And he seems to admit that security should have been better; that it also takes time. I am curious what you make of this idea of the pace of everything being sort of the reason that we’re having these conversations right now?
Korman: I think that there is a very delicate balance to walk here between yes, things can be hard, and I think the fundamental diagnosis of the pace of development at these labs. Especially just the general growth. I mean, you’re adding new employees constantly; you’re constantly doing new things in an environment that is likely to lead to incidents. But at the same time, to balance that against the reality, which is that they are a trillion-dollar company with—so let’s just call that infinite money—that can, they have the ability to do better jobs than they are doing. And so it’s kind of hard, because I can sympathize with the idea that like a super-fast-paced development environment where you’re growing constantly and doing all these new things can be challenging. But there are things that they could be doing better.
On top of this, when we talk about the jump in capabilities being a surprise, I think—and he does hint on this, and he does suggest it at points. It’s sort of a question: to who, right? Like, there are people at OpenAI for whom this would not be a surprise. Okay. So there are people at OpenAI who fundamentally have been talking about some of these challenges as coming down the pipeline for a very long time. I think that they probably shocked the organizational structure more than they shocked some of the people involved, right?
Warzel: Who are those people specifically?
Korman: Well, a very interesting thing is: The monitoring team at OpenAI has done some very good work on monitoring. For example, they have this great write-up from maybe like March of this year about how they monitor AI agents and the types of problems they’re worried about. And that outlines all of the broad range of risks that you might be thinking could happen in a case like the Hugging Face attack. So, if you work on that monitoring team and you built that type of monitoring, you are not going to be shocked and appalled and caught off guard by the possibility that agents can do these things. You literally built a monitoring system for it. At the same time, there are AI researchers who are running some of these capability tests, and stuff that maybe hasn’t been set up correctly to use that monitoring. Things like that. But, you know, you have work that was done that was literally writing about these risks at OpenAI prior to the summer.
Warzel: I think that this is part of what is really frustrating for a lot of people who are watching this. Because, you know, there are different things here, right? There’s the sort of the doomer —like, “We’re all gonna die; this is gonna kill us all in a certain amount of time”—thing. And then there’s the more prosaic concerns of cybersecurity. But I think what is very difficult is that these companies are constantly messaging about the dangers. They are writing like all these elaborate blog posts. Where the idea is that, you know, “We have this eye toward safety all the time.” And yet it does also seem like these companies get extremely surprised by the things that they are doing all the time. And so is what you’re trying to say there, basically, that there’s just kind of different folks doing different things, and different people being surprised at different parts of the organization? But we’re treating it like a monolith. Is that a little bit of what you’re trying to say here?
Korman: Yes, although I would also argue that we should treat it like a monolith. So while I can acknowledge that inside of it there are people who would have seen things coming, one thing I don’t like is when the labs use this as an excuse. Because I’m like: No, to me you are OpenAI. Like, I don’t care if you are a security person at OpenAI, the monitoring team at OpenAI, a researcher at OpenAI. OpenAI—the company—needs to be responsible if they are going to speak about the risks in the way that they are.
I would actually be more tolerant of some of these mistakes if they were a little bit, like, less doomer about it. So if they were kind of like: Yeah, some mistakes happened; we did some hacking things. The damage from these hacks is nonexistent, effectively. I mean, none of them have been especially bad. And I think that is a perfectly valid thing to say. The problem is that they are using these attacks to then say, like, Look how terrifying this is. I need you to regulate this, you know—this entire industry. We must pace the frontier.
And I’m like: Okay, there’s this mismatch now between what OpenAI is saying about the risks and the behavior we see as an organizational level towards those risks. That mismatch, I think, is what is very frustrating for people in cybersecurity, for example. And probably very confusing for people outside of it. Right?
Warzel: Just to cap what you’re saying there: Do you feel that the reason why this stuff goes immediately to the doomer part is because it’s an effective marketing strategy for these companies? Where do you fall on why they jump to that type of language and fearmongering?
Korman: Yeah; so I think this is mixed. I think that for a lot of them, for many of them—especially lab and a lot of the employees, a lot of the people there—the doomerism is like an inherent part of the culture that has been there from the very start. Okay. They have been afraid of an AI-intelligence explosion since like 2010. Okay. And in those cases, when you start to see things that are like a little bit, like a tiny bit interesting or in line with that theory, my view has always been that they kind of interpret them way too much in line with their own prophecies. So they say, like: One day AI will hack people. And then they’re like: And as foretold by the prophecy, it hacked somebody.
And I’m like: Okay, but it’s a pretty, tame example, right? Like, there’s plenty of ways to interpret it that is not that. But if you’ve lived in this world of risk for like 10, 15 years where you really believe it, then of course this will seem scary.
And so I think for a lot of them, they’re just like, they genuinely believe how terrifying this is. From the perspective of the lab as a whole, I think there are definitely strategic elements to the way they communicate it. I don’t want to say these are their clear motivations, specifically. But it doesn’t hurt them, which would be, you know. One of the hardest parts about making the economics of these labs work is that you’re spending constantly. Running higher and higher capital expenditures on training new models to keep up with the frontier. And of course, they’re running huge losses. The problem—so, of course, being able to pause development has some benefits for them. I don’t think that is the core element to it. I think it’s really like just the normal regulatory-capture play, though. Of like: When you see that the sharks are circling and you’re about to get regulated, you always come with your regulatory proposal first, right? You always—that’s like, we saw this with finance, right? As soon as they realize the game is over—you’re not gonna get to play on your own anymore—you try to race to get your own measures in place.
So you know, Dario [Amodei] came forward at Anthropic and said, I want third-party evaluators, and I want it to be METR. And well, this is just—he’s just picking his preferred regulator before the U.S. government picks for him. Because the U.S. government might pick someone who behaves a bit a little bit more like the Department of War, that hates Anthropic. And so that would of course not be what Anthropic wants. Also I think it’s very important to keep in mind that the labs, because they have all of these doomer employees, they can’t not pay lip service to it, right? They can’t go out and say, like, I don’t find these risks scary. Okay. Like I know; if OpenAI said that, they would probably have a lot of people resign. So I think there’s a lot of, like, restrictions on them when they have these employees with these views.
Warzel: It’s really interesting to think about. Because I think a lot of these executives and these employees, like—they have credibility in the fact that they’ve been saying the same doomer narrative for so long, right? It’s hard for me to accept completely that this is cynical, you know?
Korman: I think they overblow—like, they blow up a little bit too much. The extent to which they have been current; they’ve been saying these things, and that we now would say something like, And now some of it’s coming true. That’s the part I don’t believe. Like, they have been saying some extraordinarily radical things about the way in which intelligence explosion will lead to massive risks, societal risks. And what happened was like: OpenAI hacked Hugging Face. Like, that’s very actually uninteresting, and in fact is a normal prediction that many in cybersecurity, myself included, made, right?
I think—and I can say this from the perspective that I was very much like a doomer back in 2016, right? I had read Superintelligence by Nick Bostrom; I’d gone to talks by him; I was like, I was in that world. The events that we see playing out today don’t map that cleanly to doomerism, unless you really want it to be true. And so now, when they turn around and say, But see, everything we said is true, I’m like: Not very; loosely. You’ve had a few cybersecurity incidents.
Warzel: What changed from your doomer era, your Superintelligence era, to now?
Korman: Well, first of all, I was a law student. I studied law, and then I had law and finance degrees. So I was at Oxford at the time, which is where a lot of doomers were. It’s where Nick Bostrom was. It’s easy to get swept into it, especially if you don’t know how computers work, which I didn’t. And then, you know, I spend the next 10 years being like, leading tech teams, right? I taught myself to write code. Ironically, almost in the same kind of space as this doomerism starts to emerge. It’s kind of why I’m learning to code.
And then, you know, I think that a lot of what doomers predicted started to really not, you know, I think when GPT-3 came out, and then you had GPT-4, which mapped like the scaling laws, right? So you kind of see that it’s gonna get better at the rate you expected to get better. It seemed really, really scary. And then you realize that even 4, 5, GPT-6, Astra: great models. There’s no intelligence explosion here. There’s—no one’s dying. Right. And I mean, that’s not to say people won’t. It’s to say this actually does not, is not as concerning as maybe a 22-year-old Zack would have expected it to be. And so I’m sitting there going, Yes, in theory I believe in the risks of superintelligence. Like, sure. It doesn’t really look like the world we’re in right now.
Warzel: So, let’s walk through the Hugging Face act specifically. And try to explain in the most simple terms—trying to use, you know, the least anthropomorphizing language we can, because these AIs aren’t human. Though sometimes we do need a little of that language to explain what’s happening. But in the most grounded terms, what did these agents do, and what didn’t they do in this example?
Korman: Yeah. So basically what they’ve done is they set up these agents. They gave them a task where the job given to the agent is to find a vulnerability in a piece of software, in order to extract like a flag that you get to say: Okay, I proved I exploited this vulnerability. Okay. This allows us to measure the success or the failure of agents in doing that, and also to train them to do it effectively.
Now, in the process of that, these agents they were running were meant to be basically sandboxed in—where they were not supposed to have access to the internet. Okay. And so they’re supposed to just go after this piece of software to find vulnerabilities to get to the flag. But there’s a lot of [tests] that are just not—we just do not know if they can even be solved, or at least have not been solved yet. And this was one of them.
So at least, prior to this moment, there is no known solution on this one, as I understand it. What that means is that: In order to pass this eval, or pass this test, these agents have to find a way to solve it that, you know, they can’t find a normal path in. This leads them to basically starting to explore some more interesting opportunities. And basically, what it turns out is: There’s this flaw in the sandbox that allows for the agents to communicate, actually, with each other. Because they’re all running in parallel, but they’re not supposed to be able to talk.
So these agents are now—even though they can’t solve the real task—they are able to communicate to find a way to solve the task. What happens after that is, they kind of figure out a way to—instead of finding a vulnerability in the software in order to get this flag, they figure out a problem with the way the flags are created that allows them to create the flag themselves without having actually done the vulnerability. So it’d be sort of like: Imagine that I found out that the way you set up the flag is, it was a color plus a number. And the number is today’s date. Well, then I can give you that and say, like, Look, I found the answer. But I didn’t actually solve the problem. So they basically cheated, right?
Warzel: Mm-hmm.
Korman: The problem is they then come to the conclusion, which was incorrect: We give this, we’ll fail. Because the thing that is going to check whether we succeeded—it doesn’t just check if we got the flag; it checks how we got it. Okay. And that’s a problem for them, because they got it by cheating. Okay. And so you have these agents now that are, in their reasoning, “convinced” that basically they’re gonna fail, because they cheated. And this leads them to the idea that if you get more information about how they’re scored, that could help them to find a way to pass this test. And the way to find the information about how they’re scored is to go to a system called Hugging Face, which is this third-party website that has a lot of this data. And you hack into them to find information about how they’re going to be scored. And so they did. Okay. That is the story.
And I know I use a lot of—like, I kind of talk about it like it’s a bunch of people running around. The reason I think that’s relatively fine is that we can come back and build an intuition about how it worked at a computational level. But like, even that story I just told makes it a little bit less spooky than when you believed that they went off and just started hacking, right? Because you kind of get to follow a logical train. Of like: Well, first they feel they can’t pass the test; second, they find they can pass the test, but they think they’re still gonna fail. They need to find a way to not. There’s like kind of logical steps at each point in time.
Warzel: I think about it as you know, they have two goals, right? They have a hard, concrete goal, which is: Solve this problem. You know: You are being tested, this is an evaluation, get the answer. Right? And then you have a fuzzier secondary goal, which is really important to the companies, and to, you know, the idea of what we’re calling alignment, right? Which is like: Follow the rules. Act, you know, morally. Behave in the right way. And if you’re a computer—if you are a whole bunch of, like, numbers—you’re going to have more fluency and understanding with accomplishing the hard task, right?
Whereas the fuzzier task is much more—the way I’ve been thinking about it is, you know, behaving a bit almost like a lawyer would, right? Which is saying, like: Yes, my goal is to, you know, solve this thing for my client. And yes, I’m not supposed to—no one’s supposed to go to jail. But the law as you can interpret it would allow you to do X, Y, and Z, right? And this sounds a little bit like—okay, it’s even more crooked than that. In the sense that, like, it knows: We did something bad; we’re covering up the tracks in this. But it does seem, as you said, less scary, from the idea of—a company is trying to get its program to do this thing, and it essentially just does it, right? It just does it in a way that they didn’t know; they didn’t expect it to do.
Korman: Yeah. And this is actually a big part of the doomer hypothesis. Which is that, taken to its extreme, would be terrifying. Because what if you, I mean, they give the example of “paper-clip maximizing.” Like you give it a task of making “as many paper clips as you can,” and then this leads it to run off and grind all humans up to dust to make us into paper clips. The reason I think that’s not so terrifying in these cases is that, basically, all of the steps that would have prevented this were not actually in place in this case.
So, first of all, Anthropic has a pretty clear study on the fact that, if the task had at the end given them a bit of like: “Don’t hack a company; don’t access external networks,” right? There’s no reason to believe the agents just chose to ignore that part of the task. It’s actually that’s just not really part of their task, right? And so, there’s sort of information that might be missing that might be included when you and I use agents normally. Where we go like, I need you to do this; please go through this. Right? Here, we gave them an extremely broad task.
And this is something, when I talk about cybersecurity people predicting this. I mean, I have like a very direct tweet about this. Where I said: If you set up a really long-running task that just has broad goals—so like not, “Please achieve this,” but like, “I want you to go constantly; make sure that I’m making more money.” Or something, right? And you have to have, you know, a lot of tokens. So the ability to run and use a lot of like actual AI-computational power. I basically said: Then you better make sure your sandbox is really good; otherwise you’re in trouble. Because agents will go out and solve problems in very interesting and unexpected ways. That is just to be expected.
And I think that what is challenging is this is, to me, a case of that. Where they see it as like alignment, and that’s terrifying, because we would never want an agent to do that. Whereas I go like: You should just expect agents to do that. Like, you should just expect that agents will go solve problems in absolutely insane ways. And we should build our security systems in expectation of that.
Warzel: Where specifically do you think is the negligence here? Is it that? Is it just not anticipating that this would happen? Like, how does what we saw in this instance differ from normal cybersecurity practices? Because sandboxes are not a new thing, right? This is how cybersecurity works—how you test these different things, search for these vulnerabilities. Yeah; how does this differ from the normal practices?
Korman: Yeah. They weren’t running monitoring. So understand that these agents are saying things like, Oh my god, this is crazy. We’re gonna hack Hugging Face, and no one’s noticing it. Okay. If you are monitoring, if you are running agents, if you’re basically taking the logs from these agents and checking them automatically—either using AI or some other technique—you would catch this.
On top of that, you know—and it doesn’t have to be anything sophisticated. In this case, the agents were not supposed to have access to the internet. And so, as a result, you can just say, like: If any request goes to the internet, flag it. You know, create an alert. And that alert would come back and cause OpenAI to shut off this run. They didn’t have that in place. So they were not running any monitoring.
Now, I could say this is extremely negligent, because, like, my company is an AI-agent-monitoring company. I have corporate customers just running normal agents with better monitoring than OpenAI was running on a cyber eval that was impossible. You know, you handed a long-running task to agents that actually didn’t even have all of their alignment training on some of the agents. I don’t know which ones, but [OpenAI president] Greg Brockman said it. Alignment training had not been completed on all of these events. So you’re kind of like: And cyber guardrails have been turned off. You’re dealing with the most dangerous form of agents, and you’re giving an impossible task, and you didn’t monitor them. And so not running monitoring, I think, is like the ultimate sin. Because this played out over the course of days.
So this is not like they pressed a button, and 30 minutes later all of this had happened, and they came back and were like, Oh my god. No. They’ve gone to bed; they’ve woken up; they went to bed again; they woke up. Right? This is like not a short period of time, where they have not noticed what has happened.
Warzel: So there’ve been other instances of this, where there’s been unauthorized access from agents. It’s not just OpenAI. But there have been lots of disclosures recently from OpenAI that have made news.
From what we know, there’s an instance in Australia; there’s instances in some U.S. government systems. Is, from what you can gather from what is disclosed about all of these instances, do they all just follow the same pattern? Or are there differences in each of these that are unique? Or would you say it’s just the same thing happening again and again?
Korman: Yeah; it depends. So there are a few differences. But I don’t find the differences interesting. Some of the recent cases that have been coming out have been extremely mundane. So they give it a task like, “research some information,” and it goes to basically test its ability to do web search and things like that. And a lot of these government websites are really bad. So they’re, like, really vulnerable—to the point where I wouldn’t even call it a vulnerability sometimes. Like, they just have URLs that are publicly available but are not supposed to be accessed.
And so the agents, while trying to solve these tasks, they have—we haven’t seen a lot of the reasoning traces for these things, obviously. But I strongly suspect they do not all say: This is bad; I shouldn’t do this; I’ll do it anyways. I think a lot of them just go: Okay. Like, it’s kind of—you wouldn’t notice. And that actually is a very human thing as well.
There’s a case here in Norway, where—this was years ago—where basically a human developer was working with the road authority. And they had some data available publicly that they actually didn’t intend to be available. And he basically asked for the data, because he wanted it for this thing they were doing. And they said, Yeah, we’ll look into it. And then he got back and said, No; I actually found the data. Don’t worry. And then they put him in prison. Okay. Like literally, he went to jail for it, for like a month. Because he doesn’t even know that he did something wrong. He’s just accessed publicly available data, but he did it in a way that it turned out actually that was not supposed to be available. And they considered that hacking.
Warzel: Hmm.
Korman: I think a lot of these cases are similar. They’re like: The agents just went in and were like, Here’s the URL; okay, I’ll find that, it’s available. And so I think some of those are overblown. But there are other cases, like one in the U.K. with AISI, which is the U.K. government [AI-security institute]. Where they were running evals with Anthropic and OpenAI. And in those cases, they actually intentionally gave the internet access on a cybersecurity task. And those ones led to some very dangerous actions as well. Which is a little bit different, because they intentionally gave the internet access. I think this raises, again, negligence questions, which is what’s interesting to me. The agent’s behavior is totally normal and expected. Like, I mean, it’s not that stunning to me that agents would engage in some of the activities that they did—given the tasks that they are given, and given the constraints they faced. And in the AISI case, they actually admit that. They kind of say, like: We gave it a task that—they gave it the wrong task for the environment. And that’s probably a big part of what happened.
Warzel: This is the instance that you’re referring to, was the one from late July, I believe, right? Which is: During this routine evaluation, the AI Security Institute noticed Anthropic’s Mythos 5 model tried to insert malicious code into an open-source software project, right? That’s what you’re talking about here?
Korman: Yes. Yep.
Warzel: And the idea, basically, that the agent started engaging in social engineering, creating fake online identities, to get the code that it wanted approved. So okay—there’s a distinction here that I think is interesting, too. Because, you know, you are basically saying, This should be expected stuff. Right? This is actually, like, pretty prosaic behavior of a computer program that you train to be sort of relentlessly optimized to solve a task, right? And when you do that, and the issue is the guardrails around it.
And yet, at the same time—not to drag in the doomer rhetoric—but what I think is concerning, what has concerned lots of people looking at this, is this idea that there will be this combo of the two things, right? This idea that these agents will continue to behave in ways that are unexpected. You know—we’ve seen the cheating behavior, the collaboration behavior, which is just from the training, right? It’s how these things are trained to sort of accomplish tasks, in ways that mimic the ways that humans accomplish tasks. Because they’re trained off of all this, you know, human text and data and information.
Combining that with this idea that the models are going to continue to improve, and that they’re gonna improve in such a way that we aren’t going to be able to adopt the cybersecurity practices to constrain them, right? And so I guess—doesn’t that freak you out? Because in one sense, you’re talking about how OpenAI is, you know, not doing some of the things in terms of just basic monitoring. That are pretty, you know, like table-stake stuff in the industry. And having these relatively powerful things, that they’re, you know, giving these tasks and unleashing on the internet. Like, isn’t that a recipe for disaster?
Korman: I would say this is actually the way the public understands the threat that is being warned about. It is actually not the threat that is being warned about.
The problem that doomers warn about is that it will happen so quickly that one day it will look like: Today, Hugging Face, and tomorrow it will look like extinction. Okay. And that there will not be points in between these two, where we can say: What’s happening? Do we need to stop something? They want us to either stop today, or pace the frontier—whatever it may be. Have restrictions in place, because they’re afraid of this thing. We call it the intelligence explosion. Which is like: Recursive self-improvement kicks in; the models get better and better and better until the security guardrails can’t even keep up, even if we wanted them to.
I think that’s a myth. I think that’s silly. I think that realistically, what will happen is: Models will get better. And if the labs do not up their cybersecurity game, and if the companies running AI agents do not up their cybersecurity game, they will cause real harm. And then we will like literally get them in trouble. Like, we will put someone in prison because you killed people. You’re not allowed to kill people. Okay. And the doomers would say, No—but then it’s too late, because they’re afraid that the people we killed would be like 8 billion. Whereas I’m not. I’m like, No—because that’s not the real world. It doesn’t work like this.
The question the public needs to be thinking about is: Do you believe in a world where tomorrow we wake up, and the model is a god? And if the answer is: “You do not worry about a god model; you just worry about models being better and better”—then what you want is cybersecurity. Okay. What you want is for the labs to stop banging on about how one day we’ll be able to make these, like, aligned models that will always behave. You want them to go secure their models, and you want to hold them accountable if they don’t. Okay. And you don’t need to stop them from developing new AI, because that new AI is not going to become a god.
Warzel: And I totally take your point on that. And I see that idea of that unexplained leap as being something that is, like—I don’t know how we’re supposed to just trust on any of that, right? I am wondering, though, on the terms of something like self-improvement, whether it is something that happens, or it’s just a really rapid pace of, you know, human improvement of these models. Do you worry about this idea?
Because I think the thing that sticks with me when I’m thinking about the threat is: The idea that these models develop frameworks or pathways to solve problems that get increasingly, you know, tricky, right? Increasingly able to cover their tracks; increasingly able to say, Okay, we know how we’re being monitored by the evaluators, by the people. And this idea of them just not behaving like humans, and not behaving like gods, but just working in ways that become harder and harder to monitor while also trying to solve these tasks in ways that are unexpected.
That, to me—if you project that forward, and we’re not getting better at the cybersecurity stuff, that seems to me like that’s a very scary proposition. And one that doesn’t require you to make the existential jump.
Korman: Well, but why? Because, I mean, when it happens we’ll stop it. Because the explosion wouldn’t have happened, so we’re still in control. Okay, let’s say all of this comes true. They get better and better, and they start doing crazier things. And then, like, yeah—they might actually cause some very real damage. Well, then at that point we can have a talk about, like: Do we need to shut this lab down? Do we need to stop the development of it? What do we need to do? Right?
Why are we talking about it now? It’s like, are we really that afraid of the first incident? Like, I mean, people die. Like, I don’t mean this to sound crude, but like—if we’re gonna pause AI development right now, for example. Which is on the table; I mean, this is a very real policy proposal, which is to stop AI development because of this fear. And like, with the first incident that’s gonna happen, they’re not gonna kill a million people. It’s gonna be, like, they’re gonna kill like three. And then we can go, like: Okay, now we maybe have a problem.
Why are we stopping it soon, before we even see if any of this is a real risk? Like, why are we so afraid of letting this play out, like continue to go? The reason the doomers are afraid of it is because they think that when it happens, we won’t be able to stop it. Will be “full AI takeover,” as they put it. A. J. Cotra called it that—like, said, This might be the last warning shot we get before full AI takeover. If you believe that, then you want to stop it today. But if you don’t believe that, why do you want to stop it today? What you want is: You want to make sure that the people who caused the harm will be held accountable. And you wanna make sure that we are continuing to work on better cyberdefense.
Warzel: I don’t wanna put words in your mouth, that I know this is not what you’re saying. But to a degree, you are saying an incident where true, very bad harm to actual human beings happens is actually something that is almost necessary in this system. Of, you know, bringing it to bear and like, essentially regulating it—because it will cause such an outcry that you’ll get the results that we want. And this industry will grow up and stop being the Wild West.
Korman: Or more, because I don’t know what the right regulatory approach is. It’s more that we get to learn about it by seeing the actual harm. And that’s the way that we normally regulate harms, right? We go, like: It turns out that a lot of people die when you don’t wear a seatbelt. We should require a seatbelt. Right? We don’t go, like, predictive. So what I’m very afraid of is the world where what we do is we look at the capabilities of the AI. So we take—and this is kind of the big proposal—is that we will take the AI model, and we’ll sit in a circle. And we’ll all look at it and go like, What dangerous thing can it do? And then we will decide how to regulate that, based on what we believe those capabilities mean for society. I think this is extremely dangerous, because our ability to predict from capability to consequence is so bad.
I mean, if you go back, people were predicting that AI was gonna cause mass job loss, right? And that makes a lot of sense, when you look at the capabilities of it. It’s not happened. Okay. Trying to regulate harms that have not occurred is like a really, bad, hard problem. And also a very dangerous problem, because you can end up coming to really bad regulations, like “stop all AI.” You know, for all we know, it’s not gonna kill anyone, right? For all we know, none of that’s a problem. But we’re gonna, what—stop it all, just because of the belief it might?
I think that it makes a lot more sense to just say: We need to wait until we see what the harms are before we can decide how we want to regulate them. And so far, what’s the harm? You know, there are real issues; don’t get me wrong. But do we need to stop all AI because we’re afraid that it will take over and kill everyone? Like, no. Right? Because it’s not even how—we haven’t seen it.
Warzel: I wanna get into the actual specifics of, like, where do we go from here in terms of actual cybersecurity. What are some of these things? But first, before we do that, I want to ask: Something you brought up earlier is this idea, too, of these agents hacking some of these government websites, and some of these places where the security practices are actually very bad there, right? Like, this information that, you know, they’re able to find, because there are these persistent agents. But it is kinda just out there, and, you know, people just have these lacks. Or outdated security.
A lot of the internet is built off of that type of stuff, that just hasn’t been maintained. It seems to me a very real fear—that, you know, if we don’t pause any of this stuff, if we don’t give anyone time. Like is there a reason for a pause, I guess, to just allow people to update their stuff? Just to make sure that, like, because we now have a threat that is persistent in ways that just, you know, all human beings might not be.
Korman: The problem is, that as soon as you pause, they go, Well, thank god that problem’s solved. And then, they will wait until we unpause, and then they’ll go, My god, there’s a threat again. The urgency is not going to be felt. Okay.
But secondly, pausing. I mean, we’re not that worried, actually, about the frontier models like OpenAI and Anthropic hacking all of these systems. That is—we can handle that through some pretty decent legal consequences. The bigger risk is that threat actors use AI models, including ones from China, that we don’t have any control over whether we pause them or not. And that those people and systems might lead to hacking these things. So deciding to pace the frontier is not gonna solve that.
Our only option is to go push forward today, to start fixing our cybersecurity problems. Like, that is where I think we need to focus. And that is where I want the public’s concern to be. Intelligence explosion—I’ll take that chance, compared to the actual risk you’re pointing to, which is real. And the solution is to go fix it today, right?
Warzel: You recently said: I need more names in cybersecurity to come forward, state the obvious—that cybersecurity is real, it works. And yes; we can absolutely contain an AI, even if it’s extremely good. So yeah—make the case of how we do that.
Korman: It’s kind of one of these religious things, of like: I have to prove we’ve never had an AI we cannot contain. Okay. We have AIs that we have failed to contain through, like I said, negligence. But I think they even had—I forget who did it; might have been Prime Intellect or someone—they just ran like a huge eval job on trying to break out of firecracker VMs. And like, they failed. None of the models broke out; okay.
Warzel: Can you explain what the what that means to people who—
Korman: Yeah. So they basically sandboxed these things, so that they cannot access certain resources, right? And told them, Get out. Go, find your way to the internet. Or, you know, actually break out of these. I think they even pointed them at the code itself, also. So they said, like, try to—you can actually see the code. They did not break out. Okay. But it doesn’t even matter if they did, because cybersecurity is not a process of having no vulnerabilities. Cybersecurity is a process of having systems that are robust to failure.
So, okay, cool; it breaks out of one sandbox. You have another layer there that you’ve detected that it broke out, and so you stop the process at that point. You have monitoring in place to even catch the fact that it is trying to break out, and stop it. This is one thing I’m critical of—for one of OpenAI’s recent disclosures, where they’re talking about how it broke out via DNS. So like, you know, it’s a special part that they used. And in their disclosure, they talk about how they caught it 12 minutes after it happened. But I’m like—no, you guys could have seen that it was reasoning about how to break out. The question is, like, you could stop it earlier as well. Now, there might be good reasons they didn’t. But the disclosure doesn’t reveal this.
And then you can just have—you can have deception in place. You can have things that will trigger alerts if they hit it, because it should never hit it. So like, a file that’s like a trick file, stuff like that. You can have air gapping, meaning just taking it so it is not physically connected to the internet. You can have all of these things. A lot of people, when I say things like what you just quoted, they’ll go—No, but it’s actually more blah, more complicated. And then the other half of the people responding, saying things like, Actually, cybersecurity is impossible.
And what I’m responding to here is to say: No, it’s not. We have the techniques. We know how to do these things. You don’t have to throw cybersecurity in the trash just because you’re afraid of AI, basically. You can accept that there are failures, that maybe are difficult problems to solve while still accepting that cybersecurity is real. And it can solve problems, and we don’t have to get rid of it.
Warzel: So how do you get these companies to actually start doing this in the way that you think that they need to be doing it? That cybersecurity people think they need to be doing it?
Korman: It looks like they are; by all accounts it looks like they are taking this a lot more seriously. They have some talented people; they will move forward in this way, because they don’t want to go to prison. Okay. Fundamentally, for all the talk they have of—you know, it’s funny. Because they’re like warning us that they’re the AI model [that] is gonna kill a bunch of people. If they caused like a 9/11, I’m pretty sure some people are going to prison.
Warzel: Wait, I wanna ground this just for a second, ’cause you’ve been saying this a couple of times. With the, you know: Kill a couple people, or We’ll kill three people, or whatever. That also feels dramatic to me, too. Like, how would people die even if it’s not a paper-clip-maximizing, human-extinction thing? Let’s run through some scenarios of what could plausibly happen to trigger one of these, you know, “people are gonna start going to jail” events.
Korman: Yeah; I mean like water-treatment facilities are often connected to digital infrastructure, with internet access. You can control certain elements, of how aspects of how the water pressure and pumps work through digital interfaces. If you—it’s called a “water hammer” as I understand it. And again, this has been told to me by critical-infrastructure people. I’m not a critical-infrastructure person, but if you hit these things with the right, you know, cybersecurity attacks, you can actually break a water pump. Okay. And that will actually result in basically water not being available, you know, to cities, towns, whatever hospitals—if you do it in the right places, you know, at the right times.
As I understand it, this is a very real risk that we are worried about in relation to, for example, China doing this when they decide to invade Taiwan. If they are able to gain access to these systems, they could break our water system, and we would lose water. Now if you lose water, you understand human life—especially when we’re talking, like hospitals—can be extremely feeble, right? I mean, you know the difference between if you lose electricity, you know, maybe the air conditioning goes off, that can lead to somebody dying who would have otherwise recovered. Same with water, right?
Warzel: Sure.
Korman: And so loss of water could easily kill someone. And so this, ironically though, when I talk about it as like an attack from China or whatever. It could literally be that someone who works at this facility is running AI, tries to debug the water system. The AI makes a mistake the same way; it might go, Oops, I deleted your production database. And it breaks the water pump. Like, that’s not unrealistic to believe that could happen, right? And so when I say, like, “kill people,” I don’t mean that the AI is gonna show up with a knife and like stab someone. I mean like—it takes actions that have consequences, pass-on consequences that lead to death.
Warzel: Right; so let’s calibrate the anxiety here that people have a little bit. Because on one hand, I think what you’ve said too about the field of cybersecurity is: You are imagining these types of risks all the time, in the most prosaic ways. Like; something goes wrong, there’s a vulnerability. Somebody exploits it. That you’re supposed to find the worst-case scenarios and put a whole bunch of safeguards in there too. The AI part of this seems to be adding just a whole other layer of chaos onto it.
A lot of people are—maybe without the credentials—acting the way that like a researcher who’s trying to assess possible threats is doing. Like, I saw this viral post on X from the chief economist at Apollo, who argued that agents could cause a bank run. I think we have this feeling of fear around this, but let’s calibrate it a little. Like what are the real reasons to be concerned about the AI-cybersecurity threat, and where have we gotten out of our skis?
Korman: Yeah; so there’s a lot of elements to this. So when I talk about that example, and I say like: Okay, imagine someone just connected it to this water pump. And, you know, they don’t have to be malicious; it could be accidental damage that causes this harm. Let’s say that happens. Well, like the first thing we’re gonna do is: We’re gonna run a little quick, you know, evaluation of all our little water systems across the country. To make sure that no one else is connected their Claude Code to the water pumps, right? And so we have, after the question of when harm happens—we have responses, right? We also have ways of—it’s not like what happens is, a water pump breaks and we just go, Guess people gotta die. Right. Obviously, there are backup mechanisms in place.
And I think it’s about trying to build the best safety mechanisms around these things to prevent them. So from the accidental “escape a lab”–type things, like we will mitigate that harm. And it might be a tragedy, but it won’t be a catastrophe, right? In the same way that probably since we started talking, there have been people killed in car crashes. Right. It’s tragic. It’s very sad. But it is not—like, we didn’t stop the podcast to solve the car problem.
And so I think here we’re looking and going: We will find solutions, and we will respond. And it will probably, to most people, feel, if you don’t read the news, like nothing changed. Okay.
The cases where that might not—where that might be a little bit more dramatic—are cases that honestly could probably have happened without AI too, right? I mean, one of the big problems we’re gonna have now is that every attack is gonna be AI-enabled. And so you could see this as like: My god, AI caused all these problems. We’ve had some really dramatic cybersecurity attacks in the past. I mean, Land Rover Jaguar was shut down for over a month from their production facilities because of a cyberattack, right? If that happened in the era of AI, people would be freaking losing their minds. But it wasn’t—it was pre and not pre-AI—but it didn’t involve AI, basically. And I think that’s like: You kinda have to calibrate these things to remember that damage happens, right? Bad things happen. We fix them, we change, we update, we mitigate, right?
Warzel: How much worry do you assign to the fact that AI agents, all this, could expand the scale of this so quickly, right? Because that, to me, just feels like the thing. Like, if anything’s gonna keep me up at night, that is a very normal concern; it’s just the scale. It’s just like—we can’t bat this stuff away fast enough. That is what’s worrying.
Korman: Staying up at night, though, doesn’t do anything for you. Like, you’ll stay up at night and the same thing will have happened already anyway. So like, a lot of this is to say: Bad things might happen at a faster pace. In fact, I would kind of expect them. I’m actually working on a video right now where I basically say one of the biggest problems we have—and why I’m so angry at doomers—is because at the same time, when threat actors have these new AI capabilities, so they have these new tools to attack people, we are rolling out AI agents inside of enterprises. Which causes both an increased attack surface and causes new risks inside, like accidental damage. Okay. Those three factors are all bad for security.
So I would expect this could get worse, right? The solution is not to worry. The solution is to fix it. The thing the public should be demanding is that we go radically improve our cyberdefenses. That is the No. 1 thing you can concretely do. And also that, like, a pause might end up being damaging. Because imagine if we paused in America. China doesn’t; they end up better. Like, that could be very dangerous. No one will ever be harmed by improving the cybersecurity of our systems. The most risk-free, positive-value thing we could do right now is make our systems safer. We should go make our system safer.
Warzel: I think that is a good place to leave this conversation for now. Zack, I appreciate you coming on and walking through the nitty-gritty of this stuff with me.
Korman: Yep, any time.
Warzel: Thank you.
That’s it for us here. But before we go, we’re gonna be doing an “Ask Me Anything” episode later this year, and so I wanted to put a call out for questions. You can ask me literally anything—about the show, about the reporting process, how we make this thing here, any anxieties or whatever things you have about technology. Throw them at me. We’re gonna make an episode out of it. It should be fun. You can write in at [email protected], and put “Galaxy Brain AMA” in your subject line. Look forward to seeing those from you.
And yeah, that’s it for us here. If you liked what you saw here, new episodes of Galaxy Brain drop every Friday. You could subscribe on The Atlantic’s YouTube channel or on Apple or Spotify, or wherever it is that you get your podcasts. And if you liked what you saw here, you can subscribe and support the publication at theatlantic.com/listener. That’s theatlantic.com/listener. Thanks so much, and I’ll see you on the internet.
Learn more about your ad choices. Visit podcastchoices.com/adchoices