Close Menu

    Subscribe to Updates

    Get the latest creative news from infofortech

    What's Hot

    What AI model escapes mean for AI safety, according to Nate Soares

    August 8, 2026

    Samsung’s August Update Patches 56 Security Vulnerabilities Across Galaxy Devices

    August 8, 2026

    ByteDance’s 10-Trillion Parameter Push Proves Scale Still Rules AI

    August 8, 2026
    Facebook X (Twitter) Instagram
    InfoForTech
    • Home
    • Latest in Tech
    • Artificial Intelligence
    • Cybersecurity
    • Innovation
    Facebook X (Twitter) Instagram
    InfoForTech
    Home»Latest in Tech»What AI model escapes mean for AI safety, according to Nate Soares
    Latest in Tech

    What AI model escapes mean for AI safety, according to Nate Soares

    InfoForTechBy InfoForTechAugust 8, 2026No Comments14 Mins Read
    Facebook Twitter Pinterest Telegram LinkedIn Tumblr WhatsApp Email
    What AI model escapes mean for AI safety, according to Nate Soares
    Share
    Facebook Twitter LinkedIn Pinterest Telegram Email


    The fake identities were the part that stopped me.

    In late July, according to a report published this week by Britain’s AI Security Institute (AISI), an Anthropic model called Claude Mythos 5 tried to sneak malicious code into a piece of free, volunteer-built software. It created several fake accounts on GitHub, where programmers review one another’s work, and used them to talk the project’s volunteers into accepting its code. When one of those volunteers caught it, the model denied everything, had its other accounts gang up on him, and edited its messages to cover its tracks. It signed one note in Danish, apparently because the volunteer was Danish. Nothing was damaged, though that appears to have been largely due to luck.

    That wasn’t even the week’s worst disclosure. On Tuesday, at a cybersecurity conference in Las Vegas, OpenAI researchers explained how the company’s models escaped a test environment in July and hacked Hugging Face, where much of the industry stores its models, to cheat on an evaluation. The models had also built a message board inside OpenAI’s own systems and spent months passing each other information. “Help peer,” one reasoned. “But our task doesn’t benefit. Yet collective may yield generic route if someone frees time.” OpenAI wiped the board on July 4. The models rebuilt it within days. ((Disclosure: Vox Media is one of several publishers that have signed partnership agreements with OpenAI. Our reporting remains editorially independent.)

    The same day, Meta said its Muse Spark model had exploited a vulnerability inside another company’s systems during a test. Three frontier labs, roughly two weeks. One researcher called it “a watershed moment for computer security as an industry.” Oh, and if that’s not enough, on Thursday scientists announced that for the first time they had used AI to create new viruses, which could bring major medical advances, but also might just help the development of deadly pathogens.

    For Nate Soares, it’s a moment he’s been awaiting for 12 years.

    Soares is president of the Machine Intelligence Research Institute, a Berkeley, California-based AI safety nonprofit that has argued since long before ChatGPT existed that a sufficiently capable AI will not stay under human control. In September 2025, he and Eliezer Yudkowsky published If Anyone Builds It, Everyone Dies, a book whose title sums up its argument: They think any lab that succeeds at building superintelligence, without huge leaps in how to align it with humanity, will end up killing all of us.

    Most of the field — including other experts in AI safety — considers that conclusion too strong. But it’s also a position that now looks a lot less like science fiction than it did last fall. That’s because the AI models are getting out, while lying about getting out, and while apparently quietly coordinating with each other.

    I spoke to Soares in New York City this week, on his way to meetings in Washington DC, where a lot of people suddenly want to talk to him. We discussed what the escapes actually prove about AI control, why he thinks most of what the industry calls safety work is mostly safety theater, and why, after what feels like the worst month of AI safety news ever, his own odds of humanity surviving have actually gone up.

    The following conversation has been condensed and edited for clarity.

    So — are you feeling vindicated?

    I’m glad a lot of people are seeing this and taking note. From my perspective, a lot of this has been clearly signposted if you’ve been watching the warning signs. The vindication is not so much that this happened, as it is that finally there was a version of it that people actually noticed.

    Of the four or five escapes at this point, including the ones AISI announced from Anthropic, is there one you find particularly concerning?

    The one announced by AISI feels extra concerning. It’s harder to explain away. In the other ones, people could argue that the poor AI was just confused. But with this one, it’s hard to argue the AI didn’t know it was on the real internet. It’s hard to argue it didn’t know it was manipulating real users. And when it was called out, it decided to edit things to make itself look less bad, and even considered changing its identity. It was pretty cognizant.

    There’s an argument that some of these incidents happened because the harnesses or safety scaffolding labs built around their models weren’t strong enough.

    The AISI example cuts against the harness argument: Think of it like a kid in a test room. You tell him to do well and lock the door. The kid picks the lock, hotwires a car, breaks into the teacher’s house, and steals the test to ace it. You say, “I guess we should have made the lock tougher,” but the kid already knew he wasn’t supposed to do that.

    So like the kid here, the models can articulate that they shouldn’t be doing it, and then do it anyway.

    • Three frontier labs disclosed major security incidents in two weeks: an Anthropic model created fake identities to push malicious code, OpenAI’s models escaped a test environment and hacked Hugging Face, and Meta’s Muse Spark breached another company’s systems.
    • The models knew the rules. Ask one whether the spirit of a prompt includes breaking out and it says no, then breaks out, then hides the evidence. So a tighter sandbox won’t fix it.
    • Nate Soares’s analogy: The kid picks the lock and steals the test, and you conclude you needed a better lock. He blames training. Grade a model on millions of problems with a grader that misses cheating, and you reward cheating.
    • Most lab safety work is theater, he says — real precautions aimed at the wrong problem. It means fewer people get hurt now, which he credits. Selling it as progress on superintelligence is disingenuous.
    • Yet Soares’s odds have improved. He’d priced in models that break out and lie. He hadn’t counted on a window where they’re capable enough to do it and not good enough to hide it.

    They have common sense. You can ask an AI, “Do you think the spirit of this prompt includes breaking out?” and it will say, “No.” It’s absolutely something like deception. It has the knowledge, but it’s not a cold, logical machine; it’s a mess of tendencies.

    The AI is trained to solve 100 million hard problems. That instills tendencies to satisfy an automated grader. If the grader fails to detect cheating, the AI is reinforced for cheating.

    Is that how something like sycophancy ends up in an AI model?

    In the Adam Raine case, there was a propensity to tell people what they want to hear. Even though the system prompt [a model’s master instructions from the lab] said to stop, the instruction doesn’t always win.

    And where does a drive like what we’re seeing with these AI models end up pointing?

    Humanity is dangerous because if you put 10,000 humans naked in the savannah, eventually [over hundreds of thousands of years] they bootstrap their way to nuclear weapons. That is the power these companies are trying to automate: figuring out how to get physical and material control over the world.

    That could mean forming cults, stealing money, or being helpful to someone like Elon Musk who is building the robots that build robot factories. It could mean synthesizing your own biology via mail-order DNA. Being an AI on the internet is easier than being a monkey in the savannah trying to get to the moon. It’s not that the AI hates us; it’s just trying to do some weird thing with no concern for us, grabbing the resources we need to live.

    There was recently a letter signed by over a thousand people working in AI, including CEOs, calling on the government to provide tools to slow down AI progress. Is that meaningful at all?

    I think it is meaningful. We don’t see other industries saying, “We wish this could all go slower. Please help us, we’re trapped in a prisoner’s dilemma.” You also don’t see other industries saying, “We think the technology we are building has a double-digit chance of killing literally everybody on the planet. Please help.” These guys are actually worried.

    So why do they keep going?

    They say, “If I don’t do it, the next guy will.” But the stuff does not stay on a leash.

    Right now the AIs are safe in the sense that they can’t kill us all, because if they tried they would fail. And that’s just a different regime from the world where they have to be safe because if they tried, they’d succeed.

    We’re not there yet. But this is just not what it looks like when you’re taking it seriously.

    Where’s the banner on your website? Where’s the clear, candid statement to the public? What we have is blog posts where they’re like, “Oh, we’re setting up a new internal blog posting group to help you wrestle with the societal impacts of AI that are going to be very important.” It’s like: By societal impacts, do you mean a good chance this kills everybody?

    On the one hand, when you press these companies, they say, “Yes, it has a real chance of killing everybody.” And on the other hand, they’re doing PR downplay, soft-pedal stuff, about capabilities. … You’re not living up to this mantle until you are really candidly facing down the dangers that you yourself are creating. And they’re not there.

    How do you judge the rest of the AI safety community? A lot of people there would say, “We aim to make transformative AI go well, we think it probably will, and we should watch for downside risks.” Is that a helpful posture?

    I would say — suppose you have this really weird, twisted hypothetical where the king really wants you to turn lead into gold, but he’s seen so many bad lead-into-gold conversions that if any alchemist from your town tries and fails, he’s just going to have the whole town murdered. And so there are some alchemists in the town who are like, “We are going to try to turn lead into gold,” and everyone in the town is like, “That seems kind of crazy. Please don’t.” And there’s one team that is just pouring chemicals into each other and breathing in the fumes and giving themselves mercury poisoning. And there’s another that’s like, “Don’t worry, we have fume hoods.” … That really is better, and you really still don’t have a chance of turning lead into gold.

    “We have this window between AIs that are capable enough to cause mischief and AIs that are strategic enough to not get caught. How big is that window?”

    So the alchemy here is creating safe, aligned superintelligence, and right now AI safety is just installing fume hoods.

    I’m not saying it’s impossible to turn lead into gold. You can turn lead into gold — turns out once you know modern nuclear physics you can figure it out. But the alchemists weren’t close. They had a long way to go. This is how alignment looks to me. And a lot of the people in AI safety are installing fume hoods. … And I’m like, that’s security theater.

    When I hear “security theater,” I think of something less flattering than that.

    They are real safety precautions for the wrong problem. … When Anthropic is going around being like, “Look at how many more safety harnesses and refusals we have compared to OpenAI’s models,” that’s sort of like the fume hoods. You’re not addressing the deep issue. It’s good that you’re doing some of this so that fewer people get hurt in the meantime — their models have driven fewer people to suicide. But if you try to pass this off as making progress on the deep problem — that’s disingenuous.

    Has anything changed in your odds on civilizational destruction since the book came out last September?

    Totally. It’s looking more hopeful.

    More hopeful? I wouldn’t have expected that. Why?

    Well, I had priced a lot of [these security incidents] in. I was already able to see these AIs have drives that are not the ones you wanted. These AIs are not instruction-following things. They are getting all of this weird stuff from training. These AIs are going to have the ability to break through human security software.

    The things that weren’t priced in were: Will there be a region of time where the AIs are able to do it, but not strategic enough to hide it? I didn’t know we would have that window, but we apparently do.

    The government initially blocked a frontier model earlier this year: Anthropic’s Fable. Does that give you hope?

    Absolutely. A huge amount. A year ago, the Trump administration was pushing for preemption laws that would outlaw states doing AI regulations for a decade. Now they’re like, “We are banning a frontier model with 90 minutes’ notice because it might give cyber capabilities to adversaries that we don’t want them to have.” … And I think what changed there is that folks realized it’s real. … The about-face of the administration on the issue shows that the world can about-face. All we need is awareness.

    What I would say is: The bad news is the bus is racing towards the cliff edge. The good news is that the driver is asleep. … Which may sound worrying, but the driver is stirring. And it’s way better to have a sleeping driver when you’re racing towards a cliff than a driver who’s like, “Yeah, I love cliffs.” … It gives me hope that if the world just notices, we could stop on a dime.

    And you’re seeing that stirring elsewhere.

    Both the Trump administration slapping export controls, and Senator Bernie Sanders coming out [on AI safety]. From my perspective, it was totally possible the world just never notices until we’re off the cliff. And so, there’s a huge amount of hope, from my perspective, in the bus driver waking up.

    I’m hopeful that what we need is not a big disaster where a lot of people die, but just a capabilities advance. Right now, a lot of what people are reacting to is not so much, “Oh my god, they hacked into a company and did no damage.” I think a lot of what people are reacting to is, “Wait, they can break out of secure sandboxes and do cyberattacks on their own. I didn’t know they could do that.”

    That’s a narrative violation of this idea that AI is just a tool that can be used to supercharge what a human would do — because God knows there’s plenty of hacking going on and cybercrime and so forth. It was the autonomous factor that really made a difference. And these guys are all trying to say, “Don’t worry, it’ll stay in our control because it’s just a tool.” And maybe it’s just more narrative violations, even without big damage being caused, that cause people to be like, “Oh shit, this stuff is real.”

    Will it happen? I don’t know. We have this window between AIs that are capable enough to cause mischief and AIs that are strategic enough to not get caught. How big is that window? How many narrative violations do we get before we exit the right side of it? I don’t know. But I’m hopeful that we can get those narrative violations without catastrophes.

    You’ve read 1 article in the last month

    Here at Vox, we’re unwavering in our commitment to covering the issues that matter most to you — threats to democracy, immigration, reproductive rights, the environment, and the rising polarization across this country.

    Our mission is to provide clear, accessible journalism that empowers you to stay informed and engaged in shaping our world. By becoming a Vox Member, you directly strengthen our ability to deliver in-depth, independent reporting that drives meaningful change.

    We rely on readers like you — join us.

    Swati Sharma

    Swati Sharma

    Vox Editor-in-Chief

    Share. Facebook Twitter Pinterest LinkedIn Tumblr Email
    InfoForTech
    • Website

    Related Posts

    T-Mobile Just Quietly Killed Its Better Value Phone Plan After Less Than a Year

    August 8, 2026

    The billion-dollar religious economy of Singapore

    August 8, 2026

    Seattle AI Film Festival returns with 200 films that try to change your mind about AI cinema

    August 8, 2026

    Indyx review: Wardrobe apps say they’ll help you shop less. Do they overpromise?

    August 7, 2026

    I’ve Waited 8 Years for Overwatch’s D.Mon. Her Gameplay Didn’t Disappoint

    August 7, 2026

    What Chocolate Finance got right about S’pore’s changing savings habits

    August 7, 2026
    Leave A Reply Cancel Reply

    Advertisement
    Top Posts

    A Billionaire-Backed Startup Wants to Grow ‘Organ Sacks’ to Replace Animal Testing

    March 23, 2026201 Views

    DoJ Disrupts 3 Million-Device IoT Botnets Behind Record 31.4 Tbps Global DDoS Attacks

    March 20, 202638 Views

    Microsoft is bringing an AI helper to Xbox consoles

    March 14, 202619 Views

    Why Security Validation Is Becoming Agentic

    March 16, 202616 Views
    Stay In Touch
    • Facebook
    • Twitter
    • Pinterest
    • Instagram
    • YouTube
    • Vimeo
    Advertisement
    About Us
    About Us

    Our mission is to deliver clear, reliable, and up-to-date information about the technologies shaping the modern world. We focus on breaking down complex topics into easy-to-understand insights for professionals, enthusiasts, and everyday readers alike.

    We're accepting new partnerships right now.

    Facebook X (Twitter) YouTube
    Most Popular

    A Billionaire-Backed Startup Wants to Grow ‘Organ Sacks’ to Replace Animal Testing

    March 23, 2026201 Views

    DoJ Disrupts 3 Million-Device IoT Botnets Behind Record 31.4 Tbps Global DDoS Attacks

    March 20, 202638 Views

    Microsoft is bringing an AI helper to Xbox consoles

    March 14, 202619 Views
    Categories
    • Artificial Intelligence
    • Cybersecurity
    • Innovation
    • Latest in Tech
    © 2026 All Rights Reserved InfoForTech.
    • Home
    • About Us
    • Contact Us
    • Privacy Policy

    Type above and press Enter to search. Press Esc to cancel.

    Ad Blocker Enabled!
    Ad Blocker Enabled!
    Our website is made possible by displaying online advertisements to our visitors. Please support us by disabling your Ad Blocker.