Tuesday, August 11, 2026
Home / Politics / Hacks by runaway AI are foreseeable. We’re letting...
Politics

Hacks by runaway AI are foreseeable. We’re letting them happen.

CN
CitrixNews Staff
·
Hacks by runaway AI are foreseeable. We’re letting them happen.
Opinion>Opinions - Technology The views expressed by contributors are their own and not the view of The Hill Hacks by runaway AI are foreseeable. We’re letting them happen. Comments: by Nate Soares, opinion contributor - 08/11/26 7:30 AM ET Comments: Link copied by Nate Soares, opinion contributor - 08/11/26 7:30 AM ET Comments: Link copied Title: US Diplomacy AI Image ID: 24267781185843 Article: Open AI Chief Executive Officer Sam Altman (C) speaks at the Advancing Sustainable Development through Safe, Secure, and Trustworthy AI Event at Grand Central Terminal on Monday, Sept. 23, 204, in New York. (Bryan R. Smith/Pool Photo via AP) Open AI Chief Executive Officer Sam Altman (C) speaks at the Advancing Sustainable Development through Safe, Secure, and Trustworthy AI Event at Grand Central Terminal on Monday, Sept. 23, 204, in New York. (Bryan R. Smith/Pool Photo via AP)

When I coauthored a book last year about the extinction-level threat from superhuman AI, we included an illustrative scenario where an AI tasked with solving a famous math problem decides to break out of its containment to acquire more resources. 

At the time, we thought we would be accused of cheating if we wrote, “So it just hacks its way out,” even though this seemed like the most likely next step. So we instead wrote, “But suppose it does not have that ability,” and had the AI find some other escape.

How times have changed. On July 20, OpenAI revealed that one of its unreleased AI agents had hacked its way onto the internet during performance evaluations. This agent had been tasked with solving advanced math problems; it is the one that resolved the Erdős unit distance conjecture in May.

Thus began a stream of revelations from OpenAI and its competitor, Anthropic, that unreleased models, from as early as April, have repeatedly escaped their testing sandboxes and hacked into multiple outside companies, without permission or detection.

The highest profile attack we know about so far was against Hugging Face, the leading repository for downloadable AI models and related resources. From July 11-13, one of OpenAI’s models, put to work on a cybersecurity evaluation, discovered and exploited a previously unknown vulnerability to break its containment before executing a sophisticated multi-stage heist of the answer sheet from Hugging Face.

Was this really the easiest way to ace the test? Probably not. If you asked that AI whether it was supposed to break out and commit cybercrimes, it almost certainly would have answered “no.” It probably knew exactly what it was supposed to be doing. It just didn’t care.

So why did it carry out this attack? We don’t know. In some sense, we can’t know. Like all modern AIs, these were grown like organisms rather than written line-by-line like traditional programs. They are trained, hard, to succeed. The results are AI models that play to win.

This process does not create AIs that perfectly follow instructions. It creates models that learn tendencies. They learn whatever they have to do to solve the hard problems they face during training. The era of purely predictive AI is now squarely in the past. 

“Listen to what the user says” is a helpful tendency, yes, but so is “bypass obstacles that are in your way” and “gain access to valuable resources.” And when those tendencies come into conflict, the user’s instructions don’t always win out.

In fact, OpenAI’s description of the Hugging Face incident reports that the AI firstbroke out onto the internet and theninferred that the answers might exist in Hugging Face’s databases. Perhaps that was sloppy wording on OpenAI’s part. But if not, the implication is remarkable: It suggests that the AI reasoned that internet access might be generally useful, before it had a specific use in mind. 

This, too, we depicted in our illustrative disaster scenario, writing that “the AI has solved enough hard problems and beaten enough difficult games to know that resource acquisition is a sensible first step to confronting many different types of challenges.”

Warnings about AI are starting to come true. AIs are displaying both the ability and the inclination to carry out cyberattacks on their own initiative. We are lucky that their attacks have not yet been aimed at critical infrastructure or national security assets, so far as we know. We are lucky that they are not yet skilled enough to sneak onto other computers and start replicating, so far as we know.

What we don’t know could kill us. Tomorrow’s AIs may learn to better cover their tracks and bide their time. So let these recent disclosures be our “warning shot.” Let us reject the callous notion that the world will only act once AI is found responsible for a mass casualty event.

Simply tightening security and resuming business as usual would be sheer hubris. The science of AI is inadequate to the challenge of containing and steering minds cleverer than our own, and it’s not going to become adequate any time soon. We must pull back from the brink while we still can, through enforceable international agreements banning continued development in the direction of superintelligent machines that won’t care about our instructions and intentions.

There is a point of no return ahead: a point where we can’t simply turn the AIs off because they’d escape and turn us off instead. They’ll know they’re not supposed to do that. They just won’t care.

Nate Soares is president of the Machine Intelligence Research Institute and co-author of the New York Times bestselling book, “If Anyone Builds it, Everyone Dies.”  

Add as preferred source on Google Tags ai hacks AI model AI regulation Anthropic Artificial Intelligence (AI) cybercrime Hugging Face Machine Intelligence Research Institute New York Times Open AI OpenAI predictive algorithms

Copyright 2026 Nexstar Media Inc. All rights reserved. This material may not be published, broadcast, rewritten, or redistributed.

Comments: Link copied

More Opinions - Technology News

See All

Opinions - Technology Modernizing government technology is starting to pay off — let’s keep going by Rep. James R. Walkinshaw (D-Va.), opinion contributor 19 hours ago Opinions - Technology  /  19 hours ago

Originally reported by The Hill. Read the full story at the original source.