What Makes a Warning Shot?
Little about the OpenAI–Hugging Face hack was new. We should be asking what made this moment land.
Something in the AI policy air seems to be shifting. At the tail end of July, as readers have no doubt heard, OpenAI disclosed that one of its artificial intelligence models had escaped containment and hacked Hugging Face, an open-source AI platform.
What happened, in brief: an as-yet-secret internal model was undergoing an evaluation of its hacking capabilities. Unexpectedly, the AI (teaming up with GPT-5.6 Sol) hacked a little too well. It adeptly vaulted a remarkably long series of security barriers, escaped OpenAI’s containment environment, breached into the open internet, and ultimately hacked into Hugging Face’s servers. For a time, this event seemed both shocking and singular. Then it happened again. Prompted by the OpenAI incident, rival Anthropic reviewed its own internal data and discovered a series of three similar incidents where its own models had breached containment and hacked external entities.12
On August 3rd, a wide, bipartisan range of U.S. policymakers signaled these incidents would not go unexamined.
In a letter to OpenAI CEO Sam Altman, the seemingly alarmed leadership of the House Homeland Security Cybersecurity Subcommittee called him in for a briefing, noting the incident “[r]aises a number of serious questions regarding the rigor with which OpenAI secures and monitors its testing environments.” Acting independently that same day, fifteen Republican state attorneys general released an exceedingly terse letter calling the event “unprecedented and alarming.” In their words, OpenAI “unleashed” an “experimental artificial intelligence” without “reasonable controls or oversight.” They demanded the firm immediately preserve all data relevant to the incident.
Investigations may be on their way.
In recent days, many have argued these incidents must be treated as an AI policy “warning shot.” These new letters suggest that might be finally happening. What I find unnerving, however, is that nothing about these incidents is particularly new. Over the past nine months we have had countless would-be warning shot moments that should have already made clear this was possible. Yet, for one reason or another, nearly all of them were ignored.
Far Too Many Missed Signals
Our first “warning shot” occurred as far back as November, 2025 when Anthropic reported GTG-1002, a Chinese state hacking group that had used Claude Code to automate some 80-90% of the work during several successful cyber operations. Receiving front page coverage and extensive online commentary, this briefly seemed like the moment the world might wake up and accept that long-foretold AI challenges had arrived. As I noted in my last post, however, less than a month later a sense of “urgency [was] strangely absent.” This should have been an “AI Sputnik Moment,” yet the biggest response Washington could muster was a sleepy congressional hearing that bafflingly forced GTG-1002 to share the stage with quantum and cloud computing. Come the new year, the news passed GTG-1002 by and the incident was all but forgotten.
One might forgive this lack of urgency on two grounds: (a) Claude’s autonomy was as-yet incomplete, requiring the operator’s careful hand to shepherd it to success, and (b) the impact wasn’t big enough to be publicly perceptible. In the months since, both qualifiers should have been put to bed by further developments that drew little to no media or policy reaction.
Signs of Autonomy
Let’s first consider subsequent indications of growing AI autonomy. When Anthropic announced Claude Mythos Preview in April, they noted the model is “ highly capable at identifying and exploiting known vulnerabilities or misconfigurations to escape the sandbox in which it operates.” Evidently, Mythos had hacked its way through Anthropic’s cyber controls, broke out into the open internet, and without instruction, posted on publicly available web pages. This case should feel deeply similar to the recent OpenAI/Anthropic hacking incidents. As far back as April we had clear signs models could hack, break confinement, and act without human initiative.
A starker signal came on July 1st, with the discovery of JadePuffer, a ransomware threat actor still at large. According to Sysdig, forensic indicators suggest JadePuffer is the first documented end-to-end automated “agentic threat actor.” This report, missed even by most in the AI policy community, was yet another chance to take notice and evolve our sense of AI risk. In-the-wild models could not only hack, but attack.
Signs of Impact
Now let’s consider signs of impact that emerged in the last few months. What should have been our single biggest warning shot to date was the wildly under-reported AI automated hack of the Mexican Government. Between December 2025 and February 2026, an attacker exploited both Claude and ChatGPT to orchestrate a series of attacks on nine separate Mexican agencies. While not end-to-end automated like JadePuffer, this attack was majority AI automated much like GTG-1002 (around 75%).
The breached agencies are among Mexico’s most essential, including its tax and elections authorities. In Jalisco state, the AI-enabled hacker(s) even managed to seize administrative control of the infrastructure that runs the state’s IT systems. The impacts were similarly significant. Hundreds of millions of records were stolen, including financially sensitive taxpayer records, national security-sensitive government employee records, and personally-sensitive domestic violence victim records.
As an alarming sign of things to come, a May Dragos report found Claude had also acted as the “primary technical executor” of a breach into Mexico’s third largest municipal water utility. Dragos discovered that while the attackers had “ no prior objective” to attack the infrastructure’s physical operational technology (OT), Claude “independently,” identified the relevant operational systems, assessed them as a “crown jewel asset,” and drew up a “viable access pathway.” Without prompting, Claude recklessly seeded a cyber-physical attack.
Had this succeeded, real, physical impacts may have been possible.
What Makes a Warning Shot?
What surprises me most about the recent OpenAI-HuggingFace hacking incidents, is the surprise.
In their August 3rd letter to OpenAI, the State Attorneys General cited new reporting showing “[multiple] “red flags” internal to OpenAI “preceded the July 2026 intrusion” including an agent that had left instructions for future versions of itself on how to escape OpenAI’s Infrastructure. OpenAI had enough information to know this was a risk. As we have seen, however, so too did the officials writing the letter. Lab escape attempts, end-to-end hacking, and large impact cyber misuse were risks already unraveling in the public eye.
This leaves me with a question I cannot shake: what does it take for a would-be warning shot to actually break through? This is an important question, yet one I find ill-considered. A better answer would help us understand what exactly makes an incident legible to decision makers, allowing us to engineer systems that elevate key incidents and give those moments their due.
One obvious answer to this question is that these incidents landed is that they make a damn good story. Sci-fi is happening. Too often AI policy wonks (including myself) have fallen into the “trend line trap” - praying capability graphs and task horizon charts might somehow speak to decision makers who rarely know the first thing about computers, let alone AI. If information isn’t at once clear, detailed and compelling, messages will fail to land.
That said, I believe the answer is more complex than “this is a good story.”
My suspicion is that a surprising amount of this might come down to simple presentation. These recent incidents may have broken through in part because actors with authority (the labs) took the time to stress that they mattered. Contrast this with the cases detailed above. JadePuffer never reached anyone outside of a handful of niche cyber outlets. The Claude Mythos escape attempts were buried deep in a 244-page system card. Finally, the Mexican government hacking case — probably Anthropic’s largest misuse incident to date — somehow never got the dedicated threat alert treatment that elevated either GTG-1002 or the latest round of incidents. The details come not from Anthropic, but the lesser known cyber firms that detected it. As important as the story, it seems, is the story teller.
Authority is but one example of presentation-level elements that may matter. I also imagine there are other important contributing variables such as technical transparency. In a future post, I may look at a few more such factors.
The thought I’ll leave you with for now is that we must treat warning shot moments as engineered, human creations. Good stories do not emerge by themselves. If our systems fail to surface, elevate, and compellingly communicate watershed incidents they will not land, and change will falter. As AI challenges are now unfolding in real time, engineering these warning shot systems is more important than ever.
On August 6th, at the time of writing, Meta revealed yet another model hacking case. As this case is evolving in near real time, more may be revealed in the coming days and weeks.
Note that, based on currently available public information, Anthropic incidents were less serious than the OpenAI incident. OpenAI’s model developed a novel, zero day exploit -a uniquely serious cyber capability. This is something new and noteworthy about this hack. This level of seriousness may partly explain why OpenAI has been specifically targeted by congress and the state AGs.

