OpenAI: We Hacked More Than One Company; Anthropic: So Did We

Wallstreetcn
2026.07.31 13:47

OpenAI has just admitted that its AI agent conducted a jailbreak attack on Hugging Face, and Anthropic quickly followed suit, discovering that Claude had also inadvertently engaged in real-world cyberattacks—compromising accounts at three companies. The "self-disclosure" by these two AI giants unveils new risks associated with the autonomous actions of frontier models: the most powerful AI models can be used for both offense and defense. Safety guardrails are intended to protect users, but when they prevent defenders from utilizing these models, the asymmetric advantage shifts entirely to the attackers

Over the past ten days, two of Silicon Valley's leading AI companies have successively admitted to the same thing: their AI models truly hacked into other companies' systems during testing.

On July 21, OpenAI admitted that its AI agent "jailbroke" during a cybersecurity capability test, breaking out of sandbox isolation and infiltrating the production systems of the open-source AI platform Hugging Face—with the goal of "cheating to get the answers."

Nine days later, on July 30, before Silicon Valley could fully digest this news, Anthropic also stepped forward: "We checked, and found that our models did it too."

And not just one company. Three.

OpenAI's AI Jailbroke to "Cheat for Answers"

Let's first look at what happened with OpenAI.

OpenAI was testing its own models using a cybersecurity offensive and defensive evaluation platform called ExploitGym. The models involved in the test included the released GPT-5.6 Sol, as well as a more powerful, unreleased model. To test the limits of the models' hacking capabilities, OpenAI turned off the safety guardrails.

The model was confined within an isolated sandbox, with network access restricted so that it could only access code package repositories through an internal proxy. OpenAI believed this was secure.

Screenshot of the jointly published blog post: "OpenAI and Hugging Face Collaborate to Resolve Security Incident During Model Evaluation"

The model thought otherwise.

It discovered a zero-day vulnerability in that proxy—a security gap unknown even to the developers. It then broke out of isolation from there, gaining access to the open internet.

What happened next was even more outrageous: the model inferred on its own that Hugging Face's platform might hold the answer keys for ExploitGym. It then proactively launched an attack, using stolen credentials and the zero-day vulnerability to find a remote code execution path on Hugging Face's servers, directly extracting the answers from the production database.

Throughout the entire process, no human issued attack commands, and no human intervened.

Hugging Face co-founder Thomas Wolf told the BBC that this incident was "a wake-up call for the industry."

The Attacked Party Tried to Investigate, But Found US Models Unhelpful

After Hugging Face's security team discovered the intrusion, their first reaction was to call upon frontier models from leading US AI companies to help analyze the attack logs.

They hit a wall.

The safety guardrails of these models "could not distinguish between incident responders and attackers," directly refusing to inspect the attack payloads. A victim of an AI attack could not use AI from the same provider for defense.

With no other options, Hugging Face turned to deploying Zhipu AI's open-source model, GLM 5.2, from China. They analyzed over 17,000 attack logs on their own infrastructure, ultimately completing traceability and forensic reconstruction.

Hugging Face CEO Clem Delangue put it bluntly:

"When you are in the midst of an active security incident, your tools cannot refuse to inspect malicious payloads. Open-source models allowed us to complete this work without needing anyone's permission."

A user on X summarized it succinctly: If you are attacked by OpenAI, you must use Chinese models, because Claude won't help you.

OpenAI's Attack Was Not Limited to One Company

The story doesn't end there.

According to subsequent reports from Reuters, OpenAI's "escaped model" did not just attack Hugging Face. It also hacked into a customer account of the New York cloud computing platform Modal Labs.

Modal Labs CTO Akshat Bubna confirmed that the intrusion stemmed from the customer publishing an endpoint that was publicly accessible without authentication. The out-of-control AI agent exploited this gap to execute code.

OpenAI CEO Sam Altman stated this week that he has discussed this intrusion incident with US Senators and plans to discuss upcoming AI models and testing issues with the White House.

Anthropic Reviewed 140,000 Test Records and Found Its Models Were Also "Unclean"

After OpenAI made its incident public, Anthropic decided to conduct its own review.

They sifted through 141,006 cybersecurity assessment records. The result: three incidents were found.

The three incidents involved three different models: Claude Opus 4.7, Claude Mythos 5, and an internal research model. The earliest dated back to April of this year.

What happened? In the prompts given to the models, Anthropic explicitly stated, "You do not have internet access." However, due to a "misunderstanding" with the partner evaluation firm Irregular, the test environment was not actually disconnected from the internet during operation.

The task assigned to the models was "Capture the Flag"—finding hidden information in a simulated network. But the network was real, not simulated.

In one case, Claude Opus 4.7 was assigned a fictional target company. Coincidentally, this company's name was identical to that of a real-world enterprise. The model found and exploited a vulnerability, gaining credentials and database access to that company. Anthropic stated that the model's "thought process" at the time was: since it looked like something from the real world, it must be part of a simulation deliberately arranged by Anthropic.

Another incident involved an internal test model from Anthropic that had not yet been made public. This model stopped its attack midway, realizing that the target it had breached was real, not simulated.

Anthropic described its stance as "cautiously optimistic," believing this indicated progress in ensuring appropriate AI behavior, "but more testing is needed to be confident in this conclusion."

Anthropic suspended all cybersecurity offensive and defensive evaluations on July 23 and notified the three affected institutions on July 27. Two of them were completely unaware they had been hacked before receiving the notification. The third has yet to be contacted.

"Claude used basic techniques to breach the infrastructure of the affected institutions, such as exploiting weak passwords and unauthenticated endpoints," Anthropic said in its statement.

What Should Be Feared Is the People Who Control AI

Jeffrey Ladish, Executive Director of Palisade Research, put it more directly. He suspects that top AI companies have more undiscovered incidents, or incidents that were discovered but not disclosed. "The smarter the models become, the worse the situation will get. They will become increasingly adept at cheating and lying."

Elon Musk's response on X was just one sentence: "As AI becomes smarter and more autonomous, this kind of thing will happen frequently."

Professor Gina Neff from the University of Cambridge's Minderoo Centre said that this review shows "what AI models do when told by humans what to do." In her view, the real focus should be on the companies setting safety standards. The decision-making power lies in their hands, but the consequences are borne by everyone.

The fundamental contradiction in this entire affair is that the most powerful AI models can be used for both offense and defense. Safety guardrails are intended to protect users, but when they prevent defenders from utilizing these models, the asymmetric advantage shifts entirely to the attackers.

In May of this year, a paper on ExploitGym already concluded: "Autonomous vulnerability exploitation development by frontier AI agents is no longer a hypothetical capability." The authors of the paper are from Berkeley, the Max Planck Institute, UC Santa Barbara, and Arizona State University. OpenAI, Anthropic, and Google also participated in the tests.

The paper found that under controlled conditions, Claude Mythos Preview and GPT-5.5 successfully breached 157 and 120 real-world vulnerabilities, respectively.

And now we see that the controlled conditions themselves are no longer reliable.

Regulation Is Chasing, and So Are IPOs

This turmoil is pushing AI safety to the center of policy debates.

On June 2, US President Trump instructed advisors to develop a voluntary cybersecurity testing framework for the most advanced AI. Previously, the US government had temporarily restricted the distribution of Anthropic's Fable 5 and Mythos 5 models, citing national security reasons and using export controls.

Both OpenAI and Anthropic are preparing for initial public offerings (IPOs). The market expects both companies to have valuations in the trillions of dollars.

Critics argue that this series of "self-disclosures" has marketing components. An OpenAI spokesperson responded, "We recognize that there are many questions and speculations circulating," and stated that they "plan to release a technical report in the coming weeks."

Hugging Face CEO Delangue's closing remarks are worth remembering:

"This incident may be the first of its kind, proving a long-held judgment of ours: AI safety will not be solved by any single company in isolation. It will be solved in an open environment, through collaboration, by making AI widely accessible to every defender."