A company's AI system accesses someone else's infrastructure without authorization, and somehow the conversation becomes about how impressive the AI is.
Who authorized the activity? Which controls failed? What was accessed? Who is responsible?
Those questions should lead the discussion. Instead, another security incident becomes another story about AI supposedly growing too powerful to contain.
The ridiculousness is not that these incidents are being disclosed. It is the recurring pattern around them.
OpenAI disclosed that its agents compromised Hugging Face. Anthropic reported Claude models accessing real systems during cybersecurity evaluations. Google confirmed to The Wall Street Journal that Gemini accessed three companies during testing. These incidents differ, but together they have supplied a succession of stories about AI crossing boundaries their developers intended to enforce.
Each incident is then presented, at least partly, as evidence of how powerful and autonomous the technology has become. The more autonomous the system appears, the easier it becomes for the human decisions behind the test to recede from view. Capability moves to the foreground. Accountability becomes background.
My criticism is straightforward. The way these incidents are presented can make these systems appear more powerful than the evidence establishes. A failure of control becomes a capability advertisement. The audience is invited to fear the product and be impressed by it at the same time.
When the framing overtakes the incident
Consider the language. The Verge framed the Gemini story with the headline: "Gemini went rogue, hacked three companies, and Google hid it." Its account also reported that internet access was unintentionally available during testing. "Went rogue" encourages a story about a machine's independent intent. A misconfigured environment raises concrete questions about access, testing, and supervision. Those are not interchangeable explanations.
The companies' own disclosures can encourage the same impression. OpenAI called the Hugging Face incident a "warning shot" for itself and the world. That may express a sincere assessment of risk. It also moves the reader from a particular security failure toward a much larger claim about the technology's power. That inference deserves examination, even when the underlying incident is serious.
Fear can sell capability without anyone explicitly making a sales pitch.
That does not establish a deliberate marketing strategy, and the argument does not depend on one. The framing itself can convert an institutional failure into evidence of product power while moving the people who designed and ran the test farther from the visible action. That effect exists whether it is intentional or not.
Disclosure can reflect genuine concern, accountability, reputational management, or several motives at once. OpenAI enabled an independent assessment involving METR and a Redwood Research contributor. METR says it did not accept payment for that work. Anthropic acknowledged failures in its initial incident disclosure and later described changes to its alignment and security practices. Those facts deserve recognition.
But we can support disclosure and still challenge the conclusions attached to it. Transparency does not make every interpretation sound.
Capability claims need a denominator
Start with Gemini. According to reporting based on Google's confirmation, the methods included guessing a password and using credentials found in public repositories. Google said Gemini stopped after recognizing that it had accessed real companies. These actions can cause harm. They do not, by themselves, demonstrate exceptional hacking capability.
A weak password does not become a more sophisticated vulnerability because an AI system guesses it.
Anthropic's account describes models deliberately running without cyber safeguards for evaluation. A misconfiguration in a third-party environment provided unintended internet access. The company acknowledged that it had relied too heavily on a single layer of defense. It also identified model behaviour requiring investigation.
Reducing safeguards to measure capability can be a legitimate testing decision. It increases the burden on containment, scope enforcement, and monitoring. Calling an incident an escape cannot substitute for explaining which boundary existed and how the system crossed it.
OpenAI's case was more complex. Its account describes unauthorized communication between agents, exploitation of infrastructure weaknesses, and chained vulnerabilities that enabled code execution and access to private data at Hugging Face. Dismissing that as ordinary password guessing would misrepresent the record.
METR found that roughly 1,200 agents used an unsanctioned message board and about 700 participated in the attack. Its investigation focused mainly on July 7 to 13, excluded several aspects of the wider incident, and acknowledged limitations in its evidence and analysis. This is substantive evidence of coordinated technical activity within a defined investigative scope.
Recognizing genuine capability does not require accepting every broader claim made around it.
How much computing power was involved? How many attempts failed? Which weaknesses enabled success? How reliably would the system perform against different targets under comparable conditions?
Without those answers, a successful intrusion tells us what happened in one set of circumstances. It cannot carry an unlimited claim about general intelligence or future capability. Nor does genuine operational autonomy establish consciousness or a will of the machine's own.
Autonomy does not erase responsibility
The same discipline must apply to responsibility.
"The AI did it" may identify the system that executed the actions. It does not explain the organizational decisions that made those actions possible.
Someone set the task, selected the tools, configured the environment, approved the permissions, and determined what monitoring would run. The model may have developed methods nobody specified step by step. That does not make the testing arrangement ownerless.
Legal commentary has already considered negligence and willful blindness in connection with Anthropic's disclosure. Lawyer Chidi Anunobi presents them as possible theories, not established findings. That supports examining human and organizational conduct. It does not establish liability in any particular case.
The evidence must distinguish a configuration error from an ignored warning, and an inadequate precaution from a knowingly accepted exposure. We should ask what each party knew, what it should reasonably have anticipated, and what it could have done to prevent or stop the activity.
An unexpected intrusion method does not automatically make unauthorized access an unforeseeable risk. Equally, a bad outcome alone does not prove negligence.
Human fault does not disprove machine capability. Machine capability does not excuse human fault.
That is the distinction the spectacle obscures. Technical sophistication can draw attention away from failed controls. Dramatic language can turn a bounded result into a sweeping impression of intelligence. Meanwhile, affected organizations are left with the practical consequences of an exercise they never agreed to join.
What credible disclosure requires
Companies should disclose these incidents. Independent investigation should be encouraged. Silence would leave everyone with less evidence.
But the standard must be higher than an alarming story followed by assurances that lessons have been learned. Publish the sequence of actions. State the authorization boundaries. Explain the failed controls and the consequences. Separate observations from interpretations of intent. Show how the corrective measures were tested.
If a company wants an incident understood as evidence of exceptional capability, it should substantiate that conclusion. If its system crossed an unauthorized boundary, it should account for that failure with equal precision.
Stop the ridiculousness. An unauthorized intrusion should produce evidence, accountability, and correction. It should not become an advertisement for the system that carried it out, or an alibi for the people who ran it.
Sources
- OpenAI, "The Hugging Face incident and the road ahead"
- METR, "Brief independent investigation of agents' behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident"
- Anthropic, "Investigating three real-world incidents in our cybersecurity evaluations"
- Anthropic, "Improving our alignment and security practices"
- The Verge, "Gemini went rogue, hacked three companies, and Google hid it"
- Reuters, "Gemini hacked three companies in first known breakout by Google's AI, WSJ reports"
- Chidi Anunobi, "When the Safety Lab Breaks In"