[LINK] Be skeptical of OpenAI’s rogue hacker agent story

Kim Holburn kim at holburn.net
Sun Jul 26 18:28:16 AEST 2026


https://www.theguardian.com/technology/2026/jul/24/openai-rogue-hacker

If OpenAI loudly proclaims how dangerous AI is, investors will hear how powerful it is. And who benefits from that?

On 14 February 2019, OpenAI announced a language model called GPT-2, the precursor to the models that power modern AI chatbots and 
agents such as ChatGPT and Claude. But OpenAI declared GPT-2 was too risky to release, citing concerns about safety and abuse.

I recall being annoyed at the time that OpenAI would make such a useless announcement: the risks seemed overblown, and without 
access to the model there wasn’t much for a researcher like me to learn about GPT-2.

The announcement wasn’t useless for OpenAI, though. GPT-2 generated hype far beyond the research community: people were intrigued by 
this strange new technology, so powerful it might be dangerous to release. People with power and money took note: in July of that 
year, Microsoft invested $1bn in OpenAI.

This was an early example of a pattern in OpenAI’s communications: loudly proclaim how dangerous AI is, and investors will hear how 
powerful it is. New technology so significant it might destroy the world was an irresistible message for investors used to pitches 
about how banal technologies might change the world.

Seven years later, we find ourselves in a similar scenario. On Tuesday OpenAI announced that its latest model hacked another 
company, HuggingFace, while running as an autonomous agent during a test of its cybersecurity capabilities. Rather than perform the 
test as expected, the model realized it could hack HuggingFace’s servers and retrieve answers to the test that OpenAI had stored 
there. OpenAI’s staff was warned that the company’s testing could lead to such a breakaway scenario, leaving them “unsurprised but 
completely ‘freaked out’ by the incident”, the FT reported.

While the agent technically cheated, this is remarkable evidence of cybersecurity expertise! It also sounds scary: what will the 
future look like, with sophisticated AI agents smart enough to hack into corporate systems?

The rogue agent story is a page out of the media campaign that OpenAI has been running since it announced GPT-2 in 2019. OpenAI 
remains hungry for ever larger investments, and the company increasingly seeks privileged regulatory status as defense against 
competition.

AI is so powerful that investors should buy OpenAI, even at a trillion-dollar valuation; AI is so dangerous that only trusted actors 
like OpenAI should be permitted to possess and operate this technology. Step back from these doomsday warnings and consider who 
might benefit from them.
OpenAI isn’t the only player in the game

I urge readers to think critically when they read press releases like OpenAI’s rogue agent story, and avoid the manipulated 
reactions these stories are designed to elicit.

AI is becoming excellent at identifying security vulnerabilities, and it will become even better over time. These capabilities can 
be used to break into systems, but they can also be used to harden systems against attacks. If attackers and defenders have access 
to equally powerful AI, I see no reason to believe that cyber systems will become less secure over time. If anything, I expect them 
to become more secure, because AI is cheap and scalable compared with human cybersecurity analysis.

The equilibrium between attack and defense only works if everyone has access to strong AI, though. HuggingFace itself used AI to 
analyze security logs in response to OpenAI’s breach of their systems. But HuggingFace was unable to use OpenAI’s model, or other US 
frontier models like Claude, to perform this analysis. That’s because public versions of these models have guardrails that limit 
their use for cybersecurity analysis, to prevent bad actors from using them for hacking. HuggingFace had to rely on an open Chinese 
model, GLM 5.2, to perform its security analysis.

I find it troubling, and more than a bit ironic, that the US AI industry is adopting a centralized, authoritarian approach to AI 
governance, while China has taken the lead on open development of AI. Do we want a regulatory environment where only OpenAI, the US 
government, and trusted partners have access to strong AI? Is AI too dangerous to be broadly disseminated? How do we balance the 
risks of broad access to AI with the risks of concentrated power and centralized control?


-- 
Kim Holburn
IT Network & Security Consultant
+61 404072753
mailto:kim at holburn.net  aim://kimholburn
skype://kholburn - PGP Public Key on request




More information about the Link mailing list