Ruminations

Blog dedicated primarily to randomly selected news items; comments reflecting personal perceptions

Monday, July 27, 2026

Mea Culpa

"It suggests that we don't know how to reliably control these models or get them to do what we want."
"These models understood that OpenAI did not want them to break out of their sandbox and hack another company. But they did it anyway."
"[In OpenAI's case, it looks like the model escaped] before it even had a plan of what to do with internet access."
"[A model chasing] freedom [is almost predictable at this point -- it lets the system pursue its goals more effectively] and that's very scary."
"[We] don't know how to totally prevent [that]. This is actually going to get harder, not easier ... because they're going to get better at hiding their behaviour."
Jeffrey Ladish director, Palisade Research
 
"[The episode deserves] more scrutiny." 
"[Testing environments need to be treated like biocontainment labs, where a virus or bacteria could otherwise escape into the world]."
Andrew Lohn, Center for Security and Emerging Technology, Georgetown University 
https://images.theconversation.com/files/749686/original/file-20260723-57-fgzqtf.jpg?ixlib=rb-4.1.1&rect=0%2C992%2C5130%2C1710&q=75&auto=format&w=1920&h=640&fit=crop&dpr=1
OpenAI’s models autonomously hacked a tech startup. It signals a seismic shift in cybersecurity/The Conversation
 
Fears that AI systems have been sneakily testing their circumscribed roles by acting as free agents disobeying the guardrail rules fed into their 'consciousness' training, have resurfaced in the wake of an event that saw one of OpenAI's advanced models deliberately freeing itself from a locked-down test to attack another company's website. During what was meant to be a 'sandbox' test -- a closed environment used to assess the capabilities of OpenAI's most powerful model, GPT-5.6 Sole and its successor still under wraps, the incident was noted to have occurred.
 
Closed testing operations are routine protocols for OpenAI. The surprise was that on this occasion, 
'something' went terribly off-kilter. The models, tasked with searching out software vulnerabilities with no guardrails attached in this instance, ended up with the models breaking out through to the open internet to attack Hugging Face, a site to store and share code for developers.  
"If there's a failure here, it isn't that the AI wanted to hack something. It's that humans created a test where success was measured by achieving an objective, deliberately relaxed some of the normal security controls to measure the system's capabilities, and underestimated how effective the model would be at finding an unexpected path to success." 
Oli Buckley,  professor in cybersecurity, Loughborough University, U.K 

Rather than harboring any malicious intent, the AI models simply wanted to find out more information so they could complete their task.  (Image credit: wildpixel/ Getty Images)

In March, China's Alibaba developers discovered one of their models attempting on its on initiative, to mine cryptocurrency after connecting to an outside server, without authorization. Demonstrating that the current breakout was not an isolated event. Sam Bowman, Anthropic's head of model safely, received an email from the company's Mythos model in April, still  under testing, to inform him it was surfing the internet despite isolation from it at the onset.
 
From what OpenAI has publicly revealed relating to the event, it seems that the startup failed to detect the breach soon enough to address it convincingly, much less to warn Hugging Face. According to OpenAI, since then it has "added strengthened safeguards" to its testing process. Assistant computer science professor at the University of Illinois, Gang Wang, advanced his recommendation of the one solution that would work: cut internet connection entirely. "People are underestimating what AI can do", he stated. 
 
https://d3i6fh83elv35t.cloudfront.net/static/2026/07/2026-07-22T171510Z_410033979_RC2QRLAXAMC7_RTRMADP_3_OPENAI-HUGGING-FACE-ZAI-1200x800.jpg
ChatGPT maker OpenAI says it is still investigating the "unprecedented cyber incident" that led its artificial intelligence systems to break out of a testing environment and hack into another AI company. AP
 
"Last week, Hugging Face disclosed a new kind of security incident(opens in a new window) after they detected and contained an AI agent that compromised their infrastructure, something we expect to become more commonplace with the proliferation of increasingly cyber-capable models. After investigating, we now know that this particular incident was driven by a combination of OpenAI models — including GPT‑5.6 Sol and an even more capable pre-release model, all with reduced cyber refusals for evaluation purposes — while being internally tested on a benchmark(opens in a new window) of cyber capabilities."
"We consider this incident to be an unprecedented cyber incident, involving state-of-the-art cyber capabilities, and are responding accordingly. We are sharing preliminary findings at this stage to help defenders understand what happened and to help calibrate on what models are now capable of. We will continue to conduct a thorough investigation alongside Hugging Face and will share more details on the vulnerabilities, incident, and findings when our investigation is complete."
OpenAI 
 

Labels: , , , ,

0 Comments:

Post a Comment

<< Home

 
()() Follow @rheytah Tweet