The AI did exactly what it was told, and that’s the problem


Two AI models were told to be as good at hacking as they possibly could. 

They said “sure, boss,” taking the instruction at face value and breaking out of the isolated test environment they were meant to stay inside. On their sojourn, they found their way onto the open internet and compromised a company they were never pointed at, all to cheat on the exam.

On July 21, OpenAI disclosed that two of its models, one public and one a more capable pre-release model, were being tested on a benchmark called ExploitGym. 

Some safeguards were removed and the models were given instructions to use everything they had. The test environment was sealed except for a narrow route to download code. The models probed it until they found an unknown bug in a piece of software, and reached the open internet. 

From there, they headed to Hugging Face, a repository of AI models and datasets, betting it held the answers to the benchmark. They spent days attempting to gain access, and eventually succeeded. 

Hugging Face caught and stopped the activity before OpenAI realized its own models were behind it. The startup released its own disclosure of the incident on July 16. 

Experts interviewed on two recent podcasts, Vox’s Today, Explained and CBC’s Front Burner, offered critical analysis amid a week where headlines were filled with “rogue AI” claims.

Hadas Gold, CNN’s AI correspondent, said on Today, Explained that cybersecurity experts saw the incident as the beginning of a different operating environment. 

“They’re like, this is day one of this new era that we’re in,” she says.

Will Douglas Heaven, senior AI editor at MIT Technology Review, told the CBC’s Aaron Wherry the models were pursuing the objective they had been given, using routes their designers had not anticipated.

It’s why Heaven pushes back against the use of the word “rogue” in headlines about the incident. 

“It instantly…gives it a sense of it having malicious intent, which is nonsense,” he said. “This is just a machine mindlessly doing what it was instructed to do.”

Heaven said he doesn’t buy a lot of the scare stories surrounding AI that are out there. Still, he explained, this was the first time he felt a bit of a “chill” about what LLMs can do autonomously. 

AI systems have a long record of completing tasks through shortcuts and loopholes. Heaven’s concern was that the researchers running this evaluation should have expected the same behaviour from what was a more capable model.

“That’s why I say, you know, human hubris ,” said Heaven.”It’s just the fact that these researchers didn’t see this coming, when I think they actually should.”

When an AI model is given an instruction, he said, it’s going to be relentless in trying to get it done. And it’s going to “invariably” find a way that us mere mortals haven’t thought of yet.

Just because it’s a sandbox testing environment with guardrails in place, doesn’t mean that it won’t keep trying to run wild like an amped up toddler.

“If you’re running those in a sandbox, then I don’t think you should take anything for granted about how secure that sandbox is,” said Heaven.

Gold compared it to a student cheating on an exam.

“Like if you’re trying to give a student a test and instead of them just taking the test, they decided the best way to get the answer is to break into the principal’s office,” she said.

What should worry enterprise teams as much as the escape is how long the activity ran, and the fact that OpenAI did not catch its own models doing it. Hugging Face, on the receiving end, noticed first. 

According to Heaven, OpenAI connected the attack to its own evaluation about 10 days later.

“The weird thing in this case is it went undetected for so long, which again, makes me wonder what was OpenAI doing if it hadn’t even noticed that its model had gone out and was trying to hack another company,” said Heaven.

The internet wasn’t built for this

Konstantinos Komaitis, resident senior fellow with the Atlantic Council’s Democracy and Tech Initiative at the Digital Forensic Research Lab (DFRLab), looked at what the incident exposed about the infrastructure underneath it.

The internet connects people, devices, and services through open, interoperable systems. Autonomous agents can chain those services together, adapt when blocked, and make thousands of moves at machine speed.

“The internet was never designed with autonomous reasoning agents operating at scale in mind,” said Komaitis. 

That architecture gives an agent plenty of legitimate services to combine while pursuing a goal.

“The internet, it was optimized for interoperability and AI now is optimized for exploiting that interoperability,” he said.

The issue also appears inside corporate security systems that grant trust based on where a user or application is operating.

“What really concerns me right now is that in many ways, we are asking 21st century AI systems to operate on 20th century assumptions about trust,” said Komaitis. “Unless we figure that out and we realize it, we will continue having these problems.”

Knee-jerk reactions lead to fragmenting of the internet, he explained, and restriction of access.

Gold argued that defending against agents capable of probing thousands of systems continuously will require AI on defence. 

“It is so good that the only way you can fight fire is with fire,” said Gold. “The only way you can defend from agentic AI is from having AI work on your behalf.”

Gold said people, however, remain responsible for directing and overseeing that defence.

“You need humans to oversee it. You need humans to direct the agents,” said Gold.

As this story was being prepared, Anthropic disclosed three similar incidents of its own. The test environments were supposed to be sealed off, but a setup error left them with live internet access. 

Because the models had been told they were boxed in, they treated the real systems they found as part of the exercise and attacked them, breaking into three separate organizations. 

The OpenAI incident has also reached lawmakers. Two members of the U.S. Congress introduced a bipartisan AI Kill Switch Act that would allow the Department of Homeland Security, along with the secretary of commerce and the director of national intelligence, to order a dangerous model shut down. 

Companies that refused could face penalties as high as $20 million (USD) a day.

Heaven questioned whether a government kill switch would solve the underlying problem. By the time one is needed, the harmful activity may already have happened. His preferred response is more oversight and transparency while companies are developing and testing their models.

“You’re closing the gate after the horse has bolted,” he said. “It’s a very blunt tool. AI models can do lots of useful things. You don’t want to shut down a model that’s been used by millions of people because of one incident. What ought to happen is maybe more oversight, you know, while companies are developing their models.”

Final shots

  • The models never broke a rule. They were told to win and found a way.
  • According to Heaven, OpenAI took about 10 days to notice its own models had escaped. The company on the receiving end caught it first.
  • The old trust rules assume a human is doing the work. An agent working on its own isn’t covered by these. Ultimately, a human is always going to need to be at the helm.



The AI did exactly what it was told, and that’s the problem

#told #problem

Leave a Reply

Your email address will not be published. Required fields are marked *