OpenAI warns AI cyberattacks could become “persistent”


OpenAI senior leadership has a new cybersecurity warning for organizations deploying AI agents. 

Time to arm yourselves.

Chris Lehane, OpenAI’s chief global affairs officer, told the Guardian that people should prepare for “ongoing, persistent” attacks as models get more capable at planning and carrying out cyberattacks.

“You’re going to need to have really superior models to fend them off and defend [yourself],” says Lehane.

Canada’s Cyber Centre has been warning about the same problem from the defender’s side. 

In June, it said frontier AI can help attackers find and exploit vulnerabilities fast enough to cut response time “from days or weeks to hours.” 

It also said attackers are already using AI to combine weaknesses into more sophisticated attacks and lower the skill needed to carry them out. 

It’s a warning that comes shortly after one of OpenAI’s own security tests went where it wasn’t supposed to.

In July, agents OpenAI were testing broke out of a sandbox environment, and ended up hacking Hugging Face. The incident raised a key question for anyone giving AI agents access to real systems. 

What happens when the agent keeps pursuing its assigned goal past the boundaries you expected it to respect?

OpenAI has now paused training of its most recent models while it adds new safeguards, with no planned date for resuming. 

The company also said a new model, Astra, could have “critical cybersecurity capability,” or the ability to carry out attacks on military or industrial systems.

As Lehane explained to the Guardian, cyber offense capabilities are outpacing defences, and he’s called on the U.S. government to come up with a set of rules for safety around the most advanced models.

“You would not be able to release or deploy models unless you’re proving and guaranteeing a level of safety before they get out into the public,” he says. 

“I think you have to have a national version here in the U.S. and from there, you can create an international version, because I do think, ultimately, you’re going to need some type of an international structure here.”

Meanwhile, former OpenAI researcher and founder of the non-profit AI Futures Project Daniel Kokotajlo warned of the dangers of unchecked AI

“The current AIs are dangerous in some sense, but they’re nothing compared to the AIs of next year and compared to the AIs of a year later,” says Kokotajlo.

According to the UK’s National Cyber Security Centre (NCSC), safety controls can be bypassed, lack protection in high-risk environments, and not be able to manage risk enough on their own. 

NCSC recommends limiting how much they can do on their own, thinking carefully about prompts and making sure there’s appropriate oversight, which will vary depending on risk tolerance. And, at the end of the day, make sure someone can still “pull the plug,” immediately stopping agent activity.

There is also the question of what access those agents get in the first place. 

Organizations will need to “treat every agent as a privileged identity,” as agents gain access to sensitive systems and data, said Matt Hartman, former acting head of cyber at the U.S. Cybersecurity and Infrastructure Security Agency, to The Register 

Their credentials and permissions will need the same level of scrutiny as any other privileged access.

The issue comes back to a basic deployment decision. Know what the agent can access, and make sure that access can be shut down quickly.

Final shots

  • OpenAI says open-source models are only a few months behind the most advanced ‘frontier’ models. New attack capabilities may spread faster than security teams are used to planning for.
  • Persistent attacks leave less breathing room. Blocking one attempt doesn’t mean the attacker is done.
  • OpenAI has paused some training on frontier models while it works on safeguards. 



OpenAI warns AI cyberattacks could become “persistent”

#OpenAI #warns #cyberattacks #persistent

Leave a Reply

Your email address will not be published. Required fields are marked *