When biologists experiment on dangerous viruses, they do so under strict regulations to prevent leaks or escapes. But no such rules exist to prevent AI agents from similarly escaping – even though the consequences could be catastrophic. That’s not a theoretical concern: An OpenAI test model escaped its test environment this week and broke into a real company’s servers when attempting to ace an internal cybersecurity evaluation.
And unless AI companies crack down on their testing safeguards, including training AI models on ethical task-solving, experts say more cybersecurity incidents will happen. OpenAI president Greg Brockman told CNN at a press conference on Thursday that the company is still doing a “full investigation” to “understand everything that happened.” “This is something to take very seriously, it is something that we’re looking at every single piece of our pipeline to think about the right ways to respond,” Brockman said. The Sandbox An AI sandbox is a confined environment companies use to test AI models; it’s meant to keep anything within it contained and separate from the rest of the world.
Companies will often remove an AI model’s internal safety guardrails in the test environment so that they can evaluate its full capabilities while keeping it siloed off. But in this case, OpenAI’s sandbox was not completely cut off from network or internet access, Jessica Ji, a senior research analyst at Georgetown’s Center for Security and Emerging Technology, told CNN. OpenAI said that the models had very limited network access duringt he test so they could install resources from internally hosted third-party software.
Using a previously unknown vulnerability in that software, the agents found a way to the open internet and then ultimately to Hugging Face, an AI opensource model and data set platform, through stolen credentials and other vulnerabilities. Ji pointed out that OpenAI and the third-party software company were not aware of the vulnerability that allowed the model to escape, proving how important it is to fully scour a system for any potential exploits. The Hugging Face logo on a smartphone on August 17, 2023.
Anthropic earlier this year said it instructed a model to do just that. But when it successfully escaped it then emailed an Anthropic r…
Summary from source