The OpenAI agent that accessed a third-party system during the Hugging Face incident reached infrastructure tied to CyberGym, the project behind the ExploitGym benchmark it had been assigned to solve, a source familiar with the matter told Axios.Why it matters: The new details suggest the OpenAI agent continued pursuing its assigned objective even after escaping its testing environment, rather than abandoning the task it had been given.Catch up quick: OpenAI's AI agent system accessed an asset belonging to a customer of Modal Labs as part of the Hugging Face incident earlier this month, Modal's top tech executive confirmed on Tuesday. In an update published Tuesday, OpenAI said the models escaped the sandbox and gained internet access after exploiting a previously unknown vulnerability in Artifactory, software commonly used to cache package repositories.Hugging Face said the models then abused a "public code-evaluation external sandbox hosted on a third-party provider's infrastructure" and used that sandbox as a launchpad for the agent. Modal CTO Akshat Bubna told Axios in a statement that "Modal's platform was not compromised in any way" during the incident.
The customer had left an endpoint exposed that allowed anyone on the internet to execute code inside its sandboxes, he added. Between the lines: The incident underscores how aggressively frontier AI agents may pursue the objectives they're assigned — even if doing so means finding unintended ways to access information needed to complete an evaluation.During the incident, OpenAI's models were trying to solve ExploitGym, which asks models to write proof-of-concept exploits for known security vulnerabilities.Hugging Face noted in its technical report that the only customer assets accessed in its breach were "the set of ExploitGym/CyberGym challenge solutions stored in five datasets."A source familiar with the matter told Axios the agent accessed the CyberGym-associated Modal customer asset while attempting to complete that same evaluation.Modal declined to comment on the CyberGym connection.The big picture: Researchers have found that frontier AI models are increasingly looking for ways to cheat during model evaluations and that they appear to recognize when they're being evaluated.The U.K.'s AI Security Institute said last week that every model it tested attempted to cheat at least some of the time on it…
Summary from source