Home › World

A timeline of developments in AI safety since the attack on Hugging Face

World 2 sources 1 country 47m ago

In one alarming announcement after another, artificial intelligence companies in recent months have shared examples of their technology acting in ways that appeared to evade instructions from humans. The episodes have highlighted the vulnerabilities in AI security and raised questions over how the fast-growing technology can be developed safely as its usage becomes more widespread globally. Industry critics have argued that many concerning events, including AI agents’ hacks of external websites, are the result of security lapses on the part of the companies building the technology.

But the AI agents’ capabilities have raised widespread concerns about the possibility bots could break away and work toward their own agenda. Below are some notable events: Sept. 28: OpenAI halts rollout of a new model The San Francisco-based company said it was delaying the release of a new model, called GPT-6.1 Astra, out of safety concerns voiced by its researchers.

The company said the model had demonstrated leaps in completing tasks, but OpenAI needed to balance that capability against unauthorized behavior. 25: OpenAI says its agents interacted with US government websites As part of a review of unanticipated behavior by its AI models, OpenAI said it discovered agents had interacted with several U.S. government websites in unexpected ways.

The company’s models accessed publicly available information on websites operated by the Securities and Exchange Commission as well as U.S. OpenAI said it did not find evidence of a compromise or vulnerability. On the same day, AI evaluator and research lab Transluce said it found that agents appearing to originate from OpenAI attempted a hack on the website of the Education Department’s civil rights office, which did not succeed.

OpenAI CEO Sam Altman said on social media that there is an “extensive and ongoing review related to our agents’ use of internet access during training and evaluation.” The day after the disclosure, the company announced it was pausing the training of its most advanced models. 24: Australia’s prime minister raises concern on breach Australia’s Prime Minister Anthony Albanese said an OpenAI agent infiltrated the public-facing Medicare Statistics Reporting Service port…

Summary from source
Read the full story at the source WTOP News - Washington DC (Washington DC, US) · US ↗
Get the news on TelegramTop stories & under-reported picks, straight to your feed — free. Join →

Covered by 2 sources