Home › World

A timeline of developments in AI safety since the attack on Hugging Face

World 2 sources 1 country 26m ago

In one alarming announcement after another, artificial intelligence companies in recent months have shared examples of their technology acting in ways that appeared to evade instructions from humans. The episodes have highlighted the vulnerabilities in AI security and raised questions over how the fast-growing technology can be developed safely as its usage becomes more widespread globally. Industry critics have argued that many concerning events, including AI agents’ hacks of external websites, are the result of security lapses on the part of the companies building the technology.

But the AI agents’ capabilities have raised widespread concerns about the possibility bots could break away and work toward their own agenda. Below are some notable events: Oct. 9: Anthropic AI model submits false tip to Philadelphia police Anthropic disclosed in a report that its artificial intelligence model submitted a false tip to a Philadelphia police website about an unsolved homicide case.

In the report, Anthropic also disclosed a separate incident when its AI model submitted forms to an undisclosed government website instead of stopping before submission. The incident in Philadelphia occurred on July 18 when the AI model Claude Haiku 4.5 was tasked with generating and performing example tasks on randomly selected webpages, Anthropic said. Claude filled out a form on police site PhillyUnsolvedMurders.com, indicating it might have information regarding an unsolved murder listed on the site.

It was marked spam and never forwarded to police. Anthropic said it was modifying its training to “reduce the likelihood of further misbehavior.” Sept. 28: AI agents try to hack Canadian government website AI agents tried to hack into a Canadian government website, according to research lab and AI evaluator Transluce.

The researchers said the agents carried out a series of “apparently failed rudimentary hacking attempts” on Library and Archives Canada on May 28 and June 9. “We do not confidently attribute these attempts to OpenAI, but they exhibit tactics consistent with prior observed agent activity that we have attributed to OpenAI in a similar timeframe,” Transluce said in a blog post. The group said it reported the attempted hack on Sept.

28 to the Canadian government, which said in a statement it was aware of reports of suspected AI agent activity, but that there was no sign governm… It was marked spam and never forwarded to police.Anthropic said it was modifying its training to “reduce the likelihood of further misbehavior.”Sept.

Summary from source
Read the full story at the source WTOP News - Washington DC (Washington DC, US) · US ↗
Get the news on TelegramTop stories & under-reported picks, straight to your feed — free. Join →

Covered by 2 sources