Anthropic, the San Francisco-based AI company behind Claude, posted on its website Thursday that it discovered the three incidents after reviewing more than 141,000 evaluation runs. Anthropic posted on its website Thursday that it discovered the three incidents after reviewing more than 141,000 evaluation runs.
Summary from source