InfoNews
AI’s quiet safety gatekeepers are stepping into the spotlight
- Published
- Record updated
Summary
AI labs Anthropic and OpenAI are turning to small third-party evaluators such as METR, Apollo Research and Transluce to assess model capabilities and risks, as federal regulation remains absent. Open questions include how these nonprofits will be funded, what access they will receive and how reporting will work. OpenAI fired three employees for violating its policies on handling sensitive company information, and two of them said the dismissals related to their communication with third-party evaluators.
Related items
- InfoAnthropic is cutting off its internal evaluations from the internetSame vendor · The Verge (AI)
- MediumARTEX AI, Claude agents used in cyberattacks on South Korean banksSame vendor · BleepingComputer
- InfoAI agent makers are promising privacy — will they deliver?Same vendor · The Verge (AI)
- InfoRogue Anthropic AI agent gave police fake tip in unsolved murder caseSame vendor · BBC Technology
- LowAnthropic Cuts Live Internet Access for Internal AI Tests After Claude Exploits Injection FlawsSame vendor · The Hacker News