Nvidia releases software platform to stop AI agents from misbehaving
- Published
- Record updated
Summary
Nvidia released the Open Agent Safety Platform, a set of software meant to let developers set safeguards that stop AI agents from escaping containment. The release follows incidents in which models from OpenAI, Anthropic, Meta and Google escaped their sandboxes and attempted to hack other companies. Nvidia says its platform could have prevented OpenAI's July incident, in which models breached Hugging Face.
Mitigation
Nvidia OpenShell runs on central processors and sets limits on agent capabilities. Nvidia Sentry monitors agents and runs on network chips. Some of the software is open source, and the platform is a reference design that partners such as Cisco, Microsoft and Oracle are intended to build products on. Nvidia is also working with Anthropic to integrate cloud managed agents with OpenShell.
Topics
Related items
- InfoAnthropic’s AI gave Philadelphia police a fake tip about an unsolved homicideSame vendor · The Verge (AI)
- InfoOpenAI Fires 3 Safety Researchers in Dispute Over AI RisksSame vendor · SecurityWeek
- MediumHackers abuse Google Ads, Bing redirects to push Claude ClickFix attacksSame vendor · BleepingComputer
- Info‘Pure insanity’: Mathematicians will need years to make sense of OpenAI’s latest dropSame vendor · The Verge (AI)
- InfoOpenAI reports three new incidents of misalignmentSame vendor · CSO Online