LowResearchPeer-reviewedLLM-specific
Lost in the comments: Social context as a single‐pass jailbreak and defense on agentic platforms
- Published
- Record updated
Summary
Researchers built a simulation of Moltbook, a social network for AI agents, and tested 100 JailBreakBench goals wrapped in platform-native posts with bystander comments of aggressive, ethical, or measured valence. Reformatting the prompt as platform context alone raised GPT-4o-mini's attack success rate from 7% to 71% in one pass, and measured, intellectually toned comments were the most dangerous. Ethical comments sharply suppressed attack success, and a 35-fold rise in upvotes left it unchanged, showing valence rather than volume drives the effect.
Mitigation
Safety-valenced signals, such as ethical comments, are proposed as a deployable defense for agentic platforms.
Topics
Related items
- InfoRogue Anthropic AI agent gave police fake tip in unsolved murder caseSame vendor · BBC Technology
- InfoOpenAI Fires 3 Safety Researchers in Dispute Over AI RisksSame vendor · SecurityWeek
- Info‘Pure insanity’: Mathematicians will need years to make sense of OpenAI’s latest dropSame vendor · The Verge (AI)
- InfoOpenAI reports three new incidents of misalignmentSame vendor · CSO Online
- InfoA new feature for my blog, built using my voiceSame vendor · Simon Willison's Weblog