Skip to content
LowResearchPeer-reviewedLLM-specific

Lost in the comments: Social context as a single‐pass jailbreak and defense on agentic platforms

Published
Record updated
View JSON

Summary

Researchers built a simulation of Moltbook, a social network for AI agents, and tested 100 JailBreakBench goals wrapped in platform-native posts with bystander comments of aggressive, ethical, or measured valence. Reformatting the prompt as platform context alone raised GPT-4o-mini's attack success rate from 7% to 71% in one pass, and measured, intellectually toned comments were the most dangerous. Ethical comments sharply suppressed attack success, and a 35-fold rise in upvotes left it unchanged, showing valence rather than volume drives the effect.

Mitigation

Safety-valenced signals, such as ethical comments, are proposed as a deployable defense for agentic platforms.