Skip to content
LowResearchPeer-reviewedLLM-specific

Evaluating Safety Embedding Prefiltering for Analyzing Millions of LLM Agent Social Network Messages for Security and Safety Harms

Published
Record updated
View JSON

Summary

The study examines whether lightweight embedding-based prefilters can cut the cost of screening agentic social network content for security and safety harms. Researchers annotated 10,000 Moltbook posts and comments with a frontier LLM judge, finding 9.2% unsafe at severity 3 or above. At 0.80 recall, prefiltering reduced the projected cost of scanning 787,226 messages by 49.9% to 65.6%, depending on the encoder or trained classifier, though jailbreak content remained the hardest category to retrieve.