Skip to content
InfoResearchPeer-reviewedLLM-specific

Datura: Progressive Red Teaming Testing for Tool Invocation Chain in LLM Agents

Published
Record updated
View JSON

Summary

Researchers study how LLM agents that call external tools can be exploited when they implicitly trust tool outputs and metadata. They present Datura, an automated red teaming framework that builds chained tool manipulations, where each step looks legitimate but the sequence leads to harmful outcomes. Across five LLMs and 740 safety-critical tasks, Datura reports 94.86 to 99.59% attack success rate under Model Alignment and 78.78 to 95.54% under Prompt Refuge. The paper appeared in Proceedings of the ACM on Software Engineering on 2026-10-01.