InfoResearchPeer-reviewedLLM-specific
Datura: Progressive Red Teaming Testing for Tool Invocation Chain in LLM Agents
- Published
- Record updated
Summary
Researchers study how LLM agents that call external tools can be exploited when they implicitly trust tool outputs and metadata. They present Datura, an automated red teaming framework that builds chained tool manipulations, where each step looks legitimate but the sequence leads to harmful outcomes. Across five LLMs and 740 safety-critical tasks, Datura reports 94.86 to 99.59% attack success rate under Model Alignment and 78.78 to 95.54% under Prompt Refuge. The paper appeared in Proceedings of the ACM on Software Engineering on 2026-10-01.
Related items
- CriticalHermes Agent - PKCE Session Takeover via Redirect-URI Parser ConfusionSimilar attack · Tenable Research Advisories
- LowLost in the comments: Social context as a single‐pass jailbreak and defense on agentic platformsSimilar attack · OpenAlex (peer-reviewed AI security)
- MediumGHSA-hmq2-7hp6-7crh: Banks: User-controlled prompt input can be parsed as privileged chat messagesSimilar attack · GitHub Advisory Database
- HighGHSA-6wjp-v33h-5cvq: PraisonAI: AgentOS defaults to network-exposed no-auth mode, allowing unauthenticated agent invocation and instruction disclosureSimilar attack · GitHub Advisory Database
- HighCVE-2026-101998: Docker Sandboxes fail open when masking credentials in proxy responsesSimilar attack · NVD/CVE Database