Skip to content
InfoResearchPreprintLLM-specific

Formal Runtime Verification for Tool-Using LLM Agents: An Offline Same-Benchmark Study on AgentDojo and STAC

Published
Record updated
View JSON

Summary

This study evaluates metric first-order temporal logic (MFOTL) as a declarative guardrail for tool-using LLM agents by replaying recorded trajectories from AgentDojo, STAC and R-Judge through the unmodified MonPoly monitor, offline and without running an agent. Five generic obligations flag 71.8% of STAC attack chains and 70.1% of successful AgentDojo attacks, but also fire on 29.3% of benign runs, an imprecision the authors attribute to the corpora, which rarely record approvals and never record timestamps. A planted line evades a naive provenance check in 94-99% of the runs it would otherwise flag, and binding provenance to the lookup that produced it closes this evasion at no cost in detection or benign firing.

Mitigation

Binding provenance to the lookup that produced it closes this evasion at no cost in detection or benign firing. The authors also propose a twelve-field enforcement-ready trace schema.