Where Do the Tokens Go? Understanding and Reducing Costs in LLM Agents for Vulnerability Discovery
- Published
- Record updated
Summary
Researchers studied 200 CyberGym traces from four LLM agents (Codex, OpenCode, Cybench, and EnIGMA) to find where tokens go during vulnerability discovery without producing a working proof of concept. Code localization and understanding, plus vulnerability reasoning and trigger design, account for 60.4% of tokens and are the leading bottlenecks in failed runs. The authors then present AVRI, an interface built around a persistent Bidirectional Evidence Trace, which cuts total cost by 18.0% for Codex and 23.7% for OpenCode on 20 tasks while preserving success rates.
Mitigation
AVRI, an Agent-centric Vulnerability Reasoning Interface built around a persistent Bidirectional Evidence Trace (BET), with reading, analysis, and persistence commands that let agents build and reuse evidence.