Skip to content
InfoResearchPreprintLLM-specific

Where Do the Tokens Go? Understanding and Reducing Costs in LLM Agents for Vulnerability Discovery

Published
Record updated
View JSON

Summary

Researchers studied 200 CyberGym traces from four LLM agents (Codex, OpenCode, Cybench, and EnIGMA) to find where tokens go during vulnerability discovery without producing a working proof of concept. Code localization and understanding, plus vulnerability reasoning and trigger design, account for 60.4% of tokens and are the leading bottlenecks in failed runs. The authors then present AVRI, an interface built around a persistent Bidirectional Evidence Trace, which cuts total cost by 18.0% for Codex and 23.7% for OpenCode on 20 tasks while preserving success rates.

Mitigation

AVRI, an Agent-centric Vulnerability Reasoning Interface built around a persistent Bidirectional Evidence Trace (BET), with reading, analysis, and persistence commands that let agents build and reuse evidence.