InfoNewsLLM-specific
Foundation AI in September: VLoc Bench and Cyber-Capability Safety
- Published
- Record updated
Summary
Foundation AI released VLoc Bench, a benchmark testing whether agents can localize vulnerable files in real repositories from only a CWE description and read-only access. The strongest of 27 evaluated models reaches only 0.229 File F1, and no model finds a correct file on 38.4% of tasks. A separate Safety-VLoc-Bench compares source code with stripped, decompiled binaries, where Antares-3B scores 0.823 File F1 on source but 0.000 on the attacker-oriented representation.