Text Adversarial Attacks With Dynamic Outputs
Summary
This research describes a new attack method called TDOA (Textual Dynamic Outputs Attack) that can trick large language models by exploiting a weakness in how they handle variable outputs. Unlike older attack methods that assume a fixed set of possible answers, real-world LLMs often generate answers that go beyond predefined categories or produce different numbers of labels depending on the input, creating what researchers call 'dynamic outputs.' TDOA works by using a clustering approach (grouping similar outputs together) to convert these unpredictable outputs into a simpler form that existing attack techniques can target, achieving up to 80.8% success rates with very few queries.
Classification
Affected Vendors
Related Issues
Original source: http://ieeexplore.ieee.org/document/11653428
First tracked: September 21, 2026 at 08:04 PM
Classified by LLM (prompt v3) · confidence: 92%