Action-Level Backdoor Attacks Against Deep Reinforcement Learning Systems via Adaptive Reward Exploration
Summary
Researchers have demonstrated a new attack called Adapdoor that can inject hidden malicious behaviors into Deep Reinforcement Learning (DRL) models, which are AI systems trained to make sequential decisions in environments like robotics and autonomous vehicles. The attack works by poisoning the reward signal (the feedback that guides what the AI learns to do) during training, allowing attackers to later manipulate the model's actions when it is deployed in the real world. The paper shows this threat is more serious than previously thought because Adapdoor can work across many different tasks without requiring manual customization for each one.
Classification
Related Issues
Original source: http://ieeexplore.ieee.org/document/11675897
First tracked: September 21, 2026 at 08:04 PM
Classified by LLM (prompt v3) · confidence: 92%