Robust Reinforcement Learning via Leveraging Historically Optimal Policy With Regulation of Performance
inforesearchPeer-Reviewed
safetyresearch
Source: IEEE Xplore (Security & AI Journals)March 13, 2026
Summary
This research proposes HORP, a method to make reinforcement learning (RL, where AI systems learn by trial-and-error to maximize rewards) more resistant to adversarial attacks (manipulations designed to fool the AI). HORP improves robustness by using a previously learned optimal policy (the best strategy found so far) to guide learning, creating diverse attack scenarios, and adjusting how much uncertainty to introduce during training to help the AI defend itself better.
Classification
Attack SophisticationAdvanced
Impact (CIA+S)
safety
AI Component TargetedModel
Monthly digest — independent AI security research
Original source: http://ieeexplore.ieee.org/document/11434530
First tracked: September 3, 2026 at 08:02 PM
Classified by LLM (prompt v3) · confidence: 85%