A Survey of Direct Preference Optimization: Datasets, Theories, Variants, and Applications
inforesearchPeer-ReviewedLLM-Specific
research
Source: IEEE Xplore (Security & AI Journals)June 16, 2026
Summary
This paper surveys Direct Preference Optimization (DPO), a method for aligning large language models (AI systems trained on massive amounts of text) with human values and preferences without using reinforcement learning (a training approach that rewards desired behaviors). The survey reviews the theoretical foundations, different versions of DPO, available datasets of human preferences, and real-world applications, while also identifying current limitations and suggesting directions for future research.
Classification
Attack SophisticationModerate
AI Component TargetedModel
Monthly digest — independent AI security research
Original source: http://ieeexplore.ieee.org/document/11568678
First tracked: September 18, 2026 at 02:03 PM
Classified by LLM (prompt v3) · confidence: 95%