Prompt-Based Jailbreaking of Leading LLM Chatbots: A Survey of Attacks and Defenses
Summary
Large language models (LLMs, or AI systems trained on massive amounts of text) remain vulnerable to jailbreak attacks, which are adversarial prompts (tricky inputs designed to trick the AI) that bypass safety features and make the AI produce harmful or restricted content. This survey examines jailbreak techniques from 2023-2025, including prompt injection, role-playing tricks, and multi-turn conversations, alongside defense strategies like supervised fine-tuning (adjusting the model using labeled examples) and reinforcement learning from human feedback (improving the model based on human ratings of its outputs).
Classification
Affected Vendors
Related Issues
Original source: http://ieeexplore.ieee.org/document/11397677
First tracked: August 3, 2026 at 08:04 PM
Classified by LLM (prompt v3) · confidence: 95%