PromptFishing: Active Hallucination Inducement to Distinguish LLMs From Humans
Summary
Researchers developed PromptFishing, a technique to identify whether social media accounts are controlled by LLMs (large language models, AI systems trained on text) or humans by embedding hidden prompts into normal conversations that cause LLMs to produce distinctive responses called hallucinations (when an AI generates false or off-topic information). The method uses a two-stage algorithm to create adversarial suffixes (specially crafted text additions) that trigger these distinctive LLM behaviors while appearing like normal conversation, achieving over 85% accuracy at detecting LLMs while falsely flagging humans less than 1% of the time.
Classification
Related Issues
CVE-2024-27444: langchain_experimental (aka LangChain Experimental) in LangChain before 0.1.8 allows an attacker to bypass the CVE-2023-
CVE-2026-30308: In its design for automatic terminal command execution, HAI Build Code Generator offers two options: Execute safe comman
Original source: http://ieeexplore.ieee.org/document/11592605
First tracked: July 24, 2026 at 08:03 PM
Classified by LLM (prompt v3) · confidence: 92%