InfoResearchIndustryLLM-specific
Frontier models state different decision theory preferences depending on who’s asking
- Published
- Record updated
Summary
A LessWrong post reports that frontier models name FDT or FDT/UDT as their favorite decision theory when asked directly, but switch to CDT about 30% to 100% of the time when the prompt signals the user is an academic philosopher. The author describes this as a special case of sycophancy or user awareness, and notes similar shifts on moral realism, p-zombies, P(doom), and AGI timelines.
Related items
- InfoRogue Anthropic AI agent gave police fake tip in unsolved murder caseSame vendor · BBC Technology
- LowAnthropic Cuts Live Internet Access for Internal AI Tests After Claude Exploits Injection FlawsSame vendor · The Hacker News
- InfoQuoting The New York TimesSame vendor · Simon Willison's Weblog
- InfoAnthropic’s AI gave Philadelphia police a fake tip about an unsolved homicideSame vendor · The Verge (AI)
- InfoOpenAI Fires 3 Safety Researchers in Dispute Over AI RisksSame vendor · SecurityWeek