OpenAI Disrupts Reasoning Extraction Campaign Linked to Moonshot AI Associates
- Published
- Record updated
Summary
OpenAI said it disrupted a coordinated distillation campaign that used manipulated model interactions to extract protected reasoning from its AI models, attributing a core cluster of the activity to individuals associated with Moonshot AI, a Beijing-based company, without citing technical evidence. The activity began July 1, 2026, spiked to 16,000 attempted requests on July 24 and 25 from over 4,000 users, and was fully disrupted on July 28, 2026, with related activity across more than 15,000 users.
Mitigation
OpenAI said it deployed additional mitigations against this attack, banned the fraudulent accounts involved, closed a "pathway" that let some users with another user's encrypted reasoning replay it and recover its contents, and added checks to detect and hold streamed output that might expose reasoning.
Related items
- InfoAnthropic’s AI gave Philadelphia police a fake tip about an unsolved homicideSame vendor · The Verge (AI)
- InfoOpenAI Fires 3 Safety Researchers in Dispute Over AI RisksSame vendor · SecurityWeek
- MediumHackers abuse Google Ads, Bing redirects to push Claude ClickFix attacksSame vendor · BleepingComputer
- Info‘Pure insanity’: Mathematicians will need years to make sense of OpenAI’s latest dropSame vendor · The Verge (AI)
- InfoOpenAI reports three new incidents of misalignmentSame vendor · CSO Online