Disrupting a coordinated model-distillation campaign
- Published
- Record updated
Summary
OpenAI disrupted a coordinated campaign, first observed in early July, that sought to extract protected reasoning from its models through adversarial distillation. Operators manipulated model interactions, including copying encrypted reasoning from one conversation and asking a model in another to decrypt and transcribe it, rather than breaking encryption or accessing stored conversations. A July 24 and 25 spike reached 16,000 requests from over 4,000 users, and the wider cluster of more than 15,000 users was fully disrupted by July 28. OpenAI attributes a core cluster of the activity to individuals associated with Moonshot AI, the developer of Kimi.
Mitigation
OpenAI banned or restricted fraudulent accounts, strengthened signup and infrastructure controls, expanded monitoring for related networks, and strengthened protections for hidden reasoning across users, workspaces, organizations, and model families. It also closed a pathway that allowed someone who already possessed reasoning content to extract it further. Additional mitigation and investigation work is continuing.
Related items
- InfoOpenAI Fires 3 Safety Researchers in Dispute Over AI RisksSame vendor · SecurityWeek
- Info‘Pure insanity’: Mathematicians will need years to make sense of OpenAI’s latest dropSame vendor · The Verge (AI)
- InfoOpenAI reports three new incidents of misalignmentSame vendor · CSO Online
- InfoOpenAI's revenue scare, Delta earnings, what investors think of a Starbucks-Chipotle deal and more in Morning SquawkSame vendor · CNBC Technology
- InfoAnthropic bans users from being 'cruel' to its AI systemsSame vendor · BBC Technology