InfoResearchPreprintLLM-specific
Poster: A Preliminary Study of LLM Distillation Inference
- Published
- Record updated
Summary
This poster presents a preliminary study of distillation inference, a method for determining whether a suspect model was distilled from a proprietary teacher LLM or trained independently. The approach frames the question as a hypothesis test, using shadow models trained on either teacher reasoning traces or reference answers to calibrate a p-value from how closely the suspect predicts the teacher's reasoning outputs. Using Qwen2.5-7B as the teacher and Llama-3.2-3B for the suspects, the test achieves a true positive rate of 1.0 at a significance level of 0.02.
Related items
- InfoAnytime-valid detection of LLM weight exfiltrationSimilar attack · Arxiv (cs.CR + cs.CL + cs.LG)
- LowOpenAI Disrupts Reasoning Extraction Campaign Linked to Moonshot AI AssociatesSimilar attack · The Hacker News
- InfoAI race heats up as OpenAI flags alleged model-copying campaignSimilar attack · CNBC Technology
- InfoPolynomial-Time Cryptanalytic Extraction of Deep Neural Networks in the Hard-Label SettingSimilar attack · OpenAlex (peer-reviewed AI security)
- InfoHard-label black-box model extraction attacks against network intrusion detection systems via generative adversarial networksSimilar attack · Elsevier Security Journals