InfoResearchPreprintLLM-specific
Anytime-valid detection of LLM weight exfiltration
- Published
- Record updated
Summary
Researchers propose a prompt-level e-process to detect LLM weight exfiltration by a compromised inference server, which hides payload bits in token choices. The method calibrates whole-response mismatch events on trusted benign traffic and accumulates evidence across responses. It is evaluated on four models against a seed-blind attack and a stronger seed-aware attack, and it provides anytime false-alarm control compared with a hard per-token alarm.
Related items
- HighGHSA-6wjp-v33h-5cvq: PraisonAI: AgentOS defaults to network-exposed no-auth mode, allowing unauthenticated agent invocation and instruction disclosureSimilar attack · GitHub Advisory Database
- HighCVE-2026-101998: Docker Sandboxes fail open when masking credentials in proxy responsesSimilar attack · NVD/CVE Database
- InfoPoster: A Preliminary Study of LLM Distillation InferenceSimilar attack · Arxiv (cs.CR + cs.CL + cs.LG)
- InfoA novel privacy-preserving large language model integrating trust-weighted and ethical gradient maskingSimilar attack · OpenAlex (peer-reviewed AI security)
- InfoSystemic privacy risks of personal data exposure through conversational large language model agentsSimilar attack · OpenAlex (peer-reviewed AI security)