Skip to content
InfoResearchPreprintLLM-specific

Anytime-valid detection of LLM weight exfiltration

Published
Record updated
View JSON

Summary

Researchers propose a prompt-level e-process to detect LLM weight exfiltration by a compromised inference server, which hides payload bits in token choices. The method calibrates whole-response mismatch events on trusted benign traffic and accumulates evidence across responses. It is evaluated on four models against a seed-blind attack and a stronger seed-aware attack, and it provides anytime false-alarm control compared with a hard per-token alarm.