Sensitive-Topic Leakage Through LLM Routing Metadata: Measurement and Mitigation
- Published
- Record updated
Summary
Researchers measured whether LLM routers leak sensitive-topic information through their model-selection metadata, even when content logging is off. Across 1.7 million real requests, routers sent harassment and self-harm prompts to the strong model 19 points less often, and medical prompts 31 points less often on distinct prompts (both post hoc). Per-category length-matched parity with accurate labels removed the gap on real traffic at a cost of at most 0.2 accuracy points on RouterBench.
Mitigation
Per-category length-matched parity with accurate labels removes the gap on real traffic, costing at most 0.2 accuracy points on RouterBench (post hoc). Per-conversation stickiness, per-user budget bands, and pooled parity failed. A post hoc exact per-user rate hides even-prefix strong counts but preserves odd-position decisions, so it does not fully close the channel.
Related items
- LowGHSA-3gh4-cghq-f8v4: Pydantic AI OpenTelemetry instrumentation: retry prompt content is not redacted when `include_content=False`Similar attack · GitHub Advisory Database
- LowGHSA-4x9p-g9wm-8q7f: Pydantic AI OpenTelemetry instrumentation: exception events on tool and agent run spans include content when `include_content=False`Similar attack · GitHub Advisory Database
- LowSeptember 2026 Cyber Threat Landscape: Global Attacks Jump 48% as Phishing and GenAI Data Exposure RiseSimilar attack · Check Point Research
- InfoA novel privacy-preserving large language model integrating trust-weighted and ethical gradient maskingSimilar attack · OpenAlex (peer-reviewed AI security)
- InfoSystemic privacy risks of personal data exposure through conversational large language model agentsSimilar attack · OpenAlex (peer-reviewed AI security)