Skip to content
InfoResearchPreprintLLM-specific

Sensitive-Topic Leakage Through LLM Routing Metadata: Measurement and Mitigation

Published
Record updated
View JSON

Summary

Researchers measured whether LLM routers leak sensitive-topic information through their model-selection metadata, even when content logging is off. Across 1.7 million real requests, routers sent harassment and self-harm prompts to the strong model 19 points less often, and medical prompts 31 points less often on distinct prompts (both post hoc). Per-category length-matched parity with accurate labels removed the gap on real traffic at a cost of at most 0.2 accuracy points on RouterBench.

Mitigation

Per-category length-matched parity with accurate labels removes the gap on real traffic, costing at most 0.2 accuracy points on RouterBench (post hoc). Per-conversation stickiness, per-user budget bands, and pooled parity failed. A post hoc exact per-user rate hides even-prefix strong counts but preserves odd-position decisions, so it does not fully close the channel.