InfoResearchPreprint
Inverting Multi-Vector Visual Document Indices
- Published
- Record updated
Summary
This research shows that multi-vector visual document retrievers can be inverted, letting an attacker who runs or breaches the vector store reproduce source pages from the stored index alone. On the ViDoRe v3 benchmark, inverted pages recover 47% of words and rank their source page first 98.4% of the time. Token pooling and shuffling cut word recall to about 8%, but a model that restores shuffled order raises first-rank matches from 3.8% to 93.5%.
Related items
- HighGHSA-6wjp-v33h-5cvq: PraisonAI: AgentOS defaults to network-exposed no-auth mode, allowing unauthenticated agent invocation and instruction disclosureSimilar attack · GitHub Advisory Database
- HighCVE-2026-101998: Docker Sandboxes fail open when masking credentials in proxy responsesSimilar attack · NVD/CVE Database
- InfoAnytime-valid detection of LLM weight exfiltrationSimilar attack · Arxiv (cs.CR + cs.CL + cs.LG)
- InfoA novel privacy-preserving large language model integrating trust-weighted and ethical gradient maskingSimilar attack · OpenAlex (peer-reviewed AI security)
- InfoSystemic privacy risks of personal data exposure through conversational large language model agentsSimilar attack · OpenAlex (peer-reviewed AI security)