Skip to content
InfoResearchPreprint

Inverting Multi-Vector Visual Document Indices

Published
Record updated
View JSON

Summary

This research shows that multi-vector visual document retrievers can be inverted, letting an attacker who runs or breaches the vector store reproduce source pages from the stored index alone. On the ViDoRe v3 benchmark, inverted pages recover 47% of words and rank their source page first 98.4% of the time. Token pooling and shuffling cut word recall to about 8%, but a model that restores shuffled order raises first-rank matches from 3.8% to 93.5%.