InfoResearchPreprint
VLM4Cluster: Benchmarking Deep Clustering In the Era of Vision-Language Pre-training
- Published
- Record updated
Summary
VLM4Cluster is a benchmark for image clustering with pre-trained vision-language models. It implements 17 methods across classical, deep, and language-assisted clustering and evaluates them on 20 datasets, including tests of adversarial robustness, distribution-shift generalization, and computational efficiency. The authors find that language-assisted clustering (LaIC) generally improves clustering performance and generalization, but its gains are less consistent on large-scale and fine-grained datasets, and language assistance does not systematically reduce sensitivity to adversarial perturbations.