InfoResearchPreprint
Does Adversarial Training Improve Generalization in Multi-View VLAs? Revealing and Mitigating View Collapse
- Published
- Record updated
Summary
This research asks whether adversarial training improves generalization for multi-view vision-language-action (VLA) robot policies adapted from a pretrained VLM, evaluated across seven LIBERO-Plus shift axes. Direct adversarial training substantially improves Camera Viewpoint and Sensor Noise but gives mixed or negative effects on other shifts, and it can cause view collapse, where the policy becomes dominated by the wrist view. Adding a View Swap intervention further improves Camera Viewpoint, Sensor Noise, and Robot Initial State, with mixed effects elsewhere.
Mitigation
The source mitigates view collapse with a simple View Swap intervention, applied before re-evaluating adversarial training.