Skip to content
InfoResearchPreprint

Does Adversarial Training Improve Generalization in Multi-View VLAs? Revealing and Mitigating View Collapse

Published
Record updated
View JSON

Summary

This research asks whether adversarial training improves generalization for multi-view vision-language-action (VLA) robot policies adapted from a pretrained VLM, evaluated across seven LIBERO-Plus shift axes. Direct adversarial training substantially improves Camera Viewpoint and Sensor Noise but gives mixed or negative effects on other shifts, and it can cause view collapse, where the policy becomes dominated by the wrist view. Adding a View Swap intervention further improves Camera Viewpoint, Sensor Noise, and Robot Initial State, with mixed effects elsewhere.

Mitigation

The source mitigates view collapse with a simple View Swap intervention, applied before re-evaluating adversarial training.