A2Net: Affiliation Alignment Networks for Whole-Body Pose Estimation With Vision--Language Models
Summary
This paper introduces A2Net, a system for whole-body pose estimation (predicting the locations of keypoints on a person's face, body, hands, and feet from an image) that combines vision and language models to solve two problems: scale variation (different body parts appearing at different sizes) and semantic ambiguity in small-scale parts (difficulty identifying what small features represent). The approach uses text features alongside image features because text is not affected by image scaling issues, then aligns them using optimal transport (a mathematical method for matching distributions) to create a unified visual-language representation that improves keypoint localization accuracy.
Classification
Original source: http://ieeexplore.ieee.org/document/11373586
First tracked: August 23, 2026 at 02:01 AM
Classified by LLM (prompt v3) · confidence: 75%