EgoLAP: Learning from Egocentric Human Data through Language-Action Reasoning
TL;DR: We introduce EgoLAP, a vision-language-action pre-training framework that learns from egocentric human and robot trajectories through shared language actions and motion-level reasoning. EgoLAP transfers human experience to robot control, achieving 80.1% mean real-world task progress on an unseen robot configuration and a 2.3× gain over alternative action representations.