Abstract
We propose JoyAI-RA 0.5, a generalist Vision-Language-World-Action framework that scales robot manipulation learning from heterogeneous human, simulation, and robot data through dual action alignment. Implicit alignment learns transferable physical dynamics from action-free video, while explicit alignment maps reliable human and robot trajectories into a unified action space for executable control. An inner-outer-loop reinforcement learning stage combines rapid task adaptation with continued foundation-policy improvement. On the Real-World AgiBot Benchmark, JoyAI-RA 0.5 achieves strong performance on seen tasks and unseen variations. Performance improves consistently as human egocentric pretraining data scales, with no saturation observed at the largest tested scale, establishing human video as a primary scaling axis for real-world manipulation.