Skip to content

Hy-Embodied-0.5-VLA: From Vision-Language-Action Models to a Real-World Robot Learning Stack

Hy-Embodied-0.5-VLA (HyVLA-0.5) is a Tencent Hunyuan technical report presenting an end-to-end Vision-Language-Action system that spans the full robot learning stack: data collection, model design, continued pre-training, supervised fine-tuning, RL post-training, and real-world deployment [Abstract]. The framing positions the work as a stack rather than a single architecture or recipe — explicitly naming each pipeline component as a distinct contribution. The paper continues the Tencent “Hy-” rebrand (formerly Hunyuan) into embodied AI, putting Hunyuan alongside Qwen-Robot Suite, π0.5, and Embodied-R1.5 as a vertically-integrated open VLA stack. Only the abstract was retrievable at filing time; method/results details are not extractable here.

  • Hy-Embodied-0.5-VLA is presented as an end-to-end system spanning data collection, model design, continued pre-training and SFT, RL post-training, and real-world deployment [Abstract].
  • Each component of the stack is claimed to serve a distinct role rather than being collapsed into a single training objective [Abstract].

The retrievable content is limited to the abstract, which describes HyVLA-0.5 as an end-to-end stack rather than disclosing architectural specifics. The report covers, in order: data collection, model design, continued pre-training, SFT, RL post-training, and deployment [Abstract]. Body sections (architecture, action head, RL algorithm, benchmark setup) are not extractable from the arxiv landing page; a refresh once the full PDF is parseable would fill in the method specifics.

Numerical results are not retrievable from the abstract.

This is the third major industrial VLA technical report filed in a ~30-day window after Qwen-RobotManip Technical Report: Alignment Unlocks Scale for Robotic Manipulation Foundation Models (Alibaba Qwen-Robot Suite, Qwen2.5-VL backbone + manipulation/navigation/world-modeling stack) and π*0.6: a VLA That Learns From Experience (RECAP) (Physical Intelligence’s π0.6 with RECAP for RL on flow-matching action experts) — Tencent now joining Alibaba and PI in publishing a stack rather than a single model. The “RL post-training” component is the most relevant axis to watch given the live debate on the VLA Models page about whether RECAP-style advantage conditioning (π*0.6: a VLA That Learns From Experience (RECAP) §IV-B) is flow-matching-specific or generalizes; HyVLA-0.5’s RL recipe (once the PDF is parseable) would be a third datapoint after π0.6 and the community Evo-RL reproduction (Evo-RL: Open Real-World Offline RL on SO-101 and AgileX PiPER). The “Hy-” branding is consistent with prior Tencent embodied releases such as HY-World 2.0: A Multi-Modal World Model for Reconstructing, Generating, and Simulating 3D Worlds (HY-World 2.0 multi-modal world model) and HY-World 1.5 (WorldPlay): A Systematic Framework for Interactive World Modeling with Real-Time Latency and Geometric Consistency (HY-World 1.5 / WorldPlay interactive world model), suggesting HyVLA-0.5 is the action-policy half of an internal stack mirroring the Qwen-Robot Suite’s manip + world-model bifurcation.