Skip to content

H2R-Bench: Benchmarking Human-to-Robot Manipulation Video Generation in World Models

H2R-Bench is a benchmark for evaluating whether video world models can transform egocentric human manipulation demonstrations into robot-centric manipulation videos under a specified target embodiment. Each instance pairs a human video with target embodiment constraints and source-grounded annotations (task goal, action events, functional contacts, object responses); generated outputs are scored along five dimensions: goal-state completion, action-event completion, functional contact transfer, embodiment correctness, and general video quality. Eleven state-of-the-art video generation models are evaluated across six manipulation families and two robot embodiments. The headline finding is that current video world models remain limited at cross-embodiment human-to-robot transfer — even leading models routinely fail on embodiment consistency, functional interaction, and task execution.

  • Cross-embodiment human-to-robot video generation is under-evaluated by existing benchmarks and is the specific capability H2R-Bench targets: transforming egocentric human demos into robot-centric videos under specified embodiment constraints [§Abstract].
  • Evaluation is decomposed along five dimensions — goal-state completion, action-event completion, functional contact transfer, embodiment correctness, and general video quality — enabling separation of surface fidelity from cross-embodiment transfer capability [§Abstract].
  • Each benchmark instance ships source-grounded annotations covering task goals, action events, functional contacts, and object responses, giving graders explicit targets for each of the five dimensions rather than relying on holistic VLM judgment [§Abstract].
  • Eleven state-of-the-art video generation models are evaluated across six manipulation families and two robot embodiments [§Abstract].
  • Current video world models remain limited at human-to-robot transfer: leading models often fail on embodiment consistency, functional interaction, and task execution [§Abstract].

H2R-Bench frames cross-embodiment video generation as an evaluation problem rather than a modeling one. Each item bundles: (a) an egocentric human demonstration video as the source, (b) a target embodiment specification, and (c) source-grounded annotations for task goal, action events, functional contacts, and object responses. Models under test consume the human video plus embodiment specification and emit a robot-centric video; graders score the output on five axes (goal-state completion, action-event completion, functional contact transfer, embodiment correctness, general video quality). The suite covers six manipulation families and two robot embodiments, and eleven state-of-the-art video generation models are benchmarked.

The five-dimension evaluation across eleven models × six manipulation families × two embodiments diagnoses that current video world models fail on embodiment consistency, functional interaction, and task execution — surface video quality does not translate into successful cross-embodiment human-to-robot transfer. The paper positions H2R-Bench as a systematic diagnostic framework for whether video world models can bridge the human-to-robot embodiment gap and turn human manipulation observations into robot-centric training resources.

H2R-Bench opens a new axis inside Video Generation Benchmarks that no filed entry currently occupies: cross-embodiment human-to-robot video-to-video transfer, evaluated with source-grounded contact/action annotations rather than pure VLM-as-judge. It sits adjacent to Rethinking Video Generation Model for the Embodied World (ReVidGen / RBench / RoVid-X)‘s RBench (which scores robot-video generation from text/image on 5 tasks × 4 embodiments) and to GigaWorld-1: A Roadmap to Build World Models for Robot Policy Evaluation‘s WMBench (which scores policy-evaluation-oriented rollout fidelity) but changes the input channel: the source is a human demonstration, and the evaluated capability is the embodiment-gap transfer itself. It also gives the Human-to-Robot Retargeting cluster its first dedicated video-model benchmark — most of that page’s entries are training/interface recipes (kinematic retargeting, reduced-DoF bridging actions, physics-in-loop optimization) rather than evaluation of video-generative retargeting.