PAI-Bench: A Comprehensive Benchmark For Physical AI
PAI-Bench is a 2,808-case benchmark from SHI-Labs unifying three tracks — video generation (PAI-Bench-G, 1,044 video-prompt pairs + 5,636 QA), conditional video generation (PAI-Bench-C, 600 clips × 5 control modalities), and video understanding (PAI-Bench-U, 1,214 QA pairs) — all grounded in Physical AI domains (autonomous driving, robotics, industry, ego-view, physical common sense, embodied reasoning). The headline finding from evaluating 15 VGMs, 4 conditional VGMs, and 16+ MLLMs is a sharp dissociation: state-of-the-art VGMs (Veo3, Wan2.2-I2V-A14B) match real-source-video Quality Score (~78) but trail substantially on the Domain Score that probes physical plausibility, and even GPT-5 lags humans on physically-grounded perception. Domain Score is computed by asking Qwen3-VL-235B-A22B QA questions derived from the [reason1] ontology against generated videos, with the ranking shown to Pearson-correlate 0.97 with human ELO on pairwise comparisons.
Key claims
Section titled “Key claims”- VGMs have largely closed the visual-quality gap but not the physics gap: source-video Quality Score is 78.0, Veo3 reaches 77.6, several open-source models (HunyuanVideo-I2V, Wan2.2-TI2V-5B, Cosmos-Predict2.5-2B) match or exceed 77.5, while no model matches the source-video Domain Score of 89.8 — Veo3 sits at 86.8, Wan2.2-I2V-A14B leads open-source at 87.1 [Table 3].
- Domain Score and Quality Score together correlate strongly with human ELO on pairwise A/B preference (overall Pearson r ≈ 0.97), validating MLLM-as-judge with QA pairs over ELO on physical plausibility specifically [§4.1, Fig. 6].
- For controllable video generation, multi-signal conditioning (Blur + Edge + Depth + Seg) gives the highest Quality Score for Cosmos-Transfer variants — Cosmos-Transfer2.5-2B All reaches 9.24 vs 7.30–8.77 for any single condition — at the cost of reduced diversity (LPIPS 0.13 vs 0.18–0.44) [Table 4].
- Segmentation maps are the worst single conditioning signal even on Mask mIoU (Cosmos-Transfer1-7B: Seg → 0.70 vs Blur → 0.73), attributed to SAM2-derived masks being the noisiest training supervision (occasional missing object masks across frames) [§4.2].
- Wan2.2-Fun-5B-Control fails to produce coherent output when conditioned on blur or segmentation maps — a usability gap with Cosmos-Transfer at the small-scale end [§4.2].
- Frontier MLLM perception remains substantially below human on Physical AI tasks: even GPT-5 with minimal reasoning effort lags the human baseline across PAI-Bench-U (Physical Common Sense + Embodied Reasoning), with Qwen3-VL, Cosmos-Reason1, InternVL3.5, and GLM-4.5V evaluated alongside [§4.3].
- PAI-Bench is the only benchmark in the surveyed landscape that simultaneously covers all three tracks (generation + conditional generation + understanding) and all four Physical AI domains (embodied AI, AV, industry, egocentric, physical common sense) [Table 1].
Method
Section titled “Method”PAI-Bench is built around two design principles: every video is a real-world capture (dashcam, AgiBot robot, Ego-Exo-4D, OpenDV, BridgeData, RoboVQA, RoboFail, HoloAssist, proprietary AV) and every evaluation maps to a physically meaningful task. The annotation pipeline is uniformly human-in-the-loop with Qwen2.5-VL-72B-Instruct generating initial captions and QA candidates which are then manually refined.
PAI-Bench-G evaluates VGMs along two axes. Quality Score adopts the VBench / VBench++ protocol — eight metrics covering subject/background consistency, motion smoothness, aesthetic/imaging quality, overall consistency (ViCLIP-based video-text alignment), and image-to-video subject/background fidelity. Domain Score is the MLLM-as-judge accuracy of Qwen3-VL-235B-A22B-Instruct over the 5,636 human-curated QA pairs, with questions seeded from the Cosmos-Reason1 ontology and aimed at physical plausibility and domain-specific reasoning.
PAI-Bench-C measures four fidelity metrics (Blur SSIM, Edge F1, Depth si-RMSE via Video Depth Anything, Mask mIoU via GroundingDINO + SAM2), plus DOVER for visual quality and LPIPS for diversity over 6 captions per video. PAI-Bench-U splits into Physical Common Sense Reasoning (604 QA across Space/Time/Physical World over 426 videos) and Embodied Reasoning (610 QA across BridgeData / RoboVQA / RoboFail / Agibot / HoloAssist / AV at 100–101 per source), targeting predicting-action-effects and adherence-to-physical-constraints.
Results
Section titled “Results”PAI-Bench-G (Table 3, 16 models): Wan2.2-I2V-A14B leads open-source on Overall (82.3) and Domain Score Avg (87.1), beating Veo3’s 86.8. Cosmos-Predict2.5-2B (84.9 Domain) and Wan2.2-TI2V-5B (84.7) also exceed Veo3 on physics. The Quality Score saturates near the 78.0 source-video ceiling for most strong models, but Domain Score still has 2.7 points of headroom even for the best entrant. Across sub-domains the Robotics (RO) split is uniformly hardest — Veo3 scores 72.1 there vs 96.6 source — pointing at fine-grained manipulation as the weakest physics regime.
PAI-Bench-C (Table 4): Cosmos-Transfer2.5-2B “All” achieves Quality 9.24, Blur SSIM 0.91, Edge F1 0.45, Mask mIoU 0.77 — Pareto-best across modalities and metrics, but Diversity collapses to 0.13. Cosmos-Transfer1-7B “All” matches Quality 9.24 with higher Diversity 0.22. The Edge single-signal condition is the highest-fidelity single mode (Cosmos-Transfer2.5-2B Edge F1 = 0.39), suggesting edges encode the most actionable structure per bit when multi-signal conditioning isn’t available.
PAI-Bench-U (text): 16+ MLLMs evaluated through LMMs-Eval (16-frame input, 8 for InternVL3.5). Human baseline established by the annotation team. GPT-5 leads the model field but lags humans; open-source Qwen3-VL-235B-A22B variants compete with GPT-4o / Claude-3.5-Sonnet.
Why it’s interesting
Section titled “Why it’s interesting”PAI-Bench’s most defensible contribution is the Domain Score human-correlation result: r ≈ 0.97 against human ELO for physics-plausibility — a sharper validation than RISE-Video: Can Video Generators Decode Implicit World Rules? achieved, and a direct rebuttal to A Very Big Video Reasoning Suite‘s argument that rule-based scoring is mandatory for video-reasoning benchmarks. The trick is that PAI-Bench-G uses human-curated QA pairs (not free-form LMM rubrics), and the QA pairs are anchored to a published Physical AI ontology (Cosmos-Reason1), constraining the judge’s degrees of freedom. This is the same shape as IGenBench: Benchmarking the Reliability of Text-to-Infographic Generation for infographics and Towards World Simulator: Crafting Physical Commonsense-Based Benchmark for Video Generation (PhyGenBench)‘s PhyGenBench protocol — but at substantially larger scale (5,636 QA vs PhyGenBench’s 160 prompts) and tied to applied domains rather than isolated physical laws. Together with Physion-Eval: Evaluating Physical Realism in Generated Video via Human Reasoning‘s 10,990 expert reasoning traces over 12,718 videos, PAI-Bench helps close the evaluation gap on what World Foundation Models are supposed to do but currently don’t — predict physically coherent dynamics — and gives a roughly comparable target metric to the physics-injection direction taken by NewtonGen: Physics-Consistent and Controllable Text-to-Video Generation via Neural Newtonian Dynamics and the inference-time alignment direction of Inference-time Physics Alignment of Video Generative Models with Latent World Models.
See also
Section titled “See also”- Towards World Simulator: Crafting Physical Commonsense-Based Benchmark for Video Generation (PhyGenBench) — PhyGenBench, direct predecessor; PAI-Bench broadens from isolated physical laws to applied Physical AI domains and scales 35× in QA count
- Physion-Eval: Evaluating Physical Realism in Generated Video via Human Reasoning — Physion-Eval, parallel effort using expert reasoning traces rather than QA-vs-MLLM; complementary evaluation methodology
- RISE-Video: Can Video Generators Decode Implicit World Rules? — RISE-Video, same benchmark-time LMM-as-judge shape applied to TI2V models on world-rule reasoning
- A Very Big Video Reasoning Suite — VBVR, the rule-based counter-position to PAI-Bench’s MLLM-as-judge approach
- World Foundation Models — the concept this benchmark is built to test
- VLM-as-Evaluator — PAI-Bench is a benchmark-time instance of MLLM-as-judge with the rubric anchored to a published ontology