NVIDIA Launches Cosmos 3, the Open Frontier Foundation Model for Physical AI
Cosmos 3 is NVIDIA’s headline GTC-Taipei drop: an open frontier omnimodel for physical AI that natively spans text, images, video, ambient sound, and action (numerical joint angles / gripper positions / trajectory points). It ships as Cosmos 3 Super (32B, post-training-grade) and Cosmos 3 Nano (8B, real-time policy use), with a Cosmos 3 Edge variant promised. Architecturally it’s a mixture-of-transformers that pairs a reasoning transformer with an expert generation transformer; the framing is that a single omnimodel can replace the current fragmented stack of separate VLMs, video generators, and policy heads used in robotics and AV training. Released under OpenMDW 1.1 on Hugging Face + build.nvidia.com + GitHub + NIM, with a Cosmos Coalition (Agile Robots, Black Forest Labs, Generalist, LTX, Runway, Skild AI) and reported #1 placements on Artificial Analysis open-weights, Physics-IQ, R-Bench, PAI-Bench, VANTAGE-Bench, RoboLab, RoboArena.
Key claims
Section titled “Key claims”- Cosmos 3 is positioned as the first fully open omnimodel with native multimodal reasoning across text, images, video, ambient sound, and actions, with action as a generated output modality (joint angles, gripper positions, trajectory points) [§Announcement, §What Cosmos 3 Is].
- The architecture is a mixture-of-transformers pairing a reasoning transformer with an expert generation transformer, intended to let the model “think before it acts” — reason over object interactions, motion, and spatial–temporal relationships before emitting video or action trajectories [§A New Architecture for Physical AI].
- Two variants are available at launch: Cosmos 3 Super (32B), targeted at robotics / AV post-training that needs the highest physics accuracy and generation quality, and Cosmos 3 Nano (8B), targeted at fast video and action reasoning in fractions of a second; Cosmos 3 Edge for real-time on-device inference is promised but not yet released [§Availability].
- Cosmos 3 Nano post-trained as a policy is reported to lead RoboLab (simulation, language-guided tasks) and RoboArena (real-world DROID robots) [§Cosmos 3 in robotics].
- Cosmos 3 is reported as the top-ranked open vision-language model on VANTAGE-Bench (smart-infrastructure scene understanding) and TAR (traffic anomaly reasoning) [§Vision AI agents].
- Cosmos 3 variants are reported to rank first on Artificial Analysis open-weights leaderboards and top Physics-IQ, R-Bench, and PAI-Bench for world generation [§Leaderboards].
- The release is governed by the Linux Foundation’s OpenMDW 1.1 license — a single model-centric license covering weights, architecture, documentation, datasets, benchmarks, and code [§License].
- A “Cosmos Coalition” is launched as a governance / shared-infrastructure layer with Agile Robots, Black Forest Labs, Generalist, LTX, Runway, and Skild AI as founding members, with contributions back to Cosmos training tools and DGX Cloud [§Cosmos Coalition].
- Distribution is multi-surface at launch: build.nvidia.com playground, Hugging Face weights, GitHub physical-AI agent skills, NIM microservices, and inference partners (Baseten, CoreWeave, Microsoft Azure, Nebius, Deep Infra, Classmethod) [§Availability].
Method
Section titled “Method”Architecturally the announcement names a mixture-of-transformers with two cooperating experts: a reasoning transformer that ingests interleaved multimodal context (text, images, video, ambient sound) and produces an explicit reasoning trace, and a generation transformer that emits video, action trajectories, and dense captions conditioned on that trace. The action head is numerical — joint angles, gripper positions, trajectory points — which is what differentiates Cosmos 3 from prior open WFMs like Cosmos Predict / Transfer (video out only) and from omnimodels like Emu3.5 or BAGEL (text+image+video+audio but no action). The announcement does not disclose parameter splits between the two transformers, training data composition, training compute, or context length; a technical report is implied but not linked at announcement time.
Downstream use cases the announcement explicitly calls out: (i) synthetic data generation for robotics and AV training, with the model emitting physically plausible video sequences including collisions and long-tail edge cases that are dangerous to capture in the real world; (ii) policy fine-tuning per embodiment / camera layout / workspace, with Cosmos 3 Nano serving as the post-trained policy backbone (NVIDIA’s own GEAR team uses Cosmos 3 for video action models across games, simulation, and real-world robotics); (iii) vision-AI agents in smart-city / industrial settings — Linker Vision is named as analyzing live camera streams via Cosmos 3 vision-language reasoning.
Results
Section titled “Results”Quantitative numbers are sparse — the announcement is a launch post, not a tech report. The reported placements:
- #1 on Artificial Analysis open-weights leaderboards [§Leaderboards].
- Tops Physics-IQ, R-Bench, and PAI-Bench among world-generation benchmarks [§Leaderboards].
- Cosmos 3 Nano post-trained policy leads RoboLab (simulation, language-guided tasks) and RoboArena (real-world DROID robots) [§Cosmos 3 in robotics].
- Top-ranked open VLM on VANTAGE-Bench (smart-infrastructure scene understanding) and TAR (traffic anomaly reasoning) [§Vision AI agents].
No baseline numbers are quoted in the post, no closed-flagship comparison is shown, and no per-benchmark Cosmos 3 score is given. The framing is “best open” rather than “best overall” — consistent with the open-foundation-releases pattern of Kimi K2.5 / Qwen3-VL where headline tables sometimes claim parity with closed models but here Cosmos 3 is positioned only against the open-weights cohort.
Why it’s interesting
Section titled “Why it’s interesting”Cosmos 3 is the first filed datapoint where the World Foundation Models stack collapses from “Cosmos Reason 2 (VLM) + Cosmos Predict 2.5 (video gen) + Cosmos Transfer 2.5 (sim-to-real) + Isaac GR00T N1.7 (VLA policy head)” — the architecture pattern laid out in NVIDIA Unveils New Open Models, Data and Tools to Advance AI Across Every Industry just six months prior — into a single omnimodel that does perception, generation, and action emission with a shared reasoning trace. That’s a substantial reframing of NVIDIA’s own architectural bet: the previous stack treated WFMs as a backbone with VLA task-heads bolted on; Cosmos 3 puts the action head inside the omnimodel. Whether this collapse beats the modular stack on transfer to new embodiments — the claim implicit in the RoboArena leaderboard placement — will be the headline question once a technical report and independent reproductions land.
For the Unified Multimodal Models cluster, Cosmos 3 extends the omnimodel surface to action as a generated output modality, alongside text, image, video, and audio — a strict superset of recent open omnimodels (Emu3.5, Ming-Flash-Omni, LongCat-Flash-Omni, Qwen3.5-Omni) and a direct extension of Haoxiang’s “BAGEL for … +action” framing in the slack note. It also pairs naturally with HY-World 2.0: A Multi-Modal World Model for Reconstructing, Generating, and Simulating 3D Worlds (which lacked any Cosmos Predict 2.5 / Genie 3 head-to-head comparison) and with Data-regularized Reinforcement Learning for Diffusion Models at Scale (NVIDIA Cosmos’s DDRL post-training recipe), which now has an obvious successor model to apply to.
The Cosmos Coalition is the more interesting strategic move: by enrolling Black Forest Labs, Runway, and LTX — three labs the wiki already tracks as closed-flagship video generators (FLUX.2 [klein]: Towards Interactive Visual Intelligence, Introducing Runway Labs, LTX-2: Efficient Joint Audio-Visual Foundation Model) — into a shared physical-AI ecosystem, NVIDIA is bidding to make the open WFM substrate dependent on its tooling (DGX Cloud, training stack) even when the downstream products stay proprietary. Mirrors the Cosmos Coalition framing in NVIDIA Unveils New Open Models, Data and Tools to Advance AI Across Every Industry but with more concrete commercial members.
See also
Section titled “See also”- World Foundation Models — Cosmos 3 collapses the WFM-backbone + VLA-head stack into one omnimodel; sharpens the cluster’s TL;DR axis (i) first-party WFM releases.
- Open foundation-model releases — OpenMDW 1.1 license + multi-surface release (HF/build.nvidia.com/GitHub/NIM) extends the NVIDIA multi-domain bundle pattern with a coalition-governance layer.
- Unified Multimodal Models — adds action as a generated output modality on top of the existing text/image/video/audio omnimodel surface.
- NVIDIA Unveils New Open Models, Data and Tools to Advance AI Across Every Industry — the prior NVIDIA omnibus release (Cosmos Reason 2 / Predict 2.5 / Transfer 2.5 / Isaac GR00T N1.6 / Alpamayo) that Cosmos 3 supersedes.
- Data-regularized Reinforcement Learning for Diffusion Models at Scale — DDRL post-training recipe from NVIDIA Cosmos; obvious next application is Cosmos 3.
- HY-World 2.0: A Multi-Modal World Model for Reconstructing, Generating, and Simulating 3D Worlds — Tencent’s HY-World 2.0 explicitly listed Cosmos Predict 2.5 as a comparison it didn’t run; Cosmos 3 raises that bar.
- PAN: A World Model for General, Interactable, and Long-Horizon World Simulation — PAN frames interactable + long-horizon world simulation as the open problem Cosmos 3 implicitly claims to address with its reasoning+generation expert split.
- Awesome World Models — the cross-domain WFM taxonomy in which Cosmos 3 occupies the “physical AI + embodied agents” slot.