Yann LeCun: World Models — Enabling the next AI revolution (ETH Frontiers of Embodied AI)
Recording of Yann LeCun’s talk “World Models: Enabling the next AI revolution,” given at ETH Zürich on 29 May 2026 as part of the “Frontiers of Embodied AI” distinguished lecture series (alongside Jitendra Malik, Vladlen Koltun, and Shuran Song). Hosted on the Computer Vision and Geometry Group (ETH Zürich) YouTube channel. The published page reports the talk video itself; YouTube exposes no auto-transcript at filing time, so substantive technical notes below are stub-only and reflect what is reliably known from the title, the speaker’s public research program (JEPA / V-JEPA line), and the surrounding playlist context.
Key claims
Section titled “Key claims”- The talk is one of four distinguished lectures in the ETH “Frontiers of Embodied AI” series on 29 May 2026, framed by the organizers as covering vision, world models, and robotics [event title].
- The talk is hosted on the ETH Computer Vision and Geometry Group channel and lists Yann LeCun as the speaker [video metadata].
Method
Section titled “Method”The artifact is a single lecture recording. No auto-transcript was
available at filing time, and the video page exposes no description
beyond the one-line framing “Talk given by Yann LeCun at ETH Zürich
during ‘Frontiers of Embodied AI’.” A future /bud refresh yt-72Xj8k5WQX4-yann-lecun-world-models-enabling after a transcript
or a write-up appears should fill in the actual methodological
content of the lecture.
Results
Section titled “Results”Not applicable to a lecture recording. No benchmark numbers, no trained checkpoints, no released code.
Why it’s interesting
Section titled “Why it’s interesting”LeCun’s JEPA program is the canonical reference point for the latent-predictive framing of world foundation models on the wiki — the same lineage that produced V-JEPA 2 (Introducing the V-JEPA 2 world model and new benchmarks for physical reasoning), the V-JEPA 2.1 dense-feature follow-up (V-JEPA 2.1: Unlocking Dense Features in Video Self-Supervised Learning), the inference-time physics-alignment use of VJEPA-2 as a frozen reward (Inference-time Physics Alignment of Video Generative Models with Latent World Models), and the end-to-end LeWorldModel (LeWorldModel: Stable End-to-End Joint-Embedding Predictive Architecture from Pixels). That program also anchors one pole of the open debate in World Foundation Models — predictive-latent backbones vs generative-rollout WFMs (Genie 3 / Cosmos 3) — and the lab itself (AMI Labs) was recently flagged in Saining Xie joining AMI Labs (LeCun's new world-model lab) — announcement. The talk is also a sibling to the other three ETH lectures shared in the same Slack post (Jitendra Malik on vision and robotics; Vladlen Koltun on spatial cognition in frontier models; Shuran Song on manipulation); cross-pointers below give the other two non-Malik perspectives a place to land when transcripts arrive.
See also
Section titled “See also”- World Foundation Models — predictive-latent WFMs are LeCun’s pole of the cluster’s central debate
- Introducing the V-JEPA 2 world model and new benchmarks for physical reasoning — flagship public artifact from LeCun’s program at the time of this talk
- V-JEPA 2.1: Unlocking Dense Features in Video Self-Supervised Learning — V-JEPA 2.1 dense-feature follow-up
- LeWorldModel: Stable End-to-End Joint-Embedding Predictive Architecture from Pixels — first end-to-end pixel-to-latent JEPA WFM, FAIR + Mila with LeCun co-author
- Saining Xie joining AMI Labs (LeCun's new world-model lab) — announcement — Saining Xie joining AMI Labs (LeCun’s new world-model lab) — surrounding org context
- Jitendra Malik: don't let CV researchers in robotics skip the sensorimotor level — Jitendra Malik’s other ETH-adjacent public note (sensorimotor level in robotics); plausibly related to his sibling talk in this same lecture series
- The flavor of the bitter lesson for computer vision — the “video-generative pre-training dissolves explicit 3D” position that LeCun-flavoured WFM arguments are typically read against