Skip to content

Yann LeCun: World Models — Enabling the next AI revolution (ETH Frontiers of Embodied AI)

Recording of Yann LeCun’s talk “World Models: Enabling the next AI revolution,” given at ETH Zürich on 29 May 2026 as part of the “Frontiers of Embodied AI” distinguished lecture series (alongside Jitendra Malik, Vladlen Koltun, and Shuran Song). Hosted on the Computer Vision and Geometry Group (ETH Zürich) YouTube channel. The published page reports the talk video itself; YouTube exposes no auto-transcript at filing time, so substantive technical notes below are stub-only and reflect what is reliably known from the title, the speaker’s public research program (JEPA / V-JEPA line), and the surrounding playlist context.

  • The talk is one of four distinguished lectures in the ETH “Frontiers of Embodied AI” series on 29 May 2026, framed by the organizers as covering vision, world models, and robotics [event title].
  • The talk is hosted on the ETH Computer Vision and Geometry Group channel and lists Yann LeCun as the speaker [video metadata].

The artifact is a single lecture recording. No auto-transcript was available at filing time, and the video page exposes no description beyond the one-line framing “Talk given by Yann LeCun at ETH Zürich during ‘Frontiers of Embodied AI’.” A future /bud refresh yt-72Xj8k5WQX4-yann-lecun-world-models-enabling after a transcript or a write-up appears should fill in the actual methodological content of the lecture.

Not applicable to a lecture recording. No benchmark numbers, no trained checkpoints, no released code.

LeCun’s JEPA program is the canonical reference point for the latent-predictive framing of world foundation models on the wiki — the same lineage that produced V-JEPA 2 (Introducing the V-JEPA 2 world model and new benchmarks for physical reasoning), the V-JEPA 2.1 dense-feature follow-up (V-JEPA 2.1: Unlocking Dense Features in Video Self-Supervised Learning), the inference-time physics-alignment use of VJEPA-2 as a frozen reward (Inference-time Physics Alignment of Video Generative Models with Latent World Models), and the end-to-end LeWorldModel (LeWorldModel: Stable End-to-End Joint-Embedding Predictive Architecture from Pixels). That program also anchors one pole of the open debate in World Foundation Models — predictive-latent backbones vs generative-rollout WFMs (Genie 3 / Cosmos 3) — and the lab itself (AMI Labs) was recently flagged in Saining Xie joining AMI Labs (LeCun's new world-model lab) — announcement. The talk is also a sibling to the other three ETH lectures shared in the same Slack post (Jitendra Malik on vision and robotics; Vladlen Koltun on spatial cognition in frontier models; Shuran Song on manipulation); cross-pointers below give the other two non-Malik perspectives a place to land when transcripts arrive.