Reka and Moonvalley Join Forces to Advance Models and Infrastructure for Physical AI
Reka announces it has merged with Moonvalley, absorbing a team of former DeepMind / Meta / Amazon / Microsoft / Google / Wayve / Runway researchers and engineers — most prominently Mateusz Malinowski and Mikołaj Bińkowski, both ex-Google DeepMind staff who contributed to the models underlying Veo. The combined team’s stated technical bet is the World Language Action Model (WLAM), an omnimodel trained on egocentric and other physical-world data, meant to perceive, simulate, and act in the real world via “realistic simulation for planning.” Reka is restructuring around four pillars — Labs (research), Infer (multimodal inference infra), Vision (video processing infra), and Claru (training data for physical AI: egocentric video, robotics trajectories, world-model footage, expert judgement). No models, weights, benchmarks, or architectures are disclosed.
Key claims
Section titled “Key claims”- Reka has merged with Moonvalley and added a team of researchers/engineers from DeepMind, Meta, Amazon, Microsoft, Google, Wayve, and Runway to accelerate physical-AI work [§Body].
- Mateusz Malinowski (PhD, ex-DeepMind staff research scientist) and Mikołaj Bińkowski (PhD, ex-DeepMind senior research scientist) — both attributed as key contributors to models that became Google’s Veo — join Reka’s research leadership [§Body].
- The combined team will focus on a World Language Action Model (WLAM) — an omnimodel trained on egocentric and other physical-world data, framed as perceiving and acting in the real world via realistic simulation for planning [§Body].
- WLAM is positioned to (a) generate high-fidelity multimodal outputs, (b) translate multimodal perception to robotic execution with “state-of-the-art spatial and temporal awarenesses,” and (c) perform reasoning over long-form and streaming video [§Body].
- Reka is restructuring around four named pillars: Reka Labs (foundational research), Reka Infer (multimodal inference API), Reka Vision (video processing platform), and Reka Claru (training data: egocentric video, robotics trajectories, world-model footage, expert human judgement) [§Body].
- Target verticals named: autonomous defense, intelligent robots, next-generation wearables, media [§Body].
Method
Section titled “Method”Strategic / hiring announcement. No technical content — no architecture details, parameter counts, training data sizes (beyond category names), benchmarks, or model releases. The substantive claim is that a single omnimodel can unify perception, multimodal generation (video/image/audio), and action / robotic execution, with simulation-for-planning as the inference-time mechanism. This positions WLAM in the same architectural neighborhood as NVIDIA’s Cosmos 3 (mixture-of-transformers omnimodel emitting text + video + audio + action; see Cosmos 3: Omnimodal World Models for Physical AI) but Reka does not disclose whether the unification is structural (one backbone) or stacked (reasoning + generation + action heads). The “perceive → simulate → act” framing matches the Sitzmann normative position (The flavor of the bitter lesson for computer vision) more than the latent-predictive JEPA line (Action100M: A Large-scale Video Action Dataset, LeWorldModel: Stable End-to-End Joint-Embedding Predictive Architecture from Pixels).
Results
Section titled “Results”None disclosed. The announcement is a corporate / strategic event, not a paper. Concrete reportable facts: Reka previously raised 168M at a $1B+ valuation; the combined Reka + Moonvalley headcount is not given. The Moonvalley side previously shipped Marey, a closed controllable video model (Moonvalley Marey — controllable AI video model with 360° camera, pose, and motion transfer) — the press release does not say whether Marey or its successors continue as a product line under Reka or are subsumed into the WLAM effort.
Why it’s interesting
Section titled “Why it’s interesting”This is the second filed datapoint (after Saining Xie’s move to AMI Labs, Saining Xie joining AMI Labs (LeCun's new world-model lab) — announcement) of senior researchers consolidating around a physical-AI / world-model org thesis in mid-2026, and the first filed datapoint of a merger rather than a hire. The “WLAM” label collapses three things the wiki currently tracks separately — world foundation models, VLA action policies, and unified multimodal generation — into a single product framing; this is the closed-source counterpart to NVIDIA’s open Cosmos 3 bet (Cosmos 3: Omnimodal World Models for Physical AI). The team’s video-generation provenance (Veo contributors + Moonvalley’s Marey lineage) makes this a notable test of whether a video-generation-rooted lab can pivot through simulation for planning into robotic execution — the exact recipe argued for in The flavor of the bitter lesson for computer vision and instantiated (at research scale) by Direct Video-Action Models — Causal Video Models Are Data-Efficient Robot Policy Learners. Useful primarily as a market-signal marker; pending /bud refresh once any model, paper, or benchmark drops.
See also
Section titled “See also”- World Foundation Models — WLAM positions itself as a video-generation-rooted world foundation model with action output
- VLA Models — WLAM’s claimed “translate multimodal perception to robotic execution” capability is the VLA half of the recipe
- Cosmos 3: Omnimodal World Models for Physical AI — open-source omnimodel counterpart unifying VLM + multimodal generation + action under one Mixture-of-Transformers backbone
- Moonvalley Marey — controllable AI video model with 360° camera, pose, and motion transfer — Moonvalley’s prior closed controllable video model; the team’s video-generation provenance entering Reka
- Saining Xie joining AMI Labs (LeCun's new world-model lab) — announcement — companion mid-2026 datapoint of senior research talent consolidating around world-model orgs (Xie → AMI Labs)
- The flavor of the bitter lesson for computer vision — normative position the WLAM framing aligns with: generative-rollout world models as the perception substrate for embodied AI
- Introducing Runway Labs — adjacent organizational framing: closed-flagship video lab restructuring around GWM productization