Going Beyond World Models & VLAs
Pete Florence’s April 2026 positioning post explains why Generalist AI refuses to label GEN-1 as either a VLA or a world model, and why ~99% of its parameters are trained from scratch on their >500K-hour real-interaction corpus rather than initialized from a frozen VLM or video model. The argument is three-part: (1) goals beat methods — pick a concrete milestone like “99% success with one hour of task data” and use whatever gets there rather than following method bandwagons; (2) the “A or B” framing (VLA or world model) is a category error — the real question is how far one can go once objectives and constraints are understood; (3) the current data constraint (that motivates VLM/video-model pre-training as a crutch) is temporary and Generalist has now removed it. The post is a philosophy piece, not a technical report — no benchmarks, no architecture details.
Key claims
Section titled “Key claims”- Approximately 99% of GEN-1’s parameters are trained from scratch — deliberate rather than incidental, rooted in a two-year conviction that with enough data, complete control over the fundamental model beats fine-tuning someone else’s frontier [§introduction].
- GEN-1 is explicitly not a fine-tuned vision-language model with actions bolted on, and it is not just a world model — the position is that it is a “first-class-citizen native foundation model for physical interaction” [§introduction].
- Idea-driven research (follow the latest method) is contrasted with goal-driven research (pick a concrete outcome, solve whatever stands in the way); the post argues goal-driven is typically more productive and shapes both what you build and what you don’t get distracted by [§Goals are more important than the labels on your tools].
- A concrete progressive goal is proposed: instead of demanding zero-shot, allow a small amount of task-specific robot data X and drive X down while pushing performance up — “99%+ success rates with roughly one hour of robot data” is named as a commercially-viable measurable milestone [§Goals are more important than the labels on your tools].
- Choosing concrete-yet-ambitious goals is claimed to be more productive as a springboard for branching than picking a method that feels broadly applicable; PaLM-E is cited as a case where a robotics-driven goal produced a multimodal model that later scored on medical benchmarks [§Goals are more important than the labels on your tools].
- The “A or B” framing (perception or control; per-application specialized models or cotraining; VLA or world model) is called a recurring trap; the productive move is to convert “or” → “and” → “how much of each” → deeper questions about objectives and constraints [§How far can we go?].
- The current-data-scarcity constraint that motivates VLM and video-model pretraining as “crutches” is not a long-term view; Generalist claims >500K hours of physical interaction data now removes that constraint [§Building for the world that’s coming].
- Bandwagons are named as part of the nature of academic research: world models are having their moment in early 2026, VLAs had theirs from 2023 to 2025; Generalist has never referred to their own models as either label despite co-inventing VLAs and publishing on world models in robotics since 2023 [§introduction].
Method
Section titled “Method”The post is a positioning essay, not a technical report — no architecture diagrams, training details, benchmark tables, or ablations. The methodological claim is meta: Generalist has spent over a year “combining ideas from across what you might call VLAs, world models, and beyond” and has been “revising training methods” with the philosophy that when things get combined enough they become hard to categorize. The referenced footnote to the GEN-1 blog is the pointer to what was actually built; this post is the argument for why it was built the way it was. Two prior Generalist artifacts are cited as GEN-1’s antecedents — co-invention of the VLA family (Robotics-Transformer-2) and 2023-onward publication on video-language-planning world models — used to support the claim that avoiding both labels is a considered position, not a marketing choice or ignorance of either lineage.
Results
Section titled “Results”No quantitative results. The post lists prior GEN-1 “glimpses” as evidence the from-scratch bet is working — scaling laws in robotics, generalization to new environments and embodiments “in hours,” and improvisational intelligence emerging from large-scale pretraining — but defers all numbers to the GEN-1 blog and to future releases. The “one hour of robot data → 99%+ success” milestone is proposed as a target rather than a reported result.
Why it’s interesting
Section titled “Why it’s interesting”Sharpest filed philosophical statement of the “engine-is-the-lever” position that GEN-1.5: Embodied Foundation Models are One-Shot Learners later cashes out empirically — GEN-1.5’s emergent in-context learning, cross-substrate ICL (sim demo → real prompt without any sim in pretraining), and cross-embodiment ICL (human-hand → robot-hand via ICL) all read as instances of the “how far can we go with the fundamental model under our own control” framing this post articulates in April. Direct counter-position to nearly every recipe on the VLA Models recipe-lever board: π*0.6’s RL-on-flow-matching, Spirit-v1.5’s clean teleop, Embodied-R1.5’s unified VLM with pointing, XR-1’s 100K-hour UMI pretraining, µ₀’s frozen-trace-WM + action expert, and 1XWM’s video-model-as-policy all assume some pretrained backbone or borrowed representation as the substrate; Florence’s argument is that at ≥500K hours of real interaction data the crutch is no longer needed and choosing to keep it is a bet against the near future.
Complements Towards Machines with a Thousand Hands as the philosophical prequel to the empirical follow-up — the “Thousand Hands” post shows how the from-scratch bet cashes out along the embodiment-diversity axis (~9,000 end effectors, mid-rollout hand swaps), while this one is the why behind refusing to categorize the model that produces those results. Also sits alongside Jitendra Malik: don't let CV researchers in robotics skip the sensorimotor level‘s sensorimotor counter-position but on a different axis: Malik argues the VLA framing under-weights contact/tactile dynamics, while Florence argues the VLA and world-model framings themselves are the problem — both are consumed by the labels-not-goals critique. For World Foundation Models, the position is symmetric: Florence names the early-2026 “world model moment” explicitly as a bandwagon and refuses the label from the WFM side too, treating video pretraining as another crutch to be discarded when robotics data is abundant.
See also
Section titled “See also”- GEN-1.5: Embodied Foundation Models are One-Shot Learners — empirical follow-up: emergent ICL, compositional / sim-to-real / human-to-robot physical prompting from the same from-scratch model
- Towards Machines with a Thousand Hands — same Generalist thesis cashed out along the embodiment-diversity axis (~9,000 end effectors, task-vector diagnostic)
- VLA Models — Florence argues the entire VLA framing is a category error; sits as a meta-counter-recipe to every entry on the recipe-lever board
- World Foundation Models — same meta-critique applied to the world-model label (“having their moment in early 2026”)
- Jitendra Malik: don't let CV researchers in robotics skip the sensorimotor level — parallel counter-position (VLA framing under-weights sensorimotor level); Florence’s critique is broader (labels themselves)
- Spirit-v1.5: Clean Data Is the Enemy of Great Robot Foundation Models — sibling data-is-the-lever position with a different bet (curated teleop quality vs Generalist’s real-interaction quantity)