Skip to content

Sora 2 is here

OpenAI’s Sora 2 launch announcement (Sep 30, 2025) frames the model as the “GPT-3.5 moment” for video — a flagship joint video-and-audio generative model claimed to better obey physical laws (rebounds rather than ball-snapping-to-hoop) and to model failure as well as success. It ships alongside a new social iOS app called Sora, whose central feature is “characters” — likeness-injection from a short identity-and-voice recording, with end-to-end user control over who can use the character and visibility of derivative videos. The post is product/policy-heavy and discloses no architecture, parameter count, training data, or benchmarks. Retrospective note at the top of the page indicates the Sora product was discontinued as of April 26, 2026.

  • Sora 2 is positioned as a joint video + audio generation system that can produce “sophisticated background soundscapes, speech, and sound effects with a high degree of realism” [blog body].
  • The model is claimed to follow intricate multi-shot instructions while “accurately persisting world state” — i.e. cross-shot world-state consistency is presented as a headline capability [blog body].
  • Stated physical-realism improvements are framed as failure-modeling rather than success-mimicking: a missed basketball shot rebounds off the backboard rather than snapping to the hoop; “mistakes” the model makes are framed as mistakes of an implicit internal agent rather than as physics violations [blog body].
  • A “characters” capability lets a user inject a likeness (any human, animal, or object) into a Sora-generated scene after a short one-time identity-and-voice recording in the app; the post claims accurate portrayal of both appearance and voice [blog body].
  • The launch includes the Sora iOS app — an invite-based social product whose feed is described as biased toward followed/interacted-with creators and toward videos likely to inspire user re-creation, explicitly not optimized for time-in-feed [blog body].
  • A recommender stack is described as “instructable through natural language” via OpenAI’s existing LLMs, with periodic user wellbeing polling and configurable feed controls [blog body].
  • Teen safeguards listed: daily-generation caps in feed, stricter character permissions for teens, scaled-up human moderation for bullying review, ChatGPT-based parental controls (override infinite-scroll, disable algorithmic personalization, manage DMs) [blog body].
  • Likeness control: only the character’s owner authorizes use, access is revocable, and the owner can view (and presumably remove) any video — including drafts created by others — that contains their character [blog body].
  • Initial rollout: free with “generous limits” in U.S. and Canada; ChatGPT Pro users get access to an experimental higher-quality Sora 2 Pro model on sora.com; Sora 2 in the API and Sora 1 Turbo continuing to be supported are promised [blog body].
  • No architecture, parameter count, training corpus, or quantitative benchmark numbers are disclosed in the post [blog body].
  • Retrospective notice at top of page (April 26, 2026): “the Sora product is no longer available” [blog body, banner].

Product announcement, not a technical report. The post is a mix of capability claims, framing (“GPT-3.5 moment for video”, and an assertion that further scaling of video-data neural networks moves closer to simulating reality), and policy/UX commentary on the accompanying iOS app. There is no Methods section; closest the post gets to architectural commentary is the assertion that the team is “focused on training models with more advanced world simulation capabilities” and that pre-training and post-training on large-scale video data are “in their infancy compared to language” — both rhetorical framings rather than design disclosures.

The “characters” feature is described operationally: a one-time in-app video-and-audio capture serves both as identity verification and as the conditioning signal that the model later uses to inject the user’s likeness. No technical detail on how that conditioning works (image encoder, voice encoder, learned embedding, fine-tune, or in-context) is given.

No reported metrics. The post relies entirely on described example prompts (“Olympic gymnastics routines, backflips on a paddleboard that accurately model the dynamics of buoyancy and rigidity, and triple axels while a cat holds on for dear life”) and assertions about controllability and audio realism. The only public latency anchor that exists on the wiki for the Sora 2 family is anecdotal — Sora 2 Pro takes 10-15 minutes for a 15s HD video — David Attisaas first-impressions thread reports 10–15 minutes wall-clock for a 15-second HD clip from Sora 2 Pro, suggesting the production model is either parameter-heavy, undistilled, or both.

This is the closed-flagship anchor for the joint audio-video generation cluster the wiki has been tracking from the open side: Joint audio-video generation notes that the open consensus has converged on dual-stream DiT with bidirectional cross-attention (LTX-2: Efficient Joint Audio-Visual Foundation Model, MOVA: Towards Scalable and Synchronized Video-Audio Generation, SkyReels-V4: Multi-modal Video-Audio Generation, Inpainting and Editing model), and that Sora 2, Veo 3, and Kling 2.6 are the three frontier closed reference points whose recipe is opaque. Sora 2’s announcement adds the capability surface (failure-aware physics, multi-shot world-state persistence, identity-and-voice injection) but discloses none of the recipe, so the comparison stays one-sided.

In the World Foundation Models frame, the post is the strongest first-party articulation of the “WFM as generative-rollout simulator” position — explicitly contrasting Sora 2’s failure-modeling against prior models that “morph objects and deform reality to successfully execute upon a text prompt”, and framing the work as on the path to “general-purpose simulation and AI systems that can function in the physical world”. This is the same generative-rollout framing Sitzmann argues should become the central pre-training objective for embodied AI (The flavor of the bitter lesson for computer vision), and complements the closed-flagship-then-vertical-specialization pattern Waymo demonstrates on top of Genie 3 (The Waymo World Model: A New Frontier For Autonomous Driving Simulation). The retired-product banner is the loudest single data point — a frontier closed video product reached end-of-life in ~7 months from launch, which is information about productization economics that the technical literature does not normally capture.