Skip to content

Jun Gao thread (4/5) on Cosmos2.5-2B finetuned with cleaned data + kinematic skeleton conditioning

Tweet 4/5 in a thread by Jun Gao (NVIDIA Toronto AI Lab) reporting results from finetuning the Cosmos2.5-2B video model with a curated/cleaned dataset and kinematic skeleton conditioning, trained on a single GH200 GPU. The claim is SOTA-level action-following, appearance, and motion consistency at a fraction of the compute of comparable systems. The thread sits in the same week as NVIDIA’s Cosmos 3 release and reads as a small-scale recipe demonstration showing what a 2B-parameter Cosmos2.5 backbone can do when paired with clean human-pose conditioning.

  • Cosmos2.5-2B finetuned with cleaned data + kinematic-skeleton conditioning yields superior quality on action following, appearance, and motion consistency relative to unspecified baselines [tweet body].
  • The full finetune runs on a single GH200 GPU, framed as a “fraction of the compute” claim [tweet body].

The tweet itself only describes the recipe at a high level: take Cosmos2.5-2B, condition on a kinematic skeleton (presumably per-frame 2D or 3D joint locations driving the generation), and finetune on a curated dataset on one GH200. Earlier tweets in the thread (1/5–3/5) presumably cover the data-cleaning pipeline and the skeleton-conditioning architecture but were not fetched here; a /bud refresh once the full thread is mirrored would tighten this section.

The tweet shows a side-by-side video comparison (linked, not embedded here) and claims SOTA on three axes — action following, appearance, and motion consistency — without naming the baselines or the benchmark. The “fraction of the compute” phrasing is qualitative.

Two threads of the wiki connect directly. First, this is a concrete data point on what a small Cosmos backbone can do when conditioned cleanly — relevant to the broader Cosmos 3 release filed as Cosmos 3: Omnimodal World Models for Physical AI and the NVIDIA Launches Cosmos 3, the Open Frontier Foundation Model for Physical AI launch post, which themselves frame Cosmos as a synthetic-data engine. Second, skeleton-conditioned video generation is an active cluster on the wiki — PoseAnything: Universal Pose-guided Video Generation with Part-aware Temporal Coherence argues for arbitrary skeleton inputs, Wan-Animate: Unified Character Animation and Replacement with Holistic Replication is the dominant Wan-based recipe, and SCAIL-2: Unifying Controlled Character Animation with End-to-end In-Context Conditioning drops skeleton-and-reference together in an end-to-end model. This tweet is a NVIDIA-side data point in the same race, suggesting the Cosmos backbone can compete with the Wan ecosystem on the same problem when the data is right.

A near-adjacent tweet in the same thread (status id differing by ~7 ns in Twitter’s snowflake) was shared separately; only this one carried retrievable content.