Jun Gao thread (4/5) on Cosmos2.5-2B finetuned with cleaned data + kinematic skeleton conditioning
Tweet 4/5 in a thread by Jun Gao (NVIDIA Toronto AI Lab) reporting results from finetuning the Cosmos2.5-2B video model with a curated/cleaned dataset and kinematic skeleton conditioning, trained on a single GH200 GPU. The claim is SOTA-level action-following, appearance, and motion consistency at a fraction of the compute of comparable systems. The thread sits in the same week as NVIDIA’s Cosmos 3 release and reads as a small-scale recipe demonstration showing what a 2B-parameter Cosmos2.5 backbone can do when paired with clean human-pose conditioning.
Key claims
Section titled “Key claims”- Cosmos2.5-2B finetuned with cleaned data + kinematic-skeleton conditioning yields superior quality on action following, appearance, and motion consistency relative to unspecified baselines [tweet body].
- The full finetune runs on a single GH200 GPU, framed as a “fraction of the compute” claim [tweet body].
Method
Section titled “Method”The tweet itself only describes the recipe at a high level: take Cosmos2.5-2B, condition on a kinematic skeleton (presumably per-frame 2D or 3D joint locations driving the generation), and finetune on a curated dataset on one GH200. Earlier tweets in the thread (1/5–3/5) presumably cover the data-cleaning pipeline and the skeleton-conditioning architecture but were not fetched here; a /bud refresh once the full thread is mirrored would tighten this section.
Results
Section titled “Results”The tweet shows a side-by-side video comparison (linked, not embedded here) and claims SOTA on three axes — action following, appearance, and motion consistency — without naming the baselines or the benchmark. The “fraction of the compute” phrasing is qualitative.
Why it’s interesting
Section titled “Why it’s interesting”Two threads of the wiki connect directly. First, this is a concrete data point on what a small Cosmos backbone can do when conditioned cleanly — relevant to the broader Cosmos 3 release filed as Cosmos 3: Omnimodal World Models for Physical AI and the NVIDIA Launches Cosmos 3, the Open Frontier Foundation Model for Physical AI launch post, which themselves frame Cosmos as a synthetic-data engine. Second, skeleton-conditioned video generation is an active cluster on the wiki — PoseAnything: Universal Pose-guided Video Generation with Part-aware Temporal Coherence argues for arbitrary skeleton inputs, Wan-Animate: Unified Character Animation and Replacement with Holistic Replication is the dominant Wan-based recipe, and SCAIL-2: Unifying Controlled Character Animation with End-to-end In-Context Conditioning drops skeleton-and-reference together in an end-to-end model. This tweet is a NVIDIA-side data point in the same race, suggesting the Cosmos backbone can compete with the Wan ecosystem on the same problem when the data is right.
A near-adjacent tweet in the same thread (status id differing by ~7 ns in Twitter’s snowflake) was shared separately; only this one carried retrievable content.
See also
Section titled “See also”- Cosmos 3: Omnimodal World Models for Physical AI — NVIDIA’s flagship Cosmos 3 release this week
- PoseAnything: Universal Pose-guided Video Generation with Part-aware Temporal Coherence — universal pose-guided video generation
- Wan-Animate: Unified Character Animation and Replacement with Holistic Replication — Wan-based skeleton+reference character animation
- SCAIL-2: Unifying Controlled Character Animation with End-to-end In-Context Conditioning — end-to-end controlled character animation