Scaling Spatial Intelligence with Multimodal Foundation Models
SenseNova-SI is a SenseTime / SLab family of multimodal foundation models built by post-training three existing backbones (Qwen3-VL, InternVL3, and the unified-understanding-and-generation model Bagel) on SenseNova-SI-8M, an 8M-sample dataset systematically curated under a taxonomy of spatial-intelligence capabilities. The paper’s thesis is operational rather than architectural: spatial intelligence is a data problem at this point, and a principled, taxonomy-driven curation of 8M samples is what was missing. Headline numbers — 68.8% VSI-Bench, 43.3% MMSI, 85.7% MindCube, 54.7% ViewSpatial, 47.7% SITE, 63.9% BLINK, 55.5% 3DSR, 72.0% EmbSpatial — are reported as unprecedented across this benchmark suite, with general multimodal understanding (84.9% MMBench-En) largely preserved. The release lands squarely in the VLM Perception Failures cluster as a positive counterpart to the Cambrian-S / NoSense exchange: where Cambrian-S: Towards Spatial Supersensing in Video argued spatial cognition requires new inference mechanism and Solving Spatial Supersensing Without Spatial Supersensing argued the headline benchmark admits a shortcut, SenseNova-SI argues the gap simply hasn’t been data-saturated yet.
Key claims
Section titled “Key claims”- Multimodal foundation models exhibit “surprising deficiencies in spatial intelligence” despite remarkable general progress, motivating scaling along a spatial-capability axis rather than along generic multimodal mix [Abstract].
- The SenseNova-SI family is built on three established multimodal backbones: Qwen3-VL and InternVL3 (visual-understanding-only) and Bagel (unified understanding + generation) — covering both classes of current MLLM architecture under one curation recipe [Abstract].
- SenseNova-SI-8M is constructed by systematically curating eight million samples under “a rigorous taxonomy of spatial capabilities” — the central methodological commitment is taxonomy-driven rather than scale-driven [Abstract].
- Reported scores on a broad spatial-intelligence benchmark sweep: VSI-Bench 68.8%, MMSI 43.3%, MindCube 85.7%, ViewSpatial 54.7%, SITE 47.7%, BLINK 63.9%, 3DSR 55.5%, EmbSpatial 72.0% [Abstract].
- General multimodal understanding is largely preserved: MMBench-En 84.9% [Abstract].
Method
Section titled “Method”The paper takes a clean data-engineering stance over an established post-training pipeline. Three backbones — Qwen3-VL and InternVL3 (the two strongest current open visual-understanding MLLMs in this filing cycle’s reads) plus Bagel (unified understanding + generation) — are post-trained on SenseNova-SI-8M, 8M samples organized under an explicit taxonomy of spatial capabilities. The taxonomy is the load-bearing piece: rather than a single “spatial mix” the dataset is structured by capability category, intended to ensure broad coverage rather than a single benchmark’s distribution. Evaluation spans eight spatial benchmarks (VSI-Bench, MMSI, MindCube, ViewSpatial, SITE, BLINK, 3DSR, EmbSpatial) plus MMBench-En to track general-capability retention. The fetched abstract describes the curation methodology and headline numbers; finer ablations on category mixture, per-backbone deltas, and synthetic vs. real-data splits inside the 8M corpus are in the full paper and were not retrievable at filing time (the arXiv fetcher was rate-limited).
Results
Section titled “Results”- VSI-Bench: 68.8% — substantially above the prior wiki entries’ baselines. Cambrian-S: Towards Spatial Supersensing in Video reported >+30% absolute gain on VSI-Bench over its base MLLM as a SoTA result at its release; the 68.8 number here is in the same regime and may exceed it (no direct head-to-head is in the abstract).
- MMSI 43.3%, MindCube 85.7%, ViewSpatial 54.7%, SITE 47.7%, BLINK 63.9%, 3DSR 55.5%, EmbSpatial 72.0% — a broad benchmark sweep rather than a single-benchmark optimization, which is the most credible signal here given how recent benchmark audits (e.g. Solving Spatial Supersensing Without Spatial Supersensing) have collapsed single-benchmark gains.
- MMBench-En 84.9% — general multimodal understanding largely preserved; consistent with the paper’s framing that the spatial post-training does not trade off general capability.
Caveat: per-backbone breakdowns (Qwen3-VL vs InternVL3 vs Bagel) are not in the fetched abstract; the headline numbers may correspond to a specific backbone or the best across the three. Same for the 8M-sample-mix ablations and the per-capability-category gains.
Why it’s interesting
Section titled “Why it’s interesting”Three reasons SenseNova-SI is worth reading rather than skimming. First, it is the closest thing on file to a data-centric refutation of the position in Cambrian-S: Towards Spatial Supersensing in Video that data scaling alone cannot crack hard spatial benchmarks — Cambrian-S argued the lever runs out and predictive-world-model inference is needed; SenseNova-SI argues a principled 8M-sample taxonomy clears VSI-Bench at 68.8%, well above prior open numbers. Whether the same recipe scales to VSI-Super (the long-video stage that the Cambrian-S paper specifically isolates) is the obvious next experiment and not addressed in the abstract.
Second, it is one of the few filed entries that post-trains the unified Unified Multimodal Models backbone Bagel alongside the two understanding-only backbones, providing a (potential) within-paper comparison of whether the joint understanding-and-generation architecture imports more or less spatial signal from the same curated 8M than understanding-only architectures do. This is the cleanest data point in the wiki to date on whether unified models’ joint training is net-positive for downstream spatial reasoning.
Third, it sits on the Synthetic Training Data axis as a taxonomy-driven counterpart to corpora like Video-Thinker-10K, Action100M, and InfTool — same family of “the curation gate dominates returns” claim, but at 8M samples and explicitly targeted at a capability cluster that the VLM Perception Failures page documents as a load-bearing failure mode of frontier VLMs.
See also
Section titled “See also”- Cambrian-S: Towards Spatial Supersensing in Video — the position paper whose data-scaling-runs-out thesis SenseNova-SI implicitly contests on VSI-Bench
- Solving Spatial Supersensing Without Spatial Supersensing — the NoSense critique of VSI-Super; same benchmark family, methodologically important to read together
- BOP-Ask: Object-Interaction Reasoning for Vision-Language Models — sibling spatial-reasoning benchmark targeting pixel-precise 3D output where frontier VLMs collapse; SenseNova-SI is not evaluated there
- VLM Perception Failures — concept cluster the spatial-deficiency framing belongs to
- Synthetic Training Data — concept cluster on curation-first data recipes
- Unified Multimodal Models — Bagel is one of the three SenseNova-SI backbones; only filed paper that post-trains a unified model alongside understanding-only baselines on the same spatial curriculum
- SpatialVID: A Large-Scale Video Dataset with Spatial Annotations — large-scale spatial annotation corpus in the video domain; data sibling