FilmBench: A Film-Grade Benchmark for Cinematic Video Generation
FilmBench is a text-to-video and reference-to-video benchmark co-developed with directors and faculty from the Beijing Film Academy and Hujing Digital Media, evaluating generators on the professional Cinematic Language criteria used by working filmmakers rather than the “overall VQ + coarse text alignment + temporal smoothness” taxonomy of prior web-style benchmarks. Prompts are reverse-engineered from clips of award-winning films spanning 20 genres, chosen by professional directors, and 1,056 of 1,169 prompts are multi-shot. Evaluation is a three-level taxonomy of 3 axes × 12 components × 35 T2V + 3 R2V-only sub-metrics, with an in-house expert-grade automatic evaluation agent whose core Cinematic Language operators (FilmOps) are open-sourced. Across 9 T2V and 7 R2V models, the auto-evaluator matches human model-level rankings at Spearman ρ = 0.95 (T2V) and 0.96 (R2V); scores fall well below prior web-style benchmarks with consistent gaps in dynamic aesthetics and a marked single-to-multi-shot performance drop that widens for weaker models.
Key claims
Section titled “Key claims”- Prompts are anchored to verified live-action references from award-winning films across 20 cinematic genres, chosen by professional directors, with each prompt following a real shot list [Abstract].
- 1,056 of 1,169 prompts are multi-shot, unlike prior single-clip benchmarks [Abstract].
- Evaluation follows a three-level Cinematic taxonomy of 3 axes, 12 components, and 35 T2V + 3 R2V-only sub-metrics grounded in film-academy tradition [Abstract].
- The in-house expert-grade automatic evaluation agent reproduces the human model ranking at model-level Spearman ρ = 0.95 for T2V and 0.96 for R2V across the evaluated model lineup [Abstract].
- Nine T2V and seven R2V models are benchmarked, and their absolute scores fall well below results on prior web-style benchmarks [Abstract].
- Two consistent gaps show up across models on dynamic aesthetics, and a marked single- to multi-shot performance drop widens for weaker models [Abstract].
- The core suite of Cinematic Language operators (FilmOps) is open-sourced [Abstract].
Method
Section titled “Method”FilmBench rests on three design choices. First, prompts are reverse-engineered from clips of award-winning films rather than authored from LLM templates or scraped from web captions — each prompt has a verified live-action reference, follows a real shot list, and most script multiple shots (90.3% multi-shot). Second, the evaluation taxonomy is co-designed with Beijing Film Academy directors and faculty into three levels — 3 top-level Cinematic axes, 12 components, and 35 T2V sub-metrics + 3 R2V-only sub-metrics — reflecting the vocabulary by which films are actually judged (composition, cinematography craft, dynamic aesthetics, etc.) rather than the “VQ + text alignment + temporal smoothness” of prior benchmarks. Third, evaluation is delegated to an expert-grade automatic agent whose core operators (FilmOps) are released as an open-source suite; the agent is calibrated against human ratings to Spearman ρ ≈ 0.95–0.96 at model level. Both text-to-video and reference-to-video conditioning modes are supported (the R2V-only sub-metrics score fidelity to the reference).
Results
Section titled “Results”- The automatic evaluator matches human model-level rankings at Spearman ρ = 0.95 (T2V) and ρ = 0.96 (R2V) — the strongest calibration reported by any filed video-generation benchmark that spans cinematic rather than physics-commonsense axes [Abstract].
- Absolute scores on FilmBench fall well below what the same models score on prior web-style benchmarks — the “film-grade” prompt distribution is systematically harder than web-caption prompts [Abstract].
- Two dynamic-aesthetics gaps are consistent across models, and a single-shot to multi-shot performance drop widens as model quality decreases [Abstract]. Exact numeric per-model breakdowns are not readable from the abstract alone and would need the full paper for reproduction.
Why it’s interesting
Section titled “Why it’s interesting”FilmBench is the first filed benchmark that (a) anchors prompts to real, professionally judged film craft rather than physics laws or robot tasks and (b) does so at multi-shot scale (>90% of prompts). This directly extends the axis Physion-Arc 1.0: Benchmarking Video Agents on Minute-Long Video Generation opened for the Video Generation Benchmarks cluster, but with a critical difference: Physion-Arc evaluates end-to-end agentic video systems (Runway Agent 2.0, Luma, Kling AI) on complete minute-long screenplays, whereas FilmBench probes the underlying T2V and R2V generators on shorter multi-shot prompts derived from real shot lists. Together they bracket the “narrative/cinematic capability” axis at two levels of abstraction. The taxonomy co-designed with a film academy is also a step beyond CinemaCLIP: A hybrid CLIP model and taxonomy for the visual language of cinema‘s 23 classifier-head cinematic taxonomy — that one classifies existing films; FilmBench uses the taxonomy as an evaluation rubric for generators. For Luma this matters directly: film-grade multi-shot capability is a plausible next frontier for Dream Machine class models, and FilmBench provides both a target and (via open-sourced FilmOps) tooling for evaluation.
See also
Section titled “See also”- Video Generation Benchmarks — this benchmark’s home cluster; FilmBench is the second cinematic/narrative-axis entry alongside Physion-Arc
- Physion-Arc 1.0: Benchmarking Video Agents on Minute-Long Video Generation — Physion-Arc 1.0: sibling narrative-axis benchmark, but scores agentic video systems on minute-long screenplays rather than T2V/R2V generators on multi-shot prompts
- CinemaCLIP: A hybrid CLIP model and taxonomy for the visual language of cinema — CinemaCLIP: cinematic taxonomy used for classifying films; complementary to FilmBench’s use of a taxonomy as a generation evaluation rubric
- ShotAdapter: Text-to-Multi-Shot Video Generation with Diffusion Models — ShotAdapter: multi-shot T2V generation itself (the capability FilmBench probes)
- TalkCuts: A Large-Scale Dataset for Multi-Shot Human Speech Video Generation — TalkCuts: prior multi-shot video dataset with a related cinematic-annotation flavor, but for training rather than benchmarking
- VLM-as-Evaluator — the judge machinery FilmBench’s automatic agent almost certainly builds on (VLM-driven CoT scoring against a rubric)
- Camera-Controlled Video Diffusion — related capability cluster; cinematic-language operators plausibly probe camera-motion adherence as one of the sub-metrics