Skip to content

Seedream 4.5 — ByteDance image model with up to 14 reference images (Replicate announcement)

ByteDance Seed released Seedream 4.5 on Dec 3, 2025 — a unified image generation + editing model whose headline capability is accepting up to 14 reference images in a single request for consistent multi-image composition. The model is closed (API only via ByteDance Seed, Replicate, fal, etc.), with no accompanying tech report at announcement time. Versus Seedream 4.0 (August 2025), the focus is multi-image scene consistency, aesthetic-instruction following, stronger spatial understanding, and improved typography rendering. Notable for the Luma team because the multi-reference count (10–14) is well above current open competitors (Qwen-Image, FLUX.2, HunyuanImage 3) and stresses the same subject-consistency / layered-composition problems that the team’s editing and storyboard work touches.

  • Seedream 4.5 accepts up to 14 reference images per request, targeting controllable multi-image generation with subject and detail preservation [tweet].
  • The model unifies text-to-image generation and reference-based editing in one architecture (same pattern as Seedream 4.0 / Qwen-Image / FLUX.2) [Replicate tweet; ByteDance Seed product page].
  • ByteDance reports improvements over Seedream 4.0 across prompt adherence, alignment, and aesthetics on an internal MagicBench, with no public numbers [ByteDance Seed product page].
  • The release date is December 3, 2025 [tweet timestamp].

No technical disclosure was published with the launch — Seedream 4.5 is a closed product release. The Replicate tweet is the primary public artifact at the time of this filing; ByteDance Seed’s product page (seed.bytedance.com/en/seedream4_5) describes the changes versus 4.0 in marketing terms but does not disclose architecture, training data, or evaluation methodology. The most that can be inferred from the surrounding ecosystem (Seedream 3.0 and 4.0 lineage, Qwen-Image’s MMDiT comparison points; see Qwen-Image Technical Report) is that this is almost certainly an MMDiT-class double-stream diffusion transformer; the 14-reference figure suggests an in-context conditioning scheme similar in spirit to FullDiT2: Efficient In-Context Conditioning for Video Diffusion Transformers and UNIC: Unified In-Context Video Editing in the video-editing world. None of that is confirmed by ByteDance.

No independent benchmarks are reported in the announcement. The ByteDance Seed page claims improvements over Seedream 4.0 on prompt adherence, alignment, and aesthetics, measured on their internal MagicBench — no numbers, no public comparisons against Nano Banana 2 / Qwen-Image / FLUX.2. Third-party leaderboard discussion (LM Arena, Artificial Analysis) places it competitively in the closed-image-model tier but below Nano Banana 2 (Gemini 3.1 Flash Image) on the public arenas at release.

Multi-reference image generation is one of the load-bearing capabilities for production creative workflows — character consistency across shots, product placement across angles, style transfer with constraints — and the public progression has gone Seedream 4.0 → FLUX.2 Klein (multi-ref editing, see FLUX.2 [klein]: Towards Interactive Visual Intelligence and FLUX.2 [klein] 9B-KV: KV-cache optimized variant for accelerated multi-reference editing) → Qwen-Image-Edit-2509 (community camera-control LoRAs, see Qwen-Edit-2509-Multiple-angles LoRA — community camera-control adapter for Qwen-Image-Edit-2509) → Seedream 4.5 (14-ref). It also complements BindWeave: Subject-Consistent Video Generation via Cross-Modal Integration and Phantom-Data: Towards a General Subject-Consistent Video Generation Dataset which attack the same subject-consistency problem on the video side via curated reference datasets. Worth tracking because the closed labs (Seedream, Nano Banana, GPT-Image) are now competing on a capability (10+ references) that open models have not matched, which limits open-source video/image-editing pipelines that need consistent multi-character composition.