Skip to content

Wan 2.7 Pro and Wan 2.7 enter Image Arena and Image Editing Arena leaderboards (Design Arena)

A “BREAKING” tweet from @Designarena reporting that Alibaba Wan’s Wan 2.7 Pro and Wan 2.7 entered the Design Arena leaderboards: #21 and #22 on Image Arena, and #10 and #8 on Image Editing Arena. The tweet places the Image-Arena rankings in the same performance band as Google DeepMind’s Imagen 4 Generate 001 and Black Forest Labs’ FLUX.1 Kontext Max. The post is the start of a thread; the Image-Editing-Arena commentary cuts off mid-sentence in the OP. No methodology, eval prompts, or sample counts are disclosed in the tweet itself.

  • Wan 2.7 Pro is ranked #21 and Wan 2.7 is ranked #22 on Design Arena’s Image Arena leaderboard [tweet OP].
  • On Image Editing Arena, Wan 2.7 Pro is #10 and Wan 2.7 is #8 [tweet OP] (note: tweet uses “these are #10 and #8” in that order, with the Pro variant lower than the base — the OP truncates before clarifying which model maps to which rank).
  • On Image Arena, the two Wan 2.7 entries are described as in the same performance band as Imagen 4 Generate 001 (Google DeepMind) and FLUX.1 Kontext Max (Black Forest Labs) [tweet OP].
  • Source is Design Arena, a third-party blind-vote arena evaluator (not Alibaba self-reporting) [tweet account context].

Not disclosed in the tweet. Design Arena ranks generative models via pairwise blind comparisons of outputs on user-submitted prompts; rankings shift continuously with the vote pool. The OP does not state sample sizes, prompt distributions, or whether the two Wan 2.7 entries were evaluated under matched conditions. The thread that follows would presumably elaborate but is not fetched here.

Image Arena: Wan 2.7 Pro #21, Wan 2.7 #22 — same band as Imagen 4 Generate 001 and FLUX.1 Kontext Max. Image Editing Arena: Wan 2.7 Pro #10, Wan 2.7 #8 (base outperforming Pro on the editing leaderboard per the OP’s ordering). The Editing-Arena commentary truncates before naming peer comparisons.

This is the first independent leaderboard datapoint Luma’s wiki has filed for the Wan 2.7 image stack — the feature-preview leaks (Wan 2.7 upcoming features preview (video editing, V2V, real-time)) and the launch announcement (Wan2.7-Image — unified model for image generation and editing (Alibaba Wan announcement)) both came from publisher-adjacent sources without third-party benchmarks. Two things stand out. (1) The Image-Arena placement next to Imagen 4 Generate 001 and FLUX.1 Kontext Max (FLUX.2 [klein]: Towards Interactive Visual Intelligence sibling product) puts an open-weights Alibaba model in the same band as a Google closed model and a BFL premium tier — relevant for the Open foundation-model releases thread, where Wan2.7-Image self-described as unified gen+edit+understanding but did not ship comparison numbers. (2) The Editing-Arena ordering — base Wan 2.7 above Wan 2.7 Pro (#8 vs #10) — is counter-intuitive enough to want a methodology check before treating it as signal; Design Arena’s pairwise blind votes are sensitive to prompt distribution and to whether the “Pro” tier was actually deployed in the matchups people saw.