Skip to content

Arena.ai Text-to-Image Leaderboard

Arena.ai’s public text-to-image leaderboard aggregates crowd-sourced pairwise human-preference votes into an Elo-style ranking spanning 72 model entries across ~15 labs (OpenAI, Google, Meta, Microsoft AI, xAI, Alibaba, ByteDance, Luma AI, Black Forest Labs, Ideogram, Recraft, Reve, HiDream, Krea, Runway, Tencent, Nvidia, Z.ai, Pruna, Leonardo, Stability). At the snapshot fetched on 2026-07-07 the top of the board is an OpenAI model at 1385 Elo (58,643 votes), followed by Meta’s muse-image at 1280 and Reve at 1271. Luma AI holds ranks 13 (1188 Elo, 13,460 votes) and 15 (1178 Elo, 19,420 votes). The page is the definitive external reference for how the closed-and-open T2I cohort — Nano Banana, Seedream, FLUX, Qwen-Image, MAI-Image-2, Wan-Image, Ideogram, Recraft, HiDream, Ovis-Image — currently sorts under blind human vote.

  • OpenAI holds the #1 slot at Elo 1385 ± 5 across 58,643 votes, with a rank-spread of 1–1 (no overlap with the #2 model’s confidence interval) [§leaderboard row 1].
  • Meta’s muse-image sits at #2 (Elo 1280) and Reve at #3 (Elo 1271), both proprietary [§leaderboard rows 2-3].
  • Luma AI has two entries on the board — rank 13 (Elo 1188, 13,460 votes) and rank 15 (Elo 1178, 19,420 votes) — both proprietary [§leaderboard rows 13, 15].
  • The top four positions are exclusively proprietary (OpenAI, Meta, Reve, Google); Ideogram’s open model at rank 11 (Elo 1208) is the highest-ranked open-weights entry [§leaderboard rows 1-11].
  • Vote counts on individual entries vary by more than two orders of magnitude, from 2,512 (Recraft rank 17) to 805,011 (Google rank 24) [§leaderboard rows 17, 24].
  • Alibaba’s Apache-2.0-licensed entries (Qwen-Image family) span ranks 33 (Elo 1127, 82,460 votes), 48, 51 (Elo 1060, 704,524 votes), and 55 — indicating heavy adoption but middle-of-pack human preference [§leaderboard rows 33, 48, 51, 55].

The leaderboard is a live web page rendering an Elo ranking table with rank number, rank-spread bounds (the confidence interval of the rank), lab name, license (Proprietary vs open license names), Elo score ± 95% CI, and cumulative vote count per model. No methodology page is exposed at the fetched URL; models are referenced by lab + license only, without explicit product names, making cross-referencing with individual model pages non-trivial.

Full ranked list at fetch time (2026-07-07, top 15 only for brevity — full 72-entry table lives on the page):

RankLabLicenseEloVotes
1OpenAIProprietary1385 ± 558,643
2Meta (muse-image)Proprietary1280 ± 77,715
3ReveProprietary1271 ± 712,486
4GoogleProprietary1270 ± 488,379
5Microsoft AIProprietary1257 ± 531,415
6GoogleProprietary1253 ± 97,536
7GoogleProprietary1245 ± 4126,408
8OpenAIProprietary1241 ± 3131,665
9GoogleProprietary1232 ± 582,564
10xAIProprietary1229 ± 537,621
11IdeogramOpen1208 ± 616,819
12AlibabaProprietary1193 ± 86,890
13Luma AIProprietary1188 ± 613,460
14Microsoft AIProprietary1183 ± 549,042
15Luma AIProprietary1178 ± 519,420

This is the live human-preference reference the field cites when positioning new T2I releases — Introducing MAI-Image-2: for limitless creativity anchors its entire announcement on “#3 on the Arena.ai text-to-image leaderboard”, and Introducing Recraft V4: Design Taste Meets Image Generation similarly pegs V4 to leaderboard placement. Luma’s two entries at ranks 13 and 15 sit in a dense middle-of-pack cluster (Elo 1170–1200, ~5% of the score range) where rank shifts by a handful of positions with modest Elo deltas, so the more useful signal from this page is which labs have entries at the frontier vs which have entries only in the long tail — not the absolute rank of any single model. Contrast with MiMo-V2-Pro (a.k.a. Hunter Alpha) — Artificial Analysis model card and the Video Arena referenced by LTX hits #3 Image-to-Video and #4 Text-to-Video on Artificial Analysis Video Arena (LTX announcement) and Wan 2.7 Pro and Wan 2.7 enter Image Arena and Image Editing Arena leaderboards (Design Arena): those are the parallel human-preference leaderboards for LLMs and video respectively, and the design pattern (crowd-sourced blind pairwise vote → Elo) is now the standard for cross-lab visual-generation comparison, though RoboArena rolls back evaluations after benchmark hacking observed since April (Pranav Atreya announcement) documents integrity attacks on the closely related RoboArena — a reminder that Elo aggregates are only as clean as the vote-collection pipeline.