Arena.ai Text-to-Image Leaderboard
Arena.ai’s public text-to-image leaderboard aggregates crowd-sourced pairwise human-preference votes into an Elo-style ranking spanning 72 model entries across ~15 labs (OpenAI, Google, Meta, Microsoft AI, xAI, Alibaba, ByteDance, Luma AI, Black Forest Labs, Ideogram, Recraft, Reve, HiDream, Krea, Runway, Tencent, Nvidia, Z.ai, Pruna, Leonardo, Stability). At the snapshot fetched on 2026-07-07 the top of the board is an OpenAI model at 1385 Elo (58,643 votes), followed by Meta’s muse-image at 1280 and Reve at 1271. Luma AI holds ranks 13 (1188 Elo, 13,460 votes) and 15 (1178 Elo, 19,420 votes). The page is the definitive external reference for how the closed-and-open T2I cohort — Nano Banana, Seedream, FLUX, Qwen-Image, MAI-Image-2, Wan-Image, Ideogram, Recraft, HiDream, Ovis-Image — currently sorts under blind human vote.
Key claims
Section titled “Key claims”- OpenAI holds the #1 slot at Elo 1385 ± 5 across 58,643 votes, with a rank-spread of 1–1 (no overlap with the #2 model’s confidence interval) [§leaderboard row 1].
- Meta’s
muse-imagesits at #2 (Elo 1280) and Reve at #3 (Elo 1271), both proprietary [§leaderboard rows 2-3]. - Luma AI has two entries on the board — rank 13 (Elo 1188, 13,460 votes) and rank 15 (Elo 1178, 19,420 votes) — both proprietary [§leaderboard rows 13, 15].
- The top four positions are exclusively proprietary (OpenAI, Meta, Reve, Google); Ideogram’s open model at rank 11 (Elo 1208) is the highest-ranked open-weights entry [§leaderboard rows 1-11].
- Vote counts on individual entries vary by more than two orders of magnitude, from 2,512 (Recraft rank 17) to 805,011 (Google rank 24) [§leaderboard rows 17, 24].
- Alibaba’s Apache-2.0-licensed entries (Qwen-Image family) span ranks 33 (Elo 1127, 82,460 votes), 48, 51 (Elo 1060, 704,524 votes), and 55 — indicating heavy adoption but middle-of-pack human preference [§leaderboard rows 33, 48, 51, 55].
Method
Section titled “Method”The leaderboard is a live web page rendering an Elo ranking table with rank number, rank-spread bounds (the confidence interval of the rank), lab name, license (Proprietary vs open license names), Elo score ± 95% CI, and cumulative vote count per model. No methodology page is exposed at the fetched URL; models are referenced by lab + license only, without explicit product names, making cross-referencing with individual model pages non-trivial.
Results
Section titled “Results”Full ranked list at fetch time (2026-07-07, top 15 only for brevity — full 72-entry table lives on the page):
| Rank | Lab | License | Elo | Votes |
|---|---|---|---|---|
| 1 | OpenAI | Proprietary | 1385 ± 5 | 58,643 |
| 2 | Meta (muse-image) | Proprietary | 1280 ± 7 | 7,715 |
| 3 | Reve | Proprietary | 1271 ± 7 | 12,486 |
| 4 | Proprietary | 1270 ± 4 | 88,379 | |
| 5 | Microsoft AI | Proprietary | 1257 ± 5 | 31,415 |
| 6 | Proprietary | 1253 ± 9 | 7,536 | |
| 7 | Proprietary | 1245 ± 4 | 126,408 | |
| 8 | OpenAI | Proprietary | 1241 ± 3 | 131,665 |
| 9 | Proprietary | 1232 ± 5 | 82,564 | |
| 10 | xAI | Proprietary | 1229 ± 5 | 37,621 |
| 11 | Ideogram | Open | 1208 ± 6 | 16,819 |
| 12 | Alibaba | Proprietary | 1193 ± 8 | 6,890 |
| 13 | Luma AI | Proprietary | 1188 ± 6 | 13,460 |
| 14 | Microsoft AI | Proprietary | 1183 ± 5 | 49,042 |
| 15 | Luma AI | Proprietary | 1178 ± 5 | 19,420 |
Why it’s interesting
Section titled “Why it’s interesting”This is the live human-preference reference the field cites when positioning new T2I releases — Introducing MAI-Image-2: for limitless creativity anchors its entire announcement on “#3 on the Arena.ai text-to-image leaderboard”, and Introducing Recraft V4: Design Taste Meets Image Generation similarly pegs V4 to leaderboard placement. Luma’s two entries at ranks 13 and 15 sit in a dense middle-of-pack cluster (Elo 1170–1200, ~5% of the score range) where rank shifts by a handful of positions with modest Elo deltas, so the more useful signal from this page is which labs have entries at the frontier vs which have entries only in the long tail — not the absolute rank of any single model. Contrast with MiMo-V2-Pro (a.k.a. Hunter Alpha) — Artificial Analysis model card and the Video Arena referenced by LTX hits #3 Image-to-Video and #4 Text-to-Video on Artificial Analysis Video Arena (LTX announcement) and Wan 2.7 Pro and Wan 2.7 enter Image Arena and Image Editing Arena leaderboards (Design Arena): those are the parallel human-preference leaderboards for LLMs and video respectively, and the design pattern (crowd-sourced blind pairwise vote → Elo) is now the standard for cross-lab visual-generation comparison, though RoboArena rolls back evaluations after benchmark hacking observed since April (Pranav Atreya announcement) documents integrity attacks on the closely related RoboArena — a reminder that Elo aggregates are only as clean as the vote-collection pipeline.
See also
Section titled “See also”- Open foundation-model releases — the concept page tracking the closed-vs-open T2I cohort this leaderboard ranks
- Introducing MAI-Image-2: for limitless creativity — announcement post that positions itself entirely on Arena.ai leaderboard placement
- Introducing Recraft V4: Design Taste Meets Image Generation — another T2I product post citing Arena.ai
- Sophon — AI Research Catalog of Evals, Tools, and Labs — meta-catalog of eval leaderboards including LMArena / Design Arena / Artificial Analysis
- Wan 2.7 Pro and Wan 2.7 enter Image Arena and Image Editing Arena leaderboards (Design Arena) — Wan 2.7 entering Image and Image-Editing arena leaderboards (Design Arena parallel to Arena.ai)
- LTX hits #3 Image-to-Video and #4 Text-to-Video on Artificial Analysis Video Arena (LTX announcement) — LTX video ranking on Artificial Analysis Video Arena (video-side counterpart)
- MiMo-V2-Pro (a.k.a. Hunter Alpha) — Artificial Analysis model card — Artificial Analysis LLM leaderboard (LLM-side counterpart)
- RoboArena rolls back evaluations after benchmark hacking observed since April (Pranav Atreya announcement) — RoboArena benchmark-hacking retraction, a caution on crowd-sourced eval integrity