Ultra Hard: best image by XAI Grok Imagine 2.0
1 XAI Grok Imagine 2.0 1438 Elo
Ultra Hard: best image by Takumi 1
2 Takumi 1 1420 Elo
Ultra Hard: best image by OpenAI GPT Image 2.5 Sunburst
3 OpenAI GPT Image 2.5 Sunburst 1394 Elo

Full ranking for Ultra Hard

#ModelCategory EloOverall rankBest image
1 XAI Grok Imagine 2.0 1438 ±130 #5 Ultra Hard image by XAI Grok Imagine 2.0
2 Takumi 1 1420 ±127 #4 Ultra Hard image by Takumi 1
3 OpenAI GPT Image 2.5 Sunburst 1394 ±144 #1 Ultra Hard image by OpenAI GPT Image 2.5 Sunburst
4 OpenAI GPT Image 2.5 Flare 1344 ±163 #8 Ultra Hard image by OpenAI GPT Image 2.5 Flare
5 OpenAI GPT Image 2 1334 ±176 #2 Ultra Hard image by OpenAI GPT Image 2
6 Google Nano Banana 2.1 1313 ±112 #7 Ultra Hard image by Google Nano Banana 2.1
7 Alibaba Qwen-Image-3.0-Pro 1310 ±176 #9 Ultra Hard image by Alibaba Qwen-Image-3.0-Pro
8 OpenAI GPT Image 1.5 1286 ±111 #13 Ultra Hard image by OpenAI GPT Image 1.5
9 Black Forest Labs FLUX 3 Image 1278 ±197 #6 Ultra Hard image by Black Forest Labs FLUX 3 Image
10 Microsoft MAI-Image-2.6 1273 ±147 #10 Ultra Hard image by Microsoft MAI-Image-2.6
11 XAI Grok Imagine (Quality) 1237 ±215 #18 Ultra Hard image by XAI Grok Imagine (Quality)
12 Google Nano Banana Pro 1232 ±160 #11 Ultra Hard image by Google Nano Banana Pro
13 Meta Muse Image 1214 ±66 #3 Ultra Hard image by Meta Muse Image
14 XAI Grok Imagine 1171 ±81 #14 Ultra Hard image by XAI Grok Imagine
15 Google Nano Banana 2 Lite 1170 ±156 #17 Ultra Hard image by Google Nano Banana 2 Lite
16 Reve 2.1 1151 ±159 #16 Ultra Hard image by Reve 2.1
17 Google Nano Banana 2 1128 ±145 #12 Ultra Hard image by Google Nano Banana 2
18 Bytedance Seedream 5.0 Pro 1074 ±210 #15 Ultra Hard image by Bytedance Seedream 5.0 Pro
19 Recraft V4.1 1061 ±152 #27 Ultra Hard image by Recraft V4.1
20 Ideogram 4.0 (Quality) 1020 ±68 #21 Ultra Hard image by Ideogram 4.0 (Quality)
21 OpenAI ChatGPT 4o 1004 ±214 #24 Ultra Hard image by OpenAI ChatGPT 4o
22 Black Forest Labs FLUX.2 Max 983 ±237 #19 Ultra Hard image by Black Forest Labs FLUX.2 Max
23 Google Imagen 4.0 Ultra 978 ±145 #23 Ultra Hard image by Google Imagen 4.0 Ultra
24 Bytedance Seedream 4.5 977 ±164 #25 Ultra Hard image by Bytedance Seedream 4.5
25 Black Forest Labs Flux 2 Pro 970 ±152 #26 Ultra Hard image by Black Forest Labs Flux 2 Pro
26 Bytedance Seedream 4.0 961 ±162 #22 Ultra Hard image by Bytedance Seedream 4.0
27 Recraft V3 926 ±118 #29 Ultra Hard image by Recraft V3
28 Reve Image (Halfmoon) 903 ±153 #31 Ultra Hard image by Reve Image (Halfmoon)
29 MiniMax Image-01 885 ±216 #37 Ultra Hard image by MiniMax Image-01
30 Google Imagen 3.0 856 ±131 #28 Ultra Hard image by Google Imagen 3.0
31 Google Nano Banana (2.5 Flash) 851 ±109 #20 Ultra Hard image by Google Nano Banana (2.5 Flash)
32 Black Forest Labs FLUX.1 Kontext Max 808 ±85 #33 Ultra Hard image by Black Forest Labs FLUX.1 Kontext Max
33 Black Forest Labs Flux 1.1 Pro Ultra 727 ±115 #30 Ultra Hard image by Black Forest Labs Flux 1.1 Pro Ultra
34 Krea 2 Large 697 ±64 #39 Ultra Hard image by Krea 2 Large
35 Bytedance Seedream 3.0 691 ±167 #32 Ultra Hard image by Bytedance Seedream 3.0
36 Ideogram 3.0 (Quality) 653 ±80 #34 Ultra Hard image by Ideogram 3.0 (Quality)
37 Midjourney V6.1 621 ±197 #38 Ultra Hard image by Midjourney V6.1
38 OpenAI DALL-E 3 615 ±148 #36 Ultra Hard image by OpenAI DALL-E 3
39 Ideogram V2 562 ±111 #35 Ultra Hard image by Ideogram V2
40 XAI Grok 2 Image 539 ±194 #42 Ultra Hard image by XAI Grok 2 Image
41 Alibaba Z-Image Turbo 515 ±166 #41 Ultra Hard image by Alibaba Z-Image Turbo
42 Midjourney v7 430 ±304 #40 Ultra Hard image by Midjourney v7

Summary for Ultra Hard

Welcome to the ultimate stress test for AI image generation! 🚀 The Ultra Hard category is specifically designed to push models to their absolute cognitive and rendering limits. These prompts test absurd spatial relationships, precise typography, retro UI replication, and cultural/anatomical accuracy.

Here are the key discoveries from our analysis:

  • 🥇 Top Performers: Grok Imagine 2.0 dominates the category with an impressive 1439 Elo. It is closely followed by Takumi 1 (1420 Elo) and GPT Image 2.5 Sunburst (1394 Elo). These models excel because they possess deep semantic understanding, refusing to take the "easy way out" when given contradictory or complex instructions.
  • 🧠 Intelligence over Beauty: A major trend is the triumph of prompt adherence over pure artistic polish. Models that usually top aesthetic charts, such as Midjourney v7 (975 Elo) and Midjourney V6.1 (767 Elo), struggled significantly here. When asked to draw a horse riding an astronaut, they defaulted to an astronaut riding a horse—resulting in severe penalties.
  • 😲 Surprising Results: The ability of top models to handle dense, coherent mathematics is staggering. In the OpenAI Math Lecture prompt, top models successfully generated legible Bellman equations and cross-entropy formulas instead of random gibberish.
  • 📌 Quick Takeaway: If your workflow requires precise instruction following, complex typography, or unusual spatial layouts, look to the latest generations of Grok, Takumi, and GPT Image.

The 10 Ultra Hard prompts