100 Prompts — 10 Categories — 35 Models — Judged by LLM

The state of AI image generation.

Every model. The same 100 prompts. Scored side by side by an LLM judge — so you can see the difference, image next to image.

View the full leaderboard →
Best output by Grok Imagine 2.0 (Preview)
Current Champion
XAI
8.48 / 10
1 Grok Imagine 2.0 (Preview) sample output Grok Imagine 2.0 (Preview)
XAI
8.48
2 Takumi 1 sample output Takumi 1
Takumi
8.39
3 GPT Image 2 sample output GPT Image 2
OpenAI
8.31
4 Nano Banana 2 sample output Nano Banana 2
Google
8.13
5 Nano Banana 2 Lite sample output Nano Banana 2 Lite
Google
8.10

Summary

Our deep dive into the Image Battle dataset revealed some truly fascinating trends in the AI image generation landscape! 🌟

  • The Top Contenders: Grok Imagine 2.0 (Preview) took the crown with a stellar 8.48 overall score, closely followed by Takumi 1 (8.39) and GPT Image 2 (8.31).
  • Text is the New Frontier: Models are finally getting text right! The best models smoothly handled prompts like the Evergreen Brew logo.
  • The Refusal Hurdle: Interestingly, some top models like GPT Image 2 had high refusal rates (12%) due to safety filters, affecting their overall reliability.
  • Notable Surprises: The Ultra Hard category proved its name, bringing average scores down to 6.55 as models struggled with reverse-logic and complex spatial coherence.

General Analysis & Insights

Let's break down what makes these models tick (and where they stumble)! 🔍

🏆 Superb Strengths

🚧 Common Weaknesses

  • The 'AI Gloss': Many models still struggle to escape the hyper-polished 'CGI look', holding them back in the Photorealistic People & Portraits category.
  • Anatomical Anomalies: Hands interacting (like a Handshake) or complex crowds often lead to merged limbs or missing fingers.
  • Logical Reversals: Models rely heavily on training biases. In the Astronaut ridden by a horse prompt, many models stubbornly drew an astronaut riding a horse instead! 🐴

📊 Key Differentiators

The true line between a good model and a great model is Detail Execution and Prompt Adherence. For instance, Takumi 1 consistently scores high by adhering strictly to the prompt without hallucinating extra elements.

Best Model Analysis by Use Case

Choosing the right model depends entirely on what you want to create! Here is the breakdown by category: 🎨

📸 Photorealism & Portraits

Top Pick: Grok Imagine 2.0 (Preview) and Nano Banana 2 Lite. If you need lifelike skin textures and realistic lighting, dive into the Photorealistic People & Portraits category. Models here excel at capturing minute details, like the reflections in the Elderly woman with bifocals prompt. Check out this Elderly woman portrait for a great example!

✍️ Graphic Design & Typography

Top Pick: Grok Imagine 2.0 (Preview) and Takumi 1. For logos, UI elements, and social media posts, the Graphic Design category is key. These models cleanly rendered the Evergreen Brew logo and handled the WORLD PEACE NOW billboard challenge with minimal typos.

🌸 Anime & Stylized Art

Top Pick: Grok Imagine 2.0 (Preview) and GPT Image 2. Whether it's the Anime & Cartoon Style or capturing the whimsy of the Ghibli style, these models beautifully handled prompts like the Train station with cat.

🏗️ Architecture & Complex Scenes

Top Pick: Grok Imagine 2.0 (Preview) and Seedream 5.0 Pro. When designing Architecture & Interiors or Complex Scenes, spatial reasoning is vital. These models successfully mapped out the Japanese machiya townhouse and managed large, bustling market crowds without losing coherence.

🤯 Ultra Hard Scenarios

Top Pick: Takumi 1. When you need recursive logic, like a Robot painting self-portrait, the Ultra Hard category pushes models to their absolute limits.