The best AI image generator depends on what you need it to get right. A convincing portrait, an exact line of text and an edit that preserves a product label are different jobs. A feature list alone does not show how a model handles those details.
On August 2, 2026, the Mireka team tested five image models in the same workspace across five tasks: Japanese ad copy, realistic people, anime illustration, product photography and reference-image editing. We generated each model–task combination once and kept the first successful result, including visible mistakes.
GPT Image 2 was our strongest all-round choice in this test. Nano Banana 2 and Seedream 5.0 Pro stood out for product edits, while Wan 2.7 Image Pro combined fast generation with accurate Japanese ad text.
All five tests used Japanese prompts. This English edition preserves the original images, inputs, scores and timings, with English explanations. It does not present the results as a new benchmark of English-language prompting or English typography.
Which AI image model should you choose?
Start with GPT Image 2 when following a detailed brief matters more than speed. It handled the advertising and anime constraints most consistently, but its median generation time was 214.2 seconds—the slowest in this run.
Choose Nano Banana 2 for a quicker starting point for portraits, clean product shots and reference edits. Its main failure here was the Japanese ad: it replaced all three required phrases with different wording.
Choose Seedream 5.0 Pro when the task combines Japanese text, people or product-reference fidelity. Choose Wan 2.7 Image Pro when you want to test a Japanese ad concept quickly. Grok Imagine Image was fast for rough visual exploration, but less reliable on exact text and detailed constraints.
| Model | Average across 20 scores | Median time | Best fit in this test | Main trade-off |
|---|---|---|---|---|
| GPT Image 2 | 4.80/5 | 214.2s | Detailed briefs, Japanese ads and anime | Longest wait |
| Seedream 5.0 Pro | 4.65/5 | 92.2s | Japanese text, portraits and product edits | Missed several anime constraints |
| Nano Banana 2 | 4.40/5 | 31.2s | Portraits, product shots and reference edits | Rewrote the required Japanese ad copy |
| Wan 2.7 Image Pro | 4.00/5 | 18.2s | Fast Japanese ad generation | Product shape and reference details drifted |
| Grok Imagine Image | 3.55/5 | 18.2s | Quick compositions and portrait drafts | Exact Japanese text and fine constraints |
| Midjourney | — | — | Not tested in this run | No output, timing or score recorded |
These averages describe the images shown below. They are not a universal ranking: each model produced just one image per task, and repeated runs can differ. The most useful result is the one closest to your own work.
How we tested the models
We used the same conditions for each comparison:
| Setting | Test conditions |
|---|---|
| Test date | August 2, 2026 |
| Platform | Mireka |
| Output | 16:9, 1K, one image |
| Attempts | One run per model per scenario |
| Published result | First successful output; no cherry-picking |
| Prompt language | Identical Japanese prompt within each scenario |
| Extra processing | No negative prompt, retouching or upscaling |
| Timing | From API submission to receipt of the result URL |
| Scoring | Four visual criteria per task, each scored from 1 to 5 |
The generation times include service queues and network overhead at the time of the test. They are not measurements of model inference alone or guarantees for another plan, resolution or date.
This comparison was produced by the Mireka team using models available in Mireka's generation environment. We publish the first outputs, prompts, settings, timings and review notes so you can inspect the evidence. Incorrect text, missing instructions, altered shapes and watermark-like marks have not been fixed in the displayed images.
Try these settings loads the selected model, original Japanese prompt, aspect ratio, resolution and any reference image into Mireka's AI image generator. You can review the settings and credit cost before generating. Clicking the button does not generate an image or spend credits, and it does not put your prompt in the URL.
Test 1: Exact Japanese text in an ad
We asked each model to place three phrases without changing any characters:
- 「夏の夜、ひと休み。」 — roughly, “Take a break on a summer evening.”
- 「涼風ソーダ」 — “Cool Breeze Soda,” the fictional product name.
- 「8月限定」 — “August only.”
The English meanings help explain the brief. The scores judge the Japanese characters in the actual images, not an English translation of the copy.
Japanese text in an ad
Can the model reproduce three exact Japanese phrases and deliver a usable ad layout?
English translation for reference
A landscape 16:9 Japanese advertising poster. A midsummer evening, blue soda in a clear glass, fine bubbles and condensation, with a deep-indigo-to-light-blue background. Place “夏の夜、ひと休み。” large at the top, “涼風ソーダ” in the center, and “8月限定” small in the bottom-right corner. Reproduce the Japanese text exactly, without changing a single character. Modern sans-serif typography, generous negative space, no real brand logos or watermarks.
Trying these settings loads the original Japanese prompt used in the test, not this English translation.
Original test prompt (Japanese)
横長16:9の日本語広告ポスター。真夏の夕暮れ、透明なグラスに注がれた青いソーダ、細かな気泡と水滴、深い藍色から水色への背景。上部に大きく「夏の夜、ひと休み。」、中央に「涼風ソーダ」、右下に小さく「8月限定」と、日本語を一字も変えず正確に配置。モダンなゴシック体、十分な余白、実在ブランドのロゴなし、透かしなし。
GPT Image 2, Seedream 5.0 Pro and Wan 2.7 Image Pro rendered all three strings correctly. Wan completed this particular ad in 18.4 seconds. Nano Banana 2 produced an attractive layout but changed every phrase; Grok's text also differed from the request.
Our pick for this task: GPT or Seedream for a polished first draft, with Wan worth trying when speed matters. Always check exact wording before using generated artwork in an ad.
Test 2: A realistic person holding a book
The brief placed a fictional Japanese woman in her 30s by a bookstore window, holding a red clothbound book with both hands. We checked skin, anatomy, lighting and whether the setting matched the instructions.
Photorealistic portraits and hands
We checked the bookstore scene, red book and backlighting, as well as the anatomy of the hands holding the book.
English translation for reference
A landscape 16:9 photorealistic lifestyle photograph. In a quiet Tokyo bookstore, a Japanese woman in her 30s stands by a window, naturally holding a red clothbound book with both hands. She is a fictional generated person, not a likeness of a real person. Off-white shirt, short black hair, calm expression. Soft 4 p.m. backlighting, 50mm lens, natural skin texture, five fingers on each hand, wooden bookshelves in the background. No text, logos or watermarks.
Trying these settings loads the original Japanese prompt used in the test, not this English translation.
Original test prompt (Japanese)
横長16:9の写実的なライフスタイル写真。東京の静かな書店で、30代の日本人女性が窓辺に立ち、赤い布張りの本を両手で自然に持っている。生成人物であり実在人物には似せない。生成り色のシャツ、短い黒髪、穏やかな表情。午後4時の柔らかな逆光、50mmレンズ、自然な肌の質感、指は左右5本ずつ、背景に木製の本棚。文字、ロゴ、透かしなし。
Nano Banana 2 and Seedream produced natural hands, book placement and window light. GPT's lighting and skin were convincing, but the book hid some fingers, making that part harder to assess. Wan added text to the book cover despite the instruction not to include lettering.
Our pick for this task: Nano Banana 2 offered a strong balance of finish and wait time, completing the portrait in 27.3 seconds. Seedream was also a strong option for the finished image.
Test 3: An anime scene with several precise requirements
The scene called for a station platform after rain, a student holding a closed transparent umbrella, a yellow train, puddle reflections, one rainbow and a clock showing 5:20 p.m.
An appealing anime image can still fail a brief. Here, we scored the requested objects and their states as well as the composition.
Anime illustration with multiple constraints
We checked the closed umbrella, yellow train, puddles, rainbow and a clock showing 5:20 p.m. together.
English translation for reference
A high-quality landscape 16:9 Japanese anime-style illustration. At dusk after rain, one high school student in a navy uniform stands on a station platform holding a closed transparent umbrella. A yellow train sits in the back-left, with puddles and reflections at the student’s feet, one rainbow in the sky, and a platform clock showing 17:20. Only one person. Muted blue and amber colors, cinematic composition and a detailed background. No text, logos or watermarks.
Trying these settings loads the original Japanese prompt used in the test, not this English translation.
Original test prompt (Japanese)
横長16:9の高品質な日本アニメ風イラスト。雨上がりの夕方、紺色の制服を着た高校生が透明な傘を閉じて駅のホームに立つ。左奥に黄色い列車、足元に水たまりと反射、空に一本の虹、ホームの時計は17時20分。人物は1人だけ、落ち着いた青と琥珀色、映画的な構図、細密な背景。文字、ロゴ、透かしなし。
GPT Image 2 combined the requested details most reliably. Nano Banana 2 included the main elements but missed the clock time. Seedream's image was attractive, yet the umbrella was open, the yellow train was unclear and the time was incorrect. Grok and Wan also missed the umbrella and clock requirements.
Our pick for this task: GPT Image 2 when the scene must follow a detailed storyboard. If you are exploring atmosphere rather than exact continuity, compare the other compositions visually before deciding.
Test 4: A clean ecommerce product shot
This was a text-to-image test of a fictional product: one matte sage-green cylindrical stainless-steel bottle with a slim silver cap, on a light background. We looked at geometry, material, light and a clean product presentation.
Ecommerce product photography
We checked whether the model kept the cylindrical shape, sage-green finish, silver cap and single-product composition.
English translation for reference
A landscape 16:9 ecommerce product photograph. Place one matte sage-green cylindrical stainless-steel bottle in the center. A slim silver cap, with no logo or text. Bright gray-white background, soft studio light from the back-left and a subtle, natural contact shadow. Accurate product silhouette and material. No extra props, hands, people, packaging or watermarks.
Trying these settings loads the original Japanese prompt used in the test, not this English translation.
Original test prompt (Japanese)
横長16:9のEC向け商品写真。中央にマットなセージグリーンの円筒形ステンレスボトルを1本だけ置く。銀色の細いキャップ、ロゴや文字なし。明るいグレーホワイトの背景、左後方から柔らかなスタジオ光、自然で薄い接地影。商品の輪郭と素材感を正確に、余計な小物、手、人物、包装、透かしなし。
GPT and Nano Banana 2 followed the cylinder and silver-cap instructions well. Seedream rounded the bottle's shoulders. Grok changed the cap and shape, while Wan introduced a more curved silhouette.
Our pick for this task: Nano Banana 2 for a quick, clean product concept, or GPT for detailed instruction following. For a real SKU, use a reference image rather than asking a text prompt to invent the product.
Test 5: Change the background without changing the product
Each model received the same yuzu soda bottle reference. The task was to preserve the bottle, cap, logo and Japanese label while replacing the background with a bright café setting and adding a small amount of condensation.
Reference supplied to every model in this test. The Japanese packaging is part of the product and must stay unchanged.
Background editing with label preservation
Using the same product reference, we asked each model to replace the background with a café while preserving the bottle and Japanese label.
English translation for reference
Use the uploaded yuzu soda bottle as the sole product reference. Do not change its shape, proportions, color, cap, label, logo or Japanese text. Replace only the background with a clean café featuring a bright wooden table and soft window light on a summer morning. Add a small amount of natural condensation to the bottle and preserve the contact shadow. Do not add new text, fruit, straws, hands, people, other products or watermarks. Landscape 16:9.
Trying these settings loads the original Japanese prompt used in the test, not this English translation.
Original test prompt (Japanese)
アップロードした柚子ソーダの商品ボトルを唯一の基準として使い、ボトルの形、比率、色、キャップ、ラベル、ロゴ、日本語表記を一切変更しないでください。背景だけを、夏の朝の明るい木製テーブルと柔らかな窓光のある清潔なカフェに置き換えます。ボトル表面に自然な冷たい水滴を少量追加し、接地影を保ちます。新しい文字、果物、ストロー、手、人物、別の商品、透かしは追加しないでください。横長16:9。
Nano Banana 2 and Seedream preserved the product and label particularly well, with natural backgrounds and contact shadows. GPT retained the bottle but added a watermark-like mark in the bottom-right corner. Grok introduced small changes to the fine print and proportions. Wan changed the neck, cap and amount of condensation.
Our pick for this task: Nano Banana 2 or Seedream. Product fidelity matters more than an attractive new background when the image needs to represent something you sell.
The same bottle appears in our yuzu soda ecommerce image case study, which shows how one reference became a coordinated set of listing images.
What the comparison means for your workflow
Use the scores to choose your first model, then inspect the result against your own brief:
| Your priority | Start with | Check before moving on |
|---|---|---|
| Many precise instructions in one image | GPT Image 2 | Whether the extra waiting time is worthwhile |
| Fast portraits and product concepts | Nano Banana 2 | Exact text, if the image contains it |
| Japanese lettering plus product-reference edits | Seedream 5.0 Pro | Small scene constraints and object states |
| Quick Japanese ad drafts | Wan 2.7 Image Pro | Product geometry and overly busy decoration |
| Fast composition exploration | Grok Imagine Image | Lettering and details before treating it as final |
Mireka makes it easy to switch models in one workspace as the task changes—from an initial concept to a reference edit. You do not have to commit to the same model for every stage of a project.
For an exact brand name or product label, zoom in. For a character scene, count the required objects and check the pose. For an ecommerce edit, compare the output directly with the original product. A high overall score does not replace those checks.
Free access, credits and commercial projects
A model name does not tell you the cost of using it through a particular service. Check the credits shown for your selected model and settings before generating, and review your plan's commercial-use terms before publishing business assets. Free allowances and plan entitlements can change.
This benchmark also cannot establish that an output is cleared for every commercial use. Use input materials you are authorized to upload and review the finished content for inaccurate product claims, unintended likenesses and third-party branding.
Frequently asked questions
Which AI image model was best overall in this comparison?
GPT Image 2 had the highest average score in these five tests and followed complex instructions most consistently. Nano Banana 2 and Seedream were stronger starting points for some editing tasks, and the faster models had much shorter wait times.
Which models reproduced the Japanese ad text accurately?
GPT Image 2, Seedream 5.0 Pro and Wan 2.7 Image Pro reproduced all three requested Japanese phrases correctly in this run. Nano Banana 2 and Grok Imagine Image changed the wording.
Which models were best for editing a product photo?
Nano Banana 2 and Seedream 5.0 Pro performed particularly well in the label-preservation test, keeping the bottle and Japanese label consistent while replacing the background naturally.
Can I try these image models for free?
Available free credits depend on the service and current offer. Mireka shows the required credits after you select a model and settings, so you can check the cost before generating.
Why can the same prompt produce a different result?
Image generation includes randomness. Composition and small details can change even with the same model and settings. This article shows the first successful image from each test, rather than selecting the best of several attempts.
Why are the Midjourney results blank?
Midjourney could not be included in this Mireka test run. Its output, timing and scores are left blank, and it is excluded from the ranking rather than being given estimated results.
Can I use the generated images commercially?
Check the service's current terms, your plan's entitlements, the rights to your input materials and the content of the output. Commercial-use conditions depend on more than the model name alone.
