Multi-Ref Editing
Pass up to 5 reference images in one call. The model holds subject identity across iterations, so a character or product keeps its face, colors, and proportions through successive edit rounds instead of drifting.

curl --request POST "https://gptproto.com/api/v3/grok/grok-imagine-image-2.0/text-to-image" \
--header "Authorization: Bearer $GPTPROTO_API_KEY" \
--header "Content-Type: application/json" \
--data '{
"prompt": "A tiny origami fox sailing a teacup across a moonlit puddle",
"aspect_ratio": "auto",
"resolution": "1k",
"quality": "medium",
"enable_sync_mode": false,
"response_format": "url"
}'Start from the cost of a single sample and pick a testing budget.
Released August 7, 2026, Grok Imagine Image 2.0 scores 1,439 Elo on the LMArena image-edit leaderboard — second only to GPT-Image-2's 1,463 — and 1,316 Elo on text-to-image. GPTProto serves it as grok-imagine-image-2.0 at $0.04 per image, with 1K or 2K output, up to 5 reference images per call, and 14 aspect ratios from 1:1 to 20:9. One API key and one balance cover this model and 200+ others, so there is no separate SpaceXAI account, no regional sign-up check, and no per-provider credit pool to top up.
Pass up to 5 reference images in one call. The model holds subject identity across iterations, so a character or product keeps its face, colors, and proportions through successive edit rounds instead of drifting.

Magic Wand segments an uploaded image into layers, so edits apply to one selected region and leave the rest untouched. Any isolated subject exports on a transparent background.

Recompose one image across 9 aspect ratios with generative fill on the new edges. A single 1:1 master becomes 16:9, 9:16, and 4:3 deliverables without re-prompting or manual cropping.

Trained on photography, graphic design, and illustration separately, the model plans layout before rendering. Small-point type stays readable, which is where most diffusion models break on posters and e-commerce cards.

Grok Imagine Image 2.0 is the second generation of SpaceXAI's Imagine image model, shipped as the Quality Mode inside Grok's image generator on August 7, 2026 and exposed over API on August 12. It handles two jobs in one endpoint: generating images from a text prompt, and editing existing images from up to 5 references.
The training split is unusual and worth knowing before you pick it. Rather than one general aesthetic pass, SpaceXAI trained the model separately on photography, graphic design, and illustration. That shows up in three measurable places: lighting and material response on product shots, layout planning on multi-element compositions, and small-point text that survives rendering. It is the reason the model indexes higher on the edit board (1,439 Elo) than on pure text-to-image (1,316 Elo) — it is a stronger editor than it is a from-scratch generator.
Note the provider name. xAI was absorbed into SpaceX, and this model is listed on public leaderboards under the SpaceXAI brand. If you are searching for a "SpaceXAI Grok Imagine Image 2.0 API" and an "xAI" one, they are the same model.
| Spec | Value |
|---|---|
| Model string | grok-imagine-image-2.0 |
| Provider | SpaceXAI (formerly xAI) |
| Released | Aug 7, 2026 (product) · Aug 12, 2026 (API) |
| Input modalities | Text, Image |
| Output modality | Image |
| Resolution | 1k, 2k |
| Quality tiers | low, medium |
| Aspect ratios | 1:1, 3:4, 4:3, 9:16, 16:9, 2:3, 3:2, 9:19.5, 19.5:9, 9:20, 20:9, 1:2, 2:1, auto |
| Reference images | 0–5 per request (input_image_n) |
| Images per request | 1–10 (n) |
| Sync mode | enable_sync_mode (bool) |
| Response format | response_format |
| GPTProto price | $0.04 per image |
| Arena — Image Edit | 1,439 Elo (#2, 5,931 votes) |
| Arena — Text-to-Image | 1,316 Elo (#3, 2,675 votes) |
| SuperCLUE text-to-image | 90.32 |
All four models are live on GPTProto under the same key, so this comparison is about fit, not access. Scores below are LMArena Elo from the August 25, 2026 snapshot.
| Grok Imagine Image 2.0 | Seedream 5.0 Pro | Nano Banana 2 | GPT Image 2 | |
|---|---|---|---|---|
| Model string | grok-imagine-image-2.0 |
dola-seedream-5-0-pro-260628 |
gemini-3.1-flash-image |
gpt-image-2 |
| Provider | SpaceXAI | ByteDance | OpenAI | |
| Arena text-to-image Elo | 1,316 | 1,258 | 1,263 | 1,382 |
| Arena image-edit Elo | 1,439 | — | — | 1,463 |
| GPTProto price | $0.04 / image | $0.0405 (1K) · $0.081 (2K) | $0.0402 / image | $6.40 / $24.00 per 1M tokens |
| Reference images | up to 5 | Text + Image input | Text + Image input | up to 16 |
| Max resolution | 2K | native 2K | 1K | native 2K |
| Region-level editing | Magic Wand + segmentation | — | natural-language edits | mask-based |
| In-image text | strong, small-point legible | 14 languages | multilingual | multilingual incl. CJK |
| Billing model | flat per image | flat per image | flat per image | token-metered |
Read the table by job, not by total score.
If the work is editing — retouching a product shot, swapping a background, holding one character across a series — Grok Imagine Image 2.0 is the strongest pick on this list after GPT-Image-2, and the 24-point Elo gap between them (1,439 vs 1,463) is narrower than the gap in operating cost predictability. At a flat $0.04 per image, a 1,000-image edit batch costs $40 regardless of resolution or prompt length. GPT Image 2 bills per token, so the same batch lands somewhere in a range you only know after the fact — fine for exploratory work, awkward for a quoted client job.
If the work is text-to-image from scratch, GPT Image 2 wins on raw preference (1,382 Elo) and there is no honest way to argue otherwise. Grok Imagine Image 2.0 sits at 1,316, ahead of Seedream 5.0 Pro and Nano Banana 2, but behind OpenAI.
If the deliverable carries dense in-image copy in a non-Latin script, Seedream 5.0 Pro's 14-language text rendering is the more direct tool, and its native 2K comes in at $0.081 — double Grok's 2K cost.
If throughput matters more than the last few Elo points, Nano Banana 2 at $0.0402 is effectively the same price as Grok Imagine Image 2.0 and generates faster, but caps at 1K and offers no region-mask equivalent to Magic Wand.
One caveat on the numbers: Grok Imagine Image 2.0's Arena scores carry wider confidence intervals (±12 on text-to-image) than the established models, because it has 2,675 votes against GPT-Image-2's 75,014. Treat it as a strong first-tier signal rather than a settled ranking.
If you already call SpaceXAI directly, the migration is a base_url swap and a model-string swap — the request body stays OpenAI-compatible. Three differences to plan for:
Model naming. SpaceXAI's own API exposes this generation through its quality-tier model names. On GPTProto the string is grok-imagine-image-2.0, one string covering both the low and medium quality tiers via the quality parameter rather than two separate model IDs.
Account surface. Direct access requires a SpaceXAI account and its own prepaid balance. On GPTProto the same key and the same balance also reach Seedream 5.0 Pro, Nano Banana 2, GPT Image 2, and 200+ other models, so an A/B test across four image models is a one-line model-string change instead of four sign-ups, four billing relationships, and four sets of rate limits to reconcile.
Fallback path. Because the models share one balance, a failed or filtered generation can retry against a different image model in the same code path. Keep a secondary model string in config — gemini-3.1-flash-image is the closest match on price and modality at $0.0402 per image.
Grok Imagine Image 2.0 is positioned by SpaceXAI for commercial and professional work, a deliberate repositioning after the January 2026 deepfake incident that led xAI to restrict image generation to paid tiers. Two practical consequences for API users.
First, the model ships with pre-configured templates aimed at commercial categories — product stills, professional headshots, e-commerce listings, game concept art, marketing posters. These are product-surface presets rather than API parameters, but they signal what the training was optimized toward, and prompts written in those registers land more reliably.
Second, requests are subject to SpaceXAI's upstream content filtering, which GPTProto passes through unchanged. GPTProto does not add a second moderation layer and does not remove the provider's. A filtered request is a provider decision, not a platform one. If your pipeline needs a fallback for filtered generations, route the retry to another image model on the same key rather than reworking the prompt in place.
Guides, comparisons, and updates related to this model.
All Articles
Compare 6 cheapest AI image generators in 2026, from about $0.0035 per image. See batch costs, hidden fees, and the best API for startups.

Learn how to make an AI story video for kids with GPT Image 2 and Seedance 2.5, including the full prompt, captions, sound, editing, and a real test.

Compare 8 unrestricted AI photo editors for online, local, free, and image-to-image editing, including pricing, privacy, and key limits.

Compare 7 image editing AI models for APIs, batch workflows, product photos, text edits, and brand consistency, with the best pick for each task.
Input
Output
