768p Clips in Two Fixed Durations
Generate 6- or 10-second videos at 768p. The fixed duration options make per-clip budgeting predictable for previews, product loops, social posts, and storyboard tests.
curl --request POST "https://gptproto.com/api/v3/minimax/hailuo-2.3-standard/image-to-video" \
--header "Authorization: Bearer $GPTPROTO_API_KEY" \
--header "Content-Type: application/json" \
--data '{
"prompt": "A tiny origami fox sailing a teacup across a moonlit puddle",
"image": "",
"duration": 6,
"enable_prompt_expansion": false
}'Comece pelo custo de uma única amostra e escolha um orçamento de teste. As tarifas da GPTProto são 10% abaixo do preço de tabela.
| Cenário | Lista de MiniMax | OpenRouter | GPTProto | Economia / mês |
|---|---|---|---|---|
| Anúncios curtos de produto30 × 6s / mês | $8.40 | $8.86 | $7.56 | −$0.84≈ $10.08 / ano |
| Clipes sociais60 × 6s / mês | $16.80 | $17.72 | $15.12 | −$1.68≈ $20.16 / ano |
| Geração em massa100 × 6s / mês | $28.00 | $29.54 | $25.20 | −$2.80≈ $33.60 / ano |
Use the Hailuo 2.3 Standard API for short text-to-video and image-to-video jobs without opening a separate MiniMax account or managing provider-specific video points. The 768p tier suits concept tests, ecommerce motion assets, social clips, and batch variations. The same GPTProto key and balance also reach Seedance, Kling, Wan, and other video APIs, so teams can test one brief across models and add fallback routes without separate provider billing.
768p Clips in Two Fixed Durations
Generate 6- or 10-second videos at 768p. The fixed duration options make per-clip budgeting predictable for previews, product loops, social posts, and storyboard tests.
Text or First-Frame Image Input
Start from a prompt or animate an approved image. Image-to-video accepts JPG, JPEG, PNG, or WebP files under 20 MB, with a short edge above 300 pixels.
Explicit Camera Direction
Use 15 documented commands, including pan, tilt, zoom, tracking, push, pull, truck, shake, and static shot. Combine up to three commands when a shot needs simultaneous movement.
One Key for Multi-Model Testing
Call Hailuo, Seedance, Kling, Wan, and other supported models from one GPTProto account. Keep the same balance while routing each shot to the model that fits it.
Hailuo 2.3 Standard is MiniMax's cost-focused tier for text-to-video and image-to-video. MiniMax describes improvements over Hailuo 02 in body movement, motion-command response, stylized rendering, and facial micro-expressions. It is designed for short, continuous shots rather than long multi-scene sequences.
At 768p, the API supports 6- and 10-second outputs. Image-to-video uses one source image as the first frame, while text-to-video begins from a prompt. Both modes support prompts up to 2,000 characters and the same documented camera-command syntax. The prompt optimizer is enabled by default; disabling it gives developers more literal control over the supplied wording.
| Specification | Hailuo 2.3 Standard API Details |
|---|---|
| Provider | MiniMax |
| GPTProto model string | hailuo-2.3-standard |
| Supported generation | Text-to-video and image-to-video |
| Standard output on GPTProto | 768p |
| Duration at 768p | 6 or 10 seconds |
| Text prompt limit | Up to 2,000 characters |
| Image formats | JPG, JPEG, PNG, WebP |
| Image limits | Under 20 MB; short edge over 300 px; aspect ratio from 2:5 to 5:2 |
| Camera controls | 15 bracketed commands; up to 3 simultaneous commands recommended |
| Prompt optimizer | Enabled by default; can be disabled |
| Native audio | Not exposed for this model route |
| Processing pattern | Asynchronous task submission and status check |
Standard targets lower-cost 768p work, while Pro is the 1080p route. Both are available on GPTProto, so teams can draft with Standard and reserve Pro for approved outputs.
| Decision Factor | Hailuo 2.3 Standard | Hailuo 2.3 Pro |
|---|---|---|
| GPTProto model string | hailuo-2.3-standard |
hailuo-2.3-pro |
| Primary output tier | 768p | 1080p |
| Duration shown on GPTProto | 6 or 10 seconds | 6–10 seconds |
| Current starting price | $0.252 for the selected 6-second Standard run | $0.441 per Pro generation |
| Best fit | Iteration, batch variants, social and ecommerce motion | Quality-first final shots and 1080p delivery |
Compare them with the same source, prompt, duration, and acceptance criteria; a cheaper run is not cheaper if it needs more retries.
All four alternatives are available on GPTProto. Choose by duration, resolution, audio, reference inputs, and expected retry count.
| Model | Published Output Scope on GPTProto | Native Audio | Choose It When |
|---|---|---|---|
| Hailuo 2.3 Standard | 768p; 6 or 10 seconds; text or first-frame image | No | You need affordable, short, single-shot drafts with explicit camera commands |
| Seedance 2.0 Mini | 480p or 720p; up to 14 seconds; route-dependent text, image, video, and audio inputs | Yes | You need many lower-resolution variations or longer social clips with sound |
| Seedance 2.0 | Up to 4K; up to 15 seconds; multimodal references | Yes | You need a higher-resolution hero shot, native audio, or richer reference control |
| Kling v3.0 Standard | 720p; 3–15 seconds; text and image routes, with separate motion-control options | Optional | You need flexible duration, sound, or a dedicated motion-control workflow |
For a 6- or 10-second silent image animation, Hailuo is the simplest fit. Choose Seedance Mini for volume and sound, full Seedance for higher-resolution multimodal work, or Kling for flexible duration and motion control.
Hailuo 2.3 Standard fits ecommerce teams that already have approved product photography. Use the image as the first frame, then describe the intended object movement, camera path, background behavior, and details that must remain unchanged.
For batch production, reuse one template for product data, prompt structure, duration, and acceptance checks. Submit each SKU separately, store its task ID, and start with 6-second drafts. This works for rotations, packaging reveals, fabric movement, food close-ups, and subtle lifestyle scenes—not dialogue, synchronized sound, or multi-shot stories.
Separate the action, camera, environment, and preservation constraints. For image-to-video, focus on what moves and what stays fixed rather than redescribing the source.
Prompt formula: source subject + intended motion + camera command + environmental movement + lighting behavior + details to preserve + artifacts to avoid
Ecommerce example:
The matte-black wireless speaker rotates slowly on the pedestal while small water droplets vibrate beside the grille. [Push in,Pedestal down] Soft blue edge light moves across the metal surface. Preserve the logo, button layout, proportions, and surface finish. No extra text, duplicate controls, warping, or sudden background changes.
Use commands such as [Static shot], [Tracking shot], [Pan left], [Push in], or [Zoom out]. Keep simultaneous groups to three movements or fewer. Disable prompt optimization only when literal wording matters.
One GPTProto account, key, and USD balance cover Hailuo and 200+ supported models. Compare it with Hailuo 2.3 Pro, Seedance 2.0, or Kling v3.0 Standard without separate provider billing.
Use the live API tab as the source of truth for endpoint, parameters, and price. The prompt-only route has a separate text-to-video page. See the MiniMax collection related options.
Perguntas comuns sobre este modelo de IA
Guias, comparações e atualizações relacionadas a este modelo.
Todos os artigos
Explore a ascensão da MiniMax AI, seu poderoso modelo M2.7 e a eficiente arquitetura MoE. Descubra como acessar esses recursos multimodais hoje mesmo!

Embora a higgsfield ai ofereça movimento de vídeo fluido, seus altos custos de créditos e a interface confusa frustram os profissionais. Descubra se ela se adequa ao seu fluxo de trabalho.

Descubra o Kling O1, o primeiro modelo de vídeo com IA unificado do mundo, que combina geração e edição. Conheça os recursos, casos de uso e como esse "Nano Banana do mundo dos vídeos" está transformando a criação de conteúdo.

Crie vídeos cinematográficos com IA usando as expressões naturais e o trabalho de câmera suave do Vidu Q2. Veja como ele se compara ao Sora 2 e transforme imagens em vídeo instantaneamente.
Entrada
Saída