MiniMax H3 can generate a 15-second video at 2K resolution with native stereo sound. That makes for a neat launch headline, but resolution is not the upgrade that matters most.
The more consequential change is what you can give the model before it generates anything. H3 accepts text, images, video, and audio as references, then uses them to create or modify a clip. Instead of asking only, “What new video can this model generate?”, you can ask, “What parts of this existing video should stay, and what should change?”
That makes H3 relevant to campaign revisions, product videos, motion transfer, music visuals, and other jobs that previously required several separate generation and editing steps. It does not turn H3 into a frame-accurate timeline editor, however. Its video editing is generative: the model interprets the source and renders a new result.
MiniMax H3 Examples: Two Useful Tests
The following examples are designed to test two different sides of MiniMax H3. The first tests whether H3 can preserve an existing performance while changing its art direction and adding exact typography. The second tests whether it can hold together a complex 15-second action sequence.
Example 1: Edit a K-pop Video With Kinetic Typography
Input: One existing video featuring exactly three performers
Mode: Reference-to-video
What changes: Studio background, clothing direction, graphics, typography, and audio treatment
What should remain: Performer count, identities, choreography, timing, and source-camera movement
Edit the uploaded source video into a high-fashion K-pop campaign.
Preserve exactly three female performers, their facial identities, choreography timing, body positions, expressions, and the original camera movement. Do not add, remove, duplicate, or merge any performer.
Replace the original environment with a high-contrast white cyclorama fashion studio. Restyle the performers in coordinated futuristic black, silver, and deep-red techwear with metallic panels, structured silhouettes, utility straps, and reflective accessories.
Add bold kinetic typography that reacts to the rhythm and camera cuts. Render only the exact words “FEARLESS” and “NO LIMITS” in large, correctly spelled, black condensed sans-serif letters. Show one phrase at a time. Let the typography slide, scale, and briefly pass behind the performers without covering their faces or hands.
Add torn-paper collage elements, red editorial cutouts, subtle digital scan lines, and controlled glitch-signal effects in the background. Emphasize close-ups of confident expressions, finger-pointing gestures, leg-strap details, and synchronized dance poses while following the source video’s existing shot sequence.
High-key editorial lighting, clean white highlights, sharp fashion-photography detail, crisp fabric texture, stable faces and limbs, and a dynamic but controlled visual rhythm. Preserve the original video duration and synchronization. Native stereo audio with punchy electronic percussion, sharp transition hits, and subtle glitch accents.
This is not merely a test of whether H3 can create an attractive music video. Check whether it can preserve three distinct people through fast movement, keep the choreography aligned with the source, spell both phrases correctly, and place the typography without repeatedly covering faces or hands.
Those details matter because “change the look but keep the performance” is the actual editing problem. A visually impressive output that changes the dancer count or abandons the original timing has failed the brief.
Example 2: A 15-Second Miniature Skateboard Chase
Input: Text prompt only
Mode: Text-to-video
What it tests: Long motion, scale consistency, collisions, occlusion, camera tracking, and synchronized environmental sound
A 15-second photorealistic miniature chase sequence filmed as one continuous shot with no cuts.
0–4 seconds: An ultra-low, floor-level macro tracking camera follows directly behind a tiny miniature skateboarder riding smoothly across a polished wooden floor. The skateboarder and board remain clearly visible and maintain the same scale and appearance. Warm sunlight reflects softly across the wood. Shallow depth of field, realistic wheel vibration, subtle motion blur.
4–9 seconds: Giant colorful toy blocks suddenly tumble into the path from both sides. The skateboarder leans, swerves, and narrowly passes between them as a large toy truck rolls across the foreground. The camera keeps pace without losing the subject. Blocks collide naturally and cast realistic moving shadows.
9–15 seconds: A giant cute baby crawls rapidly into view in the background and reaches one enormous hand toward the skateboarder and the camera. The skateboarder accelerates beneath the approaching hand as the camera continues tracking inches above the floor. The hand passes close to the lens without touching it. End with the skateboarder escaping beneath the toy truck as the baby reacts with a surprised laugh.
Maintain consistent miniature scale, skateboarder identity, floor layout, lighting direction, and object positions throughout the shot. Realistic physics, natural hand anatomy, detailed wood and toy textures, cinematic macro perspective, warm bright indoor daylight, and energetic but readable motion.
Native stereo sound: tiny skateboard wheels rolling across wood, blocks clattering from left and right, the toy truck rumbling across the foreground, a soft baby laugh behind the camera, and a brief low whoosh as the giant hand approaches.
The timeline gives H3 a clearer sequence than a single paragraph full of simultaneous actions. Even so, this is a hard prompt. The camera must track one tiny subject while several much larger objects enter, collide, and partially obscure the view. A useful review should check continuity at the 4-second and 9-second transitions rather than judging only the sharpest frame.
What Is MiniMax H3?
MiniMax H3 is a general-purpose multimodal video model released by MiniMax on July 31, 2026. In plain language, it can read several kinds of creative input at once—written instructions, still images, video clips, and audio—and use them to generate a new video.
That definition separates H3 from MiniMax M3. H3 is the video-generation model discussed here. M3 belongs to MiniMax’s language and coding model line. Similar names, different jobs.
The official API model ID is MiniMax-H3. MiniMax’s current documentation lists three main generation paths: text-to-video, first- and/or last-frame image-to-video, and reference-to-video. The model outputs at 24 fps, supports 4 to 15 seconds in whole-second increments, and currently exposes 2K output through the public API. MiniMax H3 video documentation
| Specification |
MiniMax H3 |
| Official API model ID |
MiniMax-H3 |
| Output resolution |
2K |
| Output duration |
4–15 seconds |
| Frame rate |
24 fps |
| Input types |
Text, image, video, audio |
| Main modes |
Text-to-video, first/last-frame image-to-video, reference-to-video |
| Reference images |
Up to 9 |
| Reference videos |
Up to 3 clips; 15 seconds total |
| Reference audio |
Up to 3 clips; 15 seconds total |
| Mixed reference files |
Up to 12 |
| Prompt length |
Up to 7,000 characters |
| Supported output ratios |
21:9, 16:9, 4:3, 1:1, 3:4, 9:16, or adaptive where supported |
You may also see H3 described as “Hailuo 3.0.” That label makes sense as an informal way to place it after the Hailuo 2.x family, but MiniMax’s current model list and API documentation use MiniMax H3, not Hailuo 3.0, as the official name.
What Is New in the MiniMax H3 Upgrade?
Hailuo 2.3 focused on better physical action, human micro-expressions, stylized visuals, and motion-command response. H3 changes the scope of the product. It is no longer only a question of generating a short clip from text or one image; the model can now use existing media as working context.
1. 2K output for clips up to 15 seconds
MiniMax H3 raises the documented ceiling from the 6- or 10-second formats available in the older Hailuo API models to a flexible 4–15 seconds. It also raises the maximum resolution from 1080p to 2K.
Longer does not automatically mean more coherent. A 15-second shot gives the model more time to change a face, lose an object, or misunderstand a later event. That is why the second example above uses a timed sequence and explicit continuity constraints.
2. Native stereo audio
H3 generates sound as part of the video rather than requiring a completely separate audio-generation pass. Reuters reports native stereo sound as one of the release capabilities. For short ads, music visuals, and action clips, this can keep sound effects closer to the events that produced them. Reuters launch report
The trade-off is predictability. Native audio reduces workflow steps, but it does not guarantee exact dialogue, perfect lip synchronization, or a production-ready mix on every attempt. If a project depends on legally cleared music, precise voiceover wording, or exact loudness standards, treat the generated track as a draft until it passes a separate review.
3. Multimodal reference generation
A reference-to-video request can combine as many as 9 images, 3 video clips, and 3 audio clips, subject to a 12-file total and a 15-second total limit for reference video or audio. Every request still requires a written prompt. Audio also cannot be submitted by itself; at least one reference image or video must accompany it.
This is useful when one asset cannot carry the entire brief. A product image can define the object, a model image can define the person, a source clip can define the motion, and an audio clip can define the rhythm. The prompt then explains which traits to preserve and how the final scene should look.
More references also create more ways for instructions to conflict. Nine images are not automatically better than three. If two product shots show different packaging or two character references show different hair, the model has to guess which one wins.
4. First- and last-frame control
H3 supports a first frame, a last frame, or both. This is useful when the opening product composition and closing brand frame are already approved but the movement between them still needs to be generated.
The first/last-frame mode and the multimodal reference mode are separate API paths. MiniMax’s current specification says they cannot be mixed in the same request. That boundary is easy to miss: you can anchor the opening and closing images, or use a collection of reference images, videos, and audio, but not both input structures at once. MiniMax H3 API reference
5. Motion transfer and reference-video control
H3 can use a reference video to guide motion, camera behavior, editing rhythm, and other temporal information. That opens a practical path for choreography transfer, product-handling gestures, camera moves, and scene timing that would be difficult to describe with text alone.
This does not mean the source motion is copied frame for frame. H3 interprets the reference and regenerates the result. For recognizable choreography or tightly approved product movement, compare the output against the source timeline rather than assuming “reference” means exact replication.
6. Generative video editing
Video editing is the defining H3 feature, but the term needs a boundary. H3 is not a conventional editor where you select a layer, replace a word, and leave every unaffected pixel untouched.
Instead, the source video becomes a reference. The prompt tells H3 what to preserve and what to change, and the model generates a revised clip. This can support large creative changes—new setting, wardrobe, graphics, sound, or style—but those same changes may introduce identity drift, altered framing, misspelled text, or small motion changes.
In one sentence: H3 edits by regeneration, not by surgically modifying the original file.
How MiniMax H3 Video Editing Actually Works
A good H3 editing instruction has two layers:
-
Preservation constraints: Who stays in the clip? Which product details, actions, camera moves, timing, and composition must remain?
-
Transformation instructions: What background, clothing, visual style, typography, sound, or atmosphere should change?
The K-pop prompt uses this split deliberately. It first locks the performer count, facial identities, choreography, timing, and camera movement. Only then does it request a new studio, fashion direction, graphics, and audio treatment.
That ordering matters. “Turn this into a futuristic K-pop video” gives the model permission to reinterpret almost everything. “Preserve exactly three performers and the source choreography; replace only the art direction” gives it a hierarchy.
For harder edits, change one major category at a time. First test the background and wardrobe. Then add typography. Then refine audio. A single request may do all three, but separating them tells you which instruction caused a failure and makes retries easier to diagnose.
There is also a cost detail that launch summaries often miss. MiniMax bills a 2K output at $0.13 per second, and a reference video is billed at the same per-second rate according to its duration. A 15-second 2K generation therefore costs about $1.95. If the request also uses a 15-second reference video, the combined input-video and output charge is about $3.90 before any extra charge for more than five reference images. MiniMax pay-as-you-go pricing
That does not make reference editing unreasonable, but it changes how you plan iterations. Ten 15-second text-to-video attempts cost about $19.50 at the listed rate; ten equally long edits with a 15-second source video cost about $39.00. Short diagnostic generations are the sensible first pass.
What Kind of Videos Can MiniMax H3 Make?
H3 can cover the familiar text-to-video and image-to-video jobs, but its stronger use cases are the ones that benefit from several reference types or an existing timeline.
Revise a campaign without reshooting everything
A brand can supply the approved source performance, then request a new location, wardrobe direction, seasonal treatment, or graphic system. This is useful for concepting regional or event-specific versions. It is less suitable when every product label, facial detail, and gesture must survive with frame-level fidelity.
Turn product images into an e-commerce video
Product shots can define shape, color, and packaging while a motion reference demonstrates how the item should rotate, open, pour, fold, or move through the scene. A closing image can be useful for a fixed end card, although it cannot currently be combined with reference-to-video inputs in the same request.
Transfer choreography or camera motion
A reference clip can communicate timing that would take hundreds of words to describe. Dance, sports movement, handheld camera behavior, and product-reveal moves are obvious candidates. Always confirm that you have the right to use the source performance and likeness.
Create short music and fashion visuals
The combination of fast visual changes, generated sound, reference motion, and art-direction prompts suits music teasers and fashion campaigns. Exact typography remains a quality-control point, not a safe assumption.
Animate game promos and cinematic title screens
Reference images can keep a character, prop, or interface closer to an approved design while the prompt adds motion, camera travel, particles, and sound. Generated footage is useful for concept trailers, but interface text and gameplay-like sequences still need human review.
Generate 15-second action or chase sequences
The longer duration gives a prompt enough room for setup, escalation, and payoff. It also makes continuity harder. Prompts should divide the action into time ranges and repeat the subject traits that must not change.
MiniMax H3 for E-commerce and Advertising
The most useful way to think about H3 for e-commerce is not “generate an ad.” Think of it as assembling a brief from several media assets.
|
|
| Input |
Job in the workflow |
| Product images |
Define the product’s shape, color, material, and packaging |
| Logo or packaging references |
Constrain brand marks and visual identity |
| Model or character images |
Keep the presenter’s appearance closer to the approved look |
| Motion reference video |
Specify a hand gesture, demonstration, camera move, or edit rhythm |
| Audio reference |
Guide pacing, voice, or sound direction |
| Written prompt |
State the final scene, preservation rules, changes, and exclusions |
For example, a cosmetics brand could provide five clean product angles, one model reference, a 6-second hand-motion clip, and a short audio reference. The prompt could then request a 10-second vertical product reveal while preserving the bottle shape, label placement, cap color, model identity, and gesture timing.
That is a better brief than “make a luxury cosmetics ad.” It is also not a guarantee. Product geometry may bend, small label text may change, and a logo can drift between frames. If the output will become a paid advertisement, compare the product against the approved photography frame by frame and replace critical brand text in a conventional editor when necessary.
H3 is most useful here as a production accelerator for concepts, variants, and short campaign assets. It is not a substitute for brand approval.
MiniMax H3 vs Hailuo 2.3 and Hailuo 02
The generational difference is larger than the resolution column suggests. MiniMax’s previous Hailuo models are short-form generators built around text and image inputs. H3 adds a reference system that can take video and audio, then use those materials for creation, motion transfer, or editing.
|
|
|
|
| Capability |
MiniMax H3 |
Hailuo 2.3 |
Hailuo 02 |
| Maximum documented resolution |
2K |
1080p |
1080p |
| Documented duration |
4–15 seconds |
6 seconds at 1080p; 6 or 10 seconds at 768p |
6 seconds at 1080p; 6 or 10 seconds at 768p/512p |
| Frame rate |
24 fps |
24 fps |
24 fps |
| Text-to-video |
Yes |
Yes |
Yes |
| Image-to-video |
Yes |
Yes |
Yes |
| First/last-frame input |
Yes |
Not listed as a primary current mode |
Yes |
| Reference video and audio |
Yes |
Not listed in the current model specification |
Not listed in the current model specification |
| Native generated audio |
Yes |
Not listed |
Not listed |
| Generative video editing |
Yes, through reference-based regeneration |
Not listed |
Not listed |
| Best fit |
Multimodal briefs, editing, motion transfer, audio, and longer 2K clips |
Short T2V/I2V with stronger action and expression handling |
Conventional short generation and lower-resolution options |
The practical choice is straightforward:
-
Choose MiniMax H3 when the task needs an existing video, several reference assets, native sound, 2K output, or more than 10 seconds.
-
Choose Hailuo 2.3 when you only need a short text-to-video or image-to-video clip and do not need H3’s reference and editing system.
-
Consider Hailuo 02 when its first/last-frame path or lower-resolution pricing fits a simple legacy workflow.
MiniMax H3 is not yet available through GPT Proto. If you need a MiniMax video model that is already available there, try Hailuo 2.3 Pro for image-to-video or Hailuo 2.3 Pro for text-to-video.
MiniMax H3 Pricing and Availability
MiniMax H3’s official API is live under the model ID MiniMax-H3. Public API documentation currently lists 2K as the available output resolution. A 768p rate is published, but MiniMax marks that option as closed beta and directs users to contact sales before using it.
|
|
| Item |
Published price |
| 2K output video |
$0.13 per second |
| 768p output video |
$0.09 per second; closed beta |
| Audio references |
Free |
| Reference images |
First 5 free; $0.04 per additional image |
| Reference video |
$0.13 per input second for a 2K output request |
| 15-second 2K text-to-video generation |
About $1.95 |
| 15-second reference video + 15-second 2K output |
About $3.90 |
The API uses an asynchronous workflow: submit a generation request, receive a task ID, poll the task status, and download the result after it succeeds. MiniMax H3 is not currently available on GPT Proto, so this article does not include a GPT Proto code example or suggest that a GPT Proto H3 endpoint is already live.
The open-weights status also needs precise wording. MiniMax presents H3 as an open general-purpose model and has announced plans to publish the weights. Reuters reported that the files were expected within days. As of this article’s July 31 check, however, the downloadable weights, license terms, model size, and realistic local hardware requirements had not yet been confirmed publicly.
“Open weights” is therefore the announced direction, not proof that you can download and run H3 locally today.
What MiniMax H3 Still Does Not Guarantee
H3 expands what can be attempted in one generation, but it does not remove the usual failure modes of generative video.
Exact text may still require retries. The first example deliberately uses only two short phrases. If those words change across frames, use the generation for motion and add final typography in a normal editor.
Aggressive edits can change identity. Replacing a background, wardrobe, graphics system, audio, and lighting in one request gives the model more opportunities to reinterpret the people in the source.
A 15-second complex shot is harder than a simple 6-second animation. The longer the clip and the more events it contains, the more continuity decisions the model must preserve.
Generated editing is not pixel-preserving editing. H3 creates a new result from references. It is not the right tool for a legal, medical, or product-compliance edit where all untouched pixels must remain identical.
Official examples do not prove repeatable production reliability. A launch reel shows what a model can produce, not how many attempts the result required. Repeatability, failure rate, and prompt sensitivity need independent testing.
Local deployment requirements are still unknown. Until the actual weights, architecture details, license, and serving guidance arrive, any precise VRAM recommendation is speculation.
One independent signal is encouraging: on July 31, MiniMax H3 ranked first on Artificial Analysis’s video-editing leaderboard with audio, with an Elo score of 1,130 from 5,043 samples. That is a useful blind-preference result, but it is a snapshot rather than a guarantee for every task. Artificial Analysis video-editing leaderboard
Final Verdict: Is MiniMax H3 a Meaningful Upgrade?
Yes. MiniMax H3 is a larger upgrade over Hailuo 2.3 than “2K instead of 1080p” suggests.
Its strongest reason to exist is reference-driven generation and editing. Text, product images, a source performance, motion guidance, and audio can now contribute to one video request. For advertising and e-commerce teams, that changes the brief from “generate something similar” to “preserve these assets and this timing, then change these specific creative elements.”
The cost is control. Because H3 regenerates the clip, edits can affect details you wanted to keep. Longer shots also create more room for continuity failures, and input video increases the price of each attempt.
For a simple 6-second image animation, Hailuo 2.3 may still be enough. For a multimodal campaign, motion transfer, native sound, an existing-video revision, or a 15-second action sequence, H3 is the more relevant model to test.
MiniMax H3 is not currently available on GPT Proto. While access is being prepared, you can explore other video-generation models in the GPT Proto model gallery or visit GPT Proto to compare the models already available through one account.
Input: Text prompt only
Mode: Text-to-video
What it tests: Long motion, scale consistency, collisions, occlusion, camera tracking, and synchronized environmental sound
A 15-second photorealistic miniature chase sequence filmed as one continuous shot with no cuts.
0–4 seconds: An ultra-low, floor-level macro tracking camera follows directly behind a tiny miniature skateboarder riding smoothly across a polished wooden floor. The skateboarder and board remain clearly visible and maintain the same scale and appearance. Warm sunlight reflects softly across the wood. Shallow depth of field, realistic wheel vibration, subtle motion blur.
4–9 seconds: Giant colorful toy blocks suddenly tumble into the path from both sides. The skateboarder leans, swerves, and narrowly passes between them as a large toy truck rolls across the foreground. The camera keeps pace without losing the subject. Blocks collide naturally and cast realistic moving shadows.
9–15 seconds: A giant cute baby crawls rapidly into view in the background and reaches one enormous hand toward the skateboarder and the camera. The skateboarder accelerates beneath the approaching hand as the camera continues tracking inches above the floor. The hand passes close to the lens without touching it. End with the skateboarder escaping beneath the toy truck as the baby reacts with a surprised laugh.
Maintain consistent miniature scale, skateboarder identity, floor layout, lighting direction, and object positions throughout the shot. Realistic physics, natural hand anatomy, detailed wood and toy textures, cinematic macro perspective, warm bright indoor daylight, and energetic but readable motion.
Native stereo sound: tiny skateboard wheels rolling across wood, blocks clattering from left and right, the toy truck rumbling across the foreground, a soft baby laugh behind the camera, and a brief low whoosh as the giant hand approaches.
The timeline gives H3 a clearer sequence than a single paragraph full of simultaneous actions. Even so, this is a hard prompt. The camera must track one tiny subject while several much larger objects enter, collide, and partially obscure the view. A useful review should check continuity at the 4-second and 9-second transitions rather than judging only the sharpest frame.
What Is MiniMax H3?
MiniMax H3 is a general-purpose multimodal video model released by MiniMax on July 31, 2026. In plain language, it can read several kinds of creative input at once—written instructions, still images, video clips, and audio—and use them to generate a new video.
That definition separates H3 from MiniMax M3. H3 is the video-generation model discussed here. M3 belongs to MiniMax’s language and coding model line. Similar names, different jobs.
The official API model ID is MiniMax-H3. MiniMax’s current documentation lists three main generation paths: text-to-video, first- and/or last-frame image-to-video, and reference-to-video. The model outputs at 24 fps, supports 4 to 15 seconds in whole-second increments, and currently exposes 2K output through the public API. MiniMax H3 video documentation
| Specification |
MiniMax H3 |
| Official API model ID |
MiniMax-H3 |
| Output resolution |
2K |
| Output duration |
4–15 seconds |
| Frame rate |
24 fps |
| Input types |
Text, image, video, audio |
| Main modes |
Text-to-video, first/last-frame image-to-video, reference-to-video |
| Reference images |
Up to 9 |
| Reference videos |
Up to 3 clips; 15 seconds total |
| Reference audio |
Up to 3 clips; 15 seconds total |
| Mixed reference files |
Up to 12 |
| Prompt length |
Up to 7,000 characters |
| Supported output ratios |
21:9, 16:9, 4:3, 1:1, 3:4, 9:16, or adaptive where supported |
You may also see H3 described as “Hailuo 3.0.” That label makes sense as an informal way to place it after the Hailuo 2.x family, but MiniMax’s current model list and API documentation use MiniMax H3, not Hailuo 3.0, as the official name.
What Is New in the MiniMax H3 Upgrade?
Hailuo 2.3 focused on better physical action, human micro-expressions, stylized visuals, and motion-command response. H3 changes the scope of the product. It is no longer only a question of generating a short clip from text or one image; the model can now use existing media as working context.
1. 2K output for clips up to 15 seconds
MiniMax H3 raises the documented ceiling from the 6- or 10-second formats available in the older Hailuo API models to a flexible 4–15 seconds. It also raises the maximum resolution from 1080p to 2K.
Longer does not automatically mean more coherent. A 15-second shot gives the model more time to change a face, lose an object, or misunderstand a later event. That is why the second example above uses a timed sequence and explicit continuity constraints.
2. Native stereo audio
H3 generates sound as part of the video rather than requiring a completely separate audio-generation pass. Reuters reports native stereo sound as one of the release capabilities. For short ads, music visuals, and action clips, this can keep sound effects closer to the events that produced them. Reuters launch report
The trade-off is predictability. Native audio reduces workflow steps, but it does not guarantee exact dialogue, perfect lip synchronization, or a production-ready mix on every attempt. If a project depends on legally cleared music, precise voiceover wording, or exact loudness standards, treat the generated track as a draft until it passes a separate review.
3. Multimodal reference generation
A reference-to-video request can combine as many as 9 images, 3 video clips, and 3 audio clips, subject to a 12-file total and a 15-second total limit for reference video or audio. Every request still requires a written prompt. Audio also cannot be submitted by itself; at least one reference image or video must accompany it.
This is useful when one asset cannot carry the entire brief. A product image can define the object, a model image can define the person, a source clip can define the motion, and an audio clip can define the rhythm. The prompt then explains which traits to preserve and how the final scene should look.
More references also create more ways for instructions to conflict. Nine images are not automatically better than three. If two product shots show different packaging or two character references show different hair, the model has to guess which one wins.
4. First- and last-frame control
H3 supports a first frame, a last frame, or both. This is useful when the opening product composition and closing brand frame are already approved but the movement between them still needs to be generated.
The first/last-frame mode and the multimodal reference mode are separate API paths. MiniMax’s current specification says they cannot be mixed in the same request. That boundary is easy to miss: you can anchor the opening and closing images, or use a collection of reference images, videos, and audio, but not both input structures at once. MiniMax H3 API reference
5. Motion transfer and reference-video control
H3 can use a reference video to guide motion, camera behavior, editing rhythm, and other temporal information. That opens a practical path for choreography transfer, product-handling gestures, camera moves, and scene timing that would be difficult to describe with text alone.
This does not mean the source motion is copied frame for frame. H3 interprets the reference and regenerates the result. For recognizable choreography or tightly approved product movement, compare the output against the source timeline rather than assuming “reference” means exact replication.
6. Generative video editing
Video editing is the defining H3 feature, but the term needs a boundary. H3 is not a conventional editor where you select a layer, replace a word, and leave every unaffected pixel untouched.
Instead, the source video becomes a reference. The prompt tells H3 what to preserve and what to change, and the model generates a revised clip. This can support large creative changes—new setting, wardrobe, graphics, sound, or style—but those same changes may introduce identity drift, altered framing, misspelled text, or small motion changes.
In one sentence: H3 edits by regeneration, not by surgically modifying the original file.
How MiniMax H3 Video Editing Actually Works
A good H3 editing instruction has two layers:
-
Preservation constraints: Who stays in the clip? Which product details, actions, camera moves, timing, and composition must remain?
-
Transformation instructions: What background, clothing, visual style, typography, sound, or atmosphere should change?
The K-pop prompt uses this split deliberately. It first locks the performer count, facial identities, choreography, timing, and camera movement. Only then does it request a new studio, fashion direction, graphics, and audio treatment.
That ordering matters. “Turn this into a futuristic K-pop video” gives the model permission to reinterpret almost everything. “Preserve exactly three performers and the source choreography; replace only the art direction” gives it a hierarchy.
For harder edits, change one major category at a time. First test the background and wardrobe. Then add typography. Then refine audio. A single request may do all three, but separating them tells you which instruction caused a failure and makes retries easier to diagnose.
There is also a cost detail that launch summaries often miss. MiniMax bills a 2K output at $0.13 per second, and a reference video is billed at the same per-second rate according to its duration. A 15-second 2K generation therefore costs about $1.95. If the request also uses a 15-second reference video, the combined input-video and output charge is about $3.90 before any extra charge for more than five reference images. MiniMax pay-as-you-go pricing
That does not make reference editing unreasonable, but it changes how you plan iterations. Ten 15-second text-to-video attempts cost about $19.50 at the listed rate; ten equally long edits with a 15-second source video cost about $39.00. Short diagnostic generations are the sensible first pass.
What Kind of Videos Can MiniMax H3 Make?
H3 can cover the familiar text-to-video and image-to-video jobs, but its stronger use cases are the ones that benefit from several reference types or an existing timeline.
Revise a campaign without reshooting everything
A brand can supply the approved source performance, then request a new location, wardrobe direction, seasonal treatment, or graphic system. This is useful for concepting regional or event-specific versions. It is less suitable when every product label, facial detail, and gesture must survive with frame-level fidelity.
Turn product images into an e-commerce video
Product shots can define shape, color, and packaging while a motion reference demonstrates how the item should rotate, open, pour, fold, or move through the scene. A closing image can be useful for a fixed end card, although it cannot currently be combined with reference-to-video inputs in the same request.
Transfer choreography or camera motion
A reference clip can communicate timing that would take hundreds of words to describe. Dance, sports movement, handheld camera behavior, and product-reveal moves are obvious candidates. Always confirm that you have the right to use the source performance and likeness.
Create short music and fashion visuals
The combination of fast visual changes, generated sound, reference motion, and art-direction prompts suits music teasers and fashion campaigns. Exact typography remains a quality-control point, not a safe assumption.
Animate game promos and cinematic title screens
Reference images can keep a character, prop, or interface closer to an approved design while the prompt adds motion, camera travel, particles, and sound. Generated footage is useful for concept trailers, but interface text and gameplay-like sequences still need human review.
Generate 15-second action or chase sequences
The longer duration gives a prompt enough room for setup, escalation, and payoff. It also makes continuity harder. Prompts should divide the action into time ranges and repeat the subject traits that must not change.
MiniMax H3 for E-commerce and Advertising
The most useful way to think about H3 for e-commerce is not “generate an ad.” Think of it as assembling a brief from several media assets.
|
|
| Input |
Job in the workflow |
| Product images |
Define the product’s shape, color, material, and packaging |
| Logo or packaging references |
Constrain brand marks and visual identity |
| Model or character images |
Keep the presenter’s appearance closer to the approved look |
| Motion reference video |
Specify a hand gesture, demonstration, camera move, or edit rhythm |
| Audio reference |
Guide pacing, voice, or sound direction |
| Written prompt |
State the final scene, preservation rules, changes, and exclusions |
For example, a cosmetics brand could provide five clean product angles, one model reference, a 6-second hand-motion clip, and a short audio reference. The prompt could then request a 10-second vertical product reveal while preserving the bottle shape, label placement, cap color, model identity, and gesture timing.
That is a better brief than “make a luxury cosmetics ad.” It is also not a guarantee. Product geometry may bend, small label text may change, and a logo can drift between frames. If the output will become a paid advertisement, compare the product against the approved photography frame by frame and replace critical brand text in a conventional editor when necessary.
H3 is most useful here as a production accelerator for concepts, variants, and short campaign assets. It is not a substitute for brand approval.
MiniMax H3 vs Hailuo 2.3 and Hailuo 02
The generational difference is larger than the resolution column suggests. MiniMax’s previous Hailuo models are short-form generators built around text and image inputs. H3 adds a reference system that can take video and audio, then use those materials for creation, motion transfer, or editing.
|
|
|
|
| Capability |
MiniMax H3 |
Hailuo 2.3 |
Hailuo 02 |
| Maximum documented resolution |
2K |
1080p |
1080p |
| Documented duration |
4–15 seconds |
6 seconds at 1080p; 6 or 10 seconds at 768p |
6 seconds at 1080p; 6 or 10 seconds at 768p/512p |
| Frame rate |
24 fps |
24 fps |
24 fps |
| Text-to-video |
Yes |
Yes |
Yes |
| Image-to-video |
Yes |
Yes |
Yes |
| First/last-frame input |
Yes |
Not listed as a primary current mode |
Yes |
| Reference video and audio |
Yes |
Not listed in the current model specification |
Not listed in the current model specification |
| Native generated audio |
Yes |
Not listed |
Not listed |
| Generative video editing |
Yes, through reference-based regeneration |
Not listed |
Not listed |
| Best fit |
Multimodal briefs, editing, motion transfer, audio, and longer 2K clips |
Short T2V/I2V with stronger action and expression handling |
Conventional short generation and lower-resolution options |
The practical choice is straightforward:
-
Choose MiniMax H3 when the task needs an existing video, several reference assets, native sound, 2K output, or more than 10 seconds.
-
Choose Hailuo 2.3 when you only need a short text-to-video or image-to-video clip and do not need H3’s reference and editing system.
-
Consider Hailuo 02 when its first/last-frame path or lower-resolution pricing fits a simple legacy workflow.
MiniMax H3 is not yet available through GPT Proto. If you need a MiniMax video model that is already available there, try Hailuo 2.3 Pro for image-to-video or Hailuo 2.3 Pro for text-to-video.
MiniMax H3 Pricing and Availability
MiniMax H3’s official API is live under the model ID MiniMax-H3. Public API documentation currently lists 2K as the available output resolution. A 768p rate is published, but MiniMax marks that option as closed beta and directs users to contact sales before using it.
|
|
| Item |
Published price |
| 2K output video |
$0.13 per second |
| 768p output video |
$0.09 per second; closed beta |
| Audio references |
Free |
| Reference images |
First 5 free; $0.04 per additional image |
| Reference video |
$0.13 per input second for a 2K output request |
| 15-second 2K text-to-video generation |
About $1.95 |
| 15-second reference video + 15-second 2K output |
About $3.90 |
The API uses an asynchronous workflow: submit a generation request, receive a task ID, poll the task status, and download the result after it succeeds. MiniMax H3 is not currently available on GPT Proto, so this article does not include a GPT Proto code example or suggest that a GPT Proto H3 endpoint is already live.
The open-weights status also needs precise wording. MiniMax presents H3 as an open general-purpose model and has announced plans to publish the weights. Reuters reported that the files were expected within days. As of this article’s July 31 check, however, the downloadable weights, license terms, model size, and realistic local hardware requirements had not yet been confirmed publicly.
“Open weights” is therefore the announced direction, not proof that you can download and run H3 locally today.
What MiniMax H3 Still Does Not Guarantee
H3 expands what can be attempted in one generation, but it does not remove the usual failure modes of generative video.
Exact text may still require retries. The first example deliberately uses only two short phrases. If those words change across frames, use the generation for motion and add final typography in a normal editor.
Aggressive edits can change identity. Replacing a background, wardrobe, graphics system, audio, and lighting in one request gives the model more opportunities to reinterpret the people in the source.
A 15-second complex shot is harder than a simple 6-second animation. The longer the clip and the more events it contains, the more continuity decisions the model must preserve.
Generated editing is not pixel-preserving editing. H3 creates a new result from references. It is not the right tool for a legal, medical, or product-compliance edit where all untouched pixels must remain identical.
Official examples do not prove repeatable production reliability. A launch reel shows what a model can produce, not how many attempts the result required. Repeatability, failure rate, and prompt sensitivity need independent testing.
Local deployment requirements are still unknown. Until the actual weights, architecture details, license, and serving guidance arrive, any precise VRAM recommendation is speculation.
One independent signal is encouraging: on July 31, MiniMax H3 ranked first on Artificial Analysis’s video-editing leaderboard with audio, with an Elo score of 1,130 from 5,043 samples. That is a useful blind-preference result, but it is a snapshot rather than a guarantee for every task. Artificial Analysis video-editing leaderboard
Final Verdict: Is MiniMax H3 a Meaningful Upgrade?
Yes. MiniMax H3 is a larger upgrade over Hailuo 2.3 than “2K instead of 1080p” suggests.
Its strongest reason to exist is reference-driven generation and editing. Text, product images, a source performance, motion guidance, and audio can now contribute to one video request. For advertising and e-commerce teams, that changes the brief from “generate something similar” to “preserve these assets and this timing, then change these specific creative elements.”
The cost is control. Because H3 regenerates the clip, edits can affect details you wanted to keep. Longer shots also create more room for continuity failures, and input video increases the price of each attempt.
For a simple 6-second image animation, Hailuo 2.3 may still be enough. For a multimodal campaign, motion transfer, native sound, an existing-video revision, or a 15-second action sequence, H3 is the more relevant model to test.
MiniMax H3 is not currently available on GPT Proto. While access is being prepared, you can explore other video-generation models in the GPT Proto model gallery or visit GPT Proto to compare the models already available through one account.