Schuyler Stacy2026-07-31

MiniMax H3 Is Here: What Its Video Editing Upgrade Actually Changes

See what MiniMax H3 changes with 2K video, 15-second clips, native audio, generative video editing, e-commerce use cases, pricing, and Hailuo comparisons.

MiniMax H3 Is Here: What Its Video Editing Upgrade Actually Changes

TL;DR: What You Need to Know

  • MiniMax H3 was officially released on July 31, 2026.

  • It generates 4- to 15-second videos at 2K and 24 fps.

  • It accepts text, images, videos, and audio as multimodal inputs.

  • It supports text-to-video, first- and last-frame control, reference-based generation, motion transfer, and generative video editing.

  • It can generate native stereo sound with the video.

  • Its official API is live, but MiniMax H3 is not yet available through GPTProto.

  • MiniMax plans to publish downloadable model weights, but they were not publicly available when this article was last checked.

  • H3 is a major workflow upgrade over Hailuo 2.3 when you need source-video references, editing, audio, or 2K output. Hailuo 2.3 remains sufficient for simpler short text-to-video and image-to-video jobs.

Table of contents

MiniMax H3 can generate a 15-second video at 2K resolution with native stereo sound. That makes for a neat launch headline, but resolution is not the upgrade that matters most.

The more consequential change is what you can give the model before it generates anything. H3 accepts text, images, video, and audio as references, then uses them to create or modify a clip. Instead of asking only, “What new video can this model generate?”, you can ask, “What parts of this existing video should stay, and what should change?”

That makes H3 relevant to campaign revisions, product videos, motion transfer, music visuals, and other jobs that previously required several separate generation and editing steps. It does not turn H3 into a frame-accurate timeline editor, however. Its video editing is generative: the model interprets the source and renders a new result.

MiniMax H3 Examples: Two Useful Tests

The following examples are designed to test two different sides of MiniMax H3. The first tests whether H3 can preserve an existing performance while changing its art direction and adding exact typography. The second tests whether it can hold together a complex 15-second action sequence.

Example 1: Edit a K-pop Video With Kinetic Typography

Input: One existing video featuring exactly three performers
Mode: Reference-to-video
What changes: Studio background, clothing direction, graphics, typography, and audio treatment
What should remain: Performer count, identities, choreography, timing, and source-camera movement

Edit the uploaded source video into a high-fashion K-pop campaign.

Preserve exactly three female performers, their facial identities, choreography timing, body positions, expressions, and the original camera movement. Do not add, remove, duplicate, or merge any performer.

Replace the original environment with a high-contrast white cyclorama fashion studio. Restyle the performers in coordinated futuristic black, silver, and deep-red techwear with metallic panels, structured silhouettes, utility straps, and reflective accessories.

Add bold kinetic typography that reacts to the rhythm and camera cuts. Render only the exact words “FEARLESS” and “NO LIMITS” in large, correctly spelled, black condensed sans-serif letters. Show one phrase at a time. Let the typography slide, scale, and briefly pass behind the performers without covering their faces or hands.

Add torn-paper collage elements, red editorial cutouts, subtle digital scan lines, and controlled glitch-signal effects in the background. Emphasize close-ups of confident expressions, finger-pointing gestures, leg-strap details, and synchronized dance poses while following the source video’s existing shot sequence.

High-key editorial lighting, clean white highlights, sharp fashion-photography detail, crisp fabric texture, stable faces and limbs, and a dynamic but controlled visual rhythm. Preserve the original video duration and synchronization. Native stereo audio with punchy electronic percussion, sharp transition hits, and subtle glitch accents.

This is not merely a test of whether H3 can create an attractive music video. Check whether it can preserve three distinct people through fast movement, keep the choreography aligned with the source, spell both phrases correctly, and place the typography without repeatedly covering faces or hands.

Those details matter because “change the look but keep the performance” is the actual editing problem. A visually impressive output that changes the dancer count or abandons the original timing has failed the brief.

Example 2: A 15-Second Miniature Skateboard Chase

Input: Text prompt only
Mode: Text-to-video
What it tests: Long motion, scale consistency, collisions, occlusion, camera tracking, and synchronized environmental sound

A 15-second photorealistic miniature chase sequence filmed as one continuous shot with no cuts.

0–4 seconds: An ultra-low, floor-level macro tracking camera follows directly behind a tiny miniature skateboarder riding smoothly across a polished wooden floor. The skateboarder and board remain clearly visible and maintain the same scale and appearance. Warm sunlight reflects softly across the wood. Shallow depth of field, realistic wheel vibration, subtle motion blur.

4–9 seconds: Giant colorful toy blocks suddenly tumble into the path from both sides. The skateboarder leans, swerves, and narrowly passes between them as a large toy truck rolls across the foreground. The camera keeps pace without losing the subject. Blocks collide naturally and cast realistic moving shadows.

9–15 seconds: A giant cute baby crawls rapidly into view in the background and reaches one enormous hand toward the skateboarder and the camera. The skateboarder accelerates beneath the approaching hand as the camera continues tracking inches above the floor. The hand passes close to the lens without touching it. End with the skateboarder escaping beneath the toy truck as the baby reacts with a surprised laugh.

Maintain consistent miniature scale, skateboarder identity, floor layout, lighting direction, and object positions throughout the shot. Realistic physics, natural hand anatomy, detailed wood and toy textures, cinematic macro perspective, warm bright indoor daylight, and energetic but readable motion.

Native stereo sound: tiny skateboard wheels rolling across wood, blocks clattering from left and right, the toy truck rumbling across the foreground, a soft baby laugh behind the camera, and a brief low whoosh as the giant hand approaches.

The timeline gives H3 a clearer sequence than a single paragraph full of simultaneous actions. Even so, this is a hard prompt. The camera must track one tiny subject while several much larger objects enter, collide, and partially obscure the view. A useful review should check continuity at the 4-second and 9-second transitions rather than judging only the sharpest frame.

What Is MiniMax H3?

MiniMax H3 is a general-purpose multimodal video model released by MiniMax on July 31, 2026. In plain language, it can read several kinds of creative input at once—written instructions, still images, video clips, and audio—and use them to generate a new video.

That definition separates H3 from MiniMax M3. H3 is the video-generation model discussed here. M3 belongs to MiniMax’s language and coding model line. Similar names, different jobs.

The official API model ID is MiniMax-H3. MiniMax’s current documentation lists three main generation paths: text-to-video, first- and/or last-frame image-to-video, and reference-to-video. The model outputs at 24 fps, supports 4 to 15 seconds in whole-second increments, and currently exposes 2K output through the public API. MiniMax H3 video documentation

Specification MiniMax H3
Official API model ID MiniMax-H3
Output resolution 2K
Output duration 4–15 seconds
Frame rate 24 fps
Input types Text, image, video, audio
Main modes Text-to-video, first/last-frame image-to-video, reference-to-video
Reference images Up to 9
Reference videos Up to 3 clips; 15 seconds total
Reference audio Up to 3 clips; 15 seconds total
Mixed reference files Up to 12
Prompt length Up to 7,000 characters
Supported output ratios 21:9, 16:9, 4:3, 1:1, 3:4, 9:16, or adaptive where supported

You may also see H3 described as “Hailuo 3.0.” That label makes sense as an informal way to place it after the Hailuo 2.x family, but MiniMax’s current model list and API documentation use MiniMax H3, not Hailuo 3.0, as the official name.

What Is New in the MiniMax H3 Upgrade?

Hailuo 2.3 focused on better physical action, human micro-expressions, stylized visuals, and motion-command response. H3 changes the scope of the product. It is no longer only a question of generating a short clip from text or one image; the model can now use existing media as working context.

1. 2K output for clips up to 15 seconds

MiniMax H3 raises the documented ceiling from the 6- or 10-second formats available in the older Hailuo API models to a flexible 4–15 seconds. It also raises the maximum resolution from 1080p to 2K.

Longer does not automatically mean more coherent. A 15-second shot gives the model more time to change a face, lose an object, or misunderstand a later event. That is why the second example above uses a timed sequence and explicit continuity constraints.

2. Native stereo audio

H3 generates sound as part of the video rather than requiring a completely separate audio-generation pass. Reuters reports native stereo sound as one of the release capabilities. For short ads, music visuals, and action clips, this can keep sound effects closer to the events that produced them. Reuters launch report

The trade-off is predictability. Native audio reduces workflow steps, but it does not guarantee exact dialogue, perfect lip synchronization, or a production-ready mix on every attempt. If a project depends on legally cleared music, precise voiceover wording, or exact loudness standards, treat the generated track as a draft until it passes a separate review.

3. Multimodal reference generation

A reference-to-video request can combine as many as 9 images, 3 video clips, and 3 audio clips, subject to a 12-file total and a 15-second total limit for reference video or audio. Every request still requires a written prompt. Audio also cannot be submitted by itself; at least one reference image or video must accompany it.

This is useful when one asset cannot carry the entire brief. A product image can define the object, a model image can define the person, a source clip can define the motion, and an audio clip can define the rhythm. The prompt then explains which traits to preserve and how the final scene should look.

More references also create more ways for instructions to conflict. Nine images are not automatically better than three. If two product shots show different packaging or two character references show different hair, the model has to guess which one wins.

4. First- and last-frame control

H3 supports a first frame, a last frame, or both. This is useful when the opening product composition and closing brand frame are already approved but the movement between them still needs to be generated.

The first/last-frame mode and the multimodal reference mode are separate API paths. MiniMax’s current specification says they cannot be mixed in the same request. That boundary is easy to miss: you can anchor the opening and closing images, or use a collection of reference images, videos, and audio, but not both input structures at once. MiniMax H3 API reference

5. Motion transfer and reference-video control

H3 can use a reference video to guide motion, camera behavior, editing rhythm, and other temporal information. That opens a practical path for choreography transfer, product-handling gestures, camera moves, and scene timing that would be difficult to describe with text alone.

This does not mean the source motion is copied frame for frame. H3 interprets the reference and regenerates the result. For recognizable choreography or tightly approved product movement, compare the output against the source timeline rather than assuming “reference” means exact replication.

6. Generative video editing

Video editing is the defining H3 feature, but the term needs a boundary. H3 is not a conventional editor where you select a layer, replace a word, and leave every unaffected pixel untouched.

Instead, the source video becomes a reference. The prompt tells H3 what to preserve and what to change, and the model generates a revised clip. This can support large creative changes—new setting, wardrobe, graphics, sound, or style—but those same changes may introduce identity drift, altered framing, misspelled text, or small motion changes.

In one sentence: H3 edits by regeneration, not by surgically modifying the original file.

How MiniMax H3 Video Editing Actually Works

A good H3 editing instruction has two layers:

  1. Preservation constraints: Who stays in the clip? Which product details, actions, camera moves, timing, and composition must remain?

  2. Transformation instructions: What background, clothing, visual style, typography, sound, or atmosphere should change?

The K-pop prompt uses this split deliberately. It first locks the performer count, facial identities, choreography, timing, and camera movement. Only then does it request a new studio, fashion direction, graphics, and audio treatment.

That ordering matters. “Turn this into a futuristic K-pop video” gives the model permission to reinterpret almost everything. “Preserve exactly three performers and the source choreography; replace only the art direction” gives it a hierarchy.

For harder edits, change one major category at a time. First test the background and wardrobe. Then add typography. Then refine audio. A single request may do all three, but separating them tells you which instruction caused a failure and makes retries easier to diagnose.

There is also a cost detail that launch summaries often miss. MiniMax bills a 2K output at $0.13 per second, and a reference video is billed at the same per-second rate according to its duration. A 15-second 2K generation therefore costs about $1.95. If the request also uses a 15-second reference video, the combined input-video and output charge is about $3.90 before any extra charge for more than five reference images. MiniMax pay-as-you-go pricing

That does not make reference editing unreasonable, but it changes how you plan iterations. Ten 15-second text-to-video attempts cost about $19.50 at the listed rate; ten equally long edits with a 15-second source video cost about $39.00. Short diagnostic generations are the sensible first pass.

What Kind of Videos Can MiniMax H3 Make?

H3 can cover the familiar text-to-video and image-to-video jobs, but its stronger use cases are the ones that benefit from several reference types or an existing timeline.

Revise a campaign without reshooting everything

A brand can supply the approved source performance, then request a new location, wardrobe direction, seasonal treatment, or graphic system. This is useful for concepting regional or event-specific versions. It is less suitable when every product label, facial detail, and gesture must survive with frame-level fidelity.

Turn product images into an e-commerce video

Product shots can define shape, color, and packaging while a motion reference demonstrates how the item should rotate, open, pour, fold, or move through the scene. A closing image can be useful for a fixed end card, although it cannot currently be combined with reference-to-video inputs in the same request.

Transfer choreography or camera motion

A reference clip can communicate timing that would take hundreds of words to describe. Dance, sports movement, handheld camera behavior, and product-reveal moves are obvious candidates. Always confirm that you have the right to use the source performance and likeness.

Create short music and fashion visuals

The combination of fast visual changes, generated sound, reference motion, and art-direction prompts suits music teasers and fashion campaigns. Exact typography remains a quality-control point, not a safe assumption.

Animate game promos and cinematic title screens

Reference images can keep a character, prop, or interface closer to an approved design while the prompt adds motion, camera travel, particles, and sound. Generated footage is useful for concept trailers, but interface text and gameplay-like sequences still need human review.

Generate 15-second action or chase sequences

The longer duration gives a prompt enough room for setup, escalation, and payoff. It also makes continuity harder. Prompts should divide the action into time ranges and repeat the subject traits that must not change.

MiniMax H3 for E-commerce and Advertising

The most useful way to think about H3 for e-commerce is not “generate an ad.” Think of it as assembling a brief from several media assets.

Input Job in the workflow
Product images Define the product’s shape, color, material, and packaging
Logo or packaging references Constrain brand marks and visual identity
Model or character images Keep the presenter’s appearance closer to the approved look
Motion reference video Specify a hand gesture, demonstration, camera move, or edit rhythm
Audio reference Guide pacing, voice, or sound direction
Written prompt State the final scene, preservation rules, changes, and exclusions

For example, a cosmetics brand could provide five clean product angles, one model reference, a 6-second hand-motion clip, and a short audio reference. The prompt could then request a 10-second vertical product reveal while preserving the bottle shape, label placement, cap color, model identity, and gesture timing.

That is a better brief than “make a luxury cosmetics ad.” It is also not a guarantee. Product geometry may bend, small label text may change, and a logo can drift between frames. If the output will become a paid advertisement, compare the product against the approved photography frame by frame and replace critical brand text in a conventional editor when necessary.

H3 is most useful here as a production accelerator for concepts, variants, and short campaign assets. It is not a substitute for brand approval.

MiniMax H3 vs Hailuo 2.3 and Hailuo 02

The generational difference is larger than the resolution column suggests. MiniMax’s previous Hailuo models are short-form generators built around text and image inputs. H3 adds a reference system that can take video and audio, then use those materials for creation, motion transfer, or editing.

Capability MiniMax H3 Hailuo 2.3 Hailuo 02
Maximum documented resolution 2K 1080p 1080p
Documented duration 4–15 seconds 6 seconds at 1080p; 6 or 10 seconds at 768p 6 seconds at 1080p; 6 or 10 seconds at 768p/512p
Frame rate 24 fps 24 fps 24 fps
Text-to-video Yes Yes Yes
Image-to-video Yes Yes Yes
First/last-frame input Yes Not listed as a primary current mode Yes
Reference video and audio Yes Not listed in the current model specification Not listed in the current model specification
Native generated audio Yes Not listed Not listed
Generative video editing Yes, through reference-based regeneration Not listed Not listed
Best fit Multimodal briefs, editing, motion transfer, audio, and longer 2K clips Short T2V/I2V with stronger action and expression handling Conventional short generation and lower-resolution options

The practical choice is straightforward:

  • Choose MiniMax H3 when the task needs an existing video, several reference assets, native sound, 2K output, or more than 10 seconds.

  • Choose Hailuo 2.3 when you only need a short text-to-video or image-to-video clip and do not need H3’s reference and editing system.

  • Consider Hailuo 02 when its first/last-frame path or lower-resolution pricing fits a simple legacy workflow.

MiniMax H3 is not yet available through GPT Proto. If you need a MiniMax video model that is already available there, try Hailuo 2.3 Pro for image-to-video or Hailuo 2.3 Pro for text-to-video.

MiniMax H3 Pricing and Availability

MiniMax H3’s official API is live under the model ID MiniMax-H3. Public API documentation currently lists 2K as the available output resolution. A 768p rate is published, but MiniMax marks that option as closed beta and directs users to contact sales before using it.

Item Published price
2K output video $0.13 per second
768p output video $0.09 per second; closed beta
Audio references Free
Reference images First 5 free; $0.04 per additional image
Reference video $0.13 per input second for a 2K output request
15-second 2K text-to-video generation About $1.95
15-second reference video + 15-second 2K output About $3.90

The API uses an asynchronous workflow: submit a generation request, receive a task ID, poll the task status, and download the result after it succeeds. MiniMax H3 is not currently available on GPT Proto, so this article does not include a GPT Proto code example or suggest that a GPT Proto H3 endpoint is already live.

The open-weights status also needs precise wording. MiniMax presents H3 as an open general-purpose model and has announced plans to publish the weights. Reuters reported that the files were expected within days. As of this article’s July 31 check, however, the downloadable weights, license terms, model size, and realistic local hardware requirements had not yet been confirmed publicly.

“Open weights” is therefore the announced direction, not proof that you can download and run H3 locally today.

What MiniMax H3 Still Does Not Guarantee

H3 expands what can be attempted in one generation, but it does not remove the usual failure modes of generative video.

Exact text may still require retries. The first example deliberately uses only two short phrases. If those words change across frames, use the generation for motion and add final typography in a normal editor.

Aggressive edits can change identity. Replacing a background, wardrobe, graphics system, audio, and lighting in one request gives the model more opportunities to reinterpret the people in the source.

A 15-second complex shot is harder than a simple 6-second animation. The longer the clip and the more events it contains, the more continuity decisions the model must preserve.

Generated editing is not pixel-preserving editing. H3 creates a new result from references. It is not the right tool for a legal, medical, or product-compliance edit where all untouched pixels must remain identical.

Official examples do not prove repeatable production reliability. A launch reel shows what a model can produce, not how many attempts the result required. Repeatability, failure rate, and prompt sensitivity need independent testing.

Local deployment requirements are still unknown. Until the actual weights, architecture details, license, and serving guidance arrive, any precise VRAM recommendation is speculation.

One independent signal is encouraging: on July 31, MiniMax H3 ranked first on Artificial Analysis’s video-editing leaderboard with audio, with an Elo score of 1,130 from 5,043 samples. That is a useful blind-preference result, but it is a snapshot rather than a guarantee for every task. Artificial Analysis video-editing leaderboard

Final Verdict: Is MiniMax H3 a Meaningful Upgrade?

Yes. MiniMax H3 is a larger upgrade over Hailuo 2.3 than “2K instead of 1080p” suggests.

Its strongest reason to exist is reference-driven generation and editing. Text, product images, a source performance, motion guidance, and audio can now contribute to one video request. For advertising and e-commerce teams, that changes the brief from “generate something similar” to “preserve these assets and this timing, then change these specific creative elements.”

The cost is control. Because H3 regenerates the clip, edits can affect details you wanted to keep. Longer shots also create more room for continuity failures, and input video increases the price of each attempt.

For a simple 6-second image animation, Hailuo 2.3 may still be enough. For a multimodal campaign, motion transfer, native sound, an existing-video revision, or a 15-second action sequence, H3 is the more relevant model to test.

MiniMax H3 is not currently available on GPT Proto. While access is being prepared, you can explore other video-generation models in the GPT Proto model gallery or visit GPT Proto to compare the models already available through one account.

Input: Text prompt only
Mode: Text-to-video
What it tests: Long motion, scale consistency, collisions, occlusion, camera tracking, and synchronized environmental sound

A 15-second photorealistic miniature chase sequence filmed as one continuous shot with no cuts.

0–4 seconds: An ultra-low, floor-level macro tracking camera follows directly behind a tiny miniature skateboarder riding smoothly across a polished wooden floor. The skateboarder and board remain clearly visible and maintain the same scale and appearance. Warm sunlight reflects softly across the wood. Shallow depth of field, realistic wheel vibration, subtle motion blur.

4–9 seconds: Giant colorful toy blocks suddenly tumble into the path from both sides. The skateboarder leans, swerves, and narrowly passes between them as a large toy truck rolls across the foreground. The camera keeps pace without losing the subject. Blocks collide naturally and cast realistic moving shadows.

9–15 seconds: A giant cute baby crawls rapidly into view in the background and reaches one enormous hand toward the skateboarder and the camera. The skateboarder accelerates beneath the approaching hand as the camera continues tracking inches above the floor. The hand passes close to the lens without touching it. End with the skateboarder escaping beneath the toy truck as the baby reacts with a surprised laugh.

Maintain consistent miniature scale, skateboarder identity, floor layout, lighting direction, and object positions throughout the shot. Realistic physics, natural hand anatomy, detailed wood and toy textures, cinematic macro perspective, warm bright indoor daylight, and energetic but readable motion.

Native stereo sound: tiny skateboard wheels rolling across wood, blocks clattering from left and right, the toy truck rumbling across the foreground, a soft baby laugh behind the camera, and a brief low whoosh as the giant hand approaches.

The timeline gives H3 a clearer sequence than a single paragraph full of simultaneous actions. Even so, this is a hard prompt. The camera must track one tiny subject while several much larger objects enter, collide, and partially obscure the view. A useful review should check continuity at the 4-second and 9-second transitions rather than judging only the sharpest frame.

What Is MiniMax H3?

MiniMax H3 is a general-purpose multimodal video model released by MiniMax on July 31, 2026. In plain language, it can read several kinds of creative input at once—written instructions, still images, video clips, and audio—and use them to generate a new video.

That definition separates H3 from MiniMax M3. H3 is the video-generation model discussed here. M3 belongs to MiniMax’s language and coding model line. Similar names, different jobs.

The official API model ID is MiniMax-H3. MiniMax’s current documentation lists three main generation paths: text-to-video, first- and/or last-frame image-to-video, and reference-to-video. The model outputs at 24 fps, supports 4 to 15 seconds in whole-second increments, and currently exposes 2K output through the public API. MiniMax H3 video documentation

Specification MiniMax H3
Official API model ID MiniMax-H3
Output resolution 2K
Output duration 4–15 seconds
Frame rate 24 fps
Input types Text, image, video, audio
Main modes Text-to-video, first/last-frame image-to-video, reference-to-video
Reference images Up to 9
Reference videos Up to 3 clips; 15 seconds total
Reference audio Up to 3 clips; 15 seconds total
Mixed reference files Up to 12
Prompt length Up to 7,000 characters
Supported output ratios 21:9, 16:9, 4:3, 1:1, 3:4, 9:16, or adaptive where supported

You may also see H3 described as “Hailuo 3.0.” That label makes sense as an informal way to place it after the Hailuo 2.x family, but MiniMax’s current model list and API documentation use MiniMax H3, not Hailuo 3.0, as the official name.

What Is New in the MiniMax H3 Upgrade?

Hailuo 2.3 focused on better physical action, human micro-expressions, stylized visuals, and motion-command response. H3 changes the scope of the product. It is no longer only a question of generating a short clip from text or one image; the model can now use existing media as working context.

1. 2K output for clips up to 15 seconds

MiniMax H3 raises the documented ceiling from the 6- or 10-second formats available in the older Hailuo API models to a flexible 4–15 seconds. It also raises the maximum resolution from 1080p to 2K.

Longer does not automatically mean more coherent. A 15-second shot gives the model more time to change a face, lose an object, or misunderstand a later event. That is why the second example above uses a timed sequence and explicit continuity constraints.

2. Native stereo audio

H3 generates sound as part of the video rather than requiring a completely separate audio-generation pass. Reuters reports native stereo sound as one of the release capabilities. For short ads, music visuals, and action clips, this can keep sound effects closer to the events that produced them. Reuters launch report

The trade-off is predictability. Native audio reduces workflow steps, but it does not guarantee exact dialogue, perfect lip synchronization, or a production-ready mix on every attempt. If a project depends on legally cleared music, precise voiceover wording, or exact loudness standards, treat the generated track as a draft until it passes a separate review.

3. Multimodal reference generation

A reference-to-video request can combine as many as 9 images, 3 video clips, and 3 audio clips, subject to a 12-file total and a 15-second total limit for reference video or audio. Every request still requires a written prompt. Audio also cannot be submitted by itself; at least one reference image or video must accompany it.

This is useful when one asset cannot carry the entire brief. A product image can define the object, a model image can define the person, a source clip can define the motion, and an audio clip can define the rhythm. The prompt then explains which traits to preserve and how the final scene should look.

More references also create more ways for instructions to conflict. Nine images are not automatically better than three. If two product shots show different packaging or two character references show different hair, the model has to guess which one wins.

4. First- and last-frame control

H3 supports a first frame, a last frame, or both. This is useful when the opening product composition and closing brand frame are already approved but the movement between them still needs to be generated.

The first/last-frame mode and the multimodal reference mode are separate API paths. MiniMax’s current specification says they cannot be mixed in the same request. That boundary is easy to miss: you can anchor the opening and closing images, or use a collection of reference images, videos, and audio, but not both input structures at once. MiniMax H3 API reference

5. Motion transfer and reference-video control

H3 can use a reference video to guide motion, camera behavior, editing rhythm, and other temporal information. That opens a practical path for choreography transfer, product-handling gestures, camera moves, and scene timing that would be difficult to describe with text alone.

This does not mean the source motion is copied frame for frame. H3 interprets the reference and regenerates the result. For recognizable choreography or tightly approved product movement, compare the output against the source timeline rather than assuming “reference” means exact replication.

6. Generative video editing

Video editing is the defining H3 feature, but the term needs a boundary. H3 is not a conventional editor where you select a layer, replace a word, and leave every unaffected pixel untouched.

Instead, the source video becomes a reference. The prompt tells H3 what to preserve and what to change, and the model generates a revised clip. This can support large creative changes—new setting, wardrobe, graphics, sound, or style—but those same changes may introduce identity drift, altered framing, misspelled text, or small motion changes.

In one sentence: H3 edits by regeneration, not by surgically modifying the original file.

How MiniMax H3 Video Editing Actually Works

A good H3 editing instruction has two layers:

  1. Preservation constraints: Who stays in the clip? Which product details, actions, camera moves, timing, and composition must remain?

  2. Transformation instructions: What background, clothing, visual style, typography, sound, or atmosphere should change?

The K-pop prompt uses this split deliberately. It first locks the performer count, facial identities, choreography, timing, and camera movement. Only then does it request a new studio, fashion direction, graphics, and audio treatment.

That ordering matters. “Turn this into a futuristic K-pop video” gives the model permission to reinterpret almost everything. “Preserve exactly three performers and the source choreography; replace only the art direction” gives it a hierarchy.

For harder edits, change one major category at a time. First test the background and wardrobe. Then add typography. Then refine audio. A single request may do all three, but separating them tells you which instruction caused a failure and makes retries easier to diagnose.

There is also a cost detail that launch summaries often miss. MiniMax bills a 2K output at $0.13 per second, and a reference video is billed at the same per-second rate according to its duration. A 15-second 2K generation therefore costs about $1.95. If the request also uses a 15-second reference video, the combined input-video and output charge is about $3.90 before any extra charge for more than five reference images. MiniMax pay-as-you-go pricing

That does not make reference editing unreasonable, but it changes how you plan iterations. Ten 15-second text-to-video attempts cost about $19.50 at the listed rate; ten equally long edits with a 15-second source video cost about $39.00. Short diagnostic generations are the sensible first pass.

What Kind of Videos Can MiniMax H3 Make?

H3 can cover the familiar text-to-video and image-to-video jobs, but its stronger use cases are the ones that benefit from several reference types or an existing timeline.

Revise a campaign without reshooting everything

A brand can supply the approved source performance, then request a new location, wardrobe direction, seasonal treatment, or graphic system. This is useful for concepting regional or event-specific versions. It is less suitable when every product label, facial detail, and gesture must survive with frame-level fidelity.

Turn product images into an e-commerce video

Product shots can define shape, color, and packaging while a motion reference demonstrates how the item should rotate, open, pour, fold, or move through the scene. A closing image can be useful for a fixed end card, although it cannot currently be combined with reference-to-video inputs in the same request.

Transfer choreography or camera motion

A reference clip can communicate timing that would take hundreds of words to describe. Dance, sports movement, handheld camera behavior, and product-reveal moves are obvious candidates. Always confirm that you have the right to use the source performance and likeness.

Create short music and fashion visuals

The combination of fast visual changes, generated sound, reference motion, and art-direction prompts suits music teasers and fashion campaigns. Exact typography remains a quality-control point, not a safe assumption.

Animate game promos and cinematic title screens

Reference images can keep a character, prop, or interface closer to an approved design while the prompt adds motion, camera travel, particles, and sound. Generated footage is useful for concept trailers, but interface text and gameplay-like sequences still need human review.

Generate 15-second action or chase sequences

The longer duration gives a prompt enough room for setup, escalation, and payoff. It also makes continuity harder. Prompts should divide the action into time ranges and repeat the subject traits that must not change.

MiniMax H3 for E-commerce and Advertising

The most useful way to think about H3 for e-commerce is not “generate an ad.” Think of it as assembling a brief from several media assets.

Input Job in the workflow
Product images Define the product’s shape, color, material, and packaging
Logo or packaging references Constrain brand marks and visual identity
Model or character images Keep the presenter’s appearance closer to the approved look
Motion reference video Specify a hand gesture, demonstration, camera move, or edit rhythm
Audio reference Guide pacing, voice, or sound direction
Written prompt State the final scene, preservation rules, changes, and exclusions

For example, a cosmetics brand could provide five clean product angles, one model reference, a 6-second hand-motion clip, and a short audio reference. The prompt could then request a 10-second vertical product reveal while preserving the bottle shape, label placement, cap color, model identity, and gesture timing.

That is a better brief than “make a luxury cosmetics ad.” It is also not a guarantee. Product geometry may bend, small label text may change, and a logo can drift between frames. If the output will become a paid advertisement, compare the product against the approved photography frame by frame and replace critical brand text in a conventional editor when necessary.

H3 is most useful here as a production accelerator for concepts, variants, and short campaign assets. It is not a substitute for brand approval.

MiniMax H3 vs Hailuo 2.3 and Hailuo 02

The generational difference is larger than the resolution column suggests. MiniMax’s previous Hailuo models are short-form generators built around text and image inputs. H3 adds a reference system that can take video and audio, then use those materials for creation, motion transfer, or editing.

Capability MiniMax H3 Hailuo 2.3 Hailuo 02
Maximum documented resolution 2K 1080p 1080p
Documented duration 4–15 seconds 6 seconds at 1080p; 6 or 10 seconds at 768p 6 seconds at 1080p; 6 or 10 seconds at 768p/512p
Frame rate 24 fps 24 fps 24 fps
Text-to-video Yes Yes Yes
Image-to-video Yes Yes Yes
First/last-frame input Yes Not listed as a primary current mode Yes
Reference video and audio Yes Not listed in the current model specification Not listed in the current model specification
Native generated audio Yes Not listed Not listed
Generative video editing Yes, through reference-based regeneration Not listed Not listed
Best fit Multimodal briefs, editing, motion transfer, audio, and longer 2K clips Short T2V/I2V with stronger action and expression handling Conventional short generation and lower-resolution options

The practical choice is straightforward:

  • Choose MiniMax H3 when the task needs an existing video, several reference assets, native sound, 2K output, or more than 10 seconds.

  • Choose Hailuo 2.3 when you only need a short text-to-video or image-to-video clip and do not need H3’s reference and editing system.

  • Consider Hailuo 02 when its first/last-frame path or lower-resolution pricing fits a simple legacy workflow.

MiniMax H3 is not yet available through GPT Proto. If you need a MiniMax video model that is already available there, try Hailuo 2.3 Pro for image-to-video or Hailuo 2.3 Pro for text-to-video.

MiniMax H3 Pricing and Availability

MiniMax H3’s official API is live under the model ID MiniMax-H3. Public API documentation currently lists 2K as the available output resolution. A 768p rate is published, but MiniMax marks that option as closed beta and directs users to contact sales before using it.

Item Published price
2K output video $0.13 per second
768p output video $0.09 per second; closed beta
Audio references Free
Reference images First 5 free; $0.04 per additional image
Reference video $0.13 per input second for a 2K output request
15-second 2K text-to-video generation About $1.95
15-second reference video + 15-second 2K output About $3.90

The API uses an asynchronous workflow: submit a generation request, receive a task ID, poll the task status, and download the result after it succeeds. MiniMax H3 is not currently available on GPT Proto, so this article does not include a GPT Proto code example or suggest that a GPT Proto H3 endpoint is already live.

The open-weights status also needs precise wording. MiniMax presents H3 as an open general-purpose model and has announced plans to publish the weights. Reuters reported that the files were expected within days. As of this article’s July 31 check, however, the downloadable weights, license terms, model size, and realistic local hardware requirements had not yet been confirmed publicly.

“Open weights” is therefore the announced direction, not proof that you can download and run H3 locally today.

What MiniMax H3 Still Does Not Guarantee

H3 expands what can be attempted in one generation, but it does not remove the usual failure modes of generative video.

Exact text may still require retries. The first example deliberately uses only two short phrases. If those words change across frames, use the generation for motion and add final typography in a normal editor.

Aggressive edits can change identity. Replacing a background, wardrobe, graphics system, audio, and lighting in one request gives the model more opportunities to reinterpret the people in the source.

A 15-second complex shot is harder than a simple 6-second animation. The longer the clip and the more events it contains, the more continuity decisions the model must preserve.

Generated editing is not pixel-preserving editing. H3 creates a new result from references. It is not the right tool for a legal, medical, or product-compliance edit where all untouched pixels must remain identical.

Official examples do not prove repeatable production reliability. A launch reel shows what a model can produce, not how many attempts the result required. Repeatability, failure rate, and prompt sensitivity need independent testing.

Local deployment requirements are still unknown. Until the actual weights, architecture details, license, and serving guidance arrive, any precise VRAM recommendation is speculation.

One independent signal is encouraging: on July 31, MiniMax H3 ranked first on Artificial Analysis’s video-editing leaderboard with audio, with an Elo score of 1,130 from 5,043 samples. That is a useful blind-preference result, but it is a snapshot rather than a guarantee for every task. Artificial Analysis video-editing leaderboard

Final Verdict: Is MiniMax H3 a Meaningful Upgrade?

Yes. MiniMax H3 is a larger upgrade over Hailuo 2.3 than “2K instead of 1080p” suggests.

Its strongest reason to exist is reference-driven generation and editing. Text, product images, a source performance, motion guidance, and audio can now contribute to one video request. For advertising and e-commerce teams, that changes the brief from “generate something similar” to “preserve these assets and this timing, then change these specific creative elements.”

The cost is control. Because H3 regenerates the clip, edits can affect details you wanted to keep. Longer shots also create more room for continuity failures, and input video increases the price of each attempt.

For a simple 6-second image animation, Hailuo 2.3 may still be enough. For a multimodal campaign, motion transfer, native sound, an existing-video revision, or a 15-second action sequence, H3 is the more relevant model to test.

MiniMax H3 is not currently available on GPT Proto. While access is being prepared, you can explore other video-generation models in the GPT Proto model gallery or visit GPT Proto to compare the models already available through one account.

Frequently Asked Questions

What is MiniMax H3?

MiniMax H3 is a multimodal AI video model from MiniMax. It can generate 4- to 15-second 2K videos from text and can also use images, videos, and audio as references for creation, motion transfer, and generative video editing.

Is MiniMax H3 officially released?

Yes. MiniMax H3 was officially released on July 31, 2026, and its hosted API is live under the model ID MiniMax-H3.

Is MiniMax H3 the same as Hailuo 3.0?

MiniMax H3 succeeds the Hailuo 2.x video family, so some third-party pages informally call it Hailuo 3.0. MiniMax’s official documentation currently uses the name MiniMax H3 and the API model ID MiniMax-H3.

Can MiniMax H3 edit an existing video?

Yes, but the editing is generative. You submit the original clip as a reference video and describe what should remain and what should change. H3 then renders a new video rather than altering only selected pixels in the source file.

How long can a MiniMax H3 video be?

The official API supports whole-second durations from 4 to 15 seconds.

Does MiniMax H3 generate audio?

Yes. H3 can generate native stereo sound with the video. It can also accept reference audio when the request includes at least one reference image or video.

Is MiniMax H3 open source?

“Open weights” is more accurate than “open source.” MiniMax has announced plans to release downloadable weights, but they were not publicly available when this article was last checked on July 31, 2026. The final license and local deployment requirements should be evaluated after the files are published.

Can MiniMax H3 run locally?

Not yet based on publicly available files. MiniMax plans to release the weights, but the model size, VRAM requirements, serving stack, and license details were still unconfirmed on July 31.

Is MiniMax H3 good for e-commerce videos?

Its multimodal reference workflow makes it well suited to product concepts, campaign variants, demonstrations, and short ads. Product shape, packaging text, logos, and colors still need manual quality control before commercial publication.

What is the difference between MiniMax H3 and Hailuo 2.3?

H3 supports up to 2K, 15-second output, native audio, video and audio references, motion transfer, and generative video editing. Hailuo 2.3 is a simpler short-form text-to-video and image-to-video model with up to 1080p output.

Is MiniMax H3 available on GPTProto?

Not yet. GPTProto currently offers other video-generation models, including existing Hailuo options, and H3 access can be added to this article after the model is integrated.