Text, Image, Video, and Audio in One Context
Combine a prompt with image, video, and audio references. MiniMax H3 can use those assets to guide subject identity, motion, camera behavior, visual style, voice, and editing rhythm.
curl --request POST "https://gptproto.com/api/v3/minimax/minimax-h3/text-to-video" \
--header "Authorization: Bearer $GPTPROTO_API_KEY" \
--header "Content-Type: application/json" \
--data '{
"prompt": "A tiny origami fox sailing a teacup across a moonlit puddle",
"aspect_ratio": "16:9",
"duration": 5,
"resolution": "768P"
}'Начните со стоимости одного образца и выберите тестовый бюджет. Тарифы GPTProto на 30% ниже прайса.
$0.455 / изобр.
Access the MiniMax-H3 API through GPTProto for cinematic video generation from prompts, keyframes, and mixed reference assets. Build 4–15-second clips with native stereo audio, use up to nine reference images, three video clips, and three audio tracks, and manage MiniMax H3 alongside 200+ AI models with one API key and one balance.
Combine a prompt with image, video, and audio references. MiniMax H3 can use those assets to guide subject identity, motion, camera behavior, visual style, voice, and editing rhythm.
Choose an integer duration from 4 to 15 seconds. Generate at 768P for drafts or request 2K output when sharper product detail, typography, or final-delivery quality matters.
MiniMax H3 is MiniMax's general-purpose, open-weight video generation system. Instead of treating text-to-video, keyframe animation, and multimodal references as unrelated tools, it interprets text, images, video, and audio within one context. The hosted API uses an asynchronous workflow: submit a task, store its task ID, check its status, and retrieve the finished video URL.
Output runs from 4 to 15 seconds at 24 fps, with 768P and 2K modes, landscape, square, and portrait ratios, and native stereo audio. The released H3-Base weights use the MiniMax H3 Community License; hosted API access and self-hosting remain separate deployment choices.
| Specification | MiniMax H3 API Details |
|---|---|
| Provider | MiniMax |
| Official API model name | MiniMax-H3 |
| GPTProto model string | MiniMax-H3 |
| Generation modes | Text-to-video, first-frame I2V, last-frame I2V, first-and-last-frame I2V, multimodal reference-to-video |
| Input types | Text, images, video, and audio |
| Output duration | 4–15 seconds, integer values |
| Output resolution | 768P or 2K |
| Frame rate | 24 fps |
| Output audio | Native 32 kHz stereo |
| Aspect ratios | 21:9, 16:9, 4:3, 1:1, 3:4, 9:16; adaptive where supported |
| Reference limits | Up to 9 images, 3 video clips, and 3 audio clips; up to 12 files combined |
| Prompt limit | Up to 7,000 characters in the official API |
| Processing | Asynchronous task workflow |
Choose the input mode before building the request. First/last-frame inputs and multimodal reference inputs are separate modes in the official API and cannot be mixed in one generation task.
| Input Mode | Use It When | Practical Example |
|---|---|---|
| Text-to-video | You need a new scene without an approved visual starting point | Generate several cinematic concepts from one campaign brief |
| First-frame image-to-video | The opening composition or product image is already approved | Animate a product hero image while retaining its initial framing |
| Last-frame image-to-video | The shot must resolve to a specific final composition | End on a logo lockup, product pack, or call-to-action frame |
| First-and-last-frame video | Both endpoints of a transition matter | Move from a closed package to the product fully revealed |
| Reference-to-video | Identity, motion, voice, music, setting, or style must come from source assets | Keep the same product and spokesperson while borrowing motion from a reference clip |
For reference-to-video, MiniMax H3 accepts up to nine images, three video clips, and three audio clips, with a maximum of 12 files in total. Each reference video or audio clip can be 2–15 seconds, and the total duration for each media type cannot exceed 15 seconds. Use public URLs for large assets because the complete request body is limited to 64 MB.
For ecommerce, supply approved product images, state which details must remain unchanged, and add a reference video only when its motion or camera path matters. First-and-last-frame mode is often simpler for controlled product reveals.
For batch generation, treat every variation as a separate asynchronous job. Store its prompt, asset URLs, settings, task ID, and status; use a bounded queue; make retries idempotent; and copy successful outputs to your own storage. Keep logo treatment, framing, lighting, camera language, and CTA fixed while injecting variable SKU data per job.
These models overlap but solve different constraints. MiniMax H3 emphasizes 2K output, open weights, stereo audio, and mixed references. Seedance 2.5 supports longer generations and larger reference sets. Veo 3.1 fits short Google Cloud workflows with a documented 4K path.
| Decision Factor | MiniMax H3 | Seedance 2.5 | Veo 3.1 Generate |
|---|---|---|---|
| Single-generation duration | 4–15 seconds | 4–30 seconds | 4, 6, or 8 seconds |
| Documented output resolution | 768P or 2K | 480P, 720P, or 1080P | 720P, 1080P, or 4K |
| Input modalities | Text, image, video, audio | Text, image, video, audio | Text and image; no audio or video input in the documented Generate endpoint |
| Reference capacity | Up to 9 images, 3 videos, 3 audio clips; 12 files total | Up to 30 images, 10 videos, and 10 audio clips | Up to 3 asset images |
| First/last-frame control | Yes | Yes | Yes |
| Native generated audio | Yes | Yes | Yes |
| Open weights | H3-Base released under a community license | No published open weights | No |
| Best fit | 2K multimodal reference work, motion transfer, editing, and self-hosting research | Longer scenes, large reference packs, and timeline-led editing | Short high-fidelity clips in Google Cloud and 4K delivery workflows |
Compare duration, reference capacity, audio, retries, and finished-clip cost—not maximum resolution alone. Seedance 2.5 may reduce scene stitching; H3 may simplify mixed-reference 2K work; Veo 3.1 may suit short Google-centered 4K workflows.
Write the prompt as a compact production brief that separates source roles, timeline, camera direction, audio, preserved details, and exclusions.
Prompt formula: subject and reference roles + scene goal + timed actions + camera movement + lighting and visual style + dialogue, sound effects, and music + details to preserve + final frame
Use Image 1 for product identity and Image 2 for lighting. Create a 10-second vertical launch video for the matte-black speaker. Begin with a macro grille shot, pull back as it rotates, then show water droplets vibrating with the bass. Preserve the logo, controls, proportions, and finish. Use cool edge light, restrained motion, low electronic ambience, and no extra text.
Start from the supplied closed package and end exactly on the supplied final frame with the product assembled. Use one continuous camera move: the box opens, components rise, and the product locks into place. Preserve package graphics and final geometry. Add mechanical clicks and a short resolved tone.
Руководства, сравнения и обновления по этой модели.
Все статьи
Compare 5 affordable AI video APIs for ecommerce and AI short drama. See current pricing, clip costs, audio fees, and the best model for each job.

Compare 6 cheapest AI image generators in 2026, from about $0.0035 per image. See batch costs, hidden fees, and the best API for startups.

Learn how to make an AI story video for kids with GPT Image 2 and Seedance 2.5, including the full prompt, captions, sound, editing, and a real test.

See Seedance 2.0 vs 2.5 in the same 15-second storyboard test. Compare cinematic quality, emotion, pricing, ecommerce use cases, and API features.
Вход
Выход