Text, Bild, Video und Audio in einem Kontext
Kombiniere einen Prompt mit Bild-, Video- und Audioreferenzen. MiniMax H3 kann diese Assets nutzen, um Identität der Motive, Bewegung, Kameraführung, visuellen Stil, Stimme und Schnittrhythmus zu steuern.
curl --request POST "https://gptproto.com/api/v3/minimax/minimax-h3/text-to-video" \
--header "Authorization: Bearer $GPTPROTO_API_KEY" \
--header "Content-Type: application/json" \
--data '{
"prompt": "A tiny origami fox sailing a teacup across a moonlit puddle",
"aspect_ratio": "16:9",
"duration": 5,
"resolution": "768P"
}'Nimm die Kosten für ein einzelnes Beispiel als Ausgangspunkt und leg ein Testbudget fest. Die Preise von GPTProto liegen 30% unter dem Listenpreis.
| Szenario | MiniMax Liste | OpenRouter | GPTProto | Du sparst / Monat |
|---|---|---|---|---|
| Kurze Produktwerbung30 × 5s / Monat | $19.50 | $20.57 | $13.65 | −$5.85≈ $70.20 / Jahr |
| Social-Media-Clips60 × 5s / Monat | $39.00 | $41.14 | $27.30 | −$11.70≈ $140.40 / Jahr |
| Massengenerierung100 × 5s / Monat | $65.00 | $68.58 | $45.50 | −$19.50≈ $234.00 / Jahr |
Nutze die MiniMax-H3-API über GPTProto für die Erstellung filmischer Videos aus Prompts, Keyframes und gemischten Referenzmaterialien. Erstelle 4–15-sekündige Clips mit nativem Stereo-Audio, verwende bis zu neun Referenzbilder, drei Videoclips und drei Audiospuren und verwalte MiniMax H3 zusammen mit mehr als 200 KI-Modellen – mit einem API-Schlüssel und einem Guthaben.
Text, Bild, Video und Audio in einem Kontext
Kombiniere einen Prompt mit Bild-, Video- und Audioreferenzen. MiniMax H3 kann diese Assets nutzen, um Identität der Motive, Bewegung, Kameraführung, visuellen Stil, Stimme und Schnittrhythmus zu steuern.
4–15 Sekunden in 768P oder 2K
Wähle eine ganzzahlige Dauer von 4 bis 15 Sekunden. Generiere Entwürfe in 768P oder fordere eine Ausgabe in 2K an, wenn schärfere Produktdetails, Typografie oder Qualität für die finale Auslieferung wichtig sind.
MiniMax H3 is MiniMax's general-purpose, open-weight video generation system. Instead of treating text-to-video, keyframe animation, and multimodal references as unrelated tools, it interprets text, images, video, and audio within one context. The hosted API uses an asynchronous workflow: submit a task, store its task ID, check its status, and retrieve the finished video URL.
Output runs from 4 to 15 seconds at 24 fps, with 768P and 2K modes, landscape, square, and portrait ratios, and native stereo audio. The released H3-Base weights use the MiniMax H3 Community License; hosted API access and self-hosting remain separate deployment choices.
| Specification | MiniMax H3 API Details |
|---|---|
| Provider | MiniMax |
| Official API model name | MiniMax-H3 |
| GPTProto model string | MiniMax-H3 |
| Generation modes | Text-to-video, first-frame I2V, last-frame I2V, first-and-last-frame I2V, multimodal reference-to-video |
| Input types | Text, images, video, and audio |
| Output duration | 4–15 seconds, integer values |
| Output resolution | 768P or 2K |
| Frame rate | 24 fps |
| Output audio | Native 32 kHz stereo |
| Aspect ratios | 21:9, 16:9, 4:3, 1:1, 3:4, 9:16; adaptive where supported |
| Reference limits | Up to 9 images, 3 video clips, and 3 audio clips; up to 12 files combined |
| Prompt limit | Up to 7,000 characters in the official API |
| Processing | Asynchronous task workflow |
Choose the input mode before building the request. First/last-frame inputs and multimodal reference inputs are separate modes in the official API and cannot be mixed in one generation task.
| Input Mode | Use It When | Practical Example |
|---|---|---|
| Text-to-video | You need a new scene without an approved visual starting point | Generate several cinematic concepts from one campaign brief |
| First-frame image-to-video | The opening composition or product image is already approved | Animate a product hero image while retaining its initial framing |
| Last-frame image-to-video | The shot must resolve to a specific final composition | End on a logo lockup, product pack, or call-to-action frame |
| First-and-last-frame video | Both endpoints of a transition matter | Move from a closed package to the product fully revealed |
| Reference-to-video | Identity, motion, voice, music, setting, or style must come from source assets | Keep the same product and spokesperson while borrowing motion from a reference clip |
For reference-to-video, MiniMax H3 accepts up to nine images, three video clips, and three audio clips, with a maximum of 12 files in total. Each reference video or audio clip can be 2–15 seconds, and the total duration for each media type cannot exceed 15 seconds. Use public URLs for large assets because the complete request body is limited to 64 MB.
For ecommerce, supply approved product images, state which details must remain unchanged, and add a reference video only when its motion or camera path matters. First-and-last-frame mode is often simpler for controlled product reveals.
For batch generation, treat every variation as a separate asynchronous job. Store its prompt, asset URLs, settings, task ID, and status; use a bounded queue; make retries idempotent; and copy successful outputs to your own storage. Keep logo treatment, framing, lighting, camera language, and CTA fixed while injecting variable SKU data per job.
These models overlap but solve different constraints. MiniMax H3 emphasizes 2K output, open weights, stereo audio, and mixed references. Seedance 2.5 supports longer generations and larger reference sets. Veo 3.1 fits short Google Cloud workflows with a documented 4K path.
| Decision Factor | MiniMax H3 | Seedance 2.5 | Veo 3.1 Generate |
|---|---|---|---|
| Single-generation duration | 4–15 seconds | 4–30 seconds | 4, 6, or 8 seconds |
| Documented output resolution | 768P or 2K | 480P, 720P, or 1080P | 720P, 1080P, or 4K |
| Input modalities | Text, image, video, audio | Text, image, video, audio | Text and image; no audio or video input in the documented Generate endpoint |
| Reference capacity | Up to 9 images, 3 videos, 3 audio clips; 12 files total | Up to 30 images, 10 videos, and 10 audio clips | Up to 3 asset images |
| First/last-frame control | Yes | Yes | Yes |
| Native generated audio | Yes | Yes | Yes |
| Open weights | H3-Base released under a community license | No published open weights | No |
| Best fit | 2K multimodal reference work, motion transfer, editing, and self-hosting research | Longer scenes, large reference packs, and timeline-led editing | Short high-fidelity clips in Google Cloud and 4K delivery workflows |
Compare duration, reference capacity, audio, retries, and finished-clip cost—not maximum resolution alone. Seedance 2.5 may reduce scene stitching; H3 may simplify mixed-reference 2K work; Veo 3.1 may suit short Google-centered 4K workflows.
Write the prompt as a compact production brief that separates source roles, timeline, camera direction, audio, preserved details, and exclusions.
Prompt formula: subject and reference roles + scene goal + timed actions + camera movement + lighting and visual style + dialogue, sound effects, and music + details to preserve + final frame
Use Image 1 for product identity and Image 2 for lighting. Create a 10-second vertical launch video for the matte-black speaker. Begin with a macro grille shot, pull back as it rotates, then show water droplets vibrating with the bass. Preserve the logo, controls, proportions, and finish. Use cool edge light, restrained motion, low electronic ambience, and no extra text.
Start from the supplied closed package and end exactly on the supplied final frame with the product assembled. Use one continuous camera move: the box opens, components rise, and the product locks into place. Preserve package graphics and final geometry. Add mechanical clicks and a short resolved tone.
Anleitungen, Vergleiche und Updates zu diesem Modell.
Alle Artikel
Vergleiche 5 günstige KI-Video-APIs für E-Commerce und KI-Kurzdramen. Erfahre mehr über aktuelle Preise, Clip-Kosten, Audiogebühren und das beste Modell für jeden Anwendungsfall.

Vergleiche 6 der günstigsten KI-Bildgeneratoren 2026 – ab etwa 0,0035 $ pro Bild. Erfahre mehr über Batch-Kosten, versteckte Gebühren und die beste API für Start-ups.

Erfahre, wie du mit GPT Image 2 und Seedance 2.5 ein KI-Geschichtenvideo für Kinder erstellst – inklusive vollständigem Prompt, Untertiteln, Ton, Schnitt und einem Praxistest.

Sieh dir Seedance 2.0 und 2.5 im selben 15-Sekunden-Storyboard-Test an. Vergleiche Kinoqualität, Emotionen, Preise, E-Commerce-Anwendungsfälle und API-Funktionen.
Eingabe
Ausgabe