Multimodal Input, One Model
Start from text, a first frame, first and last frames, up to 10 reference images, five reference videos, five audio clips, a document, or a public webpage.
先從單次樣本成本開始,再選擇測試預算。GPTProto 費率比標價低 10%。
GPTProto · 價格估算
依價目表估算,最終費用可能有所不同。Top up
GPTProto vs official pricing.Save$11.10 (10%)vs Qwen official
Access Alibaba’s Wan3.0 API through GPTProto and use one account balance across Wan and other video models. Build prompt-only, first-frame, first-and-last-frame, reference-led, document-to-video, and webpage-to-video workflows without maintaining a separate provider account for every model in your pipeline.
Start from text, a first frame, first and last frames, up to 10 reference images, five reference videos, five audio clips, a document, or a public webpage.
Generate 2–30-second video at 480P, 720P, or 1080P and 30 fps. Use smart duration when the model should choose clip length from your prompt and references.
Guide characters, products, locations, props, and visual style with ordered references. Wan3.0 is designed to carry recognizable details across shots for branded and story-led video.
Generate audio by default, or disable it without changing the generation price. Revise visuals, plot, and dialogue through Wan3.0’s editing workflow instead of rebuilding the entire sequence.
Wan3.0 is Alibaba’s preview all-in-one video generation model. Instead of exposing a separate model for every creation path, the official wan3.0-video endpoint supports text-to-video, first-frame image-to-video, first-and-last-frame video, multimodal reference-to-video, document or webpage input, and video revision workflows.
The model can return 2–30 seconds of video in one generation. You can request 480P, 720P, or 1080P output at 30 fps, select a standard aspect ratio, or let the model adapt the ratio to the supplied media. Audio is enabled by default, while smart duration and prompt extension can help turn a short brief into a more complete sequence.
Wan3.0 is especially relevant when a video must preserve more than a general mood. Ordered references can define a character, product, location, camera language, voice, or musical direction. File and public-link inputs also let teams turn a product deck, report, article, or training document into a structured video brief.
| Spec | Wan3.0 API Details |
|---|---|
| Provider | Alibaba Cloud Model Studio |
| Official status | Preview |
| Official model ID | wan3.0-video |
| GPTProto model string | wan-3.0 |
| Generation modes | Text-to-video, first-frame, first-and-last-frame, reference-to-video, file-to-video, webpage-to-video |
| Output duration | 2–30 seconds; -1 enables smart duration |
| Output resolution | 480P, 720P, or 1080P |
| Frame rate | 30 fps |
| Aspect ratios | Adaptive, 16:9, 4:3, 1:1, 3:4, or 9:16 |
| Reference limits | Up to 10 images, 5 video clips, and 5 audio clips; total reference video and audio duration is capped separately at 15 seconds |
| File input | One DOC/DOCX, XLS/XLSX, PPT/PPTX, PDF, TXT, KEY, Pages, Numbers, or Markdown file; up to 100MB and 50 pages |
| Public webpage input | One publicly accessible HTTP or HTTPS page that does not require login |
| Prompt limit | 20,000 characters; Chinese and English supported |
| Audio | Enabled by default; can be disabled without changing the official generation rate |
| Watermark parameter | Disabled by default in Alibaba’s API; verify the live GPTProto implementation before making a blanket output-policy claim |
Choosing the correct mode matters because frame-control inputs and rich reference inputs are mutually exclusive in the official API. Do not combine first_frame or last_frame with reference images, reference videos, reference audio, a file, or a webpage in the same request.
| Input Mode | Use It When | Practical Example |
|---|---|---|
| Text-to-video | You want a new scene with no fixed visual asset | Generate several campaign concepts from structured shot prompts |
| First-frame video | The opening composition or product image is already approved | Animate a catalog hero image while preserving its initial framing |
| First-and-last-frame video | Both endpoints of the motion or transformation matter | Move from a closed package to a fully assembled product reveal |
| Reference-to-video | Identity, motion, voice, music, setting, or style must come from source assets | Keep the same product, spokesperson, location, and audio direction across a branded clip |
| File or webpage to video | The source idea lives in a deck, report, product page, article, or training document | Turn a product PPT into a launch video or a public article into a narrated visual summary |
Wan3.0 can reduce the gap between an approved commerce asset and a usable video variation. For a single SKU, start with an approved hero image when exact opening composition matters. Use reference mode when the product must remain recognizable across several shots, or supply a product presentation when features, positioning, and scene order already exist in a document.
For batch video generation, build a template for product data, references, ratio, duration, camera, audio, and ending CTA. Submit each variation asynchronously, store its task ID, and handle status separately from submission to reduce duplicate jobs. One approved source set can feed 9:16 social, 1:1 marketplace, and 16:9 product-page versions. Save successful outputs to your own storage.
These models overlap, but their strongest use cases are different. The best choice depends on required duration, reference volume, output resolution, document input, editing controls, and the cost of a finished usable clip—not only the advertised price per second.
| Decision Factor | Wan3.0 | Seedance 2.5 | MiniMax H3 |
|---|---|---|---|
| Maximum single generation | Up to 30 seconds | Up to 30 seconds | Up to 15 seconds |
| Published reference capacity | Up to 10 images, 5 videos, and 5 audio clips | Up to 30 images, 10 videos, and 10 audio clips | Up to 9 images, 3 videos, and 3 audio clips; 12 mixed files total |
| Highest documented API resolution | 1080P | Verify the resolution exposed by the selected endpoint | 2K |
| Document or public webpage input | Yes | Not listed as a core input in ByteDance’s official launch overview | Not listed in MiniMax’s current API input table |
| Audio | Optional audio output | Joint audio-video generation | Native stereo audio |
| Editing emphasis | Visual, plot, and dialogue revision; video extension | Timestamp-level audiovisual editing, green screen, camera perspective, and multi-round extension | Multimodal reference generation, editing, and motion transfer |
| Best fit | Long multimodal or document-led videos with a clearly documented 1080P API path | Reference-heavy 30-second storytelling that needs the largest published asset set | Shorter 2K clips, native stereo sound, and workflows that value an open model ecosystem |
Choose Wan3.0 over Seedance 2.5 when document or public-webpage input and a documented 480P-to-1080P ladder matter more than maximum reference volume. Choose MiniMax H3 when a 15-second ceiling is sufficient and 2K output or native stereo audio matters more. Compare cost using the same prompt, assets, duration, ratio, and acceptance criteria, including retries and rejected generations.
Use a prompt structure that separates the subject from the timeline and camera direction:
Prompt formula: subject and source references + scene goal + timed actions or story beats + camera movement + lighting and visual style + audio direction + details to preserve + details to avoid + final frame or CTA
Use Image 1 as the product reference and Image 2 as the lighting reference. Create a 12-second vertical launch video for a matte-black wireless speaker. Begin with a macro grille shot, pull back as the speaker rotates, then show bass vibrations moving nearby droplets. Preserve the logo, controls, proportions, and finish. Cool edge lighting, restrained camera motion, low electronic ambience, no extra text or deformation.
Turn the supplied product presentation into a 20-second 16:9 launch film. Arrange its three main benefits as problem, feature demonstration, and final value statement. Preserve documented colors and materials. Use clean studio lighting, measured camera moves, concise voiceover, and no unverified claims. End with the product centered against the brand background.
Alibaba states that audio texture and on-screen text rendering accuracy are still improving. For brand-safe output, add titles, legal copy, prices, and small labels in post-production instead of asking the model to render them perfectly inside the scene. Review dialogue and music closely before publishing, especially when the final asset will run as a paid advertisement.
First-frame and first-and-last-frame jobs cannot be mixed with reference, file, or webpage inputs. Validate rich media before submission, save task IDs, make retries idempotent, and store completed files promptly.
與本模型相關的指南、對比與更新。
所有文章
Learn how to turn an anime image into a realistic transformation video with matched first and last frames, prompts, sound, and editing tips.

See Seedance 2.0 vs 2.5 in the same 15-second storyboard test. Compare cinematic quality, emotion, pricing, ecommerce use cases, and API features.

Learn how to prompt realistic emotional facial expressions in Seedance 2.5 with timed micro-expressions, cinematic techniques, and two complete video examples.

Compare the best Chinese AI video models of 2026, including Seedance, MiniMax H3, Wan 3.0, Kling and Vidu, for text, image and reference video.