Native Audio in 1–16s Clips
Generate video and synchronized dialogue or sound effects in one request. Choose any duration from 1 to 16 seconds for text-, image-, or start/end-frame workflows.
curl --request POST "https://gptproto.com/api/v3/vidu/viduq3-turbo/text-to-video" \
--header "Authorization: Bearer $GPTPROTO_API_KEY" \
--header "Content-Type: application/json" \
--data '{
"prompt": "A tiny origami fox sailing a teacup across a moonlit puddle",
"style": "general",
"resolution": "720p",
"duration": 5,
"aspect_ratio": "4:3",
"movement_amplitude": "auto",
"audio": true,
"bgm": true,
"seed": 1
}'先從單次樣本成本開始,再選擇測試預算。GPTProto 費率比標價低 20%。
$0.28 / 秒
Use the Vidu Q3 Turbo API for text-, image-, first/last-frame, or reference-led video workflows. Generate 1–16-second clips up to 1080p with native audio, using the same GPTProto key and balance across Vidu, Seedance, Veo, and other models.
Generate video and synchronized dialogue or sound effects in one request. Choose any duration from 1 to 16 seconds for text-, image-, or start/end-frame workflows.
Use text-to-video, image-to-video, first/last-frame, or reference-to-video inputs. Select the workflow that matches whether you are creating a scene, animating an asset, or preserving a subject.
Deliver 24 fps output in common landscape, portrait, square, and editorial aspect ratios. Use 540p for testing, 720p for balanced previews, and 1080p for final short-form assets.
GPTProto rates start at $0.032 per second, with resolution-based billing and no separate Vidu credit wallet. Use one balance across 200+ supported models.
Vidu Q3 Turbo is the speed-focused model in Vidu’s Q3 video generation family. It is designed for faster iteration than Vidu Q3 Pro while retaining the Q3 series’ direct audio-video generation and short-form storytelling features. The model can generate a new scene from text, animate an existing image, connect a defined first and last frame, or use visual references to preserve a subject or scene direction.
The Vidu Q3 Turbo API supports 540p, 720p, and 1080p output at 24 fps. Text-, image-, and start/end-frame jobs can run from 1 to 16 seconds; reference-to-video jobs are listed from 3 to 16 seconds. For text-to-video, developers can choose 16:9, 9:16, 3:4, 4:3, or 1:1. Direct audio generation can add dialogue and sound effects to the clip, but Vidu’s current documentation states that the separate bgm option is unavailable for Q3 models.
| Spec | Vidu Q3 Turbo API details |
|---|---|
| Provider | Vidu |
| Model string | viduq3-turbo |
| Generation modes | Text-to-video, image-to-video, first/last-frame video, reference-to-video |
| Duration | 1–16 seconds; reference-to-video is listed at 3–16 seconds |
| Resolution | 540p, 720p, or 1080p |
| Frame rate | 24 fps |
| Text-to-video aspect ratios | 16:9, 9:16, 3:4, 4:3, 1:1 |
| Text prompt limit | Up to 5,000 characters for the official text-to-video route |
| Audio | Direct audio-video output supported; silent output can be requested |
| Task format | Asynchronous generation with a task ID, status retrieval, and optional callback |
Use a source asset whenever composition, identity, or the ending state is already defined.
| Input mode | Use it when | Practical example |
|---|---|---|
| Text-to-video | You need a new scene without a fixed visual asset | Generate several concepts for a short social ad |
| Image-to-video | The opening product, character, or composition is already approved | Animate a product hero image while preserving its shape and branding |
| First/last-frame video | Both endpoints of a movement or transformation matter | Move from a closed package to a fully assembled product display |
| Reference-to-video | Subject or scene consistency must come from source images | Reuse the same product, mascot, or location across several campaign clips |
For ecommerce video, begin with an approved product image rather than rebuilding the product from text. Use image-to-video for one catalog asset and reference-to-video for repeated product or mascot details. Specify the camera move, action, final frame, audio, and destination ratio.
For batch video generation, store each asynchronous task ID beside its SKU, prompt version, source asset, resolution, duration, and campaign variant. Process callbacks separately from submission, make retries idempotent, and include rejected generations when calculating campaign cost.
Vidu Q3 Turbo is the economical choice for rapid short-form iterations. Seedance 2.0 offers a broader multimodal workflow and higher-resolution options on its GPTProto route. Veo 3.1 fits projects that specifically require Google’s video model ecosystem.
| Decision factor | Vidu Q3 Turbo | Seedance 2.0 | Veo 3.1 |
|---|---|---|---|
| Current GPTProto starting rate | $0.032/second | $0.0739/second | $0.50/second |
| Listed clip duration | 1–16 seconds | 4–15 seconds on the current GPTProto route | Check the selected Veo route |
| Maximum resolution shown on GPTProto | 1080p | Up to 4K | Route-dependent |
| Core strength | Fast, lower-cost short clips with four Vidu generation modes | Multimodal and higher-resolution video workflows | Google-model video pipeline |
| Best fit | Iteration, social clips, product variants, short narrative scenes | Asset-rich campaigns and higher-resolution output | Teams standardized on Veo-specific workflows |
Compare live rates using the same duration, resolution, ratio, assets, and audio settings.
Use a prompt structure that separates what must stay stable from what should move:
Prompt formula: subject or reference + scene goal + timed action + camera movement + lighting and visual style + dialogue or sound effects + details to preserve + final frame
Use the uploaded matte-black wireless speaker as the exact reference. Create a 10-second vertical launch clip with a slow macro push-in and a 120-degree camera rotation. Preserve the logo, controls, proportions, and finish. Use cool edge lighting, subtle electronic ambience, no extra text, and end on a centered front view.
A late-night Tokyo noodle shop. Dolly toward the first speaker as she says, “I finally sent it.” Cut to her friend: “Then tomorrow starts now.” Use natural pauses, low kitchen ambience, rain outside, warm tungsten lighting, and consistent faces and wardrobe.
Because the separate bgm option is not available for Q3, describe music, ambience, dialogue, and sound effects inside the prompt and enable direct audio-video output when sound is required.
A single generation is capped at 16 seconds, so longer narratives require planned shots and post-production assembly.
The documented output ceiling is 1080p at 24 fps; do not promise native 4K or variable frame rates on this model page.
The current official documentation states that style and movement_amplitude do not take effect on Q3 models.
The separate bgm parameter is unavailable for Q3; use direct audio generation and prompt-level audio direction instead.
Generation is asynchronous. Store task IDs, handle callbacks or polling, and design retries to avoid duplicate charges.
Review identity, product geometry, small text, dialogue, and audio timing before publishing a generated clip or paid advertisement.
關於 vidu q3 AI 模型與 API 整合常見問題的專業解答。
與本模型相關的指南、對比與更新。
所有文章
別再只依賴文字提示。Vidu Q1 的參考圖轉影片功能,讓你完全掌控角色一致性。閱讀完整評測。

探索 Vidu Q3 如何透過卓越的角色一致性、原生視聽同步,以及面向全球創作者的專業級 16 秒影片,革新 AI 影片產業。

使用 Vidu Q2 自然的表情與流暢的鏡頭運動,創作電影級 AI 影片。了解它與 Sora 2 的比較,並立即將圖片轉換成影片。

探索 invideo ai 是否適合你的內容策略。我們將深入了解其龐大的素材庫,以及 AI 主導剪輯的實際表現。立即開始創作。
輸入
輸出