Native Audio in 1–16s Clips
Generate video and synchronized dialogue or sound effects in one request. Choose any duration from 1 to 16 seconds for text-, image-, or start/end-frame workflows.
curl --request POST "https://gptproto.com/api/v3/vidu/viduq3-turbo/text-to-video" \
--header "Authorization: Bearer $GPTPROTO_API_KEY" \
--header "Content-Type: application/json" \
--data '{
"prompt": "A tiny origami fox sailing a teacup across a moonlit puddle",
"style": "general",
"resolution": "720p",
"duration": 5,
"aspect_ratio": "4:3",
"movement_amplitude": "auto",
"audio": true,
"bgm": true,
"seed": 1
}'1回のサンプル費用から始め、テスト予算を選びます。GPTProto 料金は表示価格より 20% お得です。
$0.28 / 秒
Use the Vidu Q3 Turbo API for text-, image-, first/last-frame, or reference-led video workflows. Generate 1–16-second clips up to 1080p with native audio, using the same GPTProto key and balance across Vidu, Seedance, Veo, and other models.
Generate video and synchronized dialogue or sound effects in one request. Choose any duration from 1 to 16 seconds for text-, image-, or start/end-frame workflows.
Use text-to-video, image-to-video, first/last-frame, or reference-to-video inputs. Select the workflow that matches whether you are creating a scene, animating an asset, or preserving a subject.
Deliver 24 fps output in common landscape, portrait, square, and editorial aspect ratios. Use 540p for testing, 720p for balanced previews, and 1080p for final short-form assets.
GPTProto rates start at $0.032 per second, with resolution-based billing and no separate Vidu credit wallet. Use one balance across 200+ supported models.
Vidu Q3 Turbo is the speed-focused model in Vidu’s Q3 video generation family. It is designed for faster iteration than Vidu Q3 Pro while retaining the Q3 series’ direct audio-video generation and short-form storytelling features. The model can generate a new scene from text, animate an existing image, connect a defined first and last frame, or use visual references to preserve a subject or scene direction.
The Vidu Q3 Turbo API supports 540p, 720p, and 1080p output at 24 fps. Text-, image-, and start/end-frame jobs can run from 1 to 16 seconds; reference-to-video jobs are listed from 3 to 16 seconds. For text-to-video, developers can choose 16:9, 9:16, 3:4, 4:3, or 1:1. Direct audio generation can add dialogue and sound effects to the clip, but Vidu’s current documentation states that the separate bgm option is unavailable for Q3 models.
| Spec | Vidu Q3 Turbo API details |
|---|---|
| Provider | Vidu |
| Model string | viduq3-turbo |
| Generation modes | Text-to-video, image-to-video, first/last-frame video, reference-to-video |
| Duration | 1–16 seconds; reference-to-video is listed at 3–16 seconds |
| Resolution | 540p, 720p, or 1080p |
| Frame rate | 24 fps |
| Text-to-video aspect ratios | 16:9, 9:16, 3:4, 4:3, 1:1 |
| Text prompt limit | Up to 5,000 characters for the official text-to-video route |
| Audio | Direct audio-video output supported; silent output can be requested |
| Task format | Asynchronous generation with a task ID, status retrieval, and optional callback |
Use a source asset whenever composition, identity, or the ending state is already defined.
| Input mode | Use it when | Practical example |
|---|---|---|
| Text-to-video | You need a new scene without a fixed visual asset | Generate several concepts for a short social ad |
| Image-to-video | The opening product, character, or composition is already approved | Animate a product hero image while preserving its shape and branding |
| First/last-frame video | Both endpoints of a movement or transformation matter | Move from a closed package to a fully assembled product display |
| Reference-to-video | Subject or scene consistency must come from source images | Reuse the same product, mascot, or location across several campaign clips |
For ecommerce video, begin with an approved product image rather than rebuilding the product from text. Use image-to-video for one catalog asset and reference-to-video for repeated product or mascot details. Specify the camera move, action, final frame, audio, and destination ratio.
For batch video generation, store each asynchronous task ID beside its SKU, prompt version, source asset, resolution, duration, and campaign variant. Process callbacks separately from submission, make retries idempotent, and include rejected generations when calculating campaign cost.
Vidu Q3 Turbo is the economical choice for rapid short-form iterations. Seedance 2.0 offers a broader multimodal workflow and higher-resolution options on its GPTProto route. Veo 3.1 fits projects that specifically require Google’s video model ecosystem.
| Decision factor | Vidu Q3 Turbo | Seedance 2.0 | Veo 3.1 |
|---|---|---|---|
| Current GPTProto starting rate | $0.032/second | $0.0739/second | $0.50/second |
| Listed clip duration | 1–16 seconds | 4–15 seconds on the current GPTProto route | Check the selected Veo route |
| Maximum resolution shown on GPTProto | 1080p | Up to 4K | Route-dependent |
| Core strength | Fast, lower-cost short clips with four Vidu generation modes | Multimodal and higher-resolution video workflows | Google-model video pipeline |
| Best fit | Iteration, social clips, product variants, short narrative scenes | Asset-rich campaigns and higher-resolution output | Teams standardized on Veo-specific workflows |
Compare live rates using the same duration, resolution, ratio, assets, and audio settings.
Use a prompt structure that separates what must stay stable from what should move:
Prompt formula: subject or reference + scene goal + timed action + camera movement + lighting and visual style + dialogue or sound effects + details to preserve + final frame
Use the uploaded matte-black wireless speaker as the exact reference. Create a 10-second vertical launch clip with a slow macro push-in and a 120-degree camera rotation. Preserve the logo, controls, proportions, and finish. Use cool edge lighting, subtle electronic ambience, no extra text, and end on a centered front view.
A late-night Tokyo noodle shop. Dolly toward the first speaker as she says, “I finally sent it.” Cut to her friend: “Then tomorrow starts now.” Use natural pauses, low kitchen ambience, rain outside, warm tungsten lighting, and consistent faces and wardrobe.
Because the separate bgm option is not available for Q3, describe music, ambience, dialogue, and sound effects inside the prompt and enable direct audio-video output when sound is required.
A single generation is capped at 16 seconds, so longer narratives require planned shots and post-production assembly.
The documented output ceiling is 1080p at 24 fps; do not promise native 4K or variable frame rates on this model page.
The current official documentation states that style and movement_amplitude do not take effect on Q3 models.
The separate bgm parameter is unavailable for Q3; use direct audio generation and prompt-level audio direction instead.
Generation is asynchronous. Store task IDs, handle callbacks or polling, and design retries to avoid duplicate charges.
Review identity, product geometry, small text, dialogue, and audio timing before publishing a generated clip or paid advertisement.
vidu q3 AIモデルとAPI統合に関するよくある質問への専門的な回答。
このモデルに関連するガイド、比較、最新情報。
すべての記事
テキストプロンプトだけに頼るのはやめましょう。vidu q1のリファレンスから動画への変換機能なら、キャラクターの一貫性を完全にコントロールできます。レビュー全文を読む。

Vidu Q3が、優れたキャラクターの一貫性、ネイティブな音声と映像の同期、そして世界中のクリエイター向けのプロ品質な16秒クリップを提供し、AI動画業界にどのような革命を起こしているのかをご覧ください。

Vidu Q2の自然な表情と滑らかなカメラワークで、映画のようなAI動画を作成できます。Sora 2との違いを確認し、画像を瞬時に動画へ変換しましょう。

invideo aiがコンテンツ戦略に適したツールかどうかを確認しましょう。豊富なストックライブラリと、AI主導の編集の実態を検証します。今すぐ作成を始めましょう。
入力
出力