Native Audio in 1–16s Clips
Generate video and synchronized dialogue or sound effects in one request. Choose any duration from 1 to 16 seconds for text-, image-, or start/end-frame workflows.
curl --request POST "https://gptproto.com/api/v3/vidu/viduq3-turbo/text-to-video" \
--header "Authorization: Bearer $GPTPROTO_API_KEY" \
--header "Content-Type: application/json" \
--data '{
"prompt": "A tiny origami fox sailing a teacup across a moonlit puddle",
"style": "general",
"resolution": "720p",
"duration": 5,
"aspect_ratio": "4:3",
"movement_amplitude": "auto",
"audio": true,
"bgm": true,
"seed": 1
}'Start from the cost of a single sample and pick a testing budget. GPTProto rates are 20% below list price.
Use the Vidu Q3 Turbo API for text-, image-, first/last-frame, or reference-led video workflows. Generate 1–16-second clips up to 1080p with native audio, using the same GPTProto key and balance across Vidu, Seedance, Veo, and other models.
Generate video and synchronized dialogue or sound effects in one request. Choose any duration from 1 to 16 seconds for text-, image-, or start/end-frame workflows.
Use text-to-video, image-to-video, first/last-frame, or reference-to-video inputs. Select the workflow that matches whether you are creating a scene, animating an asset, or preserving a subject.
Deliver 24 fps output in common landscape, portrait, square, and editorial aspect ratios. Use 540p for testing, 720p for balanced previews, and 1080p for final short-form assets.
GPTProto rates start at $0.032 per second, with resolution-based billing and no separate Vidu credit wallet. Use one balance across 200+ supported models.
Vidu Q3 Turbo is the speed-focused model in Vidu’s Q3 video generation family. It is designed for faster iteration than Vidu Q3 Pro while retaining the Q3 series’ direct audio-video generation and short-form storytelling features. The model can generate a new scene from text, animate an existing image, connect a defined first and last frame, or use visual references to preserve a subject or scene direction.
The Vidu Q3 Turbo API supports 540p, 720p, and 1080p output at 24 fps. Text-, image-, and start/end-frame jobs can run from 1 to 16 seconds; reference-to-video jobs are listed from 3 to 16 seconds. For text-to-video, developers can choose 16:9, 9:16, 3:4, 4:3, or 1:1. Direct audio generation can add dialogue and sound effects to the clip, but Vidu’s current documentation states that the separate bgm option is unavailable for Q3 models.
| Spec | Vidu Q3 Turbo API details |
|---|---|
| Provider | Vidu |
| Model string | viduq3-turbo |
| Generation modes | Text-to-video, image-to-video, first/last-frame video, reference-to-video |
| Duration | 1–16 seconds; reference-to-video is listed at 3–16 seconds |
| Resolution | 540p, 720p, or 1080p |
| Frame rate | 24 fps |
| Text-to-video aspect ratios | 16:9, 9:16, 3:4, 4:3, 1:1 |
| Text prompt limit | Up to 5,000 characters for the official text-to-video route |
| Audio | Direct audio-video output supported; silent output can be requested |
| Task format | Asynchronous generation with a task ID, status retrieval, and optional callback |
Use a source asset whenever composition, identity, or the ending state is already defined.
| Input mode | Use it when | Practical example |
|---|---|---|
| Text-to-video | You need a new scene without a fixed visual asset | Generate several concepts for a short social ad |
| Image-to-video | The opening product, character, or composition is already approved | Animate a product hero image while preserving its shape and branding |
| First/last-frame video | Both endpoints of a movement or transformation matter | Move from a closed package to a fully assembled product display |
| Reference-to-video | Subject or scene consistency must come from source images | Reuse the same product, mascot, or location across several campaign clips |
For ecommerce video, begin with an approved product image rather than rebuilding the product from text. Use image-to-video for one catalog asset and reference-to-video for repeated product or mascot details. Specify the camera move, action, final frame, audio, and destination ratio.
For batch video generation, store each asynchronous task ID beside its SKU, prompt version, source asset, resolution, duration, and campaign variant. Process callbacks separately from submission, make retries idempotent, and include rejected generations when calculating campaign cost.
Vidu Q3 Turbo is the economical choice for rapid short-form iterations. Seedance 2.0 offers a broader multimodal workflow and higher-resolution options on its GPTProto route. Veo 3.1 fits projects that specifically require Google’s video model ecosystem.
| Decision factor | Vidu Q3 Turbo | Seedance 2.0 | Veo 3.1 |
|---|---|---|---|
| Current GPTProto starting rate | $0.032/second | $0.0739/second | $0.50/second |
| Listed clip duration | 1–16 seconds | 4–15 seconds on the current GPTProto route | Check the selected Veo route |
| Maximum resolution shown on GPTProto | 1080p | Up to 4K | Route-dependent |
| Core strength | Fast, lower-cost short clips with four Vidu generation modes | Multimodal and higher-resolution video workflows | Google-model video pipeline |
| Best fit | Iteration, social clips, product variants, short narrative scenes | Asset-rich campaigns and higher-resolution output | Teams standardized on Veo-specific workflows |
Compare live rates using the same duration, resolution, ratio, assets, and audio settings.
Use a prompt structure that separates what must stay stable from what should move:
Prompt formula: subject or reference + scene goal + timed action + camera movement + lighting and visual style + dialogue or sound effects + details to preserve + final frame
Use the uploaded matte-black wireless speaker as the exact reference. Create a 10-second vertical launch clip with a slow macro push-in and a 120-degree camera rotation. Preserve the logo, controls, proportions, and finish. Use cool edge lighting, subtle electronic ambience, no extra text, and end on a centered front view.
A late-night Tokyo noodle shop. Dolly toward the first speaker as she says, “I finally sent it.” Cut to her friend: “Then tomorrow starts now.” Use natural pauses, low kitchen ambience, rain outside, warm tungsten lighting, and consistent faces and wardrobe.
Because the separate bgm option is not available for Q3, describe music, ambience, dialogue, and sound effects inside the prompt and enable direct audio-video output when sound is required.
A single generation is capped at 16 seconds, so longer narratives require planned shots and post-production assembly.
The documented output ceiling is 1080p at 24 fps; do not promise native 4K or variable frame rates on this model page.
The current official documentation states that style and movement_amplitude do not take effect on Q3 models.
The separate bgm parameter is unavailable for Q3; use direct audio generation and prompt-level audio direction instead.
Generation is asynchronous. Store task IDs, handle callbacks or polling, and design retries to avoid duplicate charges.
Review identity, product geometry, small text, dialogue, and audio timing before publishing a generated clip or paid advertisement.
Expert answers to common questions regarding the vidu q3 AI model and API integration.
Guides, comparisons, and updates related to this model.
All Articles
Learn how to make an AI story video for kids with GPT Image 2 and Seedance 2.5, including the full prompt, captions, sound, editing, and a real test.

Turn one product image into a storyboard-led 25-second AI commercial using GPT Image 2 and Seedance 2.5, with prompts, shot timing, and QA tips.

See Seedance 2.0 vs 2.5 in the same 15-second storyboard test. Compare cinematic quality, emotion, pricing, ecommerce use cases, and API features.

Learn how to prompt realistic emotional facial expressions in Seedance 2.5 with timed micro-expressions, cinematic techniques, and two complete video examples.
Input
Output