Native Audio in 1–16s Clips
Generate video and synchronized dialogue or sound effects in one request. Choose any duration from 1 to 16 seconds for text-, image-, or start/end-frame workflows.
curl --request POST "https://gptproto.com/api/v3/vidu/viduq3-turbo/text-to-video" \
--header "Authorization: Bearer $GPTPROTO_API_KEY" \
--header "Content-Type: application/json" \
--data '{
"prompt": "A tiny origami fox sailing a teacup across a moonlit puddle",
"style": "general",
"resolution": "720p",
"duration": 5,
"aspect_ratio": "4:3",
"movement_amplitude": "auto",
"audio": true,
"bgm": true,
"seed": 1
}'샘플 1회 비용부터 시작하고 테스트 예산을 고르세요. GPTProto 요금은 표시가보다 20% 저렴합니다.
$0.28 / 초
Use the Vidu Q3 Turbo API for text-, image-, first/last-frame, or reference-led video workflows. Generate 1–16-second clips up to 1080p with native audio, using the same GPTProto key and balance across Vidu, Seedance, Veo, and other models.
Generate video and synchronized dialogue or sound effects in one request. Choose any duration from 1 to 16 seconds for text-, image-, or start/end-frame workflows.
Use text-to-video, image-to-video, first/last-frame, or reference-to-video inputs. Select the workflow that matches whether you are creating a scene, animating an asset, or preserving a subject.
Deliver 24 fps output in common landscape, portrait, square, and editorial aspect ratios. Use 540p for testing, 720p for balanced previews, and 1080p for final short-form assets.
GPTProto rates start at $0.032 per second, with resolution-based billing and no separate Vidu credit wallet. Use one balance across 200+ supported models.
Vidu Q3 Turbo is the speed-focused model in Vidu’s Q3 video generation family. It is designed for faster iteration than Vidu Q3 Pro while retaining the Q3 series’ direct audio-video generation and short-form storytelling features. The model can generate a new scene from text, animate an existing image, connect a defined first and last frame, or use visual references to preserve a subject or scene direction.
The Vidu Q3 Turbo API supports 540p, 720p, and 1080p output at 24 fps. Text-, image-, and start/end-frame jobs can run from 1 to 16 seconds; reference-to-video jobs are listed from 3 to 16 seconds. For text-to-video, developers can choose 16:9, 9:16, 3:4, 4:3, or 1:1. Direct audio generation can add dialogue and sound effects to the clip, but Vidu’s current documentation states that the separate bgm option is unavailable for Q3 models.
| Spec | Vidu Q3 Turbo API details |
|---|---|
| Provider | Vidu |
| Model string | viduq3-turbo |
| Generation modes | Text-to-video, image-to-video, first/last-frame video, reference-to-video |
| Duration | 1–16 seconds; reference-to-video is listed at 3–16 seconds |
| Resolution | 540p, 720p, or 1080p |
| Frame rate | 24 fps |
| Text-to-video aspect ratios | 16:9, 9:16, 3:4, 4:3, 1:1 |
| Text prompt limit | Up to 5,000 characters for the official text-to-video route |
| Audio | Direct audio-video output supported; silent output can be requested |
| Task format | Asynchronous generation with a task ID, status retrieval, and optional callback |
Use a source asset whenever composition, identity, or the ending state is already defined.
| Input mode | Use it when | Practical example |
|---|---|---|
| Text-to-video | You need a new scene without a fixed visual asset | Generate several concepts for a short social ad |
| Image-to-video | The opening product, character, or composition is already approved | Animate a product hero image while preserving its shape and branding |
| First/last-frame video | Both endpoints of a movement or transformation matter | Move from a closed package to a fully assembled product display |
| Reference-to-video | Subject or scene consistency must come from source images | Reuse the same product, mascot, or location across several campaign clips |
For ecommerce video, begin with an approved product image rather than rebuilding the product from text. Use image-to-video for one catalog asset and reference-to-video for repeated product or mascot details. Specify the camera move, action, final frame, audio, and destination ratio.
For batch video generation, store each asynchronous task ID beside its SKU, prompt version, source asset, resolution, duration, and campaign variant. Process callbacks separately from submission, make retries idempotent, and include rejected generations when calculating campaign cost.
Vidu Q3 Turbo is the economical choice for rapid short-form iterations. Seedance 2.0 offers a broader multimodal workflow and higher-resolution options on its GPTProto route. Veo 3.1 fits projects that specifically require Google’s video model ecosystem.
| Decision factor | Vidu Q3 Turbo | Seedance 2.0 | Veo 3.1 |
|---|---|---|---|
| Current GPTProto starting rate | $0.032/second | $0.0739/second | $0.50/second |
| Listed clip duration | 1–16 seconds | 4–15 seconds on the current GPTProto route | Check the selected Veo route |
| Maximum resolution shown on GPTProto | 1080p | Up to 4K | Route-dependent |
| Core strength | Fast, lower-cost short clips with four Vidu generation modes | Multimodal and higher-resolution video workflows | Google-model video pipeline |
| Best fit | Iteration, social clips, product variants, short narrative scenes | Asset-rich campaigns and higher-resolution output | Teams standardized on Veo-specific workflows |
Compare live rates using the same duration, resolution, ratio, assets, and audio settings.
Use a prompt structure that separates what must stay stable from what should move:
Prompt formula: subject or reference + scene goal + timed action + camera movement + lighting and visual style + dialogue or sound effects + details to preserve + final frame
Use the uploaded matte-black wireless speaker as the exact reference. Create a 10-second vertical launch clip with a slow macro push-in and a 120-degree camera rotation. Preserve the logo, controls, proportions, and finish. Use cool edge lighting, subtle electronic ambience, no extra text, and end on a centered front view.
A late-night Tokyo noodle shop. Dolly toward the first speaker as she says, “I finally sent it.” Cut to her friend: “Then tomorrow starts now.” Use natural pauses, low kitchen ambience, rain outside, warm tungsten lighting, and consistent faces and wardrobe.
Because the separate bgm option is not available for Q3, describe music, ambience, dialogue, and sound effects inside the prompt and enable direct audio-video output when sound is required.
A single generation is capped at 16 seconds, so longer narratives require planned shots and post-production assembly.
The documented output ceiling is 1080p at 24 fps; do not promise native 4K or variable frame rates on this model page.
The current official documentation states that style and movement_amplitude do not take effect on Q3 models.
The separate bgm parameter is unavailable for Q3; use direct audio generation and prompt-level audio direction instead.
Generation is asynchronous. Store task IDs, handle callbacks or polling, and design retries to avoid duplicate charges.
Review identity, product geometry, small text, dialogue, and audio timing before publishing a generated clip or paid advertisement.
vidu q3 AI 모델 및 API 통합에 관한 일반적인 질문에 대한 전문가 답변.
이 모델과 관련된 가이드, 비교, 업데이트입니다.
모든 글
텍스트 프롬프트에만 의존하지 마세요. vidu q1의 참조-투-비디오 기능은 캐릭터 일관성을 완벽하게 제어할 수 있게 해줍니다. 전체 리뷰를 읽어보세요.

Vidu Q3가 뛰어난 캐릭터 일관성, 기본 오디오-비주얼 동기화, 전문가 수준의 16초 클립을 전 세계 크리에이터에게 제공함으로써 AI 비디오 업계를 어떻게 혁신하고 있는지 알아보세요.

Vidu Q2의 자연스러운 표정과 부드러운 카메라 워크로 영화 같은 AI 비디오를 만들어 보세요. Sora 2와 어떻게 비교되는지 확인하고 이미지를 즉시 비디오로 변환하세요.

invideo ai가 콘텐츠 전략에 적합한 도구인지 알아보세요. 방대한 스톡 라이브러리와 AI 주도 편집의 현실을 살펴봅니다. 지금 시작하세요.
입력
출력