1–16 Second Clips
Choose any duration from 1 to 16 seconds. The default is 5 seconds, giving teams room to test short shots before rendering longer narrative beats.
curl --request POST "https://gptproto.com/api/v3/vidu/viduq3-pro/text-to-video" \
--header "Authorization: Bearer $GPTPROTO_API_KEY" \
--header "Content-Type: application/json" \
--data '{
"prompt": "A tiny origami fox sailing a teacup across a moonlit puddle",
"resolution": "720p",
"duration": 5,
"aspect_ratio": "4:3",
"movement_amplitude": "auto",
"audio": true,
"bgm": true,
"seed": 1
}'샘플 1회 비용부터 시작하고 테스트 예산을 고르세요. GPTProto 요금은 표시가보다 20% 저렴합니다.
$0.6 / 초
Use the ViduQ3 Pro API on GPTProto to create 1–16 second text-to-video clips in 540p, 720p, or 1080p. The model combines video, dialogue, and sound effects in one generation and supports smart scene changes for short narratives, ads, anime-inspired clips, and cinematic concepts. GPTProto provides access from $0.04 per second through one key and a shared balance for 200+ AI models.
Choose any duration from 1 to 16 seconds. The default is 5 seconds, giving teams room to test short shots before rendering longer narrative beats.
Use 720p by default, 540p for lower-cost iteration, or 1080p for final delivery. Vidu Q3 Pro outputs all three resolutions at 24 fps.
Set audio to true for synchronized dialogue and sound effects in the generated clip. Set it to false when your workflow needs silent footage.
Generate shot changes inside one clip instead of stitching separate renders. Prompt the sequence, camera language, pacing, dialogue, and sound cues in chronological order.
The ViduQ3 Pro API provides hosted access to Vidu’s Q3 Pro video model from ShengShu Technology. It is the quality-focused Q3 option for text-to-video, image-to-video, and first/last-frame generation. This page focuses on the text-to-video route, which converts a prompt of up to 5,000 characters into a video task.
Compared with the Q2 generation, Vidu Q3 Pro extends the maximum clip length from 10 to 16 seconds and adds direct synchronized audio plus smart scene cuts. Its longer single-pass duration is useful when a shot needs a setup, action, and payoff without joining several separately generated clips.
| Specification | Vidu Q3 Pro |
|---|---|
| Provider | Vidu / ShengShu Technology |
| Model ID | viduq3-pro |
| Official generation modes | Text-to-video, image-to-video, first/last-frame-to-video |
| Prompt length | Up to 5,000 characters for text-to-video |
| Duration | 1–16 seconds; 5 seconds by default |
| Resolution | 540p, 720p, or 1080p; 720p by default |
| Frame rate | 24 fps |
| Aspect ratios | 16:9, 9:16, 3:4, 4:3, and 1:1 |
| Audio | Direct audio-video generation; audio accepts true or false |
| Scene structure | Smart scene cuts are supported |
| Seed | Optional integer seed |
Short narrative ads: Fit a product reveal, action, reaction, and closing frame into one clip with aligned audio cues.
Dialogue-led scenes: Generate spoken lines and matching effects instead of adding every sound in post.
Anime-inspired videos: Describe cel shading, line work, motion, and framing in the prompt rather than the generic style field.
Previsualization: Test camera order, pacing, transitions, and audio direction before manual production.
These models overlap in high-resolution short-video generation, but they are built around different input and control workflows.
| Decision factor | Vidu Q3 Pro | Vidu Q2 Pro | Seedance 2.0 |
|---|---|---|---|
| Main workflow | Text, image, or first/last frame to video | Image, reference, or first/last frame to video | Unified text, image, audio, and video references plus editing |
| Maximum single generation | 16 seconds | 10 seconds | 15 seconds |
| Maximum listed resolution | 1080p | 1080p | Route-dependent; the current GPTProto page lists up to 4K |
| Audio approach | Direct synchronized audio for Q3 generation | Audio behavior varies by Q2 task; it is not the same Q3 text-to-video workflow | Joint audio-video generation with multimodal audio reference |
| Distinctive control | Smart scene cuts and a simple prompt-first workflow | Reference-based generation and video editing | Broader multimodal reference and editing controls |
| Best fit | Longest Vidu clip, synced dialogue/SFX, multi-shot text-to-video | Image-led or reference-led Vidu workflows | Projects that need several reference modalities or deeper editing control |
Choose Vidu Q3 Pro when you want the longest single-pass Vidu generation and direct audio from a text prompt. Choose Vidu Q2 Pro for reference-led Q2 workflows. Seedance 2.0 is the stronger fit when images, clips, and audio must all guide one result. For faster, lower-cost Vidu iteration, also compare Vidu Q3 Turbo.
Migrating from Vidu’s official API is a small integration change, but it is not a literal base-URL swap. The official text-to-video API uses POST https://api.vidu.com/ent/v2/text2video with an Authorization: Token header. GPTProto uses its own route and Bearer authentication.
| Integration item | Official Vidu API | GPTProto |
|---|---|---|
| Submit endpoint | /ent/v2/text2video |
/api/v3/vidu/viduq3-pro/text-to-video |
| Authentication | Authorization: Token ... |
Authorization: Bearer ... |
| Model selection | model: viduq3-pro in the request |
Model and task are identified in the route |
| Account setup | Separate Vidu account and credit balance | One GPTProto key and balance for 200+ models |
Both services use asynchronous tasks. Test request validation, status polling, callbacks, output URLs, and errors before moving traffic. Copy the current schema from the API Usage tab instead of assuming every field behaves identically.
Several fields appear in Vidu’s general text-to-video schema but do not behave as developers may expect with Q3:
| Parameter | Current Q3 behavior | Recommended handling |
|---|---|---|
audio |
Supported; true requests synchronized audio and false returns silent video |
Set it explicitly rather than relying on a default |
bgm |
The separate BGM option is not available for Q3 | Do not send bgm; describe the intended sound or music in the prompt and test the result |
style |
Vidu’s documentation says this generic field does not take effect for Q2 or Q3 | Put “anime,” “cinematic,” or another visual treatment in the prompt |
movement_amplitude |
The generic field does not take effect for Q2 or Q3 | Describe camera movement and subject motion in natural language |
seed |
An integer seed is accepted | Use it for controlled iteration, but do not promise frame-identical output |
Structure prompts in chronological order: subject, action, shot progression, visual treatment, dialogue, then sound. Avoid packing unrelated events into one short clip.
Cinematic dialogue: “Inside a rain-streaked late-night train, a detective says, ‘We missed one detail.’ Slow push-in, city lights crossing both faces, blue lighting, carriage hum, and rain against glass.”
Anime-inspired action: “Hand-drawn 2D anime. A courier races across tiled rooftops at sunset, leaps an alley, then looks back as a drone rises. Dynamic tracking shot, crisp line art, footsteps, wind, and an alarm tone.”
Product sequence: “A matte-black watch rests on stone. Macro shot of the crown, orbit to reveal the dial, then cut to the watch on a runner’s wrist at dawn. Controlled highlights, mechanism clicks, and light city ambience.”
For anime output, use descriptive language inside the prompt. Although the generic schema lists general and anime, Vidu’s current documentation states that the style field does not take effect on Q3 models.
Choose Vidu Q3 Pro for a complete 10–16 second narrative beat, synchronized dialogue or effects, and automatic shot changes. Test prompts at 540p, then move to 720p or 1080p once timing is stable. Choose Vidu Q3 Turbo for faster, lower-cost iteration or Seedance 2.0 for broader multimodal reference and editing controls.
이 모델과 관련된 가이드, 비교, 업데이트입니다.
모든 글
Vidu Q3가 뛰어난 캐릭터 일관성, 기본 제공 오디오-비디오 동기화, 전문가 수준의 16초 클립을 통해 전 세계 크리에이터를 위한 AI 비디오 업계를 어떻게 혁신하고 있는지 알아보세요.

Vidu Q2의 자연스러운 표정과 부드러운 카메라 워크로 시네마틱 AI 비디오를 만들어 보세요. Sora 2와 어떻게 비교되는지 확인하고 이미지를 즉시 비디오로 변환하세요.

카메라를 건너뛰고 싶으신가요? heygen은 텍스트로 매우 사실적인 AI 비디오를 만들 수 있지만, 크레딧 시스템이 빠르게 소진됩니다. ROI가 본인에게 맞는지 확인해 보세요.

sora.chatgpt가 AI로 비디오 제작을 어떻게 혁신하고 있는지 살펴보세요. 오늘날 산업 전반에 미치는 사용 사례, 기술적 한계, 경제적 영향을 알아보세요.
입력
출력