1–16 Second Clips
Choose any duration from 1 to 16 seconds. The default is 5 seconds, giving teams room to test short shots before rendering longer narrative beats.
curl --request POST "https://gptproto.com/api/v3/vidu/viduq3-pro/text-to-video" \
--header "Authorization: Bearer $GPTPROTO_API_KEY" \
--header "Content-Type: application/json" \
--data '{
"prompt": "A tiny origami fox sailing a teacup across a moonlit puddle",
"resolution": "720p",
"duration": 5,
"aspect_ratio": "4:3",
"movement_amplitude": "auto",
"audio": true,
"bgm": true,
"seed": 1
}'Начните со стоимости одного образца и выберите тестовый бюджет. Тарифы GPTProto на 20% ниже прайса.
$0.6 / сек
Use the ViduQ3 Pro API on GPTProto to create 1–16 second text-to-video clips in 540p, 720p, or 1080p. The model combines video, dialogue, and sound effects in one generation and supports smart scene changes for short narratives, ads, anime-inspired clips, and cinematic concepts. GPTProto provides access from $0.04 per second through one key and a shared balance for 200+ AI models.
Choose any duration from 1 to 16 seconds. The default is 5 seconds, giving teams room to test short shots before rendering longer narrative beats.
Use 720p by default, 540p for lower-cost iteration, or 1080p for final delivery. Vidu Q3 Pro outputs all three resolutions at 24 fps.
Set audio to true for synchronized dialogue and sound effects in the generated clip. Set it to false when your workflow needs silent footage.
Generate shot changes inside one clip instead of stitching separate renders. Prompt the sequence, camera language, pacing, dialogue, and sound cues in chronological order.
The ViduQ3 Pro API provides hosted access to Vidu’s Q3 Pro video model from ShengShu Technology. It is the quality-focused Q3 option for text-to-video, image-to-video, and first/last-frame generation. This page focuses on the text-to-video route, which converts a prompt of up to 5,000 characters into a video task.
Compared with the Q2 generation, Vidu Q3 Pro extends the maximum clip length from 10 to 16 seconds and adds direct synchronized audio plus smart scene cuts. Its longer single-pass duration is useful when a shot needs a setup, action, and payoff without joining several separately generated clips.
| Specification | Vidu Q3 Pro |
|---|---|
| Provider | Vidu / ShengShu Technology |
| Model ID | viduq3-pro |
| Official generation modes | Text-to-video, image-to-video, first/last-frame-to-video |
| Prompt length | Up to 5,000 characters for text-to-video |
| Duration | 1–16 seconds; 5 seconds by default |
| Resolution | 540p, 720p, or 1080p; 720p by default |
| Frame rate | 24 fps |
| Aspect ratios | 16:9, 9:16, 3:4, 4:3, and 1:1 |
| Audio | Direct audio-video generation; audio accepts true or false |
| Scene structure | Smart scene cuts are supported |
| Seed | Optional integer seed |
Short narrative ads: Fit a product reveal, action, reaction, and closing frame into one clip with aligned audio cues.
Dialogue-led scenes: Generate spoken lines and matching effects instead of adding every sound in post.
Anime-inspired videos: Describe cel shading, line work, motion, and framing in the prompt rather than the generic style field.
Previsualization: Test camera order, pacing, transitions, and audio direction before manual production.
These models overlap in high-resolution short-video generation, but they are built around different input and control workflows.
| Decision factor | Vidu Q3 Pro | Vidu Q2 Pro | Seedance 2.0 |
|---|---|---|---|
| Main workflow | Text, image, or first/last frame to video | Image, reference, or first/last frame to video | Unified text, image, audio, and video references plus editing |
| Maximum single generation | 16 seconds | 10 seconds | 15 seconds |
| Maximum listed resolution | 1080p | 1080p | Route-dependent; the current GPTProto page lists up to 4K |
| Audio approach | Direct synchronized audio for Q3 generation | Audio behavior varies by Q2 task; it is not the same Q3 text-to-video workflow | Joint audio-video generation with multimodal audio reference |
| Distinctive control | Smart scene cuts and a simple prompt-first workflow | Reference-based generation and video editing | Broader multimodal reference and editing controls |
| Best fit | Longest Vidu clip, synced dialogue/SFX, multi-shot text-to-video | Image-led or reference-led Vidu workflows | Projects that need several reference modalities or deeper editing control |
Choose Vidu Q3 Pro when you want the longest single-pass Vidu generation and direct audio from a text prompt. Choose Vidu Q2 Pro for reference-led Q2 workflows. Seedance 2.0 is the stronger fit when images, clips, and audio must all guide one result. For faster, lower-cost Vidu iteration, also compare Vidu Q3 Turbo.
Migrating from Vidu’s official API is a small integration change, but it is not a literal base-URL swap. The official text-to-video API uses POST https://api.vidu.com/ent/v2/text2video with an Authorization: Token header. GPTProto uses its own route and Bearer authentication.
| Integration item | Official Vidu API | GPTProto |
|---|---|---|
| Submit endpoint | /ent/v2/text2video |
/api/v3/vidu/viduq3-pro/text-to-video |
| Authentication | Authorization: Token ... |
Authorization: Bearer ... |
| Model selection | model: viduq3-pro in the request |
Model and task are identified in the route |
| Account setup | Separate Vidu account and credit balance | One GPTProto key and balance for 200+ models |
Both services use asynchronous tasks. Test request validation, status polling, callbacks, output URLs, and errors before moving traffic. Copy the current schema from the API Usage tab instead of assuming every field behaves identically.
Several fields appear in Vidu’s general text-to-video schema but do not behave as developers may expect with Q3:
| Parameter | Current Q3 behavior | Recommended handling |
|---|---|---|
audio |
Supported; true requests synchronized audio and false returns silent video |
Set it explicitly rather than relying on a default |
bgm |
The separate BGM option is not available for Q3 | Do not send bgm; describe the intended sound or music in the prompt and test the result |
style |
Vidu’s documentation says this generic field does not take effect for Q2 or Q3 | Put “anime,” “cinematic,” or another visual treatment in the prompt |
movement_amplitude |
The generic field does not take effect for Q2 or Q3 | Describe camera movement and subject motion in natural language |
seed |
An integer seed is accepted | Use it for controlled iteration, but do not promise frame-identical output |
Structure prompts in chronological order: subject, action, shot progression, visual treatment, dialogue, then sound. Avoid packing unrelated events into one short clip.
Cinematic dialogue: “Inside a rain-streaked late-night train, a detective says, ‘We missed one detail.’ Slow push-in, city lights crossing both faces, blue lighting, carriage hum, and rain against glass.”
Anime-inspired action: “Hand-drawn 2D anime. A courier races across tiled rooftops at sunset, leaps an alley, then looks back as a drone rises. Dynamic tracking shot, crisp line art, footsteps, wind, and an alarm tone.”
Product sequence: “A matte-black watch rests on stone. Macro shot of the crown, orbit to reveal the dial, then cut to the watch on a runner’s wrist at dawn. Controlled highlights, mechanism clicks, and light city ambience.”
For anime output, use descriptive language inside the prompt. Although the generic schema lists general and anime, Vidu’s current documentation states that the style field does not take effect on Q3 models.
Choose Vidu Q3 Pro for a complete 10–16 second narrative beat, synchronized dialogue or effects, and automatic shot changes. Test prompts at 540p, then move to 720p or 1080p once timing is stable. Choose Vidu Q3 Turbo for faster, lower-cost iteration or Seedance 2.0 for broader multimodal reference and editing controls.
Руководства, сравнения и обновления по этой модели.
Все статьи
Узнайте, как Vidu Q3 меняет индустрию видео с помощью ИИ, обеспечивая превосходную согласованность персонажей, встроенную синхронизацию аудио и видео и профессиональные 16-секундные клипы для авторов по всему миру.

Создавайте кинематографичные видео с помощью ИИ с естественной мимикой и плавной работой камеры в Vidu Q2. Узнайте, как он сравнивается с Sora 2, и мгновенно превращайте изображения в видео.

Хотите обойтись без камеры? heygen создает чрезвычайно реалистичные видео с помощью ИИ на основе текста, но кредиты быстро расходуются. Узнайте, оправдывает ли это вложения именно для вас.

Узнайте, как sora.chatgpt меняет производство видео с помощью ИИ. Изучите варианты использования, технические ограничения и экономическое влияние на современные отрасли.
Вход
Выход