ViduQ3 pro API — 20% Cheaper Than Going Direct
Call the viduq3-pro API on GPTProto and generate up to 16 seconds of 1080p text-to-video with synchronized audio in a single request. viduq3-pro is Vidu's Q3 text-to-video model — it renders dialogue, sound effects, and music alongside the footage, cuts between shots automatically, and supports realistic and anime styles. On GPTProto you reach it at $0.04 per second, 20% under the $0.05 market rate, through one API key and one balance that also covers 200+ other text, image, and video models — no separate Vidu account, no regional sign-up, no per-provider credits.
What Is viduq3-pro
viduq3-pro is the Pro tier of Vidu's Q3 text-to-video model, built by Shengshu Technology. From a text prompt it produces a short cinematic clip with motion, camera control, and native audio in one pass. Three things carry the model: it generates up to 16 seconds of continuous footage (double the Q2 generation); it produces synchronized audio — dialogue with lip-sync, sound effects, and music — described directly in your prompt rather than added in post; and it runs Smart Cuts, switching camera angles and locations within a single generation so one viduq3-pro API call can return an edited multi-shot sequence. On the Artificial Analysis Video Arena it ranks #2 globally in text-to-video, ahead of Runway Gen-4.5 and Kling 2.5 Turbo.
| Spec | viduq3-pro (text-to-video) |
|---|---|
| Provider | Vidu (Shengshu Technology) |
| Model string | viduq3-pro |
| Input | Text prompt (up to 5,000 characters ) |
| Output | Video, up to 1080p, 24 fps |
| Resolution | 540p / 720p / 1080p (default 720p) |
| Duration | 1–16 s (default 5) |
| Aspect ratio | 16:9 / 9:16 / 3:4 / 4:3 / 1:1 |
| Audio | Native dialogue, SFX & BGM (on by default) |
| Styles | General (realistic) and Anime |
| Reproducibility | Seed supported |
| Price on GPTProto | From $0.04 / second (20% under $0.05 market) |
viduq3-pro vs viduq2-pro vs Seedance 2.0
| viduq3-pro | viduq2-pro | Seedance 2.0 | |
|---|---|---|---|
| Max duration | 16 s | 8 s | 15 s |
| Max resolution | 1080p | 1080p | 1080p |
| Native audio | Yes | Yes | Yes |
| Anime style | Yes | Yes | Yes |
| Arena rank (T2V) | #2 | — | #1 |
| Price on GPTProto | from $0.04/s | from $0.032/s | from $0.1663/s |
Which to pick: viduq3-pro gives the longest single-pass clip (16 s) and a dedicated anime style — reach for it when you need a full narrative beat or stylized 2D output. viduq2-pro is the cheaper option at $0.032/s when 8 seconds is enough. Seedance 2.0 currently tops the Arena for raw text-to-video quality; choose it when benchmark fidelity outranks clip length.
Billing on GPTProto
GPTProto bills viduq3-pro per generation from your prepaid balance — no credit packs, no expiring tiers. At $0.04 per second (20% under the $0.05 market rate), a 5-second 540p clip runs about $0.20, and the dashboard shows each request's exact cost. (Per-resolution rates are in the pricing table above.)
Switching From the Official Vidu API
Already calling Vidu directly? Moving to viduq3-pro on GPTProto is a drop-in change: point the base URL at GPTProto, swap in your GPTProto API key, and set the model string to viduq3-pro. Prompt, duration, resolution, and aspect-ratio parameters carry over. You gain access without a regional Vidu sign-up, one balance instead of per-provider credits, and the same key for 200+ other models.
ViduQ3 pro Video Prompt Recipes
viduq3-pro reads prompts up to 5,000 characters. Describe subject, action, camera, lighting, and — since audio is native — the sound you want.
- Cinematic: "Slow dolly-in on a lone astronaut at a cracked space-station window, Earth rising behind her; low engine hum, distant alarm beeps." (
duration: 8,resolution: 1080p) - Anime: "Anime style, a girl sprints across a rooftop at sunset, hair and scarf trailing; upbeat synth, footsteps on metal." (style: Anime)
- Product: "Overhead push-in on a matte-black watch rotating on marble; soft studio ambience, a subtle click as the second hand moves."
Be explicit about camera moves and sound — vague prompts default to static shots and generic audio.







