GPT Proto

GPTProto

  • Dashboard
  • LLM

    • claude
      Claude Opus 5New
    • google
      Gemini 3.6 Flash
    • google
      Gemini 3.5 Flash Lite
    • moonshotai
      Kimi K3
    • openai
      GPT 5.6 Luna

    Image

    • bytedance
      Dola Seedream 5.0 Pro 260628New
    • google
      Gemini 3.1 Flash Lite Image
    • google
      Gemini 3.1 Flash Image
    • openai
      GPT Image 2
    • google
      Gemini 3.1 Flash Image Preview

    Video

    • kling
      Kling v3.0 4kNew
    • bytedance
      Dreamina Seedance 2.0 Mini 260615
    • kling
      Kling v3 Omni 4k
    • bytedance
      Dreamina Seedance 2.0 Fast 260128
    • bytedance
      Dreamina Seedance 2.0 260128
    Explore 214+ Models >
  • Generator

    • Create Image
    • Create Video
    • Edit in Canvas

    Features

    • Anime to Real Life AINew
    • Anime AI Art Generator
    • AI Object Remover
    • AI Image Editor
    • Unrestricted AI Image Generator
    • AI Motion Transfer
    • AI Clothes Remover
    • AI Watermark Remover
    • AI Image Enhancer Online
    • Online Background Remover Tool
    Explore All >

    Prompts

    • Seedance 2.0 PromptsNew
    • GPT Image 2 Prompts
    • Nano Banana Pro Prompts
    • Seedream 5.0 Pro Prompts
  • AI Blog

    • GLM 5.2 vs MiniMax M3: Which Is Better for Coding and Frontend Work?
    • How to Create Your Own AI Character With an API—No Coding Required
    • Kimi K3 vs Claude Opus 5: Which Is Better for Coding and AI Agents?
    • 20 Free Seedream 5.0 Pro Packaging Design Prompts for Products and E-commerce
    • GLM-5.2 vs Kimi K3 for Coding: Which Is Better for Developers in 2026?
    Explore All >

    AI Insight

    • What Is Emochi AI—and Why Is It Growing So Fast? (2026)
    • What Is Kimi K3—and Is It Really Close to GPT-5.6 and Fable 5?
    • 12 Best AI Video Generation Tools in 2026 for YouTube, TikTok, Text and Images
    • What Is Qwen 3.8 Max? Release Date, 2.4T Preview, Pricing, and Early Benchmarks
    • Gemini 3.6 Flash and Gemini 3.5 Flash-Lite Explained: Which One Should You Use?
    Explore All >

    AI Docs

    • gpt-image-2
    • gpt-5.4
    • kimi-k2.5
    • claude-opus-4-6
    • kling-v3.0-pro
    Explore All >

    AI Skills

    • browser-use
    • claude-to-im
    • competitive-ads-extractor
    • content-creator
    • data-storytelling
    Explore All >
Pricing
English繁體中文한국어日本語EspañolРусский
Get Started Now
  1. Home
  2. /Model
  3. /Qwen
  4. /wan-2.5 / text-to-video
Qwen
wan-2.5 / text-to-video
Documentation
Documentation
Wan 2.5 Text to Video creates cinematic videos up to 10 seconds long at 1080p from textual descriptions, with realistic motion, lighting, and rich temporal details. It also generates synchronized audio, including voice and ambient sound, ideal for storytelling and marketing. Call it on GPTProto with one balance across 200+ models, from $0.225 per run.

$ 0.027
$ 0.03

text

video

$ 0.027
$ 0.03

text

video

Playground
JSON
API

Input

Your browser does not support the video tag.
Your request will cost$0per run, for$100you can run this model approximately0times
Related Models
All Models
Kling
Kling
kling-v3.0-4k
$ 1.008
$ 1.26
Bytedance
Bytedance
dreamina-seedance-2-0-mini-260615
$ 0.2365
Vidu
Vidu
viduq3-turbo
$ 0.032
$ 0.04
Qwen
Qwen
wan-2.6
$ 0.45
$ 0.5
Google
Google
veo-3.1-fast-generate-preview
$ 1.2
MiniMax
MiniMax
hailuo-2.3-pro
$ 0.441
$ 0.49
Examples
Studio Ghibli anime style, a bustling ancient Chinese market, streets are crowded with people, vendors are shouting their wares, and children are chasing each other playfully. The background features traditional architecture and waving banners. The camera moves through the crowd in a first-person perspective. Sound: A cacophony of human voices, including vendor calls, customer haggling, and children's laughter. In the background, there are sounds of gongs, distant opera music, and the general din of footsteps and objects. The background music is a lively and festive traditional Chinese folk tune.
A handsome, muscular man with well-defined abs is catching his breath after an intense workout. Sweat drips down his torso. He is shirtless, wearing only black athletic shorts, and is leaning against gym equipment. The lighting comes from the upper side, highlighting the contours of his chest and arms. The scene is filled with a raw, masculine energy, hyper-realistic, high-contrast lighting.
A middle-aged man sitting at a wooden desk in a cozy study room, surrounded by bookshelves and a warm lamp glow. He opens an old book and reads aloud with a calm, deep voice: 'History teaches us more than just facts… it shows us who we are.' The room has subtle background sounds: pages turning, the faint ticking of a clock, and distant rain against the window.
A vibrant young woman in her early 20s runs toward the camera in Times Square at night, ecstatic and wide-eyed, shouting passionately into a black microphone. She wears a neon green windbreaker and black headphones around her neck. She yells: “Yo, Wan2.5 just dropped on WaveSpeedAI — sound and texture are next level, try it right now!” Wet reflective streets, glowing blue-white-magenta billboards, blurred pedestrians, dynamic handheld follow-shot, sharp face focus, shallow depth of field. 4K UHD, saturated colors, viral UGC style.

What is Wan 2.5?

Wan 2.5 is Alibaba's (Tongyi / Wan team) video generation model, released in September 2025. It turns a text prompt — or a still image — into a clip up to 10 seconds long at up to 1080p (24fps), and generates the audio in the same pass: dialogue with lip-sync, ambient sound effects, and music. That one-pass audio-visual output is what separates it from earlier Wan releases and from most open video models, which produce silent video only.

Wan 2.5 ships as an API-only preview — its weights are not publicly downloadable. (Wan 2.1 and 2.2 are the open-weight, Apache-2.0 versions you can self-host; 2.5 and 2.6 are API-only.) On GPTProto, you call it with the model string wan-2.5, billed per generation by resolution and duration, on the same balance as 200+ other models.

Spec Value
Provider Alibaba (Tongyi / Wan)
Modalities Text-to-video (this page), image-to-video
Audio Native synchronized — voice + lip-sync + SFX + music, one pass
Resolution 480p / 720p / 1080p (24fps)
Max duration ~10s
Languages Multilingual (strong Chinese)
Weights API-only preview (not open-source)
Model string wan-2.5
Endpoint https://gptproto.com/api/v3/alibaba/wan-2.5/text-to-video

What you can build with the Wan 2.5 API

Short-form social and ads — 5–10s vertical or landscape clips with built-in voiceover and sound, no separate audio editing. At $0.45 per 720p/5s run, A/B-testing a dozen ad variants stays cheap.

Localized video at scale — Wan 2.5 handles multilingual prompts and generates lip-synced speech, so one storyboard becomes versions in several languages without re-shooting or re-dubbing. Strong Chinese support.

Storytelling and explainers — Dialogue, ambient sound, and music generated together keep narration and visuals aligned across the clip — useful for YouTube intros, product explainers, and training segments.

Animating a still image — Feed a single image as the first frame and a motion prompt to animate it.

 

How Wan 2.5 generates video and audio together

Most video models output silent frames; sound is a separate job you bolt on later. Wan 2.5 was built the other way around. Its generation step produces the visual frames and a matching audio track jointly, so speech lands on the right lip movements, footsteps line up with steps, and music sits under the cut without you editing anything. In practice this collapses a two-tool pipeline (video model + TTS/foley) into one API call.

Two things follow from that. First, your prompt should describe sound as well as motion — name the dialogue line, the ambient bed, and whether you want music (see §6). Second, because audio and video come from one pass, a single run is one billable unit; you are not paying twice or stitching tracks. If you need to supply your own music or voiceover instead, the audio parameter accepts it and the model syncs motion to it.

 

Choosing inside the Wan family

GPTProto carries several Wan models on one balance. Pick by what the job needs, not by version number:

  • Wan 2.5 — text-to-video (this page): 1080p, up to 10s, native audio. The default for a prompt-to-clip with sound.
  • Wan 2.5 — image-to-video: start from a still image as the first frame and animate it (the "wan 2.5 animate" case). 
  • Wan 2.2 (wan-2.2-plus): silent, 720p, ~5s — but open-weight, so it's the pick when you must self-host.
  • Wan 2.6 (wan-2.6): 1080p, up to 15s, plus multi-shot scene planning and reference-based identity retention — step up when you need longer or multi-cut narratives.

This map is the kind of thing competitors leave out: they list models without telling you which to reach for. Use 2.5 for single-shot clips with audio, 2.2 if you need weights, 2.6 for length and multi-shot.

 

Wan 2.5 vs Sora 2

Both generate video with synchronized audio from text or an image. The practical differences:

  Wan 2.5 Sora 2
Audio Native synced (voice / SFX / music) Native synced
Resolution up to 1080p up to 1080p
Max duration ~10s longer — up to ~25s (Pro)
Languages Multilingual, strong Chinese Multilingual
Content / IP rules Lighter Heavier — blocks IP & photorealistic likeness on some hosts
Open weights No (API-only) No
GPTProto price $0.225–$1.35 / run (by res × duration) $0.4 / run

Honest take: Sora 2 is a flat $0.4/run and supports longer clips (up to ~25s on Pro) — reach for it when duration matters. Wan 2.5 starts lower at $0.225/run (480p/5s) and lets you trade resolution for cost, with strong multilingual speech and fewer content restrictions. Both run on one GPTProto balance.

Related: Sora 2 → · Wan 2.6 → · Wan 2.2 →

Wan 2.5 prompt recipes

Wan 2.5 generates audio with the video, so good prompts describe sound as well as motion. The platform's own examples put audio inline with a Sound: cue — follow that convention. Structure: subject + action + camera move + lighting + Sound: cue + style. Paste and adjust.

1. Dialogue + lip-sync (shows the audio engine)

Medium close-up of an elderly fisherman on a wooden boat at dawn, soft golden
light, gentle water lapping. He looks to camera and says warmly, "The sea always
keeps its promises." Slow push-in. Sound: soft waves, distant seagulls, a calm
spoken line, no music.

2. Ambient / SFX, no dialogue

Rain on a neon-lit Tokyo alley at night, a cat darts under an awning, reflections
shimmer on wet pavement. Static wide shot, cinematic. Sound: steady rain, distant
traffic, no music, no speech.

3. Product / marketing

Close-up of an iced coffee on a marble counter, condensation beading, morning
kitchen light. A spoon stirs the ice. Slow 180-degree orbit around the glass.
Sound: soft clink of ice, quiet ambient kitchen tone, light background music.

4. Motion / tracking

A skateboarder rolls across an empty concrete plaza at sunset, low-angle tracking
shot alongside, subtle lens flare. Sound: wheels rumbling on pavement, wind
ambience, no dialogue.

Tips: keep dialogue short enough for the clip length (about one spoken line per 5s); name the camera move explicitly; state "no music" if you only want ambient sound; or pass your own track via the audio parameter.

How to Get a wan-2.5 API Key

Getting a wan-2.5 API key takes four steps and a few minutes. Create a free GPTProto account, add credits, generate your key, and make your first call — at $0.027 it's a cheaper wan-2.5 API key than going direct, and one key works across every model on the platform. Full wan-2.5 Documentation is in the docs.

Sign up

Sign up

Create your free GPT Proto account to begin. You can set up an organization for your team at any time.

Top up

Top up

Your balance can be used across all models on the platform, including wan-2.5, giving you the flexibility to experiment and scale as needed.

Generate your API key

Generate your API key

In your dashboard, create an API key — you'll need it to authenticate when making requests to wan-2.5.

Make your first API call

Make your first API call

Use your API key with our sample code to send a request to wan-2.5 via GPT Proto and see instant AI-powered results.

Get API Key

wan-2.5/text-to-video FAQ

Common developer questions and answers regarding wan-2.5/text-to-video model features, use, and performance.

What is wan-2.5/text-to-video?

Alibaba's video model that turns a text prompt into a clip up to 10 seconds at up to 1080p, with synchronized audio (voice, sound effects, music) generated in the same pass.

Is Wan 2.5 open source?

No. Wan 2.5 ships as an API-only preview and its weights are not publicly downloadable. Wan 2.1 and 2.2 are the open-weight (Apache-2.0) versions you can self-host; 2.5 and 2.6 are API-only.

Wan 2.5 vs Wan 2.2 — what changed?

Wan 2.5 adds native synchronized audio, raises standard resolution to 1080p (2.2 tops out at 720p), and extends clips to about 10s (2.2 ≈ 5s). Wan 2.2 stays the choice if you need open weights for local deployment.

Wan 2.5 vs Sora 2?

Both generate video with synced audio. Sora 2 ($0.4/run) supports longer clips (up to ~25s on Pro) but carries heavier content/IP restrictions; Wan 2.5 starts at $0.225/run, is multilingual, and runs on the same balance.

Does Wan 2.5 generate audio?

Yes — voice with lip-sync, sound effects, and music are generated together with the video in one pass. You can also supply your own track via the audio parameter, or describe sound inline with a Sound: cue in the prompt.

Can I animate a still image with Wan 2.5?

Yes, via image-to-video: supply an image as the first frame plus a motion prompt.

Is Wan 2.5 uncensored / does it allow NSFW?

wan-2.5/text-to-video excels in creative fidelity and output consistency. Its multimodal understanding allows for refined video details, accurate storytelling, and smooth transitions. Output quality is suitable for professional use, supporting a range of visual styles and storytelling techniques.

How much does the Wan 2.5 API cost on GPTProto?

From $0.225 (480p/5s) to $1.35 (1080p/10s) per run — see the pricing matrix above. Billed per generation on a single balance.

How do I call Wan 2.5 via API?

Submit a POST to https://gptproto.com/api/v3/alibaba/wan-2.5/text-to-video with your prompt and settings, then poll /api/v3/predictions/{id}/result for the finished video. See the Quick Start above.

Does wan-2.5/text-to-video support multi-modal input like images or audio?

wan-2.5/text-to-video is designed for strong text-to-video conversion and may support certain multimodal input extensions. The core workflow relies on textual input but advanced developers can explore image or audio prompts, based on evolving official documentation.

Are there copyright risks when using wan-2.5/text-to-video generated content?

wan-2.5/text-to-video generates original videos based on user prompts. Typical use does not pose copyright risk if you supply unique ideas. However, commercial users should avoid prompts that directly replicate well-known protected works. Always check local copyright laws before public distribution.

More Related GPTProto AI Tools

Wan AI Video Generator

Wan AI Video Generator

Transform simple text prompts into dynamic moving stories using Wan AI. Experience seamless lip-sync, synchronized audio, and realistic camera motion in seconds.

InVideo AI

InVideo AI

Transform text into stunning visuals instantly with our advanced invideo AI maker.

Picsart AI Video Generator

Transform your ideas into viral content with an advanced ai generator that takes your picsart workflow to the next level. Access premium models for seamless animation.

Image to Video AI

Transform static media into dynamic sequences using our cutting-edge image to video ai. Build your own online photo to video maker API today.

Related Articles

More Blogs
Wan 2.5: The Future of 4K AI Video Generation

Wan 2.5: The Future of 4K AI Video Generation

Discover how Wan 2.5 transforms images into 4K cinematic videos with AI. Learn features, pricing, and GPT Proto API options.

Best AI Video Generation Models 2025: Top 5 Ranked

Best AI Video Generation Models 2025: Top 5 Ranked

Discover the top AI video generation models in 2025. Compare pricing, explore free APIs, and find the best AI video generator for your needs with our guide.

wan2.2-animate: Quality Over Speed in AI

wan2.2-animate: Quality Over Speed in AI

While other models rush output, wan2.2-animate focuses on cinematic visual fidelity and prompt adherence. Learn how to optimize your video workflow today.

Wan 2.2 Animate: Real Image-to-Video

Wan 2.2 Animate: Real Image-to-Video

Stop struggling with AI video flickering. Learn how to configure wan 2.2 animate for precise character consistency and fluid motion. Read the full guide.

GPT Proto

Empowering AI Innovation with Global Scale and Stability:

With our flagship product GPT Proto, we offer a unified interface to access and combine APIs from the world's leading AI providers—spanning text, vision, speech, and beyond. We empower developers and enterprises to simplify integration and accelerate innovation without limits.

Global Infrastructure, Local Compliance:

To ensure enterprise-grade reliability and compliance, Talent Tech Global Limited operates specifically as our global Billing and Contracting Entity. Meanwhile, our core technical infrastructure and R&D teams are strategically distributed across global innovation hubs, including Silicon Valley, Singapore, and Hong Kong.

Built to Scale:

We understand that stability is paramount. Our platform is built on a robust, decentralized architecture supporting dynamic Auto-scaling. Whether you are running a pilot or handling millions of concurrent requests, our system expands instantly to meet demand—guaranteeing that your business never outgrows our infrastructure.

Navigation

  • Dashboard
  • Models
  • Create Image
  • AI Image Upscale
  • AI Background Remover
  • Create Video
  • Edit in Canvas
  • Features
  • Pricing
  • AI Docs
  • AI Blog
  • AI Insight
  • AI Skills

Features

  • Anime to Real Life AI
  • Anime AI Art Generator
  • AI Object Remover
  • AI Image Editor
  • Unrestricted AI Image Generator
  • AI Motion Transfer
  • AI Clothes Remover
  • AI Watermark Remover
  • AI Image Enhancer Online
  • Online Background Remover Tool
  • AI Face Swap Image
  • AI Passport Photo Maker
  • MS Paint AI Generator
Explore all features >

LLM

  • Claude Opus 5
  • Gemini 3.6 Flash
  • Gemini 3.5 Flash Lite
  • Kimi K3
  • GPT 5.6 Luna
  • GPT 5.6 Terra
  • GPT 5.6 Sol
  • Grok 4.5
  • Claude Sonnet 5
  • Minimax M3
  • GLM 5.2
  • GPT 5.1 Chat Latest
  • Claude Fable 5
  • Qwen3.7 Max
  • Claude Opus 4.8 Thinking
  • Claude Opus 4.8
  • Gemini 3.5 Flash
  • DeepSeek v4 Flash
  • DeepSeek v4 Pro
  • Grok 4.3
Explore all models >

Image

  • Dola Seedream 5.0 Pro 260628
  • Gemini 3.1 Flash Lite Image
  • Gemini 3.1 Flash Image
  • GPT Image 2
  • Gemini 3.1 Flash Image Preview
  • Seedream 5.0 260128
  • Doubao Seedream 5.0 260128
  • Vidu Q2
  • Grok Imagine Image
  • Kling Image O1
  • GPT Image 1.5
  • Seedream 4.5 251128
  • Doubao Seedream 4.5 251128
  • Grok Imagine 0.9
  • Gemini 3 Pro Image Preview
  • Qwen Image Lora
  • Qwen Image Plus Lora
  • Qwen Image Plus
  • Grok 4 Image
  • GPT Image 1 Mini
Explore all models >

Video

  • Kling v3.0 4k
  • Dreamina Seedance 2.0 Mini 260615
  • Kling v3 Omni 4k
  • Dreamina Seedance 2.0 Fast 260128
  • Dreamina Seedance 2.0 260128
  • Vidu 2.0
  • Doubao Seedance 2.0 260128
  • Doubao Seedance 2.0 Fast 260128
  • Kling v3 Omni Pro
  • Kling v3 Omni Std
  • Vidu Q3 Turbo
  • Kling v3.0 Pro
  • Kling v3.0 Std
  • Vidu Q3 Pro
  • Kling v2.6 Std
  • Vidu Q2 Pro
  • Vidu Q2 Turbo
  • Vidu Q2 Pro Fast
  • Vidu Q2
  • Kling Video O1 Pro
Explore all models >

© 2026 Talent Tech Global Limited (Hong Kong) / Talent Tech Global LLC (US). All rights reserved.

  • About Us
  • Privacy Policy
  • Terms of Service
  • Sitemap