GPT Proto

GPTProto

  • Dashboard
  • LLM

    • claude
      Claude Opus 5New
    • google
      Gemini 3.6 Flash
    • google
      Gemini 3.5 Flash Lite
    • moonshotai
      Kimi K3
    • openai
      GPT 5.6 Luna

    Image

    • bytedance
      Dola Seedream 5.0 Pro 260628New
    • google
      Gemini 3.1 Flash Lite Image
    • google
      Gemini 3.1 Flash Image
    • openai
      GPT Image 2
    • google
      Gemini 3.1 Flash Image Preview

    Video

    • kling
      Kling v3.0 4kNew
    • bytedance
      Dreamina Seedance 2.0 Mini 260615
    • kling
      Kling v3 Omni 4k
    • bytedance
      Dreamina Seedance 2.0 Fast 260128
    • bytedance
      Dreamina Seedance 2.0 260128
    Explore 214+ Models >
  • Generator

    • Create Image
    • Create Video
    • Edit in Canvas

    Features

    • Anime to Real Life AINew
    • Anime AI Art Generator
    • AI Object Remover
    • AI Image Editor
    • Unrestricted AI Image Generator
    • AI Motion Transfer
    • AI Clothes Remover
    • AI Watermark Remover
    • AI Image Enhancer Online
    • Online Background Remover Tool
    Explore All >

    Prompts

    • Seedance 2.0 PromptsNew
    • GPT Image 2 Prompts
    • Nano Banana Pro Prompts
    • Seedream 5.0 Pro Prompts
  • AI Blog

    • GLM 5.2 vs MiniMax M3: Which Is Better for Coding and Frontend Work?
    • How to Create Your Own AI Character With an API—No Coding Required
    • Kimi K3 vs Claude Opus 5: Which Is Better for Coding and AI Agents?
    • 20 Free Seedream 5.0 Pro Packaging Design Prompts for Products and E-commerce
    • GLM-5.2 vs Kimi K3 for Coding: Which Is Better for Developers in 2026?
    Explore All >

    AI Insight

    • What Is Emochi AI—and Why Is It Growing So Fast? (2026)
    • What Is Kimi K3—and Is It Really Close to GPT-5.6 and Fable 5?
    • 12 Best AI Video Generation Tools in 2026 for YouTube, TikTok, Text and Images
    • What Is Qwen 3.8 Max? Release Date, 2.4T Preview, Pricing, and Early Benchmarks
    • Gemini 3.6 Flash and Gemini 3.5 Flash-Lite Explained: Which One Should You Use?
    Explore All >

    AI Docs

    • gpt-image-2
    • gpt-5.4
    • kimi-k2.5
    • claude-opus-4-6
    • kling-v3.0-pro
    Explore All >

    AI Skills

    • browser-use
    • claude-to-im
    • competitive-ads-extractor
    • content-creator
    • data-storytelling
    Explore All >
Pricing
English繁體中文한국어日本語EspañolРусский
Get Started Now
  1. Home
  2. /Model
  3. /Qwen
  4. /wan-2.6
Qwen
wan-2.6
Documentation
Documentation
Wan 2.6 is Alibaba's text-to-video model: a prompt becomes a clip up to 15 seconds at up to 1080p, with synchronized audio — voice, ambient sound, and music- in the same pass. It plans multi-shot scenes and holds character identity across cuts. Call it on GPTProto from $0.45 per run, on one balance shared across 200+ models.

$ 0.45
$ 0.5

text

video

$ 0.45
$ 0.5

text

video

Playground
JSON
API

Input

Your browser does not support the video tag.
Your request will cost$0per run, for$100you can run this model approximately0times
Related Models
All Models
Kling
Kling
kling-v3.0-4k
$ 1.008
$ 1.26
Bytedance
Bytedance
dreamina-seedance-2-0-mini-260615
$ 0.2365
Vidu
Vidu
viduq3-turbo
$ 0.032
$ 0.04
Google
Google
veo-3.1-fast-generate-preview
$ 1.2
MiniMax
MiniMax
hailuo-2.3-pro
$ 0.441
$ 0.49
Qwen
Qwen
wan-2.2-plus
$ 0.09
$ 0.1
Examples
A highly cinematic 10-second shot.[0–3 seconds]
A first-person POV camera rapidly pushes forward, gliding through a cold, dark stone tunnel, racing toward a bright exit ahead. The environment feels damp and shadowy. In the background audio, the sound of a woman’s heavy breathing while running can be heard.

[3–5 seconds]
The camera transitions smoothly and seamlessly into a breathtaking, fresh, lush natural forest. A crystal-clear waterfall cascades down rocky cliffs. The meadow is covered with vibrant wildflowers, and several elegant deer graze calmly on the grass. Warm golden sunlight filters through the tree canopy, illuminating soft green grass below.

[Final 5 seconds]
The shot cuts to a medium full-body shot of a young woman. Her face is filled with awe and wonder, her mouth slightly open, deeply overwhelmed by the beauty before her, almost breathless. The camera then rapidly and fluidly rotates 360 degrees around her, revealing the vast, sunlit forest in all its grandeur.
Ultra-realistic, 8K resolution, natural lighting, no glowing plants, natural color palette, cinematic masterpiece.
glowing blue fox runs across a bioluminescent forest at night. Mushrooms pulse with soft light as particles float in the air. The camera follows close behind the fox, weaving between trees. Magical atmosphere, vibrant colors, fantasy cinematic style, sense of wonder and discovery
Extremely fast paced cinematic FPV flying through a post apocalyptic world reclaimed by nature, showcasing collapsed megastructures, overgrown skyscrapers, rusted bridges, abandoned vehicles, flooded streets, broken highways, and ruined industrial zones, with dynamic low altitude fly throughs, sharp turns, dives, and accelerations, ultra realistic lighting, natural dust, fog, and atmospheric haze, cinematic depth of field, high contrast color grading, dramatic sunrays breaking through clouds, realistic decay textures, epic scale destruction, immersive motion, smooth camera stabilization, and a powerful cinematic aesthetic
In the snow-covered children's park during winter, many play equipment were in use, and several children were playing. Sunlight shone on the gleaming ice slide, making it sparkle and reflect light, surrounded by a thick blanket of white snow. The camera focused on a young British boy, dressed in a blue down jacket, red scarf, snow boots, and woolen gloves, excitedly sliding down the ice slide. He laughed and shouted, "Wow! This slide is super slippery!" He leaned forward slightly and slid down the ice slide quickly, the soft crunch of the ice and the reflection of the snow highlighting his movements. Reaching the bottom, he crouched down, patted the snow he had slid down, and exclaimed excitedly, "So fun!"

What is Wan 2.6?

Wan 2.6 is Alibaba's (Tongyi / Wan team) multimodal video generation model, released December 2025. It turns a text prompt — or images, a reference video, and audio — into a clip up to 15 seconds long at up to 1080p (24fps), with audio generated in the same pass: dialogue with lip-sync, sound effects, and music.

 

Over Wan 2.5 it adds three things that matter for real production. Multi-shot narrative planning: one prompt can lay out several cuts (wide → close-up → reaction) and the model sequences them, instead of you generating and stitching separate clips. Reference-based generation: supply reference images or a reference video and the model holds a character's identity and look across shots. Higher motion fidelity: smoother, more stable motion across the longer 15s runtime. It accepts text, image, reference-video, and audio inputs in one workflow.

 

Wan 2.6 is an API-only model — weights are not publicly downloadable. Wan 2.1 and 2.2 are the open-weight (Apache-2.0) versions; 2.5 and 2.6 are API-only. On GPTProto you call it with the model string wan-2.6, billed per run by resolution and duration.

Spec Value
Provider Alibaba (Tongyi / Wan) · released Dec 2025
Modalities Text-to-video (this page); also image-to-video, reference-to-video
Inputs Text, image, reference video, audio
Audio Native synchronized — voice + lip-sync + SFX + music, one pass
Multi-shot Yes — multi-shot planning + identity retention
Resolution up to 1080p (24fps); landscape / portrait / square
Duration 5 / 10 / 15s
Weights API-only (not open-source)
Model string wan-2.6
Endpoint https://gptproto.com/api/v3/alibaba/wan-2.6/text-to-video

 

How multi-shot and synced audio work

Two capabilities define Wan 2.6, and both change how you prompt.

Synced audio in one pass. The generation step produces frames and a matching audio track together, so speech lands on the right lip movements and ambient sound and music sit under the cut without separate editing. Describe the sound in your prompt with a Sound: cue, or pass your own track via the audio parameter and the model syncs motion to it.

Multi-shot in one run. Rather than rendering one continuous shot, Wan 2.6 can sequence several beats inside a single clip. Segment the prompt by time ([0-5s] … [5-10s] … [10-15s] …) and the model treats each as a shot, carrying characters and setting across the cuts. The shot_type parameter controls whether you get a single continuous shot or a multi-shot composition. For coherence, 8–12s clips tend to be the most stable; 15s is available when you need the length.

 

Choosing inside the Wan family

GPTProto carries the Wan models on one balance. Pick by the job:

  • Wan 2.6 — text-to-video (this page): 1080p, up to 15s, multi-shot, native audio. Use it for longer or multi-cut narratives from a prompt.
  • Wan 2.6 — image-to-video / reference-to-video: animate a still, or guide a scene with a reference clip / up to several reference images for identity. 
  • Wan 2.5 (wan-2.5): 1080p, up to 10s, native audio, single-shot — the cheaper daily driver, from $0.225/run.
  • Wan 2.2 (wan-2.2-plus): silent, 720p, ~5s — but open-weight, the pick when you must self-host.

Competitors list these models without a selection map. The short version: 2.6 for length + multi-shot + identity, 2.5 for cheaper single-shot clips with audio, 2.2 for open weights.

 

Wan 2.6 vs Wan 2.5 vs Wan 2.2

  Wan 2.6 Wan 2.5 Wan 2.2
Audio Native synced Native synced None (video only)
Max duration 15s 10s ~5s
Resolution up to 1080p up to 1080p 720p
Multi-shot / reference Yes No No
Open weights No (API only) No (API only) Yes (Apache 2.0)
GPTProto price $0.45–$2.025 / run $0.225–$1.35 / run $0.09 / run

Honest take: Wan 2.6 costs more per run, but it's the one to use when you need 15s, multi-shot scenes, or character consistency across cuts. Wan 2.5 is the cheaper daily driver for single-shot clips up to 10s with audio. Wan 2.2 is the pick only if you need open weights to self-host.

Related: Wan 2.5 → · Wan 2.2 → · Sora 2 →

 

What you can build with the Wan 2.6 API

Multi-shot short films and ads — One prompt lays out several shots with consistent characters and audio across cuts — a 15s narrative ad without manual stitching. A 20-variant test at 720p/5s ≈ $9.00.

Localized video at scale — Multilingual prompts and lip-synced speech turn one storyboard into several languages without re-shooting. Strong Chinese support.

Any aspect ratio, same price — Square (960×960, 1440×1440), portrait (720×1280, 1080×1920), and landscape are priced identically within a tier, so format is a free choice per platform.

Image / reference-to-video — Animate a still or guide a scene with a reference clip on the sibling endpoints. 

 

Wan 2.6 prompt recipes

Wan 2.6 generates audio and can plan multiple shots, so prompts work best when they describe shots, motion, and sound together. The platform's own example segments the prompt by time — use that for multi-shot. Structure each beat: shot + subject + action + camera + lighting + Sound: cue.

1. Multi-shot narrative (15s + shot planning)

[0-5s] Wide shot: a lone hiker reaches a misty mountain ridge at dawn, slow push-in.
[5-10s] Medium shot: she lifts her camera and smiles. Sound: wind, distant birds.
[10-15s] Close-up: the sun breaks over the peaks, light fills the frame. Sound: a
soft swell of ambient music, no dialogue.

2. Dialogue + lip-sync (single shot)

Medium close-up of a barista behind a wooden counter, warm morning light. She looks
to camera and says, "Your usual? Coming right up." Steam rises from the cup. Sound:
spoken line, espresso machine hiss, quiet cafe ambience.

3. Product/promo, square frame

Square frame. A sneaker rotates slowly on a matte pedestal, studio light sweeps
across it, subtle reflections. Sound: a low synth pulse, soft whoosh on the sweep,
no speech.

Working tips: for multi-shot, label each beat with its time range and keep dialogue short enough for the beat; set shot_type for single vs multi-shot; 8–12s is the sweet spot for coherence, 15s when you need length; use negative_prompt to exclude artifacts (e.g. "no text, no watermark"); prompt_extend: true lets the model expand a short prompt — turn it off for exact control; fix seed to reproduce a result.

 

How to Get a wan-2.6 API Key

Getting a wan-2.6 API key takes four steps and a few minutes. Create a free GPTProto account, add credits, generate your key, and make your first call — at $0.45 it's a cheaper wan-2.6 API key than going direct, and one key works across every model on the platform. Full wan-2.6 Documentation is in the docs.

Sign up

Sign up

Create your free GPT Proto account to begin. You can set up an organization for your team at any time.

Top up

Top up

Your balance can be used across all models on the platform, including wan-2.6, giving you the flexibility to experiment and scale as needed.

Generate your API key

Generate your API key

In your dashboard, create an API key — you'll need it to authenticate when making requests to wan-2.6.

Make your first API call

Make your first API call

Use your API key with our sample code to send a request to wan-2.6 via GPT Proto and see instant AI-powered results.

Get API Key

Frequently Asked Questions

Common questions about wan-2.6/text-to-video model

What is wan-2.6/text-to-video?

Alibaba's video model that turns a text prompt into a clip up to 15 seconds at up to 1080p, with synchronized audio and multi-shot scene planning, in one pass.

Is Wan 2.6 open source?

No — API-only, weights not downloadable. Wan 2.1 and 2.2 are the open-weight (Apache-2.0) versions; 2.5 and 2.6 are API-only.

Wan 2.6 vs Wan 2.5 — what changed?

2.6 extends clips to 15s (2.5 caps at 10s), adds multi-shot planning and reference-based identity retention, and improves motion fidelity. Both do native synced audio at up to 1080p; 2.5 is cheaper per run.

Wan 2.2 vs Wan 2.6?

2.6 adds native synced audio, 1080p, 15s, and multi-shot; 2.2 is silent, 720p, ~5s. 2.2 is the only one of the two with open weights.

Can I do image-to-video with Wan 2.6?

Yes, on the image-to-video endpoint (animate a still) or reference-to-video (guide with a clip).

Does Wan 2.6 generate audio?

Yes — voice with lip-sync, SFX, and music in one pass. Pass your own track via audio, or describe sound inline with a Sound: cue.

How much does Wan 2.6 cost?

From $0.45 (720p/5s) to $2.025 (1080p/15s) per run — see the matrix above.

Is Wan 2.6 uncensored / does it allow NSFW?

GPTProto's actual content policy — the "unfiltered API access" compliance layer only.

How do I call it via API?

POST to the endpoint above, then poll /predictions/{id}/result. See Quick Start.

How do I pay for wan-2.6/text-to-video on GPT Proto?

On GPT Proto, payment options for wan-2.6/text-to-video include credit-based usage, prepaid plans, and enterprise subscriptions. Users register, obtain model access, and select a plan that matches their expected output. Usage is tracked by duration or number of generated videos, with invoices reflecting actual consumption. API-based billing integrates with popular payment gateways for seamless account management. Free-tier limitations and premium add-ons are clearly outlined in the GPT Proto dashboard. For projects requiring large-scale video production, enterprise agreements with custom SLA and support are available.

More Related GPTProto AI Tools

Picsart AI Video Generator

Transform your ideas into viral content with an advanced ai generator that takes your picsart workflow to the next level. Access premium models for seamless animation.

Pollo AI Video Generator

Transform simple prompts into stunning, scroll-stopping visual content using the advanced Pollo AI text to video engine.

Pika Labs AI Video Generator

Harness the power of Pika Labs to transform your ideas into hyper-realistic content. Our AI models bring dynamic stories to life.

Image to Video AI

Transform static media into dynamic sequences using our cutting-edge image to video ai. Build your own online photo to video maker API today.

Related Articles

More Blogs
Wan 2.5: The Future of 4K AI Video Generation

Wan 2.5: The Future of 4K AI Video Generation

Discover how Wan 2.5 transforms images into 4K cinematic videos with AI. Learn features, pricing, and GPT Proto API options.

wan2.2-animate: Quality Over Speed in AI

wan2.2-animate: Quality Over Speed in AI

While other models rush output, wan2.2-animate focuses on cinematic visual fidelity and prompt adherence. Learn how to optimize your video workflow today.

Wan Video: Mastering Open Source Generation

Wan Video: Mastering Open Source Generation

Alibaba's wan video is an open-source alternative to pricey AI models. Discover how to build better prompts and animate your creative ideas today.

Wan 2.5: The Future of 4K AI Video Generation

Wan 2.5: The Future of 4K AI Video Generation

Discover how Wan 2.5 transforms images into 4K cinematic videos with AI. Learn features, pricing, and GPT Proto API options.

GPT Proto

Empowering AI Innovation with Global Scale and Stability:

With our flagship product GPT Proto, we offer a unified interface to access and combine APIs from the world's leading AI providers—spanning text, vision, speech, and beyond. We empower developers and enterprises to simplify integration and accelerate innovation without limits.

Global Infrastructure, Local Compliance:

To ensure enterprise-grade reliability and compliance, Talent Tech Global Limited operates specifically as our global Billing and Contracting Entity. Meanwhile, our core technical infrastructure and R&D teams are strategically distributed across global innovation hubs, including Silicon Valley, Singapore, and Hong Kong.

Built to Scale:

We understand that stability is paramount. Our platform is built on a robust, decentralized architecture supporting dynamic Auto-scaling. Whether you are running a pilot or handling millions of concurrent requests, our system expands instantly to meet demand—guaranteeing that your business never outgrows our infrastructure.

Navigation

  • Dashboard
  • Models
  • Create Image
  • AI Image Upscale
  • AI Background Remover
  • Create Video
  • Edit in Canvas
  • Features
  • Pricing
  • AI Docs
  • AI Blog
  • AI Insight
  • AI Skills

Features

  • Anime to Real Life AI
  • Anime AI Art Generator
  • AI Object Remover
  • AI Image Editor
  • Unrestricted AI Image Generator
  • AI Motion Transfer
  • AI Clothes Remover
  • AI Watermark Remover
  • AI Image Enhancer Online
  • Online Background Remover Tool
  • AI Face Swap Image
  • AI Passport Photo Maker
  • MS Paint AI Generator
Explore all features >

LLM

  • Claude Opus 5
  • Gemini 3.6 Flash
  • Gemini 3.5 Flash Lite
  • Kimi K3
  • GPT 5.6 Luna
  • GPT 5.6 Terra
  • GPT 5.6 Sol
  • Grok 4.5
  • Claude Sonnet 5
  • Minimax M3
  • GLM 5.2
  • GPT 5.1 Chat Latest
  • Claude Fable 5
  • Qwen3.7 Max
  • Claude Opus 4.8 Thinking
  • Claude Opus 4.8
  • Gemini 3.5 Flash
  • DeepSeek v4 Flash
  • DeepSeek v4 Pro
  • Grok 4.3
Explore all models >

Image

  • Dola Seedream 5.0 Pro 260628
  • Gemini 3.1 Flash Lite Image
  • Gemini 3.1 Flash Image
  • GPT Image 2
  • Gemini 3.1 Flash Image Preview
  • Seedream 5.0 260128
  • Doubao Seedream 5.0 260128
  • Vidu Q2
  • Grok Imagine Image
  • Kling Image O1
  • GPT Image 1.5
  • Seedream 4.5 251128
  • Doubao Seedream 4.5 251128
  • Grok Imagine 0.9
  • Gemini 3 Pro Image Preview
  • Qwen Image Lora
  • Qwen Image Plus Lora
  • Qwen Image Plus
  • Grok 4 Image
  • GPT Image 1 Mini
Explore all models >

Video

  • Kling v3.0 4k
  • Dreamina Seedance 2.0 Mini 260615
  • Kling v3 Omni 4k
  • Dreamina Seedance 2.0 Fast 260128
  • Dreamina Seedance 2.0 260128
  • Vidu 2.0
  • Doubao Seedance 2.0 260128
  • Doubao Seedance 2.0 Fast 260128
  • Kling v3 Omni Pro
  • Kling v3 Omni Std
  • Vidu Q3 Turbo
  • Kling v3.0 Pro
  • Kling v3.0 Std
  • Vidu Q3 Pro
  • Kling v2.6 Std
  • Vidu Q2 Pro
  • Vidu Q2 Turbo
  • Vidu Q2 Pro Fast
  • Vidu Q2
  • Kling Video O1 Pro
Explore all models >

© 2026 Talent Tech Global Limited (Hong Kong) / Talent Tech Global LLC (US). All rights reserved.

  • About Us
  • Privacy Policy
  • Terms of Service
  • Sitemap