GPT Proto

GPTProto

  • 儀表板
  • LLM

    • claude
      Claude Opus 5新功能
    • google
      Gemini 3.6 Flash
    • google
      Gemini 3.5 Flash Lite
    • moonshotai
      Kimi K3
    • openai
      GPT 5.6 Luna

    影像

    • bytedance
      Dola Seedream 5.0 Pro 260628新功能
    • google
      Gemini 3.1 Flash Lite Image
    • google
      Gemini 3.1 Flash Image
    • openai
      GPT Image 2
    • google
      Gemini 3.1 Flash Image Preview

    影片

    • kling
      Kling v3.0 4k新功能
    • bytedance
      Dreamina Seedance 2.0 Mini 260615
    • kling
      Kling v3 Omni 4k
    • bytedance
      Dreamina Seedance 2.0 Fast 260128
    • bytedance
      Dreamina Seedance 2.0 260128
    探索 214+ 種模型 >
  • 生成器

    • 建立圖片
    • 建立影片
    • 在畫布中編輯

    功能

    • 動漫轉真人 AI新功能
    • 動漫 AI 藝術生成器
    • AI 物件移除器
    • AI 圖片編輯器
    • 無限制 AI 圖像生成器
    • AI 動作轉移
    • AI 衣物移除器
    • AI 浮水印移除工具
    • 線上 AI 圖片增強器
    • 線上背景移除工具
    探索全部 >

    提示詞

    • Seedance 2.0 提示詞新功能
    • GPT Image 2 提示詞
    • Nano Banana Pro 提示詞
    • Seedream 5.0 Pro 提示詞
  • AI 部落格

    • GLM 5.2 與 MiniMax M3:哪個更適合程式設計與前端工作?
    • 如何使用 API 建立自己的 AI 角色——無需撰寫程式碼
    • Kimi K3 與 Claude Opus 5:哪個更適合程式設計與 AI 代理?
    • 20 個免費 Seedream 5.0 Pro 產品與電子商務包裝設計提示詞
    • GLM-5.2 與 Kimi K3 程式設計比較:2026 年哪個更適合開發者?
    探索全部 >

    AI 洞察

    • 什麼是 Emochi AI?以及它為什麼成長得這麼快?(2026)
    • Kimi K3 是什麼?真的接近 GPT-5.6 與 Fable 5 嗎?
    • 2026 年 YouTube、TikTok、文字與圖片適用的 12 款最佳 AI 影片生成工具
    • 什麼是 Qwen 3.8 Max?發布日期、2.4T 預覽版、價格與早期基準測試
    • Gemini 3.6 Flash 與 Gemini 3.5 Flash-Lite 詳解:您應該使用哪一個?
    探索全部 >

    AI 文件

    • gpt-image-2
    • gpt-5.4
    • kimi-k2.5
    • claude-opus-4-6
    • kling-v3.0-pro
    探索全部 >

    AI 技能

    • browser-use
    • claude-to-im
    • competitive-ads-extractor
    • content-creator
    • data-storytelling
    探索全部 >
定價
English繁體中文한국어日本語EspañolРусский
立即開始
  1. 首頁
  2. /模型
  3. /Qwen
  4. /wan-2.6
Qwen
wan-2.6
說明文件
說明文件
Wan 2.6 是阿里巴巴的文字轉影片模型:輸入提示詞即可生成最長 15 秒、最高 1080p 的影片片段,並在同一次生成中同步處理音訊——語音、環境音和音樂。它能規劃多鏡頭場景,並在剪輯之間維持角色一致性。你可以在 GPTProto 上呼叫它,每次執行只需 $0.45,並使用一個可在 200 多個模型之間共用的餘額。

$ 0.45
$ 0.5

text

video

$ 0.45
$ 0.5

text

video

Playground
JSON
API

輸入

Your browser does not support the video tag.
您的請求將花費$0每次執行,對於$100您大約可以執行此模型0次
相關模型
所有模型
Kling
Kling
kling-v3.0-4k
$ 1.008
$ 1.26
Bytedance
Bytedance
dreamina-seedance-2-0-mini-260615
$ 0.2365
Vidu
Vidu
viduq3-turbo
$ 0.032
$ 0.04
Google
Google
veo-3.1-fast-generate-preview
$ 1.2
MiniMax
MiniMax
hailuo-2.3-pro
$ 0.441
$ 0.49
Qwen
Qwen
wan-2.2-plus
$ 0.09
$ 0.1
範例
A highly cinematic 10-second shot.[0–3 seconds]
A first-person POV camera rapidly pushes forward, gliding through a cold, dark stone tunnel, racing toward a bright exit ahead. The environment feels damp and shadowy. In the background audio, the sound of a woman’s heavy breathing while running can be heard.

[3–5 seconds]
The camera transitions smoothly and seamlessly into a breathtaking, fresh, lush natural forest. A crystal-clear waterfall cascades down rocky cliffs. The meadow is covered with vibrant wildflowers, and several elegant deer graze calmly on the grass. Warm golden sunlight filters through the tree canopy, illuminating soft green grass below.

[Final 5 seconds]
The shot cuts to a medium full-body shot of a young woman. Her face is filled with awe and wonder, her mouth slightly open, deeply overwhelmed by the beauty before her, almost breathless. The camera then rapidly and fluidly rotates 360 degrees around her, revealing the vast, sunlit forest in all its grandeur.
Ultra-realistic, 8K resolution, natural lighting, no glowing plants, natural color palette, cinematic masterpiece.
glowing blue fox runs across a bioluminescent forest at night. Mushrooms pulse with soft light as particles float in the air. The camera follows close behind the fox, weaving between trees. Magical atmosphere, vibrant colors, fantasy cinematic style, sense of wonder and discovery
Extremely fast paced cinematic FPV flying through a post apocalyptic world reclaimed by nature, showcasing collapsed megastructures, overgrown skyscrapers, rusted bridges, abandoned vehicles, flooded streets, broken highways, and ruined industrial zones, with dynamic low altitude fly throughs, sharp turns, dives, and accelerations, ultra realistic lighting, natural dust, fog, and atmospheric haze, cinematic depth of field, high contrast color grading, dramatic sunrays breaking through clouds, realistic decay textures, epic scale destruction, immersive motion, smooth camera stabilization, and a powerful cinematic aesthetic
In the snow-covered children's park during winter, many play equipment were in use, and several children were playing. Sunlight shone on the gleaming ice slide, making it sparkle and reflect light, surrounded by a thick blanket of white snow. The camera focused on a young British boy, dressed in a blue down jacket, red scarf, snow boots, and woolen gloves, excitedly sliding down the ice slide. He laughed and shouted, "Wow! This slide is super slippery!" He leaned forward slightly and slid down the ice slide quickly, the soft crunch of the ice and the reflection of the snow highlighting his movements. Reaching the bottom, he crouched down, patted the snow he had slid down, and exclaimed excitedly, "So fun!"

What is Wan 2.6?

Wan 2.6 is Alibaba's (Tongyi / Wan team) multimodal video generation model, released December 2025. It turns a text prompt — or images, a reference video, and audio — into a clip up to 15 seconds long at up to 1080p (24fps), with audio generated in the same pass: dialogue with lip-sync, sound effects, and music.

 

Over Wan 2.5 it adds three things that matter for real production. Multi-shot narrative planning: one prompt can lay out several cuts (wide → close-up → reaction) and the model sequences them, instead of you generating and stitching separate clips. Reference-based generation: supply reference images or a reference video and the model holds a character's identity and look across shots. Higher motion fidelity: smoother, more stable motion across the longer 15s runtime. It accepts text, image, reference-video, and audio inputs in one workflow.

 

Wan 2.6 is an API-only model — weights are not publicly downloadable. Wan 2.1 and 2.2 are the open-weight (Apache-2.0) versions; 2.5 and 2.6 are API-only. On GPTProto you call it with the model string wan-2.6, billed per run by resolution and duration.

Spec Value
Provider Alibaba (Tongyi / Wan) · released Dec 2025
Modalities Text-to-video (this page); also image-to-video, reference-to-video
Inputs Text, image, reference video, audio
Audio Native synchronized — voice + lip-sync + SFX + music, one pass
Multi-shot Yes — multi-shot planning + identity retention
Resolution up to 1080p (24fps); landscape / portrait / square
Duration 5 / 10 / 15s
Weights API-only (not open-source)
Model string wan-2.6
Endpoint https://gptproto.com/api/v3/alibaba/wan-2.6/text-to-video

 

How multi-shot and synced audio work

Two capabilities define Wan 2.6, and both change how you prompt.

Synced audio in one pass. The generation step produces frames and a matching audio track together, so speech lands on the right lip movements and ambient sound and music sit under the cut without separate editing. Describe the sound in your prompt with a Sound: cue, or pass your own track via the audio parameter and the model syncs motion to it.

Multi-shot in one run. Rather than rendering one continuous shot, Wan 2.6 can sequence several beats inside a single clip. Segment the prompt by time ([0-5s] … [5-10s] … [10-15s] …) and the model treats each as a shot, carrying characters and setting across the cuts. The shot_type parameter controls whether you get a single continuous shot or a multi-shot composition. For coherence, 8–12s clips tend to be the most stable; 15s is available when you need the length.

 

Choosing inside the Wan family

GPTProto carries the Wan models on one balance. Pick by the job:

  • Wan 2.6 — text-to-video (this page): 1080p, up to 15s, multi-shot, native audio. Use it for longer or multi-cut narratives from a prompt.
  • Wan 2.6 — image-to-video / reference-to-video: animate a still, or guide a scene with a reference clip / up to several reference images for identity. 
  • Wan 2.5 (wan-2.5): 1080p, up to 10s, native audio, single-shot — the cheaper daily driver, from $0.225/run.
  • Wan 2.2 (wan-2.2-plus): silent, 720p, ~5s — but open-weight, the pick when you must self-host.

Competitors list these models without a selection map. The short version: 2.6 for length + multi-shot + identity, 2.5 for cheaper single-shot clips with audio, 2.2 for open weights.

 

Wan 2.6 vs Wan 2.5 vs Wan 2.2

  Wan 2.6 Wan 2.5 Wan 2.2
Audio Native synced Native synced None (video only)
Max duration 15s 10s ~5s
Resolution up to 1080p up to 1080p 720p
Multi-shot / reference Yes No No
Open weights No (API only) No (API only) Yes (Apache 2.0)
GPTProto price $0.45–$2.025 / run $0.225–$1.35 / run $0.09 / run

Honest take: Wan 2.6 costs more per run, but it's the one to use when you need 15s, multi-shot scenes, or character consistency across cuts. Wan 2.5 is the cheaper daily driver for single-shot clips up to 10s with audio. Wan 2.2 is the pick only if you need open weights to self-host.

Related: Wan 2.5 → · Wan 2.2 → · Sora 2 →

 

What you can build with the Wan 2.6 API

Multi-shot short films and ads — One prompt lays out several shots with consistent characters and audio across cuts — a 15s narrative ad without manual stitching. A 20-variant test at 720p/5s ≈ $9.00.

Localized video at scale — Multilingual prompts and lip-synced speech turn one storyboard into several languages without re-shooting. Strong Chinese support.

Any aspect ratio, same price — Square (960×960, 1440×1440), portrait (720×1280, 1080×1920), and landscape are priced identically within a tier, so format is a free choice per platform.

Image / reference-to-video — Animate a still or guide a scene with a reference clip on the sibling endpoints. 

 

Wan 2.6 prompt recipes

Wan 2.6 generates audio and can plan multiple shots, so prompts work best when they describe shots, motion, and sound together. The platform's own example segments the prompt by time — use that for multi-shot. Structure each beat: shot + subject + action + camera + lighting + Sound: cue.

1. Multi-shot narrative (15s + shot planning)

[0-5s] Wide shot: a lone hiker reaches a misty mountain ridge at dawn, slow push-in.
[5-10s] Medium shot: she lifts her camera and smiles. Sound: wind, distant birds.
[10-15s] Close-up: the sun breaks over the peaks, light fills the frame. Sound: a
soft swell of ambient music, no dialogue.

2. Dialogue + lip-sync (single shot)

Medium close-up of a barista behind a wooden counter, warm morning light. She looks
to camera and says, "Your usual? Coming right up." Steam rises from the cup. Sound:
spoken line, espresso machine hiss, quiet cafe ambience.

3. Product/promo, square frame

Square frame. A sneaker rotates slowly on a matte pedestal, studio light sweeps
across it, subtle reflections. Sound: a low synth pulse, soft whoosh on the sweep,
no speech.

Working tips: for multi-shot, label each beat with its time range and keep dialogue short enough for the beat; set shot_type for single vs multi-shot; 8–12s is the sweet spot for coherence, 15s when you need length; use negative_prompt to exclude artifacts (e.g. "no text, no watermark"); prompt_extend: true lets the model expand a short prompt — turn it off for exact control; fix seed to reproduce a result.

 

如何取得 wan-2.6 API 金鑰

取得 wan-2.6 API 金鑰只需四個步驟,僅需幾分鐘。建立免費的 GPTProto 帳號、儲值、產生金鑰,並進行第一次呼叫 — 在 $0.45 這裡能以比直連更划算的價格取得 wan-2.6 API 金鑰,且單一金鑰即可在平台上所有模型通用。完整 wan-2.6 說明文件 請參閱說明文件。

註冊

註冊

建立您的免費 GPT Proto 帳號即可開始。您可以隨時為您的團隊設定組織。

儲值

儲值

您的餘額可用於平台上的所有模型,包括 wan-2.6,讓您能靈活地進行實驗並隨需求擴充。

產生您的 API 金鑰

產生您的 API 金鑰

在您的儀表板中建立 API 金鑰 — 進行 wan-2.6 請求時,您將需要它來進行驗證。

進行第一次 API 呼叫

進行第一次 API 呼叫

使用您的 API 金鑰搭配我們的範例程式碼,透過 GPT Proto 向 wan-2.6 發送請求,並立即查看 AI 生成的結果。

取得 API 金鑰

常見問題

關於 wan-2.6/text-to-video 模型的常見問題

什麼是 wan-2.6/text-to-video?

阿里巴巴的影片模型,可將文字提示轉換為最長 15 秒、最高 1080p 的影片片段,並在單次生成中提供同步音訊與多鏡頭場景規劃。

Wan 2.6 是開源的嗎?

否 — 僅提供 API,無法下載權重。Wan 2.1 和 2.2 是採用開放權重(Apache-2.0)的版本;2.5 和 2.6 僅提供 API。

Wan 2.6 與 Wan 2.5 有何不同?

2.6 將片段延長至 15 秒(2.5 上限為 10 秒),新增多鏡頭規劃與基於參考的身分保持功能,並提升動態逼真度。兩者皆支援最高 1080p 的原生同步音訊;2.5 每次執行的費用較低。

Wan 2.2 與 Wan 2.6 有何不同?

2.6 新增原生同步音訊、1080p、15 秒與多鏡頭功能;2.2 無音訊,解析度為 720p,片段約 5 秒。這兩者之中,只有 2.2 提供開放權重。

Wan 2.6 支援圖像轉影片嗎?

可以,使用圖像轉影片端點(讓靜態影像動起來),或使用參考轉影片功能(以影片片段作為引導)。

Wan 2.6 會生成音訊嗎?

可以 — 單次生成即可加入配合嘴型同步的語音、音效與音樂。您可以透過 audio 傳入自己的音軌,或使用 Sound: 提示直接描述聲音。

Wan 2.6 的費用是多少?

每次執行的費用從 0.45 美元(720p/5 秒)到 2.025 美元(1080p/15 秒)不等 — 請參閱上方的矩陣。

Wan 2.6 未經審查嗎?/允許 NSFW 內容嗎?

GPTProto 的實際內容政策 — 僅限於「未過濾 API 存取」合規層。

如何透過 API 呼叫?

向上方的端點發送 POST,然後輪詢 /predictions/{id}/result。請參閱快速入門。

如何在 GPT Proto 上支付 wan-2.6/text-to-video 的費用?

在 GPT Proto 上,wan-2.6/text-to-video 的付款選項包括按點數計費的使用方式、預付方案與企業訂閱。使用者註冊後取得模型存取權,並選擇符合預期輸出需求的方案。使用量會依影片時長或生成影片數量追蹤,發票則反映實際使用量。基於 API 的計費可整合熱門付款閘道,實現順暢的帳戶管理。GPT Proto 儀表板會清楚列出免費方案限制與進階附加功能。對於需要大規模影片製作的專案,也提供具備客製化 SLA 與支援服務的企業協議。

更多相關 GPTProto AI 工具

Picsart AI 影片生成器

使用先進的 AI 生成器將您的創意轉化為病毒式內容,讓您的 Picsart 工作流程更上一層樓。使用高級模型,實現流暢動畫效果。

Pollo AI 影片生成器

使用先進的 Pollo AI 文字轉影片引擎,將簡單的提示轉化為令人驚豔、吸引目光的視覺內容。

Pika Labs AI 影片生成器

運用 Pika Labs 的強大功能,將您的創意轉化為超寫實內容。我們的 AI 模型讓充滿活力的故事栩栩如生。

圖片轉影片 AI

使用我們尖端的圖片轉影片 AI,將靜態媒體轉化為動態序列。立即打造您專屬的線上照片轉影片製作器 API。

相關文章

更多部落格
Wan 2.5:4K AI 影片生成的未來

Wan 2.5:4K AI 影片生成的未來

探索 Wan 2.5 如何運用 AI 將圖片轉換成 4K 電影級影片。了解功能、定價與 GPT Proto API 選項。

wan2.2-animate:AI 的品質優先於速度

wan2.2-animate:AI 的品質優先於速度

當其他模型急於產出結果時,wan2.2-animate 專注於電影級視覺保真度與提示詞遵循度。立即了解如何最佳化您的影片工作流程。

Wan Video:掌握開源生成技術

Wan Video:掌握開源生成技術

Alibaba 的 wan video 是昂貴 AI 模型的開源替代方案。立即探索如何打造更好的提示詞,並將您的創意想法製作成動畫。

Wan 2.5:4K AI 影片生成的未來

Wan 2.5:4K AI 影片生成的未來

探索 Wan 2.5 如何運用 AI 將圖片轉換成 4K 電影級影片。了解功能、定價與 GPT Proto API 選項。

GPT Proto

以全球規模與穩定性,賦能 AI 創新:

透過我們的旗艦產品 GPT Proto,我們提供統一的介面,讓您能存取並整合全球頂尖 AI 供應商的 API,涵蓋文字、視覺、語音等領域。我們協助開發者與企業簡化整合流程,並無限制地加速創新。

全球基礎設施,在地合規:

To ensure enterprise-grade reliability and compliance, Talent Tech Global Limited operates specifically as our global Billing and Contracting Entity. Meanwhile, our core technical infrastructure and R&D teams are strategically distributed across global innovation hubs, including Silicon Valley, Singapore, and Hong Kong.

為擴展而生:

我們深知穩定性至關重要。我們的平台建立在強大的去中心化架構之上,支援動態自動擴展。無論您是進行試點專案還是處理數百萬次併發請求,我們的系統都能即時擴展以滿足需求,確保您的業務永遠不會受限於基礎設施。

導覽

  • 儀表板
  • 模型
  • 建立圖片
  • AI 圖片放大
  • AI 背景移除
  • 建立影片
  • 在畫布中編輯
  • 功能
  • 定價
  • AI 文件
  • AI 部落格
  • AI 洞察
  • AI 技能

功能

  • 動漫轉真人 AI
  • 動漫 AI 藝術生成器
  • AI 物件移除器
  • AI 圖片編輯器
  • 無限制 AI 圖像生成器
  • AI 動作轉移
  • AI 衣物移除器
  • AI 浮水印移除工具
  • 線上 AI 圖片增強器
  • 線上背景移除工具
  • AI 臉部交換圖片
  • AI 護照照片製作器
  • MS Paint AI 生成器
Explore all features >

LLM

  • Claude Opus 5
  • Gemini 3.6 Flash
  • Gemini 3.5 Flash Lite
  • Kimi K3
  • GPT 5.6 Luna
  • GPT 5.6 Terra
  • GPT 5.6 Sol
  • Grok 4.5
  • Claude Sonnet 5
  • Minimax M3
  • GLM 5.2
  • GPT 5.1 Chat Latest
  • Claude Fable 5
  • Qwen3.7 Max
  • Claude Opus 4.8 Thinking
  • Claude Opus 4.8
  • Gemini 3.5 Flash
  • DeepSeek v4 Flash
  • DeepSeek v4 Pro
  • Grok 4.3
探索所有模型 >

影像

  • Dola Seedream 5.0 Pro 260628
  • Gemini 3.1 Flash Lite Image
  • Gemini 3.1 Flash Image
  • GPT Image 2
  • Gemini 3.1 Flash Image Preview
  • Seedream 5.0 260128
  • Doubao Seedream 5.0 260128
  • Vidu Q2
  • Grok Imagine Image
  • Kling Image O1
  • GPT Image 1.5
  • Seedream 4.5 251128
  • Doubao Seedream 4.5 251128
  • Grok Imagine 0.9
  • Gemini 3 Pro Image Preview
  • Qwen Image Lora
  • Qwen Image Plus Lora
  • Qwen Image Plus
  • Grok 4 Image
  • GPT Image 1 Mini
探索所有模型 >

影片

  • Kling v3.0 4k
  • Dreamina Seedance 2.0 Mini 260615
  • Kling v3 Omni 4k
  • Dreamina Seedance 2.0 Fast 260128
  • Dreamina Seedance 2.0 260128
  • Vidu 2.0
  • Doubao Seedance 2.0 260128
  • Doubao Seedance 2.0 Fast 260128
  • Kling v3 Omni Pro
  • Kling v3 Omni Std
  • Vidu Q3 Turbo
  • Kling v3.0 Pro
  • Kling v3.0 Std
  • Vidu Q3 Pro
  • Kling v2.6 Std
  • Vidu Q2 Pro
  • Vidu Q2 Turbo
  • Vidu Q2 Pro Fast
  • Vidu Q2
  • Kling Video O1 Pro
探索所有模型 >

© 2026 Talent Tech Global Limited (Hong Kong) / Talent Tech Global LLC (US). 保留所有權利。

  • 關於我們
  • 隱私權政策
  • 服務條款
  • 網站地圖