Schuyler Stacy2026-07-28

Qwen 3.8 Max vs Kimi K3: 실제 코딩 작업에 준비된 모델은 무엇일까?

코딩, API 가격, 컨텍스트, 멀티모달 지원 및 오픈 가중치 측면에서 Qwen 3.8 Max와 Kimi K3를 비교하고, 어떤 모델이 배포할 준비가 되었는지 알아보세요.

Qwen 3.8 Max vs Kimi K3: 실제 코딩 작업에 준비된 모델은 무엇일까?

업데이트 — 2026년 7월 28일: Moonshot AI가 이제 전체 Kimi K3 가중치, 모델 카드, 커스텀 라이선스 및 기술 보고서를 공개했습니다. 이번 출시로 Kimi 측의 사용 가능 여부에 대한 의문은 해소되었습니다. 하지만 2.8T 파라미터 모델을 직접 호스팅하기 쉬워진 것은 아닙니다. 공식 저장소의 용량은 약 1.56TB이며, Moonshot은 64개 이상의 가속기를 사용하는 슈퍼노드 배포를 권장합니다.

 

Qwen 3.8 Max와 Kimi K3의 비교는 두 거대 중국 AI 모델 간의 단순한 대결처럼 보입니다. Alibaba’s의 2.4조 파라미터 프리뷰 모델과 Moonshot AI’s의 2.8조 파라미터 플래그십 모델을 비교하는 구도입니다. 이 수치만 보면 간단한 결론을 내리기 쉽습니다. 더 큰 모델이 승리할 것이라는 결론입니다.

하지만 현재 확인 가능한 증거는 이를 보여주지 않으며, 개발자에게 가장 유용한 비교 방식도 아닙니다.

2026년 7월 23일 기준으로 Qwen 3.8 Max는 여전히 Alibaba’s Token Plan을 통해 배포되는 변동 중인 프리뷰 모델입니다. Kimi K3는 이미 문서화된 API, 공개된 토큰 가격, 1M 토큰 컨텍스트 윈도우, 전체 가중치 공개를 위한 예정일을 갖추고 있습니다. 성능 격차는 좁을 수 있습니다. 하지만 제품 준비도 격차는 그렇지 않습니다.

제 판단은 명확합니다. 오늘 실제 애플리케이션을 구축하고 비용을 책정해야 한다면 Kimi K3가 더 안전한 선택입니다. Qwen 3.8 Max Preview는 코딩 워크플로에서 테스트해볼 가치가 있습니다. 특히 Alibaba’s의 프로모션 Credits 덕분에 실험 비용이 저렴한 동안에는 더욱 그렇습니다. 그러나 아직 프로덕션 의사결정을 이끌 만큼 충분한 안정적 정보를 제공하지는 못했습니다.

요약: 현재 프로덕션에 더 안전한 선택은 Kimi K3

지금 기존 방식의 API, 예측 가능한 토큰당 비용, 네이티브 이미지 및 동영상 이해 기능, 또는 고객 대상 제품에 바로 적용할 수 있는 모델이 필요하다면 Kimi K3를 선택하세요. 이미 Alibaba’s 코딩 생태계를 사용하고 있고 저렴한 프로모션 비용으로 유망한 새 모델을 테스트하고 싶다면 Qwen 3.8 Max Preview를 선택하세요.

발행 시점에 이용 가능한 유일한 상세 매칭 코딩 테스트에서 Kimi K3는 83점, Qwen 3.8 Max는 80점을 받았습니다. 이 3점 차이는 유용한 증거이지만 보편적인 순위는 아닙니다. 테스트에서 Qwen은 더 깔끔한 시스템 경계와 완벽한 도구 실행을 보여주었고, Kimi는 수정 이력과 재생성을 더 완전하게 처리했습니다. 두 모델 모두 사실에 근거하지 않은 추론을 내놓아 사실 확인과 수정이 필요했습니다.

쉽게 말하면 현재 배포 의사결정에서는 Kimi가 앞섭니다. Qwen이 성능 경쟁에서 진 것은 아닙니다. 단지 아직 승리했다고 선언하기에는 너무 이른 것입니다.

목차

Qwen 3.8 Max vs Kimi K3 at a Glance

Category Qwen 3.8 Max Kimi K3
Product status Stable production API Stable production API
Total parameters 2.4T 2.8T
Active parameters 95B 104B
Context window Up to 1M tokens 1,048,576 tokens
Inputs Text, images, and video Text, images, and video
Official API price $2/M input, $6/M output $3/M input, $15/M output
GPT Proto access Available now Available now
Open weights Announced; not yet downloadable Released
Best hosted-API fit Cost-sensitive multimodal coding and long agents Revision-heavy agents and Kimi-specific workflows
Best self-hosting fit Wait for the checkpoint and license Kimi K3, with data-center-scale hardware

The Comparison Is Now More Equal—but Deployment Still Differs

Qwen3.8-Max and Kimi K3 are now both viable production API models. The main difference is no longer “Preview versus production.” It is now hosted API value versus immediate open-weight ownership.

Qwen 3.8 Max Is Now a Stable Production API

Alibaba released the stable qwen3.8-max model on August 3, 2026, replacing the earlier Preview-era positioning with normal pay-as-you-go API access.

The production release documents a 2.4-trillion-parameter Sparse Mixture-of-Experts architecture with approximately 95 billion active parameters per request. It supports up to a 1-million-token context window, up to 128K output tokens, and text, image, and video input.

This changes the practical comparison. Qwen3.8-Max can now be evaluated for customer-facing applications, coding agents, multimodal analysis, and other production workloads without relying on the earlier Credits-based Personal Token Plan.

Its official price is also lower than Kimi K3’s:

Model Input Price Output Price
Qwen3.8-Max $2 per 1M tokens $6 per 1M tokens
Kimi K3 $3 per 1M tokens $15 per 1M tokens

For teams using a hosted API, Qwen3.8-Max on GPT Proto is now a serious production option rather than an experimental endpoint.

However, Alibaba’s announced open-weight checkpoint has not yet been released. Developers should not assume that the final weights, license, or self-hosting terms are available until they are officially published.

Kimi K3 Has Released Downloadable Weights

Kimi K3 is available through both a hosted API and a downloadable checkpoint. Moonshot AI has published the model weights, model card, technical report, deployment guidance, and custom license.

Kimi K3 has 2.8 trillion total parameters and activates approximately 104 billion parameters per token. Its deployment documentation covers frameworks including vLLM, SGLang, and TokenSpeed, giving teams a clearer path to controlled or private infrastructure.

That makes Kimi K3 the more practical choice when downloadable weights, deployment ownership, or immediate self-hosting is a firm requirement.

Open weights do not make Kimi K3 easy or inexpensive to run locally. It remains a multi-trillion-parameter model that requires data-center-scale storage, memory, networking, and accelerator capacity. Its custom license may also impose conditions on very large commercial products or Model-as-a-Service deployments.

What the Deployment Difference Means

Deployment Need Better Starting Point Why
Hosted production API Qwen 3.8 Max Stable API with substantially lower official output-token pricing
Multimodal coding and visual analysis Qwen 3.8 Max Native text, image, and video input
Downloadable weights today Kimi K3 Its checkpoint and license are already public
Private or controlled deployment Kimi K3 Teams can deploy the released model on their own infrastructure
Easy local installation Neither Both models require serious infrastructure to self-host
Head-to-head API testing Test both Compare accepted tasks, retries, tool failures, latency, and total cost

The comparison is therefore no longer unequal because Qwen is “only a Preview.” Both models can serve production API workloads. The remaining difference is simpler: Qwen3.8-Max currently offers the stronger hosted cost-and-capability proposition, while Kimi K3 offers immediate access to released weights and greater deployment control.

How Developers Should Test Qwen 3.8 Max Against Kimi K3

Test Give Both Models Measure
Repository architecture review The same frozen commit, architecture question, read-only tools, and time limit Correct file citations, missed dependencies, unsupported claims, and review time
Multi-file implementation The same issue, tests, writable files, and tool permissions Tests passed, files changed, retries, regressions, and human corrections
Visual frontend repair The same screenshot, source files, browser tools, and target behavior Visual match, valid code, repair loops, and final test result

Keep the agent shell and permissions identical. Set a fixed time limit. Record the full model ID and date, especially for Qwen’s moving Preview. Then capture task completion, wall-clock time, input and output tokens, cache hits, failed tool calls, retries, human interventions, and final tests passed.

Do not score an answer because it “looks thorough.” Check whether the patch works and whether the model’s claims survive review.

For high-value architecture decisions, there is another useful pattern: run both models independently, hide their identities during review, and compare their disagreements. The 269-file test produced its strongest design only after combining Qwen’s system boundaries with Kimi’s lifecycle model. Sometimes the right answer to Qwen versus Kimi is both, followed by verification.

Qwen 3.8 Max vs Kimi K3 Pricing

Model Official Input Price Official Output Price
Qwen3.8-Max $2 per 1M tokens $6 per 1M tokens
Kimi K3 $3 per 1M tokens $15 per 1M tokens
Kimi K3 on GPT Proto $2.70 per 1M tokens $13.50 per 1M tokens

At official list price, Qwen’s output tokens cost 60% less than Kimi K3’s. That difference matters for reasoning-heavy coding agents that produce long plans, tool traces, explanations, and patches.

Token price is not the entire cost. Measure retries, failed tool calls, human corrections, latency, and accepted task completion. Kimi can still be cheaper on a specific workflow if its stronger revision and lifecycle handling prevents expensive repair loops.

Check the live Qwen3.8-Max API page for GPT Proto’s current price before calculating a production budget.

Which Model Should You Choose?

Situation Better Choice Why
Hosted production API Qwen 3.8 Max Stable access and substantially lower official output price
Complex multimodal coding Qwen 3.8 Max Text, image, and video input with strong frontend and visual-agent positioning
Architecture boundaries and tool discipline Test Qwen first The Preview completed 44 of 44 tool calls in the matched test
Revision and regeneration history Test Kimi first Kimi handled lifecycle state more completely in the matched test
Downloadable weights today Kimi K3 Full checkpoint and license are already public
Lowest self-hosting complexity Neither Both are multi-trillion-parameter models requiring serious infrastructure
One account for head-to-head testing Both through GPT Proto Run the same task, tools, reasoning settings, and evaluation criteria

Multimodal Inputs, Reasoning, and Agent Behavior

Capability Qwen 3.8 Max Preview Kimi K3
Image understanding Documented Documented
Video understanding Not clearly documented as a current model input Documented
Thinking mode Always on Always on
Reasoning levels low, high, xhigh low, high, max
Default reasoning level xhigh max
Context window Not clearly disclosed on the current product page 1,048,576 tokens
Structured output Preview capabilities require continued verification JSON mode and strict JSON Schema documented
Long tool histories Behavior may change with the Preview Full assistant messages, including reasoning content, should be preserved

Qwen’s documented vision support makes it relevant for screenshot-based debugging. Kimi goes further by accepting video, which is useful when the input is a screen recording, animation reference, or product demo that would be difficult to describe frame by frame.

Kimi’s richer documented interface also adds integration work. Its Preserved Thinking behavior means a multi-turn application should return the complete assistant message, including reasoning content and tool calls, rather than keeping only the visible answer. That history occupies context and is billed. The 1M-token window is large, but it is not free storage.

Both models default to their highest reasoning setting. For evaluation, keep those settings consistent. For production, test lower effort on simpler tasks. A model that solves 99% of requests at lower effort may be cheaper and faster than one left at maximum reasoning for every autocomplete, classification, or short transformation.

Which Model Should You Choose?

Your situation Better current choice Why
Building a customer-facing application now Kimi K3 Conventional API and forecastable token price
Running a long repository task with image or video input Kimi K3 1M context and documented image/video support
Testing inside Qwen Code, Qoder, or another supported Alibaba tool Qwen 3.8 Max Preview Low promotional Credits consumption
Need stable repeatable benchmarks Kimi K3, for now Qwen’s Preview may change between runs
Need the cleanest architecture and replay metadata Test Qwen It showed a real advantage in the matched architecture review
Need strong revision and regeneration handling Test Kimi It was more complete in the available matched test
Need downloadable weights today Kimi K3 The full checkpoint is public; Qwen 3.8 Max still has no released weights
Need to compare several model families behind one account Kimi K3 on GPT Proto One key and balance can access the broader model catalog

For an application going live this week, I would choose Kimi K3. It has enough published information to estimate cost, define integration behavior, and repeat a test against a stable model name.

For an internal coding experiment, I would not ignore Qwen. Its Preview handled a long repository analysis with zero failed tool calls and showed better architectural boundaries than Kimi in the matched test. That is a serious capability signal. It is not yet a production contract.

How to Try Kimi K3 Through GPT Proto

Qwen 3.8 Max is not yet available on GPT Proto, so the current integration example uses Kimi K3 only. You can call it through GPT Proto’s OpenAI-compatible Chat Completions endpoint with the kimi-k3 model string:

curl https://gptproto.com/v1/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: $GPTPROTO_API_KEY" \
  -d '{
    "model": "kimi-k3",
    "messages": [
      {
        "role": "user",
        "content": "Review this migration plan. Identify unsupported assumptions, missing rollback steps, and the tests required before production."
      }
    ]
  }'

The endpoint and authorization structure follow the GPT Proto API quickstart. For multi-turn Kimi workflows, preserve the complete assistant message returned by the API, including reasoning and tool-call fields, rather than storing only the final visible text.

You can try the Kimi K3 API, browse 200+ AI models, or use the GPT Proto unified AI API to compare Kimi with other text, image, video, and audio models. When Qwen 3.8 Max is added, the useful test will be the same task, prompt, tool permissions, and scoring method on both model endpoints.

Final Verdict

Kimi K3 no longer wins simply because Qwen is a Preview. Qwen3.8-Max now has a stable API, normal token billing, documented specifications, and direct availability through GPT Proto.

For most teams choosing a hosted model, Qwen3.8-Max is the stronger starting point because its official $2/$6 rate is far below Kimi K3’s $3/$15 rate, especially on output-heavy agent tasks.

Kimi K3 remains the better choice when downloadable weights and deployment ownership are non-negotiable. Its stronger revision and regeneration handling in the available matched test also makes it worth evaluating for state-heavy creative and engineering workflows.

The clean answer is now: Qwen for hosted cost and broad multimodal capability; Kimi for released weights and lifecycle-heavy tasks.

하나의 키로 더 많은 AI 모델

하나의 OpenAI 호환 API를 통해 주요 AI 모델을 합리적인 가격으로 이용해 보세요.

API 모델 둘러보기
하나의 키로 더 많은 AI 모델
관련 모델
모든 모델
MoonshotAI
10% OFF
Qwen
by Qwen
10% OFF
MiniMax
30% OFF
DeepSeek

자주 묻는 질문

코딩에서 Qwen 3.8 Max가 Kimi K3보다 우수한가요?

Qwen 3.8 Max가 전반적으로 더 우수하다고 말할 충분한 증거는 없습니다. 사용 가능한 269개 파일 아키텍처 테스트에서 Kimi K3는 83점, Qwen은 80점을 받았습니다. Qwen은 더 깔끔한 시스템 경계를 만들었고 44회의 도구 호출을 모두 성공적으로 수행했으며, Kimi는 수정과 재생성을 더 완전하게 처리했습니다. 선택하기 전에 자체 구현 및 수정 작업에서 두 모델을 모두 테스트하세요.

Qwen 3.8 Max와 Kimi K3 중 어느 쪽이 더 저렴한가요?

현재 Alibaba의 Token Plan을 통한 프로모션 실험에서는 Qwen이 더 저렴하지만, Credits가 공개된 100만 토큰당 가격으로 환산되지는 않습니다. Kimi는 프로덕션 비용을 책정하기가 더 쉽습니다. GPTProto는 현재 Kimi K3의 입력 토큰 100만 개당 가격을 $2.70, 출력 토큰 100만 개당 가격을 $13.50으로 표시하고 있습니다.

Qwen 3.8 Max와 Kimi K3는 오픈 소스인가요?

Kimi K3는 커스텀 Kimi K3 License에 따라 현재 오픈 가중치로 제공됩니다. 이 라이선스는 광범위한 사용과 수정을 허용하지만 대규모 Model-as-a-Service 사업 및 초대형 상용 제품에 대한 조건을 포함하므로 MIT 라이선스라고 설명해서는 안 됩니다. Qwen 3.8 Max는 오픈 가중치를 발표했지만 체크포인트나 라이선스를 아직 출시하지 않았습니다.

Qwen 3.8 Max를 프로덕션 API에서 사용할 수 있나요?

현재 Qwen 3.8 Max Preview는 Alibaba의 Token Plan을 통해 사용할 수 있지만 Personal Plan 약관은 자동화 스크립트, 커스텀 애플리케이션 백엔드 또는 비대화형 일괄 호출에 키를 사용하는 것을 금지합니다. 고객 대상 백엔드를 설계하기 전에 문서화된 프로덕션 경로와 일반적인 상업 가격이 제공될 때까지 기다리세요.

긴 코딩 작업에는 어떤 모델이 더 우수한가요?

Kimi K3가 현재 더 안전한 선택입니다. 문서화된 1M 토큰 컨텍스트 윈도우, 기존 방식의 API 및 사용 가능한 매칭 테스트에서 더 높은 결과를 갖추고 있기 때문입니다. 그래도 Qwen은 테스트할 가치가 있습니다. 동일한 60분 저장소 작업을 실패한 도구 호출 없이 완료했습니다. 긴 컨텍스트 용량만으로 작업 완료가 보장되지는 않으므로 통과한 테스트, 재시도 및 사람의 수정을 측정하세요.

Kimi K3는 이미지와 동영상을 이해할 수 있나요?

예. Moonshot은 Kimi K3의 네이티브 텍스트·이미지·동영상 이해를 문서화하고 있습니다. 모델은 텍스트를 반환하므로 스크린샷 기반 디버깅, 동영상 분석, 인터페이스 검토 및 기타 멀티모달 코딩 또는 지식 작업에 적합합니다.

Qwen 3.8 Max 또는 Kimi K3를 로컬에서 실행할 수 있나요?

Kimi K3 가중치는 다운로드할 수 있지만, 실제 셀프 호스팅에는 데이터센터 인프라가 필요합니다. 공식 저장소의 용량은 약 1.56TB이며 Moonshot은 64개 이상의 가속기를 권장합니다. Alibaba가 가중치를 출시하기 전까지는 Qwen 3.8 Max를 셀프 호스팅할 수 없습니다.
Kimi K3란 무엇이며, 정말 GPT-5.6 및 Fable 5에 가까운가?

Kimi K3란 무엇이며, 정말 GPT-5.6 및 Fable 5에 가까운가?

TL;DR Kimi K3는 장기 코딩, 지식 작업, 추론 및 에이전트 워크플로를 위해 Moonshot AI가 개발한 2.8조 파라미터 규모의 멀티모달 모델입니다. 독립적인 테스트에서 전반적으로 Claude Opus 4.8 및 GPT-5.5에 근접한 결과를 보였지만, GPT-5.6 Sol과 Claude Fable 5가 여전히 앞서 있습니다. K3는 에이전트 벤치마크에서 격차를 좁혔고 일부 자동화 테스트에서는 선두를 차지했지만, 측정된 환각률은 K2.6보다 증가했습니다. Kimi K3는 이제 오픈 웨이트 모델입니다. Moonshot AI는 전체 체크포인트, 모델 카드, 기술 보고서 및 자체 Kimi K3 라이선스를 공개했습니다. 공식 Hugging Face 저장소는 96개의 safetensors 샤드로 구성된 약 1.56TB 규모이며, Moonshot은 64개 이상의 가속기를 사용하는 슈퍼노드 배포를 권장합니다. 오픈 웨이트 공개로 소유권에 관한 문제는 해결되었습니다. 하지만 K3가 일반적인 로컬 모델이 되는 것은 아닙니다. 대부분의 개발자에게 호스팅 API는 여전히 실용적인 출발점입니다. GPTProto의 Kimi K3 API 는 현재 입력 토큰 100만 개당 2.70달러, 출력 토큰 100만 개당 13.50달러로 책정되어 있습니다. 데이터 제어, 맞춤형 추론 또는 모델 수정이 인프라 및 라이선스 검토 비용을 감수할 만큼 가치 있다면 웨이트를 선택하세요. 요약하면 Kimi K3는 GPT-5.6 및 Fable 5와 같은 논의의 장에 포함될 만큼 충분히 근접했으며—이제 오픈 웨이트 출시를 통해 두 폐쇄형 모델에는 없는 배포 선택지를 개발자에게 제공합니다.

Michael Johnson | 2026-07-28

코딩을 위한 GLM-5.2 vs Kimi K3: 2026년 개발자에게 더 나은 모델은?

코딩을 위한 GLM-5.2 vs Kimi K3: 2026년 개발자에게 더 나은 모델은?

TL;DR: 어렵고 장시간 실행되거나 시각적 요소가 필요한 작업에서는 Kimi K3가 더 강력한 코딩 모델입니다. Moonshot이 공개한 코딩 비교에서 GLM-5.2를 앞서며, 호스팅 서비스에서 이미지와 동영상도 입력으로 받을 수 있습니다. GLM-5.2는 일상적인 저장소 작업의 기본값으로는 여전히 더 낫습니다. 비용이 훨씬 저렴하고 운영하기 쉬우며, 허용 범위가 넓은 MIT 라이선스를 사용하기 때문입니다. Kimi K3도 이제 가중치를 공개했지만, 1.56TB 규모의 저장소, 64개 이상의 가속기를 권장하는 배포 환경, 맞춤형 라이선스로 인해 자체 호스팅에는 훨씬 더 큰 투자가 필요합니다. 역량이 병목이면 Kimi를, 비용과 운영 단순성이 매일 중요하면 GLM을 선택하세요. GLM-5.2와 Kimi K3 Code 비교에서 흥미로운 점은 두 모델 모두 React 컴포넌트를 작성하거나 짧은 알고리즘을 해결할 수 있다는 사실이 아닙니다. 이 수준의 모델은 이미 그 기준을 충족합니다. 중요한 질문은 과제가 복잡해졌을 때 어떤 일이 발생하는가입니다. 저장소 감사, 여러 파일에 걸친 마이그레이션, 스크린샷에서만 나타나는 버그, 또는 여러 시스템의 일관성을 유지해야 하는 실행 가능한 Three.js 프로토타입 같은 작업 말입니다. 가격 차이가 중요해지기 시작하는 지점도 바로 여기입니다. Kimi K3는 가장 어려운 공개 테스트에서 더 나은 성능을 보이지만, 공식 출력 가격은 GLM-5.2보다 세 배 이상 비쌉니다. 수천 건의 일반적인 리뷰를 처리하는 팀이라면 GLM을 사용할 때 달러당 더 많은 작업을 수행할 수 있습니다. 반면 하나의 까다로운 시각적 프로젝트를 해결하려는 개발자라면 K3에 기꺼이 비용을 지불할 수 있습니다.

Tiffany Layne | 2026-07-28

Kimi K3 대 GPT-5.6 Sol: 더 저렴한 토큰인가, 더 저렴한 작업인가?

Kimi K3 대 GPT-5.6 Sol: 더 저렴한 토큰인가, 더 저렴한 작업인가?

TL;DR 업데이트 — 2026년 7월 28일 : 이제 Kimi K3의 전체 가중치를 공개적으로 사용할 수 있습니다. Moonshot AI는 공식 저장소에 2.8T 체크포인트, 기술 보고서, Kimi K3 라이선스를 공개했습니다. 이번 공개로 GPT-5.6 Sol에 대한 K3의 제어권 및 배포 측면의 경쟁력이 강화되었지만, 독립 벤치마크 결과가 바뀌거나 K3를 직접 운영하는 비용이 저렴해진 것은 아닙니다. Kimi K3는 토큰당 비용이 더 저렴합니다. GPT-5.6 Sol은 중요도가 높은 프로덕션 에이전트의 기본 선택으로 더 강력합니다. 두 명제는 모두 참일 수 있습니다. 격차는 가격표가 보여 주는 것보다 작습니다. Artificial Analysis 테스트에서 GPT-5.6 Sol max의 Intelligence Index 점수는 59점으로, Kimi K3의 57점보다 높습니다. 그러나 측정된 작업당 비용은 Sol이 약 $1.04, K3가 $0.95로, 공식 출력 가격이 암시하는 2배의 격차와는 다릅니다. 짧게 답하면 다음과 같습니다. 폭넓은 안정성, 코딩 에이전트 성능, OpenAI의 호스팅 도구 스택이 가장 중요하다면 GPT-5.6 Sol 을 선택하세요. 비디오 입력, 긴 컨텍스트 작업, 더 낮은 정가 또는 공개된 오픈 가중치에 대한 접근성이 결정에 영향을 준다면 Kimi K3 를 선택하세요.

Schuyler Stacy | 2026-07-28

Qwen 3.8 Max란 무엇인가? 출시, 사양, 가격 및 오픈 가중치

Qwen 3.8 Max란 무엇인가? 출시, 사양, 가격 및 오픈 가중치

2026년 8월 7일 업데이트: 알리바바가 8월 3일 Qwen3.8-Max 프로덕션 버전을 공식 출시하면서 기존 qwen3.8-max-preview 를 현재 플래그십 API 모델로 대체했습니다. 이 글은 확정된 아키텍처, 컨텍스트 창, 가격, API 이용 가능 여부 및 오픈 가중치 일정을 반영해 업데이트되었습니다. Qwen3.8-Max는 현재까지 알리바바의 가장 강력한 Qwen 모델로, 요청당 총 2조 4천억 개의 파라미터와 950억 개의 활성 파라미터를 갖춘 멀티모달 Sparse Mixture-of-Experts 모델입니다. 안정화 버전은 텍스트, 이미지, 비디오 입력을 지원하고 텍스트 출력을 생성하며, 최대 128K 출력 토큰과 함께 최대 100만 토큰의 컨텍스트 창을 제공합니다. 복잡한 코딩, 시각 분석, 리서치, 전문 작업 및 장기 에이전트 작업을 위해 설계되었습니다. 이번 출시는 개발자에게 실질적인 답도 바꿉니다. Qwen3.8-Max는 더 이상 변경되는 프리뷰나 크레딧 전용 개인 플랜에 국한되지 않습니다. 이제 일반적인 종량제 API를 제공하며, 알리바바는 입력 토큰 100만 개당 2달러, 출력 토큰 100만 개당 6달러로 책정했습니다. 또한 GPTProto의 Qwen 3.8 Max API 를 통해서도 사용할 수 있습니다. 간단히 말하면, Qwen 3.8 Max는 완성된 제품 출시라기보다 실제적이고 이례적으로 흥미로운 프리뷰입니다. 개발자는 이를 테스트하고 모든 결과의 날짜를 기록하고, 알리바바가 아직 공개하지 않은 사양을 기준으로 프로덕션 마이그레이션을 계획하지 않는 것이 좋습니다.

Tiffany Layne | 2026-07-23