Compare to

Discover how DeepSeek's DeepSeek V4 Pro and Alibaba's Qwen3.8-Max stack up against each other in this comprehensive comparison of two leading AI language models.

Explore their capabilities, pricing, and performance metrics, with DeepSeek V4 Pro achieving 90.1% on GPQA Diamond and Qwen3.8-Max scoring 92.6%, making this comparison essential for developers and organizations seeking the right AI solution for their specific needs.

Models Overview

DeepSeek DeepSeek V4 Pro
Qwen3.8-Max

Provider

The company that provides the model.
DeepSeekAlibaba

Context Length

Maximum number of tokens the model can process
1M1M

Maximum Output

Maximum number of tokens the model can generate in one response
384K131.07K

Release Date

When the model was first released.
Unknown02-09-2026

Knowledge Cutoff

When the model's training data ends.
UnknownUnknown

Open Source

Whether the model weights are openly available.
TRUETRUE

Pricing Comparison

Compare the pricing of DeepSeek's DeepSeek V4 Pro and Alibaba's Qwen3.8-Max to determine the most cost-effective solution for your AI needs. Prices are the standard API tier per million tokens, as published by each provider as of September 2026.

DeepSeek DeepSeek V4 Pro
Qwen3.8-Max

Input Cost

Cost per million input tokens
$1.32 / 1M tokens$2 / 1M tokens

Output Cost

Cost per million tokens generated
$3.96 / 1M tokens$6 / 1M tokens

Comparing Benchmarks and Performance

Compare the performances of DeepSeek's DeepSeek V4 Pro and Alibaba's Qwen3.8-Max on industry benchmarks. Scores are the ones the providers and public leaderboards report; a benchmark neither reports is left out.

DeepSeek DeepSeek V4 Pro
Qwen3.8-Max

LMArena Elo

Crowd-sourced blind preference rating on the LMArena text leaderboard.
Benchmark not available1,481

GPQA Diamond

Graduate-level science questions written to be search-proof.
90.1%92.6%

SWE-bench Verified

Resolving real GitHub issues end to end.
80.6%Benchmark not available

SWE-bench Pro

Harder, contamination-resistant successor of SWE-bench Verified; not comparable with it.
Benchmark not available67.7%

MMLU-Pro

Broad knowledge and reasoning across 14 subjects, harder successor of MMLU.
87.5%Benchmark not available

HumanEval

Functional correctness of generated code.
76.8%Benchmark not available

Sources — DeepSeek V4 Pro: api-docs.deepseek.com, huggingface.co; Qwen3.8-Max: alibabacloud.com, huggingface.co, arena.ai.

Compare More Models