Compare to

Discover how Google's Gemini 3.8 Flash and Alibaba's Qwen3.8-Max stack up against each other in this comprehensive comparison of two leading AI language models. Released in September 2026 and September 2026 respectively, these models represent significant advancements in artificial intelligence, with Gemini 3.8 Flash offering a 1,048,576-token context window and Qwen3.8-Max offering a 1,000,000-token context window.

Explore their capabilities, pricing, and performance metrics, with Gemini 3.8 Flash achieving 1,493 on LMArena Elo and Qwen3.8-Max scoring 1,481, making this comparison essential for developers and organizations seeking the right AI solution for their specific needs.

Models Overview

Google Gemini 3.8 Flash
Qwen3.8-Max

Provider

The company that provides the model.
GoogleAlibaba

Context Length

Maximum number of tokens the model can process
1.05M1M

Maximum Output

Maximum number of tokens the model can generate in one response
65.54K131.07K

Release Date

When the model was first released.
02-09-202602-09-2026

Knowledge Cutoff

When the model's training data ends.
UnknownUnknown

Open Source

Whether the model weights are openly available.
FALSETRUE

Pricing Comparison

Compare the pricing of Google's Gemini 3.8 Flash and Alibaba's Qwen3.8-Max to determine the most cost-effective solution for your AI needs. Prices are the standard API tier per million tokens, as published by each provider as of September 2026.

Google Gemini 3.8 Flash
Qwen3.8-Max

Input Cost

Cost per million input tokens
$0.75 / 1M tokens$2 / 1M tokens

Output Cost

Cost per million tokens generated
$3.75 / 1M tokens$6 / 1M tokens

Comparing Benchmarks and Performance

Compare the performances of Google's Gemini 3.8 Flash and Alibaba's Qwen3.8-Max on industry benchmarks. Scores are the ones the providers and public leaderboards report; a benchmark neither reports is left out.

Google Gemini 3.8 Flash
Qwen3.8-Max

LMArena Elo

Crowd-sourced blind preference rating on the LMArena text leaderboard.
1,4931,481

GPQA Diamond

Graduate-level science questions written to be search-proof.
Benchmark not available92.6%

SWE-bench Pro

Harder, contamination-resistant successor of SWE-bench Verified; not comparable with it.
Benchmark not available67.7%

Terminal-Bench 2.1

Agentic tasks completed in a real terminal.
89.4%Benchmark not available

Sources — Gemini 3.8 Flash: ai.google.dev, blog.google, arena.ai, deepmind.google; Qwen3.8-Max: alibabacloud.com, huggingface.co, arena.ai.

Compare More Models