📊 Full opportunity report: Qwen3.8-Max’s AI Capabilities Unveiled: The Numbers That Matter on ThorstenMeyerAI.com — validation score, market gap, and execution plan.

TL;DR

Alibaba announced the full specifications of its flagship AI model, Qwen3.8-Max, confirming a 2.4 trillion-parameter size and strong benchmark results. The open weights will be available next week, marking a significant step in open AI models.

Alibaba has officially announced the detailed specifications of its flagship AI model, Qwen3.8-Max, confirming it has 2.4 trillion total parameters and demonstrating strong benchmark performance. This marks a significant milestone in large-scale open AI models, with open weights scheduled for release next week.

On August 3, Alibaba published the full benchmark table for Qwen3.8-Max, revealing its architecture built on Qwen3.5 with a 95 billion active parameter count per query, utilizing sparse mixture-of-experts technology. The model supports multimodal inputs — text, images, and videos — with text output. It achieved top scores on several benchmarks, including Terminal-Bench 2.1 (86.6), outperforming models like Claude Opus 4.8 and Claude Fable 5, while trailing only GPT-5.6 Sol at 88.8.

Alibaba demonstrated the model’s capabilities by reproducing research results and outperforming its own previous models on tasks like AIME24, showcasing its long-horizon reasoning abilities. However, it scored lower on software engineering benchmarks such as SWE-bench Pro (67.7), indicating limitations in some areas. The open weights for the 2.4 trillion-parameter model will be released next week, though they are expected to be a multi-node artifact due to their size. A smaller, 27B checkpoint, Qwen3.8-27B, designed for local deployment, will also be available, suitable for high-memory single machines.

At a glance
updateWhen: announced August 3, 2023; full details…
The developmentAlibaba officially released detailed specs and benchmark results for Qwen3.8-Max, confirming its 2.4 trillion parameters and upcoming open weights.
AI DISPATCH · REALITY CHECK Released 3 Aug 2026
Alibaba’s Qwen3.8-Max leaves preview
Second Only to Fable 5?

For fifteen days the claim ran without a benchmark table. Today Alibaba published the table, the active-parameter count, and a weights timeline. The numbers are genuinely strong on the rows Alibaba chose — and twelve to fifteen points behind on the rows it didn’t.

▲ All performance figures: Alibaba’s own harness
2.4T / 95B
Total / active parameters (MoE)
~1M
Context window · 131K max output
Text+Img+Video
Multimodal in · text out
„Next week“
Open weights · licence unpublished
01
Fifteen days from slogan to spec sheet

The claim shipped on a Sunday. The evidence shipped two weeks later. In between, the claim did its work.

17 Jul
Moonshot releases Kimi K3
2.8T parameters; rattles US tech stocks, later suspends new subscriptions under demand.
18 Jul
„kaleb“ appears on Code Arena
Anonymous model introduces itself as „Claude“ — a distillation artifact — and is identified within a day by a Qwen tokenizer quirk.
19 Jul
WAIC preview: „second only to Fable 5“
No benchmark table, no model card, no licence, no active-parameter count. Paid preview at 10% of standard pricing.
20 Jul
Shares rise as much as 5.4%
The market prices the claim, not the table.
3 Aug
General availability + full benchmark table
95B active confirmed; 2.4T weights and a Qwen3.8-27B checkpoint promised for next week. Licence still unwritten.
02
The table, both halves

„Second only to Fable 5“ is true on the rows Alibaba chose and false on the rows it didn’t. Both halves below are from the same release.

Where it leads
Terminal-Bench 2.1 · agentic terminal work
Qwen3.8-Max
86.6
GPT-5.6 Sol
88.8
Fable 5
84.6
OSWorld-Verified · computer use — plus PaperBench 93.0, CAD Bench 91.5
Qwen3.8-Max
86.1
Where it trails — the rows the slogan skips
SWE-bench Pro · deep software engineering
Qwen3.8-Max
67.7
Fable 5
80.0
FrontierSWE · frontier coding agents
Qwen3.8-Max
73.5
Fable 5
88.8
The real jump: one generation of agentic gains vs Qwen3.7-Max
DeepSWE 1.1
21.6 → 56.6
FrontierSWE
40.7 → 73.5
JobBench
31.3 → 53.4
03
Three artifacts, three different facts

„Qwen3.8 is going open-weight“ describes three things with very different deployment realities.

Hosted API
Live today

OpenAI- and DashScope-compatible — a base-URL change to A/B against your current backend.

2.4T weights
„Next week“ · no licence yet

A multi-node datacenter artifact. At 95B active, no single machine serves it. A flag planted, not a deployment option.

Qwen3.8-27B
Announced · no benchmarks yet

The checkpoint that fits real hardware. Whether the agentic gains survive distillation is the question that decides whether next week matters.

04
Bull and bear

Three Chinese frontier releases in seventeen days, each measured against the same export-controlled model. The contest is real; it is not the same thing as your workload.

Bull
  • The generation jump is real and consistent across a dozen agentic rows, with a stated mechanism: RL-environment scaling.
  • More disclosure than Kimi K3 shipped — full table, active-parameter count, weights timeline.
  • If 2.4T lands under a permissive licence, the ceiling of „open weight“ moves permanently.
  • The 27B sibling could become the best local agent model on hardware people already own.
Bear
  • Every number is Alibaba’s harness. Independent testing already tempered Kimi K3’s launch claims substantially.
  • The paying use case still belongs to Fable 5 — twelve to fifteen points on deep software engineering.
  • „Next week“ comes from a company that sat on a finished benchmark table for fifteen days.
  • Until the licence text exists, „going open-weight“ is a press strategy, not a property of the model.
The claim ran for fifteen days without evidence. Now the evidence exists —
and it says „second only“ depends entirely on which row you read.

Implications of Alibaba’s Largest Open-Weight Model

The announcement confirms Alibaba’s position as a major player in large-scale AI development, with the largest open-weight model ever shipped if the 2.4 trillion-parameter weights are released. This could accelerate research and deployment by enabling wider access to powerful models, particularly through the upcoming open weights. The release also demonstrates Alibaba’s focus on agentic reasoning and multimodal capabilities, which are crucial for advanced AI applications.

For the AI community, the detailed benchmark results and open weights represent a significant step toward more accessible and transparent large models. However, the substantial size of the model indicates that only large-scale data centers will host it, limiting immediate local deployment. The smaller 27B model offers a practical alternative for local use, but its performance relative to the flagship remains to be seen.

Local LLM Inference Optimization: A Comprehensive Guide to Quantization, Hardware Acceleration, and Efficient Private AI Deployment

Local LLM Inference Optimization: A Comprehensive Guide to Quantization, Hardware Acceleration, and Efficient Private AI Deployment

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Alibaba’s AI Model Development Timeline

Alibaba’s AI model journey has been marked by strategic previews and stealth releases, culminating in the full reveal of Qwen3.8-Max on August 3. Prior to this, models like Kimi K3 and the anonymous kaleb surfaced briefly, hinting at Alibaba’s ongoing efforts. The company’s approach involved staged announcements, starting with a stealth preview in July and culminating in a detailed spec release, aligning with industry practices for high-profile AI launches.

The model’s architecture builds on the Qwen3.5 foundation, with innovations in sparse mixture-of-experts and multimodal processing. Benchmarking on proprietary and public tests highlights its competitive performance, particularly in agentic reasoning tasks, though some software engineering benchmarks reveal gaps. The upcoming open weights will enable broader testing and deployment, especially for the 27B variant designed for local use.

"We are committed to providing open access to our largest models, with full weights available next week for research and development."

— Alibaba spokesperson

ASRock Intel Arc Pro B70 Creator 32GB Workstation Graphics Card, Xe2-HPG, 32GB GDDR6, PCIe 5.0, 4X DP 2.1, Blower Fan, Vapor Chamber, Honeywell PTM7950

ASRock Intel Arc Pro B70 Creator 32GB Workstation Graphics Card, Xe2-HPG, 32GB GDDR6, PCIe 5.0, 4X DP 2.1, Blower Fan, Vapor Chamber, Honeywell PTM7950

  • System Compatibility: Measures 271 x 112 x 39 mm
  • Power Requirements: Requires 12V-2x6-pin connector
  • Customer Support: Direct Amazon contact for assistance

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Outstanding Questions About Model Deployment and Licensing

While Alibaba has announced the upcoming release of the 2.4 trillion-parameter weights, details about the licensing terms remain unpublished, raising questions about usage rights and restrictions. It is also unclear whether the open weights will be fully functional for all types of deployment or limited to research purposes. The performance of the smaller 27B checkpoint in real-world applications compared to the flagship model is still to be validated.

Further, the actual infrastructure requirements for hosting the full model are not specified, leaving uncertainty about accessibility for smaller institutions or independent developers.

Building MCP Servers for AI Agents: Scalable Architecture Patterns, Security Design, and Production-Ready AI Infrastructure for Large Language Models

Building MCP Servers for AI Agents: Scalable Architecture Patterns, Security Design, and Production-Ready AI Infrastructure for Large Language Models

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Next Steps for Access and Evaluation of Qwen3.8-Max

The open weights for Qwen3.8-Max are scheduled to be released next week, enabling researchers and developers to evaluate its capabilities firsthand. Alibaba is expected to publish licensing details alongside the weights, clarifying usage rights. Additionally, the smaller 27B model will become available for local deployment, offering an accessible option for practical applications. Industry analysts will closely monitor how the model performs across diverse benchmarks and real-world tasks, especially in agentic and multimodal contexts.

ASRock Intel Arc Pro B70 Creator 32GB Workstation Graphics Card, Xe2-HPG, 32GB GDDR6, PCIe 5.0, 4X DP 2.1, Blower Fan, Vapor Chamber, Honeywell PTM7950

ASRock Intel Arc Pro B70 Creator 32GB Workstation Graphics Card, Xe2-HPG, 32GB GDDR6, PCIe 5.0, 4X DP 2.1, Blower Fan, Vapor Chamber, Honeywell PTM7950

  • System Compatibility: Measures 271 x 112 x 39 mm
  • Power Requirements: Requires 12V-2x6-pin connector
  • Customer Support: Direct Amazon contact for assistance

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

When will Alibaba release the open weights for Qwen3.8-Max?

The open weights are scheduled for release next week, on August 10, 2023, according to Alibaba’s announcement.

What are the main capabilities of Qwen3.8-Max?

Qwen3.8-Max supports multimodal inputs, demonstrates strong benchmark performance in reasoning and agentic tasks, and is built on sparse mixture-of-experts architecture with 2.4 trillion total parameters.

Will the open weights be usable for local deployment?

The full 2.4 trillion-parameter weights are expected to be a multi-node artifact, unsuitable for local deployment. However, a 27B checkpoint, Qwen3.8-27B, will be available for local use on high-memory machines.

What are the licensing terms for the open weights?

Licensing details are not yet published, but they are considered important given Alibaba’s historical licensing practices. Further information is expected alongside the weight release.

How does Qwen3.8-Max compare to other models like GPT-5 or Fable 5?

In benchmark tests, Qwen3.8-Max outperforms models like Claude Opus 4.8 and Fable 5 in several areas but trails behind GPT-5.6 at maximum effort. Its agentic reasoning capabilities have shown significant improvement over previous Alibaba models.

Source: ThorstenMeyerAI.com

This content is for general information only and is not financial, tax or legal advice. Consult a qualified professional for decisions about your money.
You May Also Like

The gigawatt gap. Why China is structurally positioned for AI power and the US is engineering around its grid.

Analysis of how China’s centralized infrastructure and renewable buildout give it an edge in AI power capacity, contrasting US fragmentation and grid constraints.

Every Benchmark Launched 2023-2024 Has Fallen — The METR / SWE-Bench / CORE-Bench / MLE-Bench / PostTrainBench Sequence

Every major AI research benchmark launched in 2023-2024 has been saturated or is nearing saturation, signaling rapid progress in AI capabilities.

Will The Lowest Temperature In Hong Kong Be 26°C On August 3?

Speculation surrounds Hong Kong’s weather with a new market suggesting a 20% chance of lowest temperature hitting 26°C on August 3. Details are still emerging.

Évian and the Fallout: What Europe Actually Wants From Amodei, Hassabis, and Altman

Europe pushes for reliable AI access, sovereignty, and safety at G7 summit with Amodei, Hassabis, and Altman, amid U.S. export controls and geopolitical tensions.