Qwen3.8-Max's AI Capabilities: The Data That Might Redefine The Race

📊 Full opportunity report: Qwen3.8-Max's AI Capabilities: The Data That Might Redefine The Race on ThorstenMeyerAI.com — validation score, market gap, and execution plan.

TL;DR

Alibaba has officially released details of Qwen3.8-Max, a 2.4 trillion-parameter AI model with impressive benchmark scores. The open weights will be available next week, marking a significant development in AI model transparency and capability.

Alibaba has officially confirmed the specifications and benchmark results of Qwen3.8-Max, a 2.4 trillion-parameter AI model, and announced that open weights will be available next week. This marks a significant step in AI model transparency and competitiveness, with the model outperforming several rivals on key benchmarks.

On August 3, Alibaba published the full benchmark table for Qwen3.8-Max, revealing it as a 2.4 trillion-parameter sparse mixture-of-experts model built on the Qwen3.5 architecture. The model features approximately 95 billion active parameters per query and supports multimodal inputs, including text, images, and videos, with text output. The benchmark results show it surpassing many competitors in various tests, such as Terminal-Bench 2.1 with a score of 86.6, ahead of Claude Opus 4.8 and Claude Fable 5, but behind GPT-5.6 Sol at 88.8.

Alibaba also confirmed that full open weights will ship next week, with the release including a 27B checkpoint optimized for local deployment. The company emphasized that the 2.4T model is primarily a data artifact, with the open weights representing a multi-node datacenter deployment rather than a single-machine solution. The announcement was preceded by a stealth preview and a staged reveal, which generated significant media coverage and a 5.4% rise in Alibaba’s stock price.

At a glance
updateWhen: announced August 3, 2023; open weights…
The developmentAlibaba announced the broad availability of Qwen3.8-Max, confirming its specifications and benchmark performance, and revealed that open weights will be shipped next week.
AI DISPATCH · REALITY CHECK Released 3 Aug 2026
Alibaba’s Qwen3.8-Max leaves preview
Second Only to Fable 5?

For fifteen days the claim ran without a benchmark table. Today Alibaba published the table, the active-parameter count, and a weights timeline. The numbers are genuinely strong on the rows Alibaba chose — and twelve to fifteen points behind on the rows it didn’t.

▲ All performance figures: Alibaba’s own harness
2.4T / 95B
Total / active parameters (MoE)
~1M
Context window · 131K max output
Text+Img+Video
Multimodal in · text out
“Next week”
Open weights · licence unpublished
01
Fifteen days from slogan to spec sheet

The claim shipped on a Sunday. The evidence shipped two weeks later. In between, the claim did its work.

17 Jul
Moonshot releases Kimi K3
2.8T parameters; rattles US tech stocks, later suspends new subscriptions under demand.
18 Jul
“kaleb” appears on Code Arena
Anonymous model introduces itself as “Claude” — a distillation artifact — and is identified within a day by a Qwen tokenizer quirk.
19 Jul
WAIC preview: “second only to Fable 5”
No benchmark table, no model card, no licence, no active-parameter count. Paid preview at 10% of standard pricing.
20 Jul
Shares rise as much as 5.4%
The market prices the claim, not the table.
3 Aug
General availability + full benchmark table
95B active confirmed; 2.4T weights and a Qwen3.8-27B checkpoint promised for next week. Licence still unwritten.
02
The table, both halves

“Second only to Fable 5” is true on the rows Alibaba chose and false on the rows it didn’t. Both halves below are from the same release.

Where it leads
Terminal-Bench 2.1 · agentic terminal work
Qwen3.8-Max
86.6
GPT-5.6 Sol
88.8
Fable 5
84.6
OSWorld-Verified · computer use — plus PaperBench 93.0, CAD Bench 91.5
Qwen3.8-Max
86.1
Where it trails — the rows the slogan skips
SWE-bench Pro · deep software engineering
Qwen3.8-Max
67.7
Fable 5
80.0
FrontierSWE · frontier coding agents
Qwen3.8-Max
73.5
Fable 5
88.8
The real jump: one generation of agentic gains vs Qwen3.7-Max
DeepSWE 1.1
21.6 → 56.6
FrontierSWE
40.7 → 73.5
JobBench
31.3 → 53.4
03
Three artifacts, three different facts

“Qwen3.8 is going open-weight” describes three things with very different deployment realities.

Hosted API
Live today

OpenAI- and DashScope-compatible — a base-URL change to A/B against your current backend.

2.4T weights
“Next week” · no licence yet

A multi-node datacenter artifact. At 95B active, no single machine serves it. A flag planted, not a deployment option.

Qwen3.8-27B
Announced · no benchmarks yet

The checkpoint that fits real hardware. Whether the agentic gains survive distillation is the question that decides whether next week matters.

04
Bull and bear

Three Chinese frontier releases in seventeen days, each measured against the same export-controlled model. The contest is real; it is not the same thing as your workload.

Bull
  • The generation jump is real and consistent across a dozen agentic rows, with a stated mechanism: RL-environment scaling.
  • More disclosure than Kimi K3 shipped — full table, active-parameter count, weights timeline.
  • If 2.4T lands under a permissive licence, the ceiling of “open weight” moves permanently.
  • The 27B sibling could become the best local agent model on hardware people already own.
Bear
  • Every number is Alibaba’s harness. Independent testing already tempered Kimi K3’s launch claims substantially.
  • The paying use case still belongs to Fable 5 — twelve to fifteen points on deep software engineering.
  • “Next week” comes from a company that sat on a finished benchmark table for fifteen days.
  • Until the licence text exists, “going open-weight” is a press strategy, not a property of the model.
The claim ran for fifteen days without evidence. Now the evidence exists —
and it says “second only” depends entirely on which row you read.

Implications of Alibaba’s Open-Weight AI Model Release

The release of Qwen3.8-Max’s benchmark data and upcoming open weights could significantly influence the AI landscape by setting new standards for transparency and performance. The model’s high benchmark scores, especially in multimodal and agentic tasks, demonstrate progress in AI capabilities, potentially impacting commercial applications and research. The availability of open weights, particularly the 27B checkpoint suitable for local deployment, may democratize access to advanced AI, sparking increased competition and innovation. However, the large size of the 2.4T model also underscores ongoing challenges in deploying such models at scale, raising questions about accessibility and licensing.

The FPGA Programming Handbook: An essential guide to FPGA design for transforming ideas into hardware using SystemVerilog and VHDL

The FPGA Programming Handbook: An essential guide to FPGA design for transforming ideas into hardware using SystemVerilog and VHDL

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Background on Alibaba’s AI Model Development Timeline

Over the past two weeks, Alibaba’s AI model developments have been characterized by staged reveals and strategic disclosures. The company’s previous model, Kimi K3, with 2.8 trillion parameters, briefly impacted US tech stocks before demand exceeded capacity. An anonymous model called 'kaleb' appeared on the Code Arena leaderboard, later confirmed as Qwen3.8-Max during the World AI Conference in Shanghai. The staged rollout included a preview endpoint and limited access at discounted pricing, building anticipation for the full release. Alibaba’s approach contrasts with other major AI players, emphasizing staged disclosures and selective benchmarking.

"We are committed to advancing AI capabilities and making them accessible through open weights, starting next week."

— Alibaba spokesperson

Unresolved Aspects of Qwen3.8-Max’s Deployment and Licensing

While Alibaba has confirmed the specifications and benchmark results, several details remain unclear. The exact licensing terms for the open weights are still unpublished, raising questions about usage rights and restrictions. Additionally, the performance of the 27B checkpoint in real-world deployment, especially regarding agentic and long-horizon tasks, has yet to be tested outside controlled benchmarks. The long-term impact of the model’s size and complexity on accessibility and deployment costs also remains uncertain.

Next Steps for Alibaba’s AI Model Rollout and Community Engagement

Next week, Alibaba will release the open weights for Qwen3.8-Max, enabling broader access for researchers and developers. The company is expected to publish licensing details and provide documentation on deployment practices. Observers will closely monitor how the community adopts and adapts the model, especially the 27B checkpoint designed for local use. Further benchmark tests and real-world applications will likely follow, shaping the competitive dynamics among large AI models in the coming months.

Key Questions

What are the main capabilities of Qwen3.8-Max?

Qwen3.8-Max features 2.4 trillion parameters, supports multimodal inputs (text, images, videos), and excels in agentic tasks and benchmark performance, especially in multimodal and research-oriented benchmarks.

When will the open weights be available?

Alibaba has announced that the open weights for Qwen3.8-Max will be shipped next week, with the 27B checkpoint targeted for local deployment available immediately.

What are the licensing implications of the open weights?

The licensing terms are still unpublished. It is unclear whether the weights will be available under open licenses like Apache 2.0 or under more restrictive terms, which could impact usage and commercialization.

How does Qwen3.8-Max compare to other models like GPT-5.6 or Fable 5?

In benchmark tests, Qwen3.8-Max outperforms models like Claude Opus 4.8 and Fable 5 in several areas but trails behind GPT-5.6 at the highest effort scores. Its strengths lie particularly in multimodal and agentic tasks.

What does this development mean for AI research and industry?

The transparency and scale of Qwen3.8-Max could accelerate AI innovation, democratize access, and intensify competition among major AI labs, influencing both research and commercial deployment strategies.

Source: ThorstenMeyerAI.com

This content is for general information only and is not financial, tax or legal advice. Consult a qualified professional for decisions about your money.
You May Also Like

The 4.8 Staircase: What the Market Actually Believes About Claude’s Next Release

Market probabilities suggest a Claude 4.8 release by mid-June, but official confirmation is still pending. Here’s what is known and what remains uncertain.

The SSD Squeeze: Why Storage Joined the Party

Storage prices are rising sharply due to supply constraints, AI-driven demand, and factory competition, impacting enterprise and consumer markets.

The Power Outage Problem That Quietly Corrupts Office Gear

Find out how unseen power fluctuations can silently damage your office equipment and what steps you can take to prevent costly failures.

Why AI-Enhanced 4K Monitors Are A Must-Have In 2026

Exploring how AI-enhanced 4K monitors are transforming productivity, gaming, and creative work in 2026, with confirmed features and future implications.