Qwen3.8-Max: a new AI model from China – AI News – #1 August 2026

3min.

Comments:0

03 August 2026

Qwen3.8-Max: a new AI model from China – AI News – #1 August 2026d-tags
Alibaba has officially released Qwen3.8-Max—its most capable model to date built on a Mixture of Experts (MoE) architecture, scaling to 2.4 trillion total parameters with 95 billion active per token. Built upon the Qwen 3.5 foundation, the new flagship model sets higher benchmarks for multi-day autonomous coding, complex agentic workflows, and real-time multimodal reasoning. Alibaba announced that open weights will be released next week.

3min.

Comments:0

03 August 2026

Qwen3.8-Max: A new bar for coding and AI autonomy from China

The large language model (LLM) landscape is evolving at a breakneck pace, with the battle for frontier capabilities shifting heavily toward Asia. Following the recent launch of Kimi K3—which we previously analyzed on our blog—China’s AI sector has demonstrated its momentum once again. Alibaba has officially made its flagship Qwen3.8-Max available, proving that open-weight architectures and massive scalability can directly challenge closed-source models from Silicon Valley like GPT-5.6 or Claude Opus 4.8.

Qwen3.8-Max is not merely an incremental update to Qwen3.7-Max. It represents an architectural leap designed specifically to execute complex, long-horizon tasks from start to finish (end-to-end) without requiring continuous human intervention.

What makes Qwen3.8-Max stand out?

Alibaba’s new flagship is built around two primary architectural pillars: sustained, multi-day systemic planning and deep multimodal integration.

Autonomous coding across multiple days

Traditional code-generation models typically excel at writing isolated functions or refactoring short snippets of code. Qwen3.8-Max pushes this boundary into true long-horizon autonomy. In official evaluation cases, the model demonstrated:

  • 10+ days of autonomous software engineering: In the oh-my-cli project, the model spent over 16 days operating completely on its own—ingesting community issues, generating code, running E2E tests, and merging pull requests on GitHub.
  • Reproducing and advancing research papers: Starting from an empty environment with GPU access, the model wrote ~7,600 lines of code across 125 hours, reproduced a research paper’s findings, and subsequently developed its own data selection method that improved math benchmark performance (AIME24) by +2.7 points over the original paper.
  • Designing silicon hardware: Operating inside a sandboxed hardware design environment (RTL/OpenROAD), the model executed hundreds of feedback loops, cutting the synthesized gate count of a cryptographic hardware accelerator from 8,298 down to just 678 logic gates.

Multimodal capabilities and real-time feedback loops

Qwen3.8-Max is also a capable Hybrid Agent that bridges backend code execution with direct graphical user interface (GUI) control. The model can process 100+ hour video streams to build continuous memory graphs, as well as digest multi-hundred-page financial reports into structured data.

Crucially, vision in Qwen3.8-Max acts as a native feedback loop. While executing visual or frontend tasks—such as rendering 3D interior scenes in Blender or generating web applications—the model continuously observes intermediate outputs, identifies visual flaws or UI misalignments, and corrects them autonomously mid-run.

Qwen3.8-Max vs. Kimi K3: Comparing China’s frontier models

Comparing the two most prominent Chinese models of recent weeks highlights clear differences in engineering philosophy. While both Kimi K3 (Moonshot AI) and Qwen3.8-Max (Alibaba) represent the pinnacle of open-weight performance, each is optimized for distinct use cases.

FeatureQwen3.8-Max (Alibaba)Kimi K3 (Moonshot AI)
Total parameters2.4T (MoE)2.8T (MoE)
Active parameters~95B / token~104B / token
Context windowUp to 1M tokensUp to 1M tokens
Release strategyOpen weights announced (releasing next week)Open weights available at launch
Primary focusAll-around versatility, multimodal intelligence, enterprise coworkAgentic autonomy, coding efficiency, long-context reasoning
Deployment modelAPI via QwenCloud + upcoming local weightsImmediate local deployment (vLLM, SGLang)

Release strategy and weight availability

Kimi K3 prioritized developer accessibility from day one, publishing its open weights immediately for local execution. Alibaba opted for a hybrid release for Qwen3.8-Max: launching first on QwenCloud with native compatibility for OpenAI and Anthropic API protocols (enabling direct integration with tools like Claude Code and Codex), with open-weight releases scheduled for the following week.

Performance in coding and agentic tasks

Kimi K3 features custom architectural innovations (such as Kimi Delta Attention) that deliver impressive cost efficiency and speed when processing massive codebases. Conversely, Qwen3.8-Max excels in multi-step business logic and systematic execution. In real-world simulation benchmarks (such as the 365-day E-Commerce Bench simulation), Qwen achieved a 4.16x return on initial capital, outperforming competing models through strategic risk control and adaptive supplier negotiations.

Ecosystem and cost optimization

Qwen3.8-Max introduces fine-grained reasoning control via the reasoning_effort parameter (xhigh, medium, low), allowing developers to balance latency, cost, and reasoning depth dynamically. Meanwhile, Kimi K3 remains a formidable rival in pure coding benchmarks, offering high performance with low token consumption.

What Qwen3.8-Max means for the future of LLMs

The launch of Qwen3.8-Max underscores a broader shift in AI development: transitioning from simple conversational assistants to fully autonomous agentic systems capable of executing multi-day workflows. Alibaba’s flagship model proves that open-weight architectures from China are not only matching Western closed-source benchmarks, but are actively pioneering new capabilities in hybrid agent control, long-horizon decision making, and autonomous hardware engineering.

The upcoming open-weight release of a 2.4T parameter model will give research teams and enterprises worldwide access to frontier-class capabilities previously restricted to proprietary API endpoints.

Stay updated on the latest in AI

The world of artificial intelligence and digital marketing changes daily. To receive the latest analyses, model breakdowns, and actionable AI insights delivered straight to your inbox, subscribe to the Delante newsletter today!

Source: https://qwen.ai/blog?id=qwen3.8

Author
Maciej Jakubiec - Junior SEO Specialist
Author
Maciej Jakubiec

SEO Specialist

A marketing graduate specializing in e-commerce from the University of Economics in Kraków – part of Delante’s SEO team since 2022. A firm believer in the importance of well-crafted content, and apart from being an SEO, a passionate music producer crafting sounds since his early teens.

Author
Maciej Jakubiec - Junior SEO Specialist
Author
Maciej Jakubiec

SEO Specialist

A marketing graduate specializing in e-commerce from the University of Economics in Kraków – part of Delante’s SEO team since 2022. A firm believer in the importance of well-crafted content, and apart from being an SEO, a passionate music producer crafting sounds since his early teens.