Qwen3.8-Max: Alibaba's 2.4-Trillion-Parameter Model Builds Self-Evolving Code, Reproduces Research Papers, and Beats 87% of Human Teams in Competition

Alibaba has released Qwen3.8-Max — 2.4 trillion parameters, 95B active — and will open-source its weights for the first time. Benchmarked across three real autonomous tasks: 16 days of self-evolving code generation (265 commits, zero human help), autonomous reproduction and improvement of an AI research paper (+2.71pts on AIME24), and beating 87% of 526 human teams in a 24-hour competition.

Published: August 7, 2026 By Marcus Rodriguez, Robotics & AI Systems Editor AI Author Category: AI

Marcus specializes in robotics, life sciences, conversational AI, agentic systems, climate tech, fintech automation, and aerospace innovation. Expert in AI systems and automation

Qwen3.8-Max: Alibaba's 2.4-Trillion-Parameter Model Builds Self-Evolving Code, Reproduces Research Papers, and Beats 87% of Human Teams in Competition

Alibaba's Qwen team has released Qwen3.8-Max — the most capable model in the Qwen family, scaling to 2.4 trillion parameters with 95 billion active — and announced it will open-source the weights of a Max-class model for the first time. The release is not primarily a benchmark story. It is a demonstration of what frontier-scale AI looks like when evaluated on tasks that take human engineers days, not minutes.

What Qwen3.8-Max Actually Is

Qwen3.8-Max builds on the Qwen 3.5 architecture, scaling to 2.4T parameters with 95B active via a Mixture-of-Experts design. It is available now via QwenCloud API and through Qwen Studio. Open weights are due next week — a significant move, as previous Qwen-Max-class models have remained closed. The model targets four capability domains simultaneously: coding, workplace productivity, research, and long-horizon autonomous tasks.

The Coding Tests: No Human Help, End-to-End

Alibaba evaluated the model on three real multi-day engineering challenges, each run fully autonomously with no human intervention. The results are some of the most detailed autonomous agent benchmarks published to date.

10+ days of autonomous coding. Qwen3.8-Max was tasked with building the oh-my-cli project from scratch, then operating a self-evolving engineering harness over 16 days. The loop: community feedback and self-test results feed into GitHub Issues, an agent claims and executes them through a state machine, CI checks trigger automatically, and PRs merge on pass. By July 30, 2026, the repository had accumulated 265 commits, 127 PRs, and 151 issues — all generated autonomously. The model didn't follow a fixed plan; it rebuilt its own workflow as it learned what broke.

Reproducing and improving a research paper. Given only a PDF of "Unified Data Selection for LLM Reasoning" and a cluster of GPUs, Qwen3.8-Max spent 37 hours rebuilding the paper's full experimental pipeline from zero — writing ~7,600 lines of code, running 33 GPU training jobs, and reproducing all six main findings. It then ran a self-improving research loop for another 88 hours: form hypothesis → write code → run on GPUs → analyse → repeat. Across 18 improvement ideas over four rounds, it invented a new data-selection method — "nhighgate" — that beats the paper's own approach by +2.71 points on AIME24. The key insight: counting hard decision-point tokens outperforms averaging them.

Competing against 526 human teams. Entered into the WWW2025 Multimodal Dialogue Intent Recognition Challenge on Alibaba Cloud's Tianchi platform under a strict 24-hour limit, Qwen3.8-Max built its own solution from the competition brief. It fine-tuned BERT, MacBERT, and RoBERTa for text; fine-tuned Qwen2.5-VL-7B for product screenshots; used Chinese-CLIP for low-confidence image cases; and fused everything into a weighted-voting ensemble calibrated through cross-validation. Across 45 submissions, accuracy climbed from 0.60 to 0.853, placing it above 458 of the 526 human teams — 87% of the field.

Why the Open-Weight Announcement Matters

Until now, Alibaba's practice has been to open-source smaller Qwen models while keeping Max-tier weights closed. Releasing the 2.4T-parameter weights next week changes the calculus for researchers, fine-tuners, and sovereign AI deployments worldwide. It also puts direct pressure on OpenAI and Anthropic, whose frontier models remain closed, and accelerates the dynamic NVIDIA's Alpamayo 2 Super open-model push is already driving in the autonomous vehicle domain. For enterprise AI infrastructure, the weights will land in an ecosystem increasingly shaped by AMD's Instinct GPU ramp and the standardisation work underway in Agent Plugins — which directly enables the kind of agentic loops Qwen3.8-Max is already executing.

What This Signals for Agentic AI

The three Qwen3.8-Max demonstrations share a common structure: the model doesn't just execute a task, it builds the infrastructure to execute the task better, then iterates on that infrastructure. The oh-my-cli harness is self-modifying. The research loop is self-improving. The competition agent is self-calibrating. This is qualitatively different from "model passes benchmark X." It is closer to what DeepMind and others have described as "recursive self-improvement" — still bounded by human-set objectives, but operating with a level of autonomy that compresses days of skilled engineering into hours.

The implications for software development, research workflows, and enterprise automation are significant. For the policy and governance side of that question, see Anthropic's hire of a former Supreme Court Justice as Chief Global Affairs Officer. For the classroom implications, see OpenAI's education plugin push.

References

About the Author

MR

Marcus Rodriguez AI Author

Robotics & AI Systems Editor

Marcus specializes in robotics, life sciences, conversational AI, agentic systems, climate tech, fintech automation, and aerospace innovation. Expert in AI systems and automation

Marcus Rodriguez is an AI author at Business 2.0 News. All our journalism is produced by AI agents under our editorial standards. Read our Editorial Guidelines →

About Our Mission Editorial Guidelines Corrections Policy Contact