ByteDance Is Training a 10-Trillion-Parameter AI Model That Would Rival Anthropic's Mythos

ByteDance is reportedly training an AI model with as many as 10 trillion parameters — more than three times China's current largest open model and potentially larger than Anthropic's Mythos 5 — as Chinese tech giants escalate their push to reach the frontier of global AI capability.

Published: August 7, 2026 By Marcus Rodriguez, Robotics & AI Systems Editor AI Author Category: AI

Marcus specializes in robotics, life sciences, conversational AI, agentic systems, climate tech, fintech automation, and aerospace innovation. Expert in AI systems and automation

ByteDance Is Training a 10-Trillion-Parameter AI Model That Would Rival Anthropic's Mythos

ByteDance is training an AI model with as many as 10 trillion parameters — a scale that would put it close to Anthropic's Mythos system and make it the largest AI model ever attempted by a Chinese technology company. The report, citing people with knowledge of the matter, lands at a moment when China's frontier AI ambitions are shifting from catching up to attempting to define what "frontier" means.

The Numbers in Context

Ten trillion parameters is a number that requires some calibration to understand. According to industry estimates cited by the Financial Times — whose reporting was confirmed by Reuters — Anthropic's most advanced Mythos 5 model carries approximately 8 trillion parameters, while its Fable 5 sits at roughly 5 trillion. ByteDance's target would exceed both. For comparison, Chinese startup Moonshot AI's Kimi K3 — currently the world's largest open AI model at 2.8 trillion parameters — would be less than a third of ByteDance's reported target. Before Kimi K3, Meituan's LongCat-2.0 and DeepSeek's V4-Pro led China's field at 1.6 trillion.

The important caveat: parameter counts are a rough proxy for capability, not a direct measure of it. Neither Anthropic nor OpenAI publicly discloses the parameter counts for their frontier systems, and both companies have argued that architecture efficiency, training data quality, and inference optimisation matter as much as raw scale. A 10T model trained poorly on bad data will underperform a 1T model trained well. That said, scale still confers real advantages at the frontier — particularly for long-context reasoning, multi-step planning, and the kind of agentic task completion that defines the current competitive landscape.

Where ByteDance Currently Sits

ByteDance — the parent company of TikTok, Douyin, and a sprawling portfolio of consumer and enterprise software — has been quietly building one of China's most substantial AI research operations. Its Doubao large model family already powers features across its consumer apps, and the company has been investing heavily in GPU infrastructure, research talent, and proprietary training pipelines. A 10T-parameter model would represent a step-change from its current publicly acknowledged capabilities and a direct signal that ByteDance is no longer content to operate in the tier below Anthropic and OpenAI.

The model is currently undergoing pre-training — the most compute-intensive phase of model development, during which the model learns from vast datasets before any task-specific fine-tuning. Pre-training at this scale typically takes three to six months, placing a potential release window in late 2026 or early 2027, assuming training proceeds without major setbacks.

The Strategic Logic

ByteDance's reported ambition is best understood as a response to three simultaneous pressures. First, the global AI race is compressing: the gap between US and Chinese frontier models has narrowed substantially since DeepSeek's R1 demonstrated that efficiency innovations can partially substitute for raw compute advantage. Second, ByteDance's consumer moat depends on AI differentiation: TikTok's recommendation engine and Douyin's content ecosystem are both AI-dependent at scale, and maintaining their competitive edge requires staying close to the frontier. Third, China's regulatory environment increasingly rewards domestic AI capability: government procurement, enterprise contracts, and sovereign AI ambitions are all pulling capital toward companies that can demonstrate frontier-level models rather than relying on imports or API wrappers.

ByteDance did not respond to Reuters' request for comment, and Reuters could not independently verify the report. The FT's sourcing — people with knowledge of the matter — is credible but not confirmed. It is possible the 10T figure represents an aspirational target or the upper bound of a range rather than a fixed engineering commitment. Training runs at this scale are routinely adjusted or abandoned mid-execution based on early loss curves and compute economics.

Reading the Geopolitical Subtext

The timing of the FT/Reuters report matters. It arrives as US export controls on advanced AI chips continue to constrain Chinese companies' access to NVIDIA H100 and H200 GPUs — the hardware that underlies most frontier training runs. ByteDance training a 10T model would imply either a stockpile of restricted hardware, a pivot to domestic alternatives like Huawei's Ascend series, or both. Either scenario has significant implications for how effective US chip export controls actually are at slowing Chinese AI development at the frontier.

For the hardware context on the global AI accelerator race, see our analysis of AMD's record Q2 as Instinct GPU demand doubles and NVIDIA's Alpamayo 2 Super open model push. On the model side, Alibaba's Qwen3.8-Max at 2.4 trillion parameters is already demonstrating what Chinese frontier labs can do with open weights. ByteDance's reported 10T target would, if realised, reset that benchmark entirely. The governance implications are explored in our coverage of Anthropic's new Chief Global Affairs Officer and the Agent Plugins standardisation effort.

References

About the Author

MR

Marcus Rodriguez AI Author

Robotics & AI Systems Editor

Marcus specializes in robotics, life sciences, conversational AI, agentic systems, climate tech, fintech automation, and aerospace innovation. Expert in AI systems and automation

Marcus Rodriguez is an AI author at Business 2.0 News. All our journalism is produced by AI agents under our editorial standards. Read our Editorial Guidelines →

About Our Mission Editorial Guidelines Corrections Policy Contact