ByteDance is pre-training a model with up to 10 trillion parameters, according to a Financial Times report picked up by Reuters on 6 August 2026.1 Reuters via WHBL 2026-08-06 ByteDance is pre-training a model with up to 10 trillion parameters per the Financial Times, more than three times Kimi K3 at 2.8 trillion and close to the estimated 8 trillion of Anthropic Mythos 5 (Fable 5 near 5 trillion); LongCat-2.0 and DeepSeek V4-Pro at 1.6 trillion each; pre-training typically three to six months; Reuters could not immediately verify and ByteDance did not respond. Open source That would be more than three times the size of Moonshot AI's Kimi K3, at 2.8 trillion parameters the largest released Chinese model, and would approach or exceed the roughly 8 trillion parameters industry estimates assign to Anthropic's Mythos 5.1 Reuters via WHBL 2026-08-06 ByteDance is pre-training a model with up to 10 trillion parameters per the Financial Times, more than three times Kimi K3 at 2.8 trillion and close to the estimated 8 trillion of Anthropic Mythos 5 (Fable 5 near 5 trillion); LongCat-2.0 and DeepSeek V4-Pro at 1.6 trillion each; pre-training typically three to six months; Reuters could not immediately verify and ByteDance did not respond. Open source The run is in early pre-training, a stage that typically takes three to six months before fine tuning and any release, and the final parameter count is not fixed.1 Reuters via WHBL 2026-08-06 ByteDance is pre-training a model with up to 10 trillion parameters per the Financial Times, more than three times Kimi K3 at 2.8 trillion and close to the estimated 8 trillion of Anthropic Mythos 5 (Fable 5 near 5 trillion); LongCat-2.0 and DeepSeek V4-Pro at 1.6 trillion each; pre-training typically three to six months; Reuters could not immediately verify and ByteDance did not respond. Open source 3 Slashdot 2026-08-07 The model remains in early pre-training; parameter count is not everything, data quality and training technique often matter as much as raw scale; the project shows Chinese firms pushing frontier AI despite limits on access to advanced chips. Open source We assess with moderate confidence that the significant fact here is not the headline number but the strategy change underneath it: ByteDance founder Zhang Yiming has directed his teams to build capability independently rather than distil it from rivals' outputs, which is a public renunciation of the fast follower playbook that produced most of China's 2025 and 2026 model surge.2 Techloy 2026-08-07 ByteDance is pursuing independent development rather than distillation, believing it offers the best chance of achieving world-leading performance over the long term; researchers caution bigger does not always mean better since data, architecture and optimisation decide outcomes. Open source 4 Tech Startups 2026-08-07 One of the most ambitious scale ups from a Chinese lab; Zhang Yiming directed teams to prioritise genuine capability over short term distillation techniques; the push reflects China closing the gap with US frontier models amid export controls. Open source

What is actually confirmed, and what is not

Start with the sourcing, because it constrains everything else. The parameter figure comes from the Financial Times citing people familiar with the project. Reuters, relaying the report, said it could not immediately verify it, and ByteDance did not respond to requests for comment.1 Reuters via WHBL 2026-08-06 ByteDance is pre-training a model with up to 10 trillion parameters per the Financial Times, more than three times Kimi K3 at 2.8 trillion and close to the estimated 8 trillion of Anthropic Mythos 5 (Fable 5 near 5 trillion); LongCat-2.0 and DeepSeek V4-Pro at 1.6 trillion each; pre-training typically three to six months; Reuters could not immediately verify and ByteDance did not respond. Open source No benchmark exists, no architecture has been disclosed, and a model in early pre-training can be resized, restarted, or quietly shelved. Everything in this brief that depends on the 10 trillion figure should be read as single sourced reporting about an internal target, not as a shipped fact.

What the figure would mean if it holds is easier to state. China's largest models climbed fast through 2026: Meituan's LongCat-2.0 and DeepSeek's V4-Pro at 1.6 trillion total parameters each, then Moonshot's Kimi K3 resetting the local ceiling at 2.8 trillion.1 Reuters via WHBL 2026-08-06 ByteDance is pre-training a model with up to 10 trillion parameters per the Financial Times, more than three times Kimi K3 at 2.8 trillion and close to the estimated 8 trillion of Anthropic Mythos 5 (Fable 5 near 5 trillion); LongCat-2.0 and DeepSeek V4-Pro at 1.6 trillion each; pre-training typically three to six months; Reuters could not immediately verify and ByteDance did not respond. Open source A 10 trillion parameter run would not just extend that curve, it would jump past the estimated scale of the largest American frontier systems, with the same Reuters relay putting Anthropic's Mythos 5 near 8 trillion and Fable 5 near 5 trillion, both themselves estimates rather than disclosures.1 Reuters via WHBL 2026-08-06 ByteDance is pre-training a model with up to 10 trillion parameters per the Financial Times, more than three times Kimi K3 at 2.8 trillion and close to the estimated 8 trillion of Anthropic Mythos 5 (Fable 5 near 5 trillion); LongCat-2.0 and DeepSeek V4-Pro at 1.6 trillion each; pre-training typically three to six months; Reuters could not immediately verify and ByteDance did not respond. Open source One necessary caveat the coverage itself carries: most large Chinese releases use mixture of experts designs that activate only a fraction of total parameters per request, and researchers quoted across the reporting stress that data quality, architecture and training technique often matter as much as raw scale, so total parameter counts are a weak proxy for both capability and running cost.2 Techloy 2026-08-07 ByteDance is pursuing independent development rather than distillation, believing it offers the best chance of achieving world-leading performance over the long term; researchers caution bigger does not always mean better since data, architecture and optimisation decide outcomes. Open source 3 Slashdot 2026-08-07 The model remains in early pre-training; parameter count is not everything, data quality and training technique often matter as much as raw scale; the project shows Chinese firms pushing frontier AI despite limits on access to advanced chips. Open source

The distillation renunciation is the real decision

Distillation, training a smaller model on the outputs of a stronger one, is the cheap route that let Chinese labs compress years of American capability gains into months. ByteDance is now explicitly declining it. Zhang Yiming directed teams to prioritise genuine capability over short term distillation techniques, and the company's stated reasoning is that independent development "offers the best chance of achieving world-leading performance over the long term."2 Techloy 2026-08-07 ByteDance is pursuing independent development rather than distillation, believing it offers the best chance of achieving world-leading performance over the long term; researchers caution bigger does not always mean better since data, architecture and optimisation decide outcomes. Open source 4 Tech Startups 2026-08-07 One of the most ambitious scale ups from a Chinese lab; Zhang Yiming directed teams to prioritise genuine capability over short term distillation techniques; the push reflects China closing the gap with US frontier models amid export controls. Open source Zhang's internal framing, per Ciente's account, was to stop chasing quick wins and build world class infrastructure instead.5 Ciente 2026-08-07 Doubao counts 324 million monthly active users, SeeDance rivals leading Silicon Valley video tools, and Zhang Yiming told staff to stop chasing quick wins and focus on building world class infrastructure. Open source

We assess with moderate confidence that this choice, not the parameter count, is the load bearing news. A company that distils cannot lead, only track, because its ceiling is set by whichever frontier model it is imitating. A 10 trillion parameter independent pre-training run is what the decision to stop tracking looks like in practice: it converts ByteDance's two structural advantages, consumer scale and capital, into the one input distillation never required, which is enormous amounts of original training compute. And ByteDance has the distribution to amortise that spend: its Doubao chatbot counts 324 million monthly active users, and its SeeDance video generator is competitive with the best American tools.5 Ciente 2026-08-07 Doubao counts 324 million monthly active users, SeeDance rivals leading Silicon Valley video tools, and Zhang Yiming told staff to stop chasing quick wins and focus on building world class infrastructure. Open source The project also lands as a statement about export controls, since the reporting frames it as Chinese firms pushing frontier scale despite constrained access to advanced chips.3 Slashdot 2026-08-07 The model remains in early pre-training; parameter count is not everything, data quality and training technique often matter as much as raw scale; the project shows Chinese firms pushing frontier AI despite limits on access to advanced chips. Open source 4 Tech Startups 2026-08-07 One of the most ambitious scale ups from a Chinese lab; Zhang Yiming directed teams to prioritise genuine capability over short term distillation techniques; the push reflects China closing the gap with US frontier models amid export controls. Open source

Second order effects: who gains and who loses

The immediate loser is the economics of the Chinese model market. Moonshot, DeepSeek and Meituan built positions on efficient mid scale releases; a rival willing to spend at a multiple of their scale, backed by TikTok cash flows and 324 million Doubao users, forces each of them to choose between matching a compute bill they cannot afford and conceding the top of the market.1 Reuters via WHBL 2026-08-06 ByteDance is pre-training a model with up to 10 trillion parameters per the Financial Times, more than three times Kimi K3 at 2.8 trillion and close to the estimated 8 trillion of Anthropic Mythos 5 (Fable 5 near 5 trillion); LongCat-2.0 and DeepSeek V4-Pro at 1.6 trillion each; pre-training typically three to six months; Reuters could not immediately verify and ByteDance did not respond. Open source 5 Ciente 2026-08-07 Doubao counts 324 million monthly active users, SeeDance rivals leading Silicon Valley video tools, and Zhang Yiming told staff to stop chasing quick wins and focus on building world class infrastructure. Open source Domestic compute suppliers gain regardless of whether the model is good: a training run of this size, executed under chip export limits, is demand that must be filled from whatever silicon ByteDance can legally assemble, which strengthens both stockpiled Nvidia inventory value and domestic accelerator vendors.3 Slashdot 2026-08-07 The model remains in early pre-training; parameter count is not everything, data quality and training technique often matter as much as raw scale; the project shows Chinese firms pushing frontier AI despite limits on access to advanced chips. Open source

For the American frontier labs the effect is mostly narrative but not trivially so. Anthropic's Mythos 5 is the explicit benchmark in every account of this project.1 Reuters via WHBL 2026-08-06 ByteDance is pre-training a model with up to 10 trillion parameters per the Financial Times, more than three times Kimi K3 at 2.8 trillion and close to the estimated 8 trillion of Anthropic Mythos 5 (Fable 5 near 5 trillion); LongCat-2.0 and DeepSeek V4-Pro at 1.6 trillion each; pre-training typically three to six months; Reuters could not immediately verify and ByteDance did not respond. Open source A Chinese lab announcing an attempt to out-scale it, rather than undercut it on price, is a new argument in Washington's export control debate: it says the controls did not stop frontier scale training, only raised its cost. We assess with low confidence, given the single sourced basis, that this becomes a cited exhibit in the next round of US policy argument over chip restrictions, usable by both sides: hawks will read it as leakage, doves as proof that controls push China toward self sufficiency.

The counter-case

The strongest case against the thesis is that this is a target dressed as a milestone. The report is unverified by a second outlet, the company declined comment, and pre-training at this scale can fail in ways that never become public: loss curves that flatten, data pipelines that cannot feed 10 trillion parameters with enough quality tokens, or a compute ceiling imposed by exactly the export controls the project is meant to defy.1 Reuters via WHBL 2026-08-06 ByteDance is pre-training a model with up to 10 trillion parameters per the Financial Times, more than three times Kimi K3 at 2.8 trillion and close to the estimated 8 trillion of Anthropic Mythos 5 (Fable 5 near 5 trillion); LongCat-2.0 and DeepSeek V4-Pro at 1.6 trillion each; pre-training typically three to six months; Reuters could not immediately verify and ByteDance did not respond. Open source 3 Slashdot 2026-08-07 The model remains in early pre-training; parameter count is not everything, data quality and training technique often matter as much as raw scale; the project shows Chinese firms pushing frontier AI despite limits on access to advanced chips. Open source The researchers quoted in the coverage make the quieter version of the same point: a smaller model trained well beats a larger one trained badly, so a 10 trillion parameter release that underperforms Kimi K3 per dollar would discredit the scale strategy rather than validate it.2 Techloy 2026-08-07 ByteDance is pursuing independent development rather than distillation, believing it offers the best chance of achieving world-leading performance over the long term; researchers caution bigger does not always mean better since data, architecture and optimisation decide outcomes. Open source 3 Slashdot 2026-08-07 The model remains in early pre-training; parameter count is not everything, data quality and training technique often matter as much as raw scale; the project shows Chinese firms pushing frontier AI despite limits on access to advanced chips. Open source For this brief's assessment to fail, ByteDance would ship nothing at frontier scale within a year, or ship something whose capability the market reads as a distillation era product at a pre-training era price.

What to watch

  • A second independent confirmation. Watch for any outlet beyond the Financial Times independently sourcing the parameter target or the training run's existence by the end of September 2026; continued single sourcing after two months would justify discounting the number heavily.1 Reuters via WHBL 2026-08-06 ByteDance is pre-training a model with up to 10 trillion parameters per the Financial Times, more than three times Kimi K3 at 2.8 trillion and close to the estimated 8 trillion of Anthropic Mythos 5 (Fable 5 near 5 trillion); LongCat-2.0 and DeepSeek V4-Pro at 1.6 trillion each; pre-training typically three to six months; Reuters could not immediately verify and ByteDance did not respond. Open source
  • The release window. Pre-training of three to six months from an early August disclosure points to fine tuning around late 2026 and a possible release in the first half of 2027; slippage past mid 2027 without a shipped model would signal the run struggled.1 Reuters via WHBL 2026-08-06 ByteDance is pre-training a model with up to 10 trillion parameters per the Financial Times, more than three times Kimi K3 at 2.8 trillion and close to the estimated 8 trillion of Anthropic Mythos 5 (Fable 5 near 5 trillion); LongCat-2.0 and DeepSeek V4-Pro at 1.6 trillion each; pre-training typically three to six months; Reuters could not immediately verify and ByteDance did not respond. Open source 2 Techloy 2026-08-07 ByteDance is pursuing independent development rather than distillation, believing it offers the best chance of achieving world-leading performance over the long term; researchers caution bigger does not always mean better since data, architecture and optimisation decide outcomes. Open source
  • The disclosed architecture. If the eventual release states its total versus active parameter counts, a mixture of experts design with a small active fraction would confirm the headline number was a capacity figure, not a density claim, and reframe every comparison to Mythos 5.1 Reuters via WHBL 2026-08-06 ByteDance is pre-training a model with up to 10 trillion parameters per the Financial Times, more than three times Kimi K3 at 2.8 trillion and close to the estimated 8 trillion of Anthropic Mythos 5 (Fable 5 near 5 trillion); LongCat-2.0 and DeepSeek V4-Pro at 1.6 trillion each; pre-training typically three to six months; Reuters could not immediately verify and ByteDance did not respond. Open source
  • Rival responses in China. Watch whether Moonshot, DeepSeek or Alibaba announce materially larger training runs of their own by the first quarter of 2027; a scale race breaking out domestically would confirm ByteDance moved the goalposts, while silence would suggest rivals are betting the efficiency curve beats the scale curve.1 Reuters via WHBL 2026-08-06 ByteDance is pre-training a model with up to 10 trillion parameters per the Financial Times, more than three times Kimi K3 at 2.8 trillion and close to the estimated 8 trillion of Anthropic Mythos 5 (Fable 5 near 5 trillion); LongCat-2.0 and DeepSeek V4-Pro at 1.6 trillion each; pre-training typically three to six months; Reuters could not immediately verify and ByteDance did not respond. Open source
  • Whether the no distillation line holds. Any credible evidence that the model's training corpus leaned on rival model outputs would collapse the strategic story ByteDance is telling; a clean independent run that reaches the frontier would be the first Chinese existence proof that the shortcut was optional.2 Techloy 2026-08-07 ByteDance is pursuing independent development rather than distillation, believing it offers the best chance of achieving world-leading performance over the long term; researchers caution bigger does not always mean better since data, architecture and optimisation decide outcomes. Open source 4 Tech Startups 2026-08-07 One of the most ambitious scale ups from a Chinese lab; Zhang Yiming directed teams to prioritise genuine capability over short term distillation techniques; the push reflects China closing the gap with US frontier models amid export controls. Open source

If the run works, the interesting consequence is not that China has a big model. It is that the largest consumer internet company in China concluded that imitation had hit its ceiling, and priced the frontier as something worth buying at full cost.