ByteDance, the company behind TikTok, is pre-training a language model with up to 10 trillion parameters — the largest ever built by a Chinese lab, and a scale that puts it in the same conversation as Anthropic, according to Financial Times reporting published August 7, based on three people with direct knowledge of the project.
The model is being built by ByteDance's Seed team, roughly 2,000 researchers strong, and would be more than three times larger than Moonshot AI's Kimi K3 — until now the biggest Chinese model publicly released, at 2.8 trillion parameters. The comparison making the rounds internally is direct: the declared target is to close in on Claude Mythos, Anthropic's frontier model, estimated at roughly 8 trillion parameters.
It's worth being precise about what the reporting does not confirm. The project is in the pretraining phase — a stage that typically takes three to six months to complete on its own, before any fine-tuning, safety evaluation, or release. There's no disclosed active-parameter architecture, no specified chips, no release date, and not a single public benchmark. ByteDance founder Zhang Yiming has told the team to avoid distilling from rival models — building independent capability even if that means falling behind in the short term. This is a long-term bet, not a product ready for evaluation.
Why "trillions of parameters" tells you little about what to buy
Here's the part that matters for anyone deciding on AI adoption at a Brazilian company: parameter count is an engineering metric, not a procurement metric. Anthropic itself doesn't disclose parameter counts for any of its models — including Mythos, which became the comparison benchmark in this story. That means much of the "gap" referenced is a third-party estimate, not a number Anthropic confirms. Comparing models by parameter count is like comparing cars by engine displacement without knowing fuel consumption, maintenance cost, or whether the engine even fits the application you need.
More relevant to IT budgets: bigger models cost more to run. Every additional trillion parameters pushes up inference cost — the bill your company pays per API call, not just the one-time training cost the lab absorbs. We've already shown here how AI pricing can swing significantly by time of day and demand; raw model scale is another variable pushing that cost upward, without necessarily delivering proportional gains on the task your company actually needs solved.
What this means for decision-makers at Brazilian companies
Before letting a scale announcement influence contract timing or AI budget, three questions matter more than parameter count:
- Does the model solve your specific task, on the right benchmark? General reasoning performance doesn't guarantee performance on customer service, tax-document extraction, or code generation compliant with internal policy.
- What's the cost per outcome, not per token? A smaller, cheaper model that gets it right 90% of the time can cost less overall than a giant model that gets it right 94% of the time — especially at high volume.
- Where does the data live, and under which jurisdiction? Chinese and American frontier models carry different data-sovereignty implications — a topic we covered here when Kimi K3 opened its weights.
The scale race between ByteDance, Anthropic, OpenAI, and Google will keep generating headlines in the months ahead. For anyone leading AI adoption in Brazil, the practical advice is to resist treating "bigger model" as shorthand for "better purchasing decision" — and keep evaluation criteria anchored to task, cost, and governance, not announced size.