On September 3, MBZUAI's Institute of Foundation Models, an AI university in Abu Dhabi, released K2 Horizon: six language models, ranging from 0.9 billion to 375 billion parameters, under the Apache 2.0 license. The difference that matters isn't the size — it's that, unlike almost every "open" release on the market, K2 Horizon makes available not just the model weights, but also the code, the training data, and the full methodology.
What Changed
The six models (0.9B, 3.7B, 7B, 32B, 36B-A4B, and 375B-A23B) are already available via Hugging Face, vLLM, and SGLang. According to MBZUAI, the three smaller models (0.9B, 3.7B, and 7B) set a new state of the art for their respective scales in reasoning, math, code, and agentic tasks. The Apache 2.0 license allows commercial use, modification, and redistribution, with no obligation to open-source any derivative — but the real differentiator lies elsewhere: most "open-weight" releases on the market (Meta, DeepSeek, and others) only release the trained weights, not the data used to train them. K2 Horizon releases the data and methodology alongside, which MBZUAI calls the largest fully open AI release in history — a characterization from the company itself, but one that points to a genuinely uncommon transparency practice in the industry.
Why It Matters
Training-data transparency has stopped being just an abstract open-source community value — it's become a concrete legal-risk question. We covered here last week the lawsuit Sony Music Publishing and Warner Chappell filed against Anthropic over alleged unauthorized use of copyrighted lyrics to train Claude — one of at least five similar suits against the company in under a year. A model whose training data is publicly auditable doesn't automatically eliminate intellectual-property risk, but it gives anyone evaluating the model a concrete basis for verification, instead of relying solely on the vendor's contractual assurance. The size range — from 0.9B, viable for edge or on-device deployment, up to 375B, frontier scale — also provides real cost optionality under a single permissive license.
The Impact for Brazil
For Brazilian companies already using open models as part of a cost-reduction strategy — a theme we've returned to several times this month, from DeepSeek to Hugging Face — K2 Horizon adds another well-funded option (MBZUAI has direct backing from the Abu Dhabi government) to the basket of alternatives outside the major US and Chinese labs. For anyone operating in a regulated sector or sensitive to intellectual-property risk, auditable training data is one more concrete criterion in vendor evaluation — it's worth explicitly asking, in your next model negotiation, whether the training data is known and verifiable, not just whether the weights are "open."
Entercast's Take
This adds a third geopolitical pole to the map we've been building here — Pax Silica, led by the US, and WAICO, led by China — with the United Arab Emirates investing heavily in sovereign, open AI through MBZUAI. For whoever leads AI adoption at your company, the practical point isn't picking a geopolitical side: it's recognizing that the basket of open-model options keeps growing, with vendors from different origins offering different trade-offs on cost, transparency, and legal risk. The same verification discipline we've applied to what an AI processes (Forcepoint) and to who has access to an agent (Google Cloud) applies to the model's own origin: auditable training data isn't a technical footnote, it's risk information.