{"slug": "k2-horizon-6-open-ai-models-with-full-training-data", "title": "K2 Horizon: 6 Open AI Models With Full Training Data", "summary": "On September 3, the Institute of Foundation Models (IFM) at Mohamed bin Zayed University of AI released K2 Horizon, a fleet of six open-source AI models ranging from 0.9B to 375B parameters, including full training code, datasets, intermediate checkpoints, and evaluation logs under Apache 2.0. IFM Founder Eric Xing said, 'Open source is much more than open weights. Science works when others can see the data, follow the method, reproduce the result, and improve on it.' The release, which coincided with Nvidia's $12.9 billion acquisition of Hugging Face, aims to enable auditing of training data for bias and compliance, a capability lacking in open-weights-only models like Meta's Llama.", "body_md": "On September 3, the Institute of Foundation Models (IFM) — run out of Abu Dhabi’s Mohamed bin Zayed University of AI — released [K2 Horizon](https://ifm.ai/blog/k2), a fleet of six foundation models spanning 0.9B to 375B parameters. All under Apache 2.0. What makes the release remarkable is not the parameter counts but what comes with them: the training code, training datasets, intermediate checkpoints, and evaluation logs. That is not what the AI industry calls “open.” That is what open source actually means.\n\n## Open Weights Is Not Open Source\n\nAlmost every AI model marketed as “open” is, in practice, open-weights only. Meta’s Llama, Alibaba’s Qwen, Mistral, DeepSeek — you can download the finished model and fine-tune it. You cannot inspect the training data, reproduce the training run, or audit what the model was built on. The [Open Source Initiative and the Free Software Foundation both state](https://futureagi.com/blog/open-source-vs-open-weight/) plainly that Llama is not open source. Meta’s Llama Community License restricts users above 700 million monthly active users and releases no training data. “Open source” is what Meta calls it. “Open weights” is what it actually is.\n\nK2 Horizon does not play that game. IFM releases weights plus training code and recipes, training datasets where redistribution licenses permit, intermediate checkpoints, fine-grained training logs, and evaluation results. For restricted datasets, it publishes source descriptions, filtering procedures, and mixture composition so the pipeline can still be reproduced. As IFM Founder Eric Xing put it: “Open source is much more than open weights. Science works when others can see the data, follow the method, reproduce the result, and improve on it.” That is a pointed critique of the status quo, and the release backs it up.\n\nRelated:[Nvidia Buys Hugging Face for $12.9B: What Devs Must Know]\n\nThe practical consequence is significant. Regulated industries — finance, healthcare, legal — can now audit a frontier-scale model’s training data for bias, IP violations, and compliance gaps. That auditing is impossible with Llama, Qwen, or any other open-weights-only model. If your organization has legal requirements around data provenance, K2 Horizon is the first frontier-scale option that can satisfy them.\n\n## Six Models, One Architecture\n\nThe fleet covers every deployment tier with a shared core architecture and vocabulary. The 0.9B model targets wearables. The 3.7B and 7B models serve edge and mobile deployments. The 32B is designed for on-premise servers. The 36B-A4B and 375B-A23B use IFM’s Mixture of Value Attention (MoVA) — a sparse attention mechanism that activates only a fraction of stored parameters per token. The 375B model stores 375 billion parameters but activates roughly 23 billion per token, cutting per-token compute substantially while preserving reasoning quality.\n\nThe flagship 375B carries a 512K token context window — enough for entire codebases, extended agent histories, or large document collections in a single pass. Day-zero support ships for vLLM, SGLang, and Ollama. Models are available now on [Hugging Face](https://huggingface.co/IFM/K2-Horizon-375B-A23B) with inference APIs from Compass, Cerebras, AWS, and Nebius. IFM also ships Uno, a diffusion-distillation LoRA adapter that generates token blocks in parallel, delivering 2.5–3× faster inference without quality loss — meaningful if you are running inference at scale and watching operating costs.\n\n## Same Day as the Nvidia–Hugging Face Deal\n\nThe timing is not subtle. [Nvidia confirmed its $12.9 billion Hugging Face acquisition](https://techcrunch.com/2026/09/03/nvidia-confirms-it-will-buy-hugging-face-for-12-9-billion/) on September 3 — the same day K2 Horizon launched. Hugging Face currently hosts K2 Horizon’s weights. Nvidia now owns the platform most developers use to access and distribute open models. The AI community’s “neutral” infrastructure is no longer neutral.\n\nHowever, K2 Horizon’s Apache 2.0 license means Nvidia cannot lock it down. Any party can mirror the weights, training data, and code without restriction. The fully open release is, in practical terms, a hedge against exactly this kind of platform consolidation. An open-weights-only release under a restrictive license would not survive the same scenario intact.\n\n## Key Takeaways\n\n- K2 Horizon is the largest fully open AI release since BLOOM — weights, training data, training code, and checkpoints, all under Apache 2.0\n- “Open weights” and “open source” are not the same thing; Meta’s Llama, Qwen, and Mistral are open weights only\n- The six-model fleet covers wearables through enterprise, with a shared architecture and day-zero support for vLLM, SGLang, and Ollama\n- Uno diffusion distillation provides 2.5–3× faster inference at no quality cost — useful if you are running K2 Horizon in production\n- K2 Horizon’s Apache 2.0 license is a direct hedge against platform consolidation — the Nvidia–Hugging Face deal makes that hedge immediately relevant", "url": "https://wpnews.pro/news/k2-horizon-6-open-ai-models-with-full-training-data", "canonical_source": "https://byteiota.com/k2-horizon-6-open-ai-models-with-full-training-data/", "published_at": "2026-09-04 01:15:10+00:00", "updated_at": "2026-09-04 01:23:42.745369+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-research", "ai-policy"], "entities": ["Institute of Foundation Models", "Mohamed bin Zayed University of AI", "K2 Horizon", "Eric Xing", "Meta", "Nvidia", "Hugging Face", "Apache 2.0"], "alternates": {"html": "https://wpnews.pro/news/k2-horizon-6-open-ai-models-with-full-training-data", "markdown": "https://wpnews.pro/news/k2-horizon-6-open-ai-models-with-full-training-data.md", "text": "https://wpnews.pro/news/k2-horizon-6-open-ai-models-with-full-training-data.txt", "jsonld": "https://wpnews.pro/news/k2-horizon-6-open-ai-models-with-full-training-data.jsonld"}}