{"slug": "announcing-ling-3-0-flash-free-on-kilo-for-a-limited-time", "title": "Announcing Ling 3.0 Flash: Free on Kilo for a Limited Time", "summary": "Kilo announced that Ling 3.0 Flash, the latest model from inclusionAI (Ant Group), is now available on its platform and free for a limited time. The model features 124B total parameters with only ~5.1B activated per token, a 256K context window extendable to 1M, and combines fast general-purpose capabilities with deep reasoning from the Ring series. This release continues the industry trend toward efficient open-weight models for production-grade agentic tasks.", "body_md": "# Announcing Ling 3.0 Flash: Free on Kilo for a Limited Time\n\n### Welcome to the flash model wars for intelligent efficiency\n\nWe are excited to announce that **Ling 3.0 Flash**, the latest breakthrough from inclusionAI (Ant Group), is now available on Kilo. To celebrate the rollout, developers can access and build with Ling 3.0 Flash **100% free for a limited time**.\n\nWhether you are running multi-turn coding agents or building daily agentic workflows on tight token budgets, Ling 3.0 Flash brings high capabilities at lightning speed. Give it a spin today in the [Kilo CLI and IDE extensions](https://kilo.ai/) and help it rise up on the [Kilo Leaderboard](https://kilo.ai/leaderboard) like previous Ring and Ling models have done!\n\n## The Rise of Highly Efficient Open-Weight Models\n\nThe open-weight landscape has shifted dramatically. The industry is no longer solely chasing sheer parameter bloat or the most advanced reasoning on the planet. Instead, the focus has moved to **architectural efficiency**.\n\nI first highlighted this trend with a focus on [Gemini and StepFun flash models](https://blog.kilo.ai/p/the-age-of-the-flash-model-gemini) back in May (a year ago in AI time…), and this new Ling model continues the trend with a focus on autonomy and efficiency. Early testing has found the model remarkably capable of following a path of thought and taking action.\n\nSparse Mixture-of-Experts (MoE) architectures and hybrid attention stacks have proven that you do not need to fire 70B+ dense parameters on every token to achieve top-tier performance. Hot on the heels of the latest model from Google DeepMind, [Gemini 3.6 Flash](https://kilo.ai/models/google-gemini-3-6-flash), **Ling 3.0 Flash** represents the pinnacle of this new wave:\n\n**Massive Capacity, Minimal Footprint:** Houses**124B total parameters** while activating only**~5.1B parameters** per token.**Ultra-Low Latency:** Optimized for high-throughput, token-efficient inference, drastically cutting down execution time and compute costs.**Pretty Big Context:** Features a native**256K context window**(extendable up to 1M), perfect for long-horizon task stability and deep repository reasoning.\n\nThis rise of hyper-efficient open-weight models means developers can finally run production-grade agentic tasks without sacrificing speed or burning through API budgets.\n\n## Standing on the Shoulders of Ling & Ring\n\nIn [the world of inclusionAI](https://kilo.ai/models/by/inclusionai), **Ling** models serve as fast, general-purpose base models designed for high-throughput production, while **Ring** models are specialized, deep-reasoning “thinking” variant built on top of Ling for more complex workflows. But this new model seems to combine both families.\n\nLing 3.0 Flash does not exist in isolation; it fuses the best innovations from inclusionAI’s established model families:\n\n**The Ling Series (e.g., Ling-2.6-1T, Ling-2.6-Flash):** Known for non-thinking flagship throughput, rapid execution, and massive linear attention stacks built for ultra-fast context processing.**The Ring Series (e.g., Ring-2.6-1T):** Ant Group’s dedicated reasoning line, engineered specifically for deep step-by-step logic and complex problem-solving.\n\n**Ling 3.0 Flash bridges these two worlds.** By combining the raw speed and low active-parameter overhead of the Flash line with a **hybrid Reasoning mode** inspired by the Ring series, Ling 3.0 Flash dynamically scales its thinking effort depending on task difficulty—delivering logic precision when needed without wasting tokens on simpler prompts.\n\n## New Ways to Control Your AI Coding Spend\n\nCost-efficiency is not just about using lighter models; it is also about having the right infrastructure controls in place. As our team wrote recently in [More Ways to Control AI Coding Spend](https://blog.kilo.ai/p/more-ways-to-control-ai-coding-spend), we have rolled out several new platform features to help you manage your API budgets effectively. The new Ling model is free for a limited time, but pricing is expected to be extremely affordable—and the model is designed to be used in tandem with heavyweights from frontier labs too.\n\nBy taking advantage of custom usage limits, prompt caching, and intelligent model routing, you can confidently deploy autonomous agents and coding assistants at scale. Pairing these granular spend-control features with a free-to-use, ultra-efficient model like Ling 3.0 Flash ensures you maximize your development runway without unexpected bills. It’s a model can take you further by following its own logic across turns.\n\n## Try Ling 3.0 Flash Today\n\nTake advantage of the limited-time FREE availability and test Ling 3.0 Flash on your toughest agentic workloads today, whether you’re coding for fun or driving a major startup.\n\n**Start building:** Access the model[wherever you use Kilo Code](https://kilo.ai)**Compare benchmarks:** See where it ranks on the live[Kilo AI Leaderboard](https://kilo.ai/leaderboard)\n\nAnd don’t forget to check out the latest open-weight releases on the [Kilo Open-Weight Model Tracker](https://kilo.ai/new-open-weight-models). InclusionAI hasn’t yet released the weights for this new release, but they are expected to release them at some point as they typically do on [Hugging Face](https://huggingface.co/inclusionAI).", "url": "https://wpnews.pro/news/announcing-ling-3-0-flash-free-on-kilo-for-a-limited-time", "canonical_source": "https://blog.kilo.ai/p/announcing-ling-30-flash-free-on", "published_at": "2026-07-23 20:54:16+00:00", "updated_at": "2026-07-23 20:57:27.014238+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-products", "ai-tools", "ai-infrastructure"], "entities": ["Kilo", "inclusionAI", "Ant Group", "Ling 3.0 Flash", "Google DeepMind", "Gemini 3.6 Flash", "Ling-2.6-1T", "Ring-2.6-1T"], "alternates": {"html": "https://wpnews.pro/news/announcing-ling-3-0-flash-free-on-kilo-for-a-limited-time", "markdown": "https://wpnews.pro/news/announcing-ling-3-0-flash-free-on-kilo-for-a-limited-time.md", "text": "https://wpnews.pro/news/announcing-ling-3-0-flash-free-on-kilo-for-a-limited-time.txt", "jsonld": "https://wpnews.pro/news/announcing-ling-3-0-flash-free-on-kilo-for-a-limited-time.jsonld"}}