Announcing Ling 3.0 Flash: Free on Kilo for a Limited Time Kilo announced that Ling 3.0 Flash, the latest model from inclusionAI (Ant Group), is now available on its platform and free for a limited time. The model features 124B total parameters with only ~5.1B activated per token, a 256K context window extendable to 1M, and combines fast general-purpose capabilities with deep reasoning from the Ring series. This release continues the industry trend toward efficient open-weight models for production-grade agentic tasks. Announcing Ling 3.0 Flash: Free on Kilo for a Limited Time Welcome to the flash model wars for intelligent efficiency We are excited to announce that Ling 3.0 Flash , the latest breakthrough from inclusionAI Ant Group , is now available on Kilo. To celebrate the rollout, developers can access and build with Ling 3.0 Flash 100% free for a limited time . Whether you are running multi-turn coding agents or building daily agentic workflows on tight token budgets, Ling 3.0 Flash brings high capabilities at lightning speed. Give it a spin today in the Kilo CLI and IDE extensions https://kilo.ai/ and help it rise up on the Kilo Leaderboard https://kilo.ai/leaderboard like previous Ring and Ling models have done The Rise of Highly Efficient Open-Weight Models The open-weight landscape has shifted dramatically. The industry is no longer solely chasing sheer parameter bloat or the most advanced reasoning on the planet. Instead, the focus has moved to architectural efficiency . I first highlighted this trend with a focus on Gemini and StepFun flash models https://blog.kilo.ai/p/the-age-of-the-flash-model-gemini back in May a year ago in AI time… , and this new Ling model continues the trend with a focus on autonomy and efficiency. Early testing has found the model remarkably capable of following a path of thought and taking action. Sparse Mixture-of-Experts MoE architectures and hybrid attention stacks have proven that you do not need to fire 70B+ dense parameters on every token to achieve top-tier performance. Hot on the heels of the latest model from Google DeepMind, Gemini 3.6 Flash https://kilo.ai/models/google-gemini-3-6-flash , Ling 3.0 Flash represents the pinnacle of this new wave: Massive Capacity, Minimal Footprint: Houses 124B total parameters while activating only ~5.1B parameters per token. Ultra-Low Latency: Optimized for high-throughput, token-efficient inference, drastically cutting down execution time and compute costs. Pretty Big Context: Features a native 256K context window extendable up to 1M , perfect for long-horizon task stability and deep repository reasoning. This rise of hyper-efficient open-weight models means developers can finally run production-grade agentic tasks without sacrificing speed or burning through API budgets. Standing on the Shoulders of Ling & Ring In the world of inclusionAI https://kilo.ai/models/by/inclusionai , Ling models serve as fast, general-purpose base models designed for high-throughput production, while Ring models are specialized, deep-reasoning “thinking” variant built on top of Ling for more complex workflows. But this new model seems to combine both families. Ling 3.0 Flash does not exist in isolation; it fuses the best innovations from inclusionAI’s established model families: The Ling Series e.g., Ling-2.6-1T, Ling-2.6-Flash : Known for non-thinking flagship throughput, rapid execution, and massive linear attention stacks built for ultra-fast context processing. The Ring Series e.g., Ring-2.6-1T : Ant Group’s dedicated reasoning line, engineered specifically for deep step-by-step logic and complex problem-solving. Ling 3.0 Flash bridges these two worlds. By combining the raw speed and low active-parameter overhead of the Flash line with a hybrid Reasoning mode inspired by the Ring series, Ling 3.0 Flash dynamically scales its thinking effort depending on task difficulty—delivering logic precision when needed without wasting tokens on simpler prompts. New Ways to Control Your AI Coding Spend Cost-efficiency is not just about using lighter models; it is also about having the right infrastructure controls in place. As our team wrote recently in More Ways to Control AI Coding Spend https://blog.kilo.ai/p/more-ways-to-control-ai-coding-spend , we have rolled out several new platform features to help you manage your API budgets effectively. The new Ling model is free for a limited time, but pricing is expected to be extremely affordable—and the model is designed to be used in tandem with heavyweights from frontier labs too. By taking advantage of custom usage limits, prompt caching, and intelligent model routing, you can confidently deploy autonomous agents and coding assistants at scale. Pairing these granular spend-control features with a free-to-use, ultra-efficient model like Ling 3.0 Flash ensures you maximize your development runway without unexpected bills. It’s a model can take you further by following its own logic across turns. Try Ling 3.0 Flash Today Take advantage of the limited-time FREE availability and test Ling 3.0 Flash on your toughest agentic workloads today, whether you’re coding for fun or driving a major startup. Start building: Access the model wherever you use Kilo Code https://kilo.ai Compare benchmarks: See where it ranks on the live Kilo AI Leaderboard https://kilo.ai/leaderboard And don’t forget to check out the latest open-weight releases on the Kilo Open-Weight Model Tracker https://kilo.ai/new-open-weight-models . InclusionAI hasn’t yet released the weights for this new release, but they are expected to release them at some point as they typically do on Hugging Face https://huggingface.co/inclusionAI .