Xiaomi’s MiMo-V3 to adopt new architecture as HySparse2 cuts long-context costs Xiaomi MiMo lead Fuli Luo said the upcoming MiMo-V3 model will adopt a new architecture built on HySparse2, a technique that at a context length of 1 million tokens reduces prefill computation by 5.02 times and cuts the KV cache by 4.5 times. Xiaomi reported higher MRCRv2 and RULER-v2 scores along with lower AgentPPL and LongPPL results, positioning HySparse2 as a way to lower long-context inference costs and improve retrieval for agentic workloads. Xiaomi MiMo lead Fuli Luo said MiMo-V3 will adopt a new architecture. Its core technology, HySparse2, is designed to reduce the cost of long-context inference and improve retrieval for agentic workloads. At a context length of 1 million tokens, HySparse2 reduces prefill computation by 5.02 times and cuts the KV cache by 4.5 times. Xiaomi also reported higher MRCRv2 and RULER-v2 scores, along with lower AgentPPL and LongPPL results. arXiv https://arxiv.org/abs/2609.26368