Xiaomi MiMo lead Fuli Luo said MiMo-V3 will adopt a new architecture. Its core technology, HySparse2, is designed to reduce the cost of long-context inference and improve retrieval for agentic workloads.
At a context length of 1 million tokens, HySparse2 reduces prefill computation by 5.02 times and cuts the KV cache by 4.5 times. Xiaomi also reported higher MRCRv2 and RULER-v2 scores, along with lower AgentPPL and LongPPL results. [arXiv]