MiMo-V2.6 Pro Architecture and Training Notes
Xiaomi's MiMo-V2.6 Pro ranks No. 1 on open-weight benchmarks by weighted average despite using a classic Grouped Query Attention architecture with Sliding Window Attention at a 128-token window, accor…
Xiaomi's MiMo-V2.6 Pro ranks No. 1 on open-weight benchmarks by weighted average despite using a classic Grouped Query Attention architecture with Sliding Window Attention at a 128-token window, accor…
The MiMo-V2.5 model family introduces Hybrid Sliding Window Attention (Hybrid SWA) to reduce KVCache storage to roughly 1/7 of Full Attention, sparse MoE activation to cut per-token compute, and multi…
JetBrains released Mellum2, a 12-billion-parameter Mixture-of-Experts model with 2.5 billion active parameters per token, under the Apache 2.0 license. The model is specialized for software engineerin…
DeepSeek-V4's million-token context capability stems from a hybrid attention architecture that compresses context before KV storage, reducing cache pressure. Together's early bring-up on NVIDIA HGX B2…