Zhipu's Mythos benchmark leak suggests GLM-4.
Zhipu AI's leaked Mythos benchmark suggests its upcoming GLM-4 model scores 87.2% on MMLU-Pro, 78.5% on GPQA-Diamond, and 99.1% on a 128k-context retrieval test, potentially outperforming GPT-4o on MM…
Zhipu AI's leaked Mythos benchmark suggests its upcoming GLM-4 model scores 87.2% on MMLU-Pro, 78.5% on GPQA-Diamond, and 99.1% on a 128k-context retrieval test, potentially outperforming GPT-4o on MM…
Moonshot AI's Kimi K3 technical report details a 2.8-trillion-parameter model with 104 billion activated parameters, a 1-million-token context window, and native vision, achieving 2.5x scaling efficie…
Chinese AI startup DeepSeek has halted its second fundraising round, Bloomberg News reported, citing people familiar with the matter. The decision pauses what was expected to be a significant capital …
**Summary:** Multi-Head Latent Attention (MLA) is an attention mechanism used in DeepSeek-V2/V3 and Kimi K2.x models that compresses the Key-Value (KV) cache by projecting full KV pairs into a shared,…