What Is Index Share? How GLM 5.2 Achieves 2.9x Fewer Compute Operations at 1M Token Context
Zhipu AI's GLM 5.2 uses Index Share, a technique that reuses sparse attention indexers across four layers, reducing compute operations by 2.9x at 1M token context. This optimization makes long-context inference more affo…