MoME: Mixture-of-Memory Embeddings for Context-Aware Sparse Lookup Researchers introduced MoME, a Mixture-of-Memory Embeddings method for context-aware sparse lookup that addresses a limitation in existing memory-embedding approaches, which retrieve via a deterministic mechanism. The work targets efficient scaling of large language models by combining sparse capacity mechanisms such as Mixture-of-Experts with token-indexed embedding tables that augment the backbone through cheap parametric lookups. Scaling large language models efficiently has motivated sparse capacity mechanisms such as Mixture-of-Experts and, more recently, conditional memory: token-indexed embedding tables that augment the backbone with cheap parametric lookups. Existing memory-embedding methods retrieve via a deterministic