04:00
2026-09-07
machinebrief.com
large-language-models
Compact-Memory LLM Agents via Online Max-Member Clustering and Atom-Aware Packing
Researchers introduced RSM-full, an online clustered-memory pipeline for long-horizon LLM deployments, achieving 83% of Full-Context quality at 32% of the token cost on the AMA-Bench benchmark at a 4k…