13:42
2026-08-12
aws.amazon.com
artificial-intelligence
Tiered KV cache for large LLMs on Amazon SageMaker HyperPod with Curvine
Amazon Web Services (AWS) introduced a tiered KV cache architecture on Amazon SageMaker HyperPod that extends the cache hierarchy beyond GPU and CPU memory into a shared, distributed NVMe pool using Cโฆ