21:58
2026-09-10
aws.amazon.com
large-language-models
Reduce LLM latency with prefix-aware routing on Amazon SageMaker Inference
Amazon SageMaker Inference launched prefix-aware routing, a routing strategy that directs requests sharing the same prompt prefix to the same instance so cached key-value pairs are reused. Benchmarks …