Disaggregated prefill and decode for LLM inference on SageMaker HyperPod
AWS introduced disaggregated prefill and decode (DPD) for LLM inference on SageMaker HyperPod, separating compute-bound prefill and memory-bound decode onto different GPU pools connected via EFA RDMA …