15:30
2026-07-27
systems.seas.harvard.edu
large-language-models
MorphServe: Making Model Precision Elastic for Bursty LLM Serving
MorphServe, a system developed by researchers and detailed in a paper on arXiv, enables elastic model precision for bursty LLM serving by temporarily swapping low-impact FP16 layers to INT4 quantizatiβ¦