vLLM: An Efficient Inference Engine for Large Language Models [pdf]
Researchers have released vLLM, a new inference engine designed to efficiently serve large language models by optimizing memory management and batching. The system achieves up to 24x higher throughput than existing solut…