04:00
2026-09-17
arxiv.org
large-language-models
The Inference Engineering Pareto Atlas: Which Optimizations Dominate the Cost, Quality, and Latency Frontier?
A study measuring 54 configurations of Qwen2.5-7B-Instruct on vLLM 0.12 across L4, A100, and H100 GPUs found that 18 of 36 configurations reach the cost, quality, and latency Pareto frontier, with com…