22:41
2026-09-26
modal.com
ai-infrastructure
>1B tokens/minute/GPU by combining query planner and inference engine
Quail, an inference engine built by the QUery-Aware Inference Layer project, processes over 1 billion tokens per minute per H100 GPU on a multi-join AI-SQL query, more than 10x faster than the team's …