{"slug": "datakernelbench-can-llms-optimize-database-queries-on-gpus", "title": "DataKernelBench: Can LLMs Optimize Database Queries on GPUs?", "summary": "A new benchmark, DataKernelBench, evaluates whether large language models can optimize database queries on GPUs, achieving up to 2.11x speedup over torch.compile on TPC-H SF10 with an H100 GPU. The benchmark translates SQL into PyTorch TorchPlan programs and tests ten proprietary and open-weight models, finding that stronger models benefit most from full-query specialization and that kernel fusion is key. On TPC-H SF100 with four H100 GPUs, the approach achieves 2.54x speedup using Dask-cuDF for on-demand partition loading.", "body_md": "arXiv:2608.25061v1 Announce Type: new\nAbstract: GPUs increasingly accelerate database systems, but query-specific peak performance still often relies on hand-written kernels. Existing LLM kernel benchmarks focus on machine learning operators, leaving irregular, heterogeneous, data-movement-heavy database-style operators untested. We introduce DataKernelBench, which translates SQL into validated PyTorch TorchPlan programs and evaluates LLMs that optimize either the core tensor-bounded snippet or the full query in CUDA or Triton through execution-guided repair. Across ten proprietary and open-weight models on TPC-H SF10 with an H100 GPU, the strongest full-query CUDA configuration achieves $2.11\\times$ speedup over torch.compile at full pass rate. We find that higher-performing implementations commonly use kernel fusion and execution-strategy changes, stronger models benefit most from full-query specialization, and workload context matters more than hardware context. To handle data larger than GPU memory, we extend TorchPlan with Dask-cuDF for on-demand partition loading on TPC-H SF100 with four H100 GPUs, achieving $2.54\\times$ speedup", "url": "https://wpnews.pro/news/datakernelbench-can-llms-optimize-database-queries-on-gpus", "canonical_source": "https://arxiv.org/abs/2608.25061", "published_at": "2026-08-27 04:00:00+00:00", "updated_at": "2026-08-27 04:20:00.255132+00:00", "lang": "en", "topics": ["large-language-models", "ai-research", "ai-infrastructure"], "entities": ["DataKernelBench", "PyTorch", "TorchPlan", "CUDA", "Triton", "TPC-H", "H100", "Dask-cuDF"], "alternates": {"html": "https://wpnews.pro/news/datakernelbench-can-llms-optimize-database-queries-on-gpus", "markdown": "https://wpnews.pro/news/datakernelbench-can-llms-optimize-database-queries-on-gpus.md", "text": "https://wpnews.pro/news/datakernelbench-can-llms-optimize-database-queries-on-gpus.txt", "jsonld": "https://wpnews.pro/news/datakernelbench-can-llms-optimize-database-queries-on-gpus.jsonld"}}