{"slug": "cuda-harness-harnessing-agentic-cuda-kernel-generation-and-optimization-from", "title": "CUDA-Harness: Harnessing Agentic CUDA Kernel Generation and Optimization from Natural Language", "summary": "Researchers introduced CUDA-Harness, a framework for generating and optimizing CUDA kernels directly from natural language (Text2CUDA), addressing the expertise barrier in high-performance kernel development. The framework introduces Intermediate-Structured Generation to bridge semantic understanding with low-level kernel generation, Synthesis-Based Verification to prevent reward hacking, and Feedback-Adaptive Evolution to prioritize correctness while optimizing performance. Experiments demonstrate effectiveness across LLMs, hardware platforms, and C-to-CUDA transpilation.", "body_md": "arXiv:2609.00058v1 Announce Type: new\nAbstract: Developing high-performance CUDA kernels demands specialized knowledge in algorithm implementation, correctness validation, and hardware-aware parallel optimization, creating a substantial expertise barrier and making generating CUDA kernels directly from natural language (Text2CUDA) essential. Meanwhile, the general-purpose code generation capability of Large Language Models (LLMs) prompts a series of works exploring LLM-based CUDA kernel generation. They mainly focus on transpilation from high-level frameworks such as PyTorch to CUDA (Torch2CUDA) rather than Text2CUDA, where models must understand the high-level input semantics and handle low-level kernel implementation and validation. Additionally, these methods are vulnerable to reward hacking due to reliance on predefined test inputs. In this paper, we propose CUDA-Harness, a framework for harnessing agentic CUDA kernel generation and optimization from natural language. Specifically, we introduce Intermediate-Structured Generation to connect high-level semantic understanding with low-level kernel generation. To dilute reward hacking in Text2CUDA, we construct Synthesis-Based Verification to provide isolated test data and progressive validation. Furthermore, we propose Feedback-Adaptive Evolution, a kernel evolution strategy that prioritizes correctness while optimizing performance. Finally, through extensive experiments, we demonstrate the effectiveness of CUDA-Harness, with further evaluations illustrating generalization across LLMs, hardware platforms, and to C-to-CUDA transpilation.", "url": "https://wpnews.pro/news/cuda-harness-harnessing-agentic-cuda-kernel-generation-and-optimization-from", "canonical_source": "https://arxiv.org/abs/2609.00058", "published_at": "2026-09-02 04:00:00+00:00", "updated_at": "2026-09-02 04:25:49.946826+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-research", "ai-tools"], "entities": ["CUDA-Harness", "arXiv", "Large Language Models (LLMs)", "PyTorch"], "alternates": {"html": "https://wpnews.pro/news/cuda-harness-harnessing-agentic-cuda-kernel-generation-and-optimization-from", "markdown": "https://wpnews.pro/news/cuda-harness-harnessing-agentic-cuda-kernel-generation-and-optimization-from.md", "text": "https://wpnews.pro/news/cuda-harness-harnessing-agentic-cuda-kernel-generation-and-optimization-from.txt", "jsonld": "https://wpnews.pro/news/cuda-harness-harnessing-agentic-cuda-kernel-generation-and-optimization-from.jsonld"}}