{"slug": "perfreasoning-how-well-do-llms-reason-on-hardware-performance", "title": "PerfReasoning: How Well Do LLMs Reason on Hardware Performance?", "summary": "Researchers introduced PerfReasoning, a benchmark evaluating large language models on hardware performance reasoning and analytical performance-model code generation. The strongest closed-source models exceed 90% on reasoning-based Q&A, while the best open-weight model reaches 82.4%, but model construction is harder: GPT-5.6 Sol exceeds 80% pass rate, whereas all other configurations average below 15%. Task-specific RL improved a 4B model's mapping-reasoning accuracy by 15.7 points, but feedback-free self-revision prompting proved unreliable.", "body_md": "arXiv:2609.04476v1 Announce Type: new \nAbstract: Performance modeling is central to hardware design and software optimization, yet constructing these models requires structured reasoning about computation, data reuse, storage, and movement. We introduce PerfReasoning, a benchmark that evaluates LLMs both as direct performance reasoners and as generators of analytical performance-model code. Given workload, architecture, and mapping specifications, models compare mappings and predict off-chip traffic and buffer requirements. The strongest closed-source models exceed 90% on reasoning-based Q&A, and the best open-weight model reaches 82.4%. However, model construction is substantially harder: while GPT-5.6 Sol exceeds 80% pass rate, all other model configurations average below 15% and vary markedly across runs. Task-specific RL raises a 4B model's mapping-reasoning accuracy by 15.7 points, whereas feedback-free multi-round self-revision prompting is not reliably effective. PerfReasoning exposes the gap between plausible architectural reasoning and reliable performance-model construction. We will publicly release the benchmark to support reproducible evaluation and track future progress.", "url": "https://wpnews.pro/news/perfreasoning-how-well-do-llms-reason-on-hardware-performance", "canonical_source": "https://arxiv.org/abs/2609.04476", "published_at": "2026-09-07 04:00:00+00:00", "updated_at": "2026-09-07 04:28:30.375871+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-research"], "entities": ["PerfReasoning", "GPT-5.6 Sol", "arXiv"], "alternates": {"html": "https://wpnews.pro/news/perfreasoning-how-well-do-llms-reason-on-hardware-performance", "markdown": "https://wpnews.pro/news/perfreasoning-how-well-do-llms-reason-on-hardware-performance.md", "text": "https://wpnews.pro/news/perfreasoning-how-well-do-llms-reason-on-hardware-performance.txt", "jsonld": "https://wpnews.pro/news/perfreasoning-how-well-do-llms-reason-on-hardware-performance.jsonld"}}