cd /news/artificial-intelligence/perfreasoning-how-well-do-llms-reaso… · home topics artificial-intelligence article
[ARTICLE · art-121919] src=arxiv.org ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

PerfReasoning: How Well Do LLMs Reason on Hardware Performance?

Researchers introduced PerfReasoning, a benchmark evaluating large language models on hardware performance reasoning and analytical performance-model code generation. The strongest closed-source models exceed 90% on reasoning-based Q&A, while the best open-weight model reaches 82.4%, but model construction is harder: GPT-5.6 Sol exceeds 80% pass rate, whereas all other configurations average below 15%. Task-specific RL improved a 4B model's mapping-reasoning accuracy by 15.7 points, but feedback-free self-revision prompting proved unreliable.

read1 min views2 publishedSep 7, 2026

arXiv:2609.04476v1 Announce Type: new Abstract: Performance modeling is central to hardware design and software optimization, yet constructing these models requires structured reasoning about computation, data reuse, storage, and movement. We introduce PerfReasoning, a benchmark that evaluates LLMs both as direct performance reasoners and as generators of analytical performance-model code. Given workload, architecture, and mapping specifications, models compare mappings and predict off-chip traffic and buffer requirements. The strongest closed-source models exceed 90% on reasoning-based Q&A, and the best open-weight model reaches 82.4%. However, model construction is substantially harder: while GPT-5.6 Sol exceeds 80% pass rate, all other model configurations average below 15% and vary markedly across runs. Task-specific RL raises a 4B model's mapping-reasoning accuracy by 15.7 points, whereas feedback-free multi-round self-revision prompting is not reliably effective. PerfReasoning exposes the gap between plausible architectural reasoning and reliable performance-model construction. We will publicly release the benchmark to support reproducible evaluation and track future progress.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @perfreasoning 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/perfreasoning-how-we…] indexed:0 read:1min 2026-09-07 ·