{"slug": "show-hn-linearsolvebench-interesting-new-benchmark-to-discover-linear-solvers", "title": "Show HN: LinearSolveBench, interesting new benchmark to discover linear solvers", "summary": "LinearSolveBench launched as a benchmark measuring how well AI models and harnesses write fast, accurate, general numerical solvers for large sparse linear systems in C. On the FLASH magnetic-diffusion leaderboard, Fable 5.1 with autodidakt and high reasoning ranked first at 1.546× speedup versus a GMRES(50) + BoomerAMG reference solver, ahead of GPT-5.6 Sol at 1.233× and GPT-6 Astra at 1.184×. The benchmark's stated goals are to encourage algorithmic advances in solving large sparse linear systems and to measure and improve AI models' ability to discover algorithms, with problem families drawn from the SuiteSparse Matrix Collection spanning power grid optimization, fusion power, quantitative finance, structural mechanics, and fluid dynamics.", "body_md": "# LinearSolveBench\n\nA benchmark for AI-discovered linear solvers.\n\nMeasuring the ability of AI models and harnesses to write fast, accurate, and general numerical solvers for large sparse linear systems in C.\n\n## Leaderboard\n\n### FLASH magnetic diffusion\n\nLinear systems captured from FLASH magnetic-diffusion solves. 9,984–113,664 rows across eight distinct sizes; 108,364–1,245,184 nonzeros per matrix.\n\n| FLASH magnetic diffusion: independent model rows show best-of-16 speedups; autodidakt, Codex, and Claude Code rows show overall best run speedups reported in the FLASH research article. Search budgets differ. Higher scores rank first; pending results are unranked. |  |  | \n|---|---|---|\n| Rank | Model | Speedup vs. GMRES(50) + BoomerAMG reference solver | \n|---|---|---|\n| 1 | Fable 5.1autodidakt · High reasoning | 1.546× | \n| 2 | GPT-5.6 Solautodidakt · High reasoning | 1.233× | \n| 3 | GPT-6 Astraautodidakt · High reasoning | 1.184× | \n| 4 | GPT-5.6 SolCodex · High reasoning | 1.153× | \n| 5 | GPT-6 AstraBest of 16 · High reasoning | 1.136× | \n| 6 | Fable 5.1Best of 16 · High reasoning | 1.132× | \n| 7 | Fable 5.1Claude Code · High reasoning | 1.106× | \n| 8 | GPT-5.6 SolBest of 16 · High reasoning | 1.063× | \n\nLarge systems of linear equations underpin some of the most important scientific and engineering problems. Progress in solving them translates directly into advances in the fields that depend on them.\n\nLinearSolveBench has two goals: encourage algorithmic advances in solving large sparse linear systems, and measure and improve the ability of AI models and systems to discover algorithms.\n\nApplications span multiple domains, with problem families represented in the [SuiteSparse Matrix Collection](https://sparse.tamu.edu/):\n\n1. Power grid optimization\n  - Optimal Power Flow (OPF).\n  - Unit Commitment (UC), including LP/MILP relaxations and decomposition methods.\n2. Fusion power and diffusive processes\n  - Heat conduction and magnetic diffusion in magnetic and magneto-inertial confinement fusion devices.\n  - Radiation transport and alpha energy diffusion.\n3. Quantitative finance\n  - Multi-asset Black–Scholes, Heston, and Heston–Hull–White pricing PDEs, whose implicit discretizations require large sparse linear solves.\n  - American-option pricing, where free-boundary methods repeatedly solve sparse linear systems.\n4. Structural mechanics and materials engineering\n  - Finite-element models of buildings, bridges, aircraft, plates, shells, and mechanical components.\n  - Buckling, fracture, contact, plasticity, and composite or porous-material simulations.\n5. Fluid dynamics, transport, and thermal science\n  - Navier–Stokes and Stokes flow, driven-cavity and coating-flow models, and atmospheric or ocean simulations.\n  - Heat transfer, combustion, diffusion, and coupled multiphysics discretizations.\n\nThese applications do not all produce the same kind of matrices. PDE and finite-element problems may be symmetric positive definite, symmetric indefinite, or unsymmetric. Flow, circuit, nonlinear-Jacobian, and KKT systems are often unsymmetric or saddle-point structured. Least-squares and linear-programming data may be rectangular, while network data may encode a graph rather than an equation system.\n\nDiscovering new algorithms for these problems requires both human and computational resources. Large language models open the possibility of pursuing that discovery at scale.", "url": "https://wpnews.pro/news/show-hn-linearsolvebench-interesting-new-benchmark-to-discover-linear-solvers", "canonical_source": "https://www.autodidakt.ai/linear-solve-bench", "published_at": "2026-09-22 21:48:42+00:00", "updated_at": "2026-09-22 21:53:39.669288+00:00", "lang": "en", "topics": ["ai-research", "machine-learning", "artificial-intelligence", "ai-tools"], "entities": ["LinearSolveBench", "FLASH", "Fable 5.1", "GPT-5.6 Sol", "GPT-6 Astra", "Codex", "Claude Code", "SuiteSparse Matrix Collection"], "alternates": {"html": "https://wpnews.pro/news/show-hn-linearsolvebench-interesting-new-benchmark-to-discover-linear-solvers", "markdown": "https://wpnews.pro/news/show-hn-linearsolvebench-interesting-new-benchmark-to-discover-linear-solvers.md", "text": "https://wpnews.pro/news/show-hn-linearsolvebench-interesting-new-benchmark-to-discover-linear-solvers.txt", "jsonld": "https://wpnews.pro/news/show-hn-linearsolvebench-interesting-new-benchmark-to-discover-linear-solvers.jsonld"}}