Show HN: I benchmarked LLMs on predicting knife steel properties A new open benchmark from the Steel-predictor project shows that Anthropic's Claude Sonnet 5 leads large language models in predicting knife-steel properties from chemical composition, achieving a mean Spearman rank correlation of 0.869 against laboratory measurements, while a reference machine-learning model trained on the same data scored 0.969. The benchmark, which tested 48 steels for edge retention via CATRA testing and 12 for toughness via Charpy impact energy, found LLMs rank edge retention well (ρ 0.85–0.92) but struggle with toughness (ρ 0.38–0.84). How well can large language models predict knife-steel properties from chemical composition alone? An open, reproducible benchmark that gives an LLM only a steel's composition e.g. C=1.45%, Cr=20%, V=4%, Mo=1%, powder-metallurgy: yes and asks it to rate two properties on a 1–10 scale, then scores those ratings against objective laboratory measurements : Edge retention ← CATRA standardized machine-cutting test total card stock cut, mm — 48 steels Toughness ← Charpy impact energy ft-lbs — 12 steels Scoring is scale-free rank correlation + pairwise ranking accuracy , so a model is judged purely on whether it orders steels correctly, not on how it calibrates the 1–10 scale. 📊 Site with charts & analysis: https://steel-predictor-project.github.io/steel-llm-eval/ https://steel-predictor-project.github.io/steel-llm-eval/ · Docs: methodology /Steel-predictor-project/steel-llm-eval/blob/main/docs/methodology.md · results & analysis /Steel-predictor-project/steel-llm-eval/blob/main/docs/results-analysis.md Ranked by mean Spearman rank correlation ρ vs. the measurements. Higher is better; 1.0 = perfect ordering, 0.0 = random. | Model | Edge ρ n | Edge pairwise | Tough ρ n | Tough pairwise | Mean ρ | |---|---|---|---|---|---| | steel-predictor reference ML † | 0.992 48 | 0.98 | 0.946 12 | 0.938 | 0.969 | | anthropic/claude-sonnet-5 | 0.894 48 | 0.918 | 0.844 12 | 0.881 | 0.869 | | google/gemini-3.6-flash | 0.918 48 | 0.913 | 0.698 12 | 0.797 | 0.808 | | openai/gpt-4o | 0.868 48 | 0.907 | 0.600 12 | 0.746 | 0.734 | | meta-llama/llama-3.3-70b-instruct | 0.864 47 | 0.964 | 0.514 12 | 0.780 | 0.689 | | deepseek/deepseek-chat-v3.1 | 0.869 48 | 0.910 | 0.380 12 | 0.661 | 0.625 | | openai/gpt-4o-mini | 0.850 48 | 0.984 | 0.385 12 | 0.689 | 0.617 | Edge retention n=48 CATRA , toughness n=12 Charpy . Zero-shot, temperature 0, one sample per steel. † Important fairness caveat: the reference ML model Steel-predictor https://github.com/Steel-predictor-project/Steel-predictor was trained on these same CATRA/Charpy measurements , so its scores here are largely in-sample and are shown as an upper-reference bar, not as a fair head-to-head with the zero-shot LLMs. The model's honest out-of-sample performance is its LOOCV MAE 0.391 , reported in that repo. The LLMs, by contrast, have never seen this labeled set. LLMs are genuinely good at ranking edge retention ρ ≈ 0.85–0.92 . Wear resistance is strongly and legibly encoded in composition carbide-forming elements — C, V, Cr, W, Mo , and frontier models clearly "know" that chemistry. Toughness is where they struggle ρ 0.38–0.84 . It depends on subtler factors carbide size/distribution, powder-metallurgy processing, matrix state that aren't obvious from a composition string, and the spread across models is large. Frontier small. Claude Sonnet and Gemini lead; the smaller/cheaper models drop off sharply on toughness while staying competitive on edge retention. git clone https://github.com/Steel-predictor-project/steel-llm-eval.git cd steel-llm-eval export OPENROUTER API KEY=sk-or-... one key → OpenAI, Anthropic, Google, Meta, DeepSeek, ... ./run benchmark.sh runs every model and rebuilds the leaderboard Run a single model, or a quick offline sanity check with no API key: python harness/run eval.py --model anthropic/claude-sonnet-5 python harness/run eval.py --provider mock deterministic heuristic, no key needed python harness/score.py Raw per-steel responses are written to results/raw