cd /news/ai-research/carat-do-materials-llms-reason-or-re… · home › topics › ai-research › article
[ARTICLE · art-142970] src=arxiv.org ↗ pub= topic=ai-research verified=true sentiment=· neutral

CARAT: Do Materials LLMs Reason or Recite?

The CARAT benchmark, detailed in arXiv paper 2609.38340v1, finds that materials LLMs often recite structural relations rather than reason from them: a frozen model quoted a link yet answered identically when the link was redirected in 95.6% of paired cases. Grounded structural views beat formula inputs by 17.3 points on the hardest benchmark families, but GraphSpace's 19.3-point edge over a plain periodic graph splits into 1.96 points where the plain rendering already carries the needed fields and 46.7 points where it omits them. After matched supervision the model reached 99.8% on the paired cases, and deleting the link dropped it to 23.4%, below the 27.0% best shortcut, showing both steps are learnable.

by read1 min views1 publishedOct 1, 2026

arXiv:2609.38340v1 Announce Type: new Abstract: When a materials LLM answers a question about crystal structure, does it reason from the structure or copy an answer already printed in its input? Accuracy cannot tell: a structural description often prints the very field it is scored against. CARAT holds question and gold answer fixed across eight matched views, names each structural relation separately in GraphSpace, and adds matched fine-tuning, answer masking, evidence injection, paired inference, and a rule that can withhold claims. First, on the benchmark's hardest families the grounded view is worth 17.3 points over formula inputs. Second, we turn that scrutiny on ourselves. GraphSpace beats a plain periodic graph by 19.3 points, but that margin is two effects at once: where the plain rendering carries everything the question needs it is 1.96 points, and where it omits those fields entirely, 46.7 points. The headline mostly measures what the baseline lacked, not how evidence is presented. Third, we attack our own benchmark. A rule that skips the link and reads the list directly answers four of seven hardened families, so we rebuilt it until eleven such shortcuts sat near chance. The frozen model quotes that link yet answers the same when we redirect it, on 95.6% of paired cases: it repeats the relation without using it. After matched supervision it reaches 99.8%, and deleting the link drops it to 23.4%, below the 27.0% the best shortcut reaches: both steps are learnable.

── more in #ai-research 4 stories · sorted by recency
── more on @carat 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/carat-do-materials-l…] indexed:0 read:1min 2026-10-01 · —