{"slug": "carat-do-materials-llms-reason-or-recite", "title": "CARAT: Do Materials LLMs Reason or Recite?", "summary": "The CARAT benchmark, detailed in arXiv paper 2609.38340v1, finds that materials LLMs often recite structural relations rather than reason from them: a frozen model quoted a link yet answered identically when the link was redirected in 95.6% of paired cases. Grounded structural views beat formula inputs by 17.3 points on the hardest benchmark families, but GraphSpace's 19.3-point edge over a plain periodic graph splits into 1.96 points where the plain rendering already carries the needed fields and 46.7 points where it omits them. After matched supervision the model reached 99.8% on the paired cases, and deleting the link dropped it to 23.4%, below the 27.0% best shortcut, showing both steps are learnable.", "body_md": "arXiv:2609.38340v1 Announce Type: new \nAbstract: When a materials LLM answers a question about crystal structure, does it reason from the structure or copy an answer already printed in its input? Accuracy cannot tell: a structural description often prints the very field it is scored against. CARAT holds question and gold answer fixed across eight matched views, names each structural relation separately in GraphSpace, and adds matched fine-tuning, answer masking, evidence injection, paired inference, and a rule that can withhold claims. First, on the benchmark's hardest families the grounded view is worth 17.3 points over formula inputs. Second, we turn that scrutiny on ourselves. GraphSpace beats a plain periodic graph by 19.3 points, but that margin is two effects at once: where the plain rendering carries everything the question needs it is 1.96 points, and where it omits those fields entirely, 46.7 points. The headline mostly measures what the baseline lacked, not how evidence is presented. Third, we attack our own benchmark. A rule that skips the link and reads the list directly answers four of seven hardened families, so we rebuilt it until eleven such shortcuts sat near chance. The frozen model quotes that link yet answers the same when we redirect it, on 95.6% of paired cases: it repeats the relation without using it. After matched supervision it reaches 99.8%, and deleting the link drops it to 23.4%, below the 27.0% the best shortcut reaches: both steps are learnable.", "url": "https://wpnews.pro/news/carat-do-materials-llms-reason-or-recite", "canonical_source": "https://arxiv.org/abs/2609.38340", "published_at": "2026-10-01 04:00:00+00:00", "updated_at": "2026-10-01 04:17:26.925673+00:00", "lang": "en", "topics": ["ai-research", "large-language-models", "machine-learning", "artificial-intelligence"], "entities": ["CARAT", "GraphSpace", "arXiv"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/carat-do-materials-llms-reason-or-recite", "markdown": "https://wpnews.pro/news/carat-do-materials-llms-reason-or-recite.md", "text": "https://wpnews.pro/news/carat-do-materials-llms-reason-or-recite.txt", "jsonld": "https://wpnews.pro/news/carat-do-materials-llms-reason-or-recite.jsonld"}}