Computational models of pragmatic reasoning with flexible generation of meaning and expression alternatives Researchers propose SAGE (ScAffolded Generative models for Explanation), a neuro-symbolic framework combining language models with cognitive models for pragmatic reasoning, achieving high accuracy across three case studies including referential expression generation and implicatures. The framework uses LM proposers to generate alternatives and LM evaluators to assess them, but component-level analyses reveal that LM proposers reliably generate alternatives while LM evaluators are better at intuitive judgments than theoretical measures. arXiv:2607.18443v1 Announce Type: new Abstract: Pragmatic language use requires reasoning about alternatives: the alternative expressions a speaker might have chosen, or the alternative interpretations a listener might entertain. Formal and computational models of pragmatics must therefore specify the sets of alternatives that interlocutors reason over, which is often done through manual specification. Here we propose a framework, ScAffolded Generative models for Explanation SAGE , that combines the explanatory transparency of cognitive models with the generative flexibility of language models LMs . SAGE decomposes a pragmatic process into three kinds of modules: proposers, which use LMs to generate an open-ended space of candidate alternatives; evaluators, which assess those alternatives e.g., their semantics, complexity, or typicality ; and selectors, which implement the rule-based computational steps of a cognitively motivated task analysis. We assess SAGE in three case studies spanning pragmatic generation and interpretation-referential expression generation, manner M- implicatures, and Gricean conversational implicatures. SAGE models are evaluated critically using established methods from computational cognitive modeling, including ablations, baseline comparisons, and quantitative fit to human data. Across studies, SAGE models achieved high accuracy and often outperformed baselines, but component-level analyses reveal an asymmetry: LM proposers reliably generated alternatives well-suited to pragmatic modeling, whereas LM evaluators are better at providing intuitive judgements rather than judgements of theoretical or formal measures. We discuss the promise and the limitations of neuro-symbolic models as candidate explanatory accounts of human pragmatic language use.