{"slug": "cross-dialect-generalization-without-retraining-benchmarks-and-evaluation-of-for", "title": "Cross-Dialect Generalization Without Retraining: Benchmarks and Evaluation of Schema-Derived Constrained Decoding for MLIR", "summary": "Researchers released four natural-language-to-MLIR benchmarks totaling 410 in-scope pairs and a three-layer schema-derived constraint stack that lets SmolLM2-1.7B match or exceed 15B-34B code LMs on structurally dominated dialects without retraining. On the linalg dialect, SmolLM2 achieved 80.0% verify-valid, beating CodeLlama-34B, Granite-Code-34B, and StarCoder2-15B by 21-44 percentage points. The work, published on arXiv, demonstrates that inference-time priors from Operation Definition Specifications can substitute for gradient-based adaptation across MLIR dialects.", "body_md": "arXiv:2607.18254v1 Announce Type: new\nAbstract: Multi-Level Intermediate Representation (MLIR) underlies modern ML compiler infrastructure (TensorFlow, JAX/StableHLO, PyTorch Inductor, IREE), yet appears only in trace amounts in code-LM pretraining corpora. MLIR is also extensible by design: new dialects ship per application domain, so a fine-tuned model per dialect does not scale. We ask whether inference-time priors derived mechanically from each dialect's Operation Definition Specification (ODS) can substitute for gradient-based adaptation. First, we release four natural-language-to-MLIR benchmarks across three dialects - MLIR-Spec-150, Linalg-Spec-30, StableHLO-Spec-30, and StableHLO-Held-Out-200 - totaling 410 in-scope NL-to-MLIR pairs, plus a 25-program out-of-grammar stress set and a hand-authored n=30 functional reference set, shipped under Apache-2.0 with Gebru datasheets and Croissant 1.0 metadata. Second, we build a three-layer schema-derived constraint stack: a CFG over op signatures(C1), type-domain splits from an ODS-extracted type lattice (C2), and an SSA-scope validator driving five-retry rejection sampling (C3). Porting from arith+func+memref+linalg to StableHLO required no new constraint-layer code. On dialects whose verifier semantics are dominated by structural constraints, schema-derived priors let SmolLM2-1.7B match or exceed 15B-34B code LMs at 8-25x the per-generation speed: on linalg, SmolLM2 reaches 80.0% verify-valid (three-seed mean, n=125), beating CodeLlama-34B, Granite-Code-34B, and StarCoder2-15B by 21-44 percentage points with non-overlapping CIs. On arith+func and on the templated parametric StableHLO-Held-Out-200, where verifier semantics turn on attribute values rather than structure, the same baselines match or beat the SLM; we scope these as non-win cells. We release benchmarks, decoder, all per-prompt generations, and a reproducibility Docker image.", "url": "https://wpnews.pro/news/cross-dialect-generalization-without-retraining-benchmarks-and-evaluation-of-for", "canonical_source": "https://arxiv.org/abs/2607.18254", "published_at": "2026-07-22 04:00:00+00:00", "updated_at": "2026-07-22 04:12:18.291051+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "ai-research", "ai-tools", "developer-tools"], "entities": ["arXiv", "SmolLM2-1.7B", "CodeLlama-34B", "Granite-Code-34B", "StarCoder2-15B", "MLIR", "TensorFlow", "JAX"], "alternates": {"html": "https://wpnews.pro/news/cross-dialect-generalization-without-retraining-benchmarks-and-evaluation-of-for", "markdown": "https://wpnews.pro/news/cross-dialect-generalization-without-retraining-benchmarks-and-evaluation-of-for.md", "text": "https://wpnews.pro/news/cross-dialect-generalization-without-retraining-benchmarks-and-evaluation-of-for.txt", "jsonld": "https://wpnews.pro/news/cross-dialect-generalization-without-retraining-benchmarks-and-evaluation-of-for.jsonld"}}