{"slug": "beyond-distribution-matching-semantics-consistent-tabular-diffusion-with-weak", "title": "Beyond Distribution Matching: Semantics-Consistent Tabular Diffusion with Weak Semantic Priors", "summary": "Researchers proposed a semantics-consistent tabular diffusion framework, described in arXiv paper 2609.16069v1, that uses large language models to extract intra-column semantics and inter-column symbolic rules from metadata and validates them on the real training split, then applies those priors as generation conditions rather than post-hoc filters. Across six real-world tabular benchmarks, the method consistently improved distributional fidelity, semantic consistency, and downstream task utility over representative VAE-, GAN-, LLM-, and diffusion-based baselines, and remained robust when semantic priors were partially unavailable.", "body_md": "arXiv:2609.16069v1 Announce Type: new \nAbstract: Synthetic tabular data can match real data distributions while still violating the semantic constraints that govern valid tabular rows. This reveals a key limitation of existing tabular generators: they mainly optimize distributional fidelity, but do not explicitly model weak semantic priors encoded in tabular schema and textual descriptions. In this paper, we propose \\ours, a semantics-consistent tabular diffusion framework for high-fidelity synthetic data generation under weakly specified semantic priors. \\ours\\ first constructs two types of priors, namely intra-column semantics and inter-column symbolic rules, with LLM-assisted extraction from metadata and validation on the real training split. These priors are then used as generation conditions rather than post-hoc filters. Specifically, \\ours\\ maps heterogeneous column values, column identities, and semantic priors into a unified semantic space, and performs column-wise forward corruption and prior-conditioned reverse denoising to preserve both marginal distributions and rule-consistent cross-column dependencies. Extensive experiments on six real-world tabular benchmarks show that \\ours\\ consistently improves distributional fidelity, semantic consistency, and downstream task utility over representative VAE-, GAN-, LLM-, and diffusion-based baselines. Additional analyses further demonstrate the robustness of \\ours\\ when semantic priors are partially unavailable.", "url": "https://wpnews.pro/news/beyond-distribution-matching-semantics-consistent-tabular-diffusion-with-weak", "canonical_source": "https://arxiv.org/abs/2609.16069", "published_at": "2026-09-16 04:00:00+00:00", "updated_at": "2026-09-16 04:07:02.180217+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "generative-ai", "large-language-models", "ai-research"], "entities": ["arXiv", "2609.16069v1"], "alternates": {"html": "https://wpnews.pro/news/beyond-distribution-matching-semantics-consistent-tabular-diffusion-with-weak", "markdown": "https://wpnews.pro/news/beyond-distribution-matching-semantics-consistent-tabular-diffusion-with-weak.md", "text": "https://wpnews.pro/news/beyond-distribution-matching-semantics-consistent-tabular-diffusion-with-weak.txt", "jsonld": "https://wpnews.pro/news/beyond-distribution-matching-semantics-consistent-tabular-diffusion-with-weak.jsonld"}}