{"slug": "synthesis-without-training-an-inference-only-pipeline-for-tabular-temporal-and", "title": "Synthesis Without Training: An Inference-Only Pipeline for Tabular, Temporal, and Relational Synthetic Data", "summary": "Researchers introduced GENSCRIPT, an inference-only synthetic data pipeline that eliminates model training by computing a deterministic statistical profile of source data and passing it to a language model to infer field semantics and integrity constraints, which a coding agent then compiles into an executable sampler. Across four single-table benchmarks, GENSCRIPT builds generators in 2 minutes and samples 50,000 rows within 6 seconds while staying within a few points of leading methods in marginal fidelity, and it is the only method that perfectly preserves a 1-to-1 column mapping in the Adult dataset. On a smart-building dataset it produced conditional time series closer to the real distribution than two baselines and perfectly preserved primary- and foreign-key relationships in the corresponding relational database.", "body_md": "arXiv:2609.38414v1 Announce Type: new \nAbstract: Synthetic data generation is dominated by the fit-then-sample paradigm: a generative model is trained on a private dataset and then sampled from. Despite its widespread adoption, this paradigm faces three challenges: (1) a new training run is required for every dataset; (2) different data modalities, such as single tables, time series, and relational databases, require task-specific models and feature engineering; and (3) the resulting model is opaque, making its behavior under data constraints difficult to inspect. We propose GENSCRIPT, an inference-only pipeline that eliminates model training. GENSCRIPT computes a deterministic statistical profile of the source data (column types, ranges, missingness, categories, correlations, etc.) and passes it--rather than raw rows--to a language model to infer field semantics and cross-column integrity constraints. A coding agent then compiles the profile and constraints into an executable, auditable sampler. This unified approach supports single-table, temporal, and relational data without task-specific modeling. Across four single-table benchmarks, GENSCRIPT builds generators in 2 minutes and samples 50k rows within 6 seconds, while remaining within a few points of leading methods in marginal fidelity. Notably, it is the only method that perfectly preserves a 1-to-1 mapping between columns in the Adult dataset. On a smart-building dataset, it produces conditional time series that more closely match the real distribution than two baselines and perfectly preserves primary- and foreign-key relationships in the corresponding relational database.", "url": "https://wpnews.pro/news/synthesis-without-training-an-inference-only-pipeline-for-tabular-temporal-and", "canonical_source": "https://arxiv.org/abs/2609.38414", "published_at": "2026-10-02 04:00:00+00:00", "updated_at": "2026-10-02 04:16:44.447772+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "ai-agents", "ai-research", "large-language-models"], "entities": ["GENSCRIPT", "Adult dataset"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/synthesis-without-training-an-inference-only-pipeline-for-tabular-temporal-and", "markdown": "https://wpnews.pro/news/synthesis-without-training-an-inference-only-pipeline-for-tabular-temporal-and.md", "text": "https://wpnews.pro/news/synthesis-without-training-an-inference-only-pipeline-for-tabular-temporal-and.txt", "jsonld": "https://wpnews.pro/news/synthesis-without-training-an-inference-only-pipeline-for-tabular-temporal-and.jsonld"}}