{"slug": "dotime-a-synthetic-benchmark-generator-for-interventional-and-counterfactual", "title": "DoTime: A Synthetic Benchmark Generator for Interventional and Counterfactual Time Series", "summary": "Researchers introduced DoTime, an open-source synthetic benchmark generator for interventional and counterfactual time series, released as the dotime PyPI package with four frozen evaluation suites. It provides 100,000 trajectories and eight named identification structures with exact ground truth, and tests a falsifiable claim that interventional training yields a measurable direction-accuracy advantage over observational models, with positive gaps in all tested structures, trajectory lengths, and seeds.", "body_md": "arXiv:2607.27263v1 Announce Type: new\nAbstract: Most benchmarks for causal inference over time series are observational, small, or domain-specific, leaving interventional and counterfactual estimation under-served exactly where it matters most, such as in healthcare, policy evaluation, and climate science. We introduce \\textbf{DoTime}, an open, scalable, and theoretically grounded generator of multivariate temporal structural causal models (TSCMs) with interventions, released as the \\code{dotime} PyPI package together with four frozen evaluation suites. Beyond existing work, it adds capabilities absent from prior generators: continuous-time intervention \\emph{windows}, counterfactual sampling modes with a positivity guard, regime-switching SCMs as a strict generalization of interrupted time series, non-stationary dynamics by construction with switching SCM parameters, and deterministic ramp and sinusoidal intervention profiles that place trends and structural breaks \\emph{inside} the evaluation window. Moreover, it demonstrates the suitability of the generator as a prior for a causal foundation model reference implementation. The released suites span a training-scale snapshot of $100{,}000$ trajectories and eight named identification structures, each with exact ground truth: paired interventional trajectories from the same SCM throughout, and shared-noise counterfactuals in the continuous-time suite. We ship reference baseline implementations with an evaluation harness, and pose a falsifiable claim: interventional training buys a measurable direction-accuracy advantage over an observational model of identical capacity. It is tested across three training seeds per arm. Under structure-matched evaluation on held-out episodes, the interventional prior-fitted network's (PFN) gap is positive in every structure, trajectory length, and seed tested.", "url": "https://wpnews.pro/news/dotime-a-synthetic-benchmark-generator-for-interventional-and-counterfactual", "canonical_source": "https://arxiv.org/abs/2607.27263", "published_at": "2026-07-31 04:00:00+00:00", "updated_at": "2026-07-31 04:34:06.836652+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "ai-research", "ai-tools"], "entities": ["DoTime", "dotime", "PyPI"], "alternates": {"html": "https://wpnews.pro/news/dotime-a-synthetic-benchmark-generator-for-interventional-and-counterfactual", "markdown": "https://wpnews.pro/news/dotime-a-synthetic-benchmark-generator-for-interventional-and-counterfactual.md", "text": "https://wpnews.pro/news/dotime-a-synthetic-benchmark-generator-for-interventional-and-counterfactual.txt", "jsonld": "https://wpnews.pro/news/dotime-a-synthetic-benchmark-generator-for-interventional-and-counterfactual.jsonld"}}