cd /news/artificial-intelligence/dotime-a-synthetic-benchmark-generat… · home topics artificial-intelligence article
[ARTICLE · art-81316] src=arxiv.org ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

DoTime: A Synthetic Benchmark Generator for Interventional and Counterfactual Time Series

Researchers introduced DoTime, an open-source synthetic benchmark generator for interventional and counterfactual time series, released as the dotime PyPI package with four frozen evaluation suites. It provides 100,000 trajectories and eight named identification structures with exact ground truth, and tests a falsifiable claim that interventional training yields a measurable direction-accuracy advantage over observational models, with positive gaps in all tested structures, trajectory lengths, and seeds.

read1 min views1 publishedJul 31, 2026

arXiv:2607.27263v1 Announce Type: new Abstract: Most benchmarks for causal inference over time series are observational, small, or domain-specific, leaving interventional and counterfactual estimation under-served exactly where it matters most, such as in healthcare, policy evaluation, and climate science. We introduce \textbf{DoTime}, an open, scalable, and theoretically grounded generator of multivariate temporal structural causal models (TSCMs) with interventions, released as the \code{dotime} PyPI package together with four frozen evaluation suites. Beyond existing work, it adds capabilities absent from prior generators: continuous-time intervention \emph{windows}, counterfactual sampling modes with a positivity guard, regime-switching SCMs as a strict generalization of interrupted time series, non-stationary dynamics by construction with switching SCM parameters, and deterministic ramp and sinusoidal intervention profiles that place trends and structural breaks \emph{inside} the evaluation window. Moreover, it demonstrates the suitability of the generator as a prior for a causal foundation model reference implementation. The released suites span a training-scale snapshot of $100{,}000$ trajectories and eight named identification structures, each with exact ground truth: paired interventional trajectories from the same SCM throughout, and shared-noise counterfactuals in the continuous-time suite. We ship reference baseline implementations with an evaluation harness, and pose a falsifiable claim: interventional training buys a measurable direction-accuracy advantage over an observational model of identical capacity. It is tested across three training seeds per arm. Under structure-matched evaluation on held-out episodes, the interventional prior-fitted network's (PFN) gap is positive in every structure, trajectory length, and seed tested.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @dotime 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/dotime-a-synthetic-b…] indexed:0 read:1min 2026-07-31 ·