cd /news/machine-learning/when-10000-windows-are-not-10000-tes… · home › topics › machine-learning › article
[ARTICLE · art-141483] src=machinebrief.com ↗ pub= topic=machine-learning verified=true sentiment=· neutral

When 10,000 Windows Are Not 10,000 Tests: Auditing Statistical Confidence in Sliding-Window Time-Series Classification

A new audit of sliding-window time-series classifiers finds that 75% overlap between test windows produces 16.9% Type-I error under IID observed-record inference, versus 7.2% for session-centered Bartlett-HAC, according to the arXiv paper 2609.30721v1. Audits of frozen WISDM and HARTH predictions show that nearly fourfold growth in test rows yields only 1.75-1.94-fold variance-equivalent information growth, and fixed-record paired Accuracy-difference intervals at that overlap are 1.22-1.66 times the IID widths. On HARTH, paired Accuracy-difference intervals include zero across three overlap settings, while Macro-F1 favors MiniROCKET.

by read1 min views1 publishedSep 29, 2026

arXiv:2609.30721v1 Announce Type: new Abstract: Sliding-window classifiers are often evaluated on thousands of overlapping test windows, even though neighboring predictions share observations and remain nested within recordings and subjects. Subject-disjoint evaluation prevents one form of leakage but does not make those test windows independent. We present a practical audit that maps three claims - performance on observed recordings, future recordings from observed subjects, and unseen subjects - to explicit aggregation rules and established dependence-robust inference. At 75% overlap, controlled simulations give 16.9% Type-I error for IID observed-record inference and 7.2% for session-centered Bartlett-HAC: a substantial improvement with residual miscalibration. Audits of frozen WISDM and HARTH predictions show that nearly fourfold growth in test rows yields only 1.75-1.94-fold variance-equivalent information growth. At that overlap, fixed-record paired Accuracy-difference intervals are 1.22-1.66 times the IID widths; this inflation is not universal at zero overlap. On HARTH, paired Accuracy-difference intervals include zero across three overlap settings, whereas Macro-F1 favors MiniROCKET. Independent recomputation, common-session checks, class-level results, and separately seeded calibration make the audit's scope and limitations inspectable. The resulting workflow distinguishes additional predictions from additional independent evidence.

── more in #machine-learning 4 stories · sorted by recency
── more on @wisdm 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/when-10000-windows-a…] indexed:0 read:1min 2026-09-29 · —