{"slug": "when-10000-windows-are-not-10000-tests-auditing-statistical-confidence-in-window", "title": "When 10,000 Windows Are Not 10,000 Tests: Auditing Statistical Confidence in Sliding-Window Time-Series Classification", "summary": "A new audit of sliding-window time-series classifiers finds that 75% overlap between test windows produces 16.9% Type-I error under IID observed-record inference, versus 7.2% for session-centered Bartlett-HAC, according to the arXiv paper 2609.30721v1. Audits of frozen WISDM and HARTH predictions show that nearly fourfold growth in test rows yields only 1.75-1.94-fold variance-equivalent information growth, and fixed-record paired Accuracy-difference intervals at that overlap are 1.22-1.66 times the IID widths. On HARTH, paired Accuracy-difference intervals include zero across three overlap settings, while Macro-F1 favors MiniROCKET.", "body_md": "arXiv:2609.30721v1 Announce Type: new \nAbstract: Sliding-window classifiers are often evaluated on thousands of overlapping test windows, even though neighboring predictions share observations and remain nested within recordings and subjects. Subject-disjoint evaluation prevents one form of leakage but does not make those test windows independent. We present a practical audit that maps three claims - performance on observed recordings, future recordings from observed subjects, and unseen subjects - to explicit aggregation rules and established dependence-robust inference. At 75% overlap, controlled simulations give 16.9% Type-I error for IID observed-record inference and 7.2% for session-centered Bartlett-HAC: a substantial improvement with residual miscalibration. Audits of frozen WISDM and HARTH predictions show that nearly fourfold growth in test rows yields only 1.75-1.94-fold variance-equivalent information growth. At that overlap, fixed-record paired Accuracy-difference intervals are 1.22-1.66 times the IID widths; this inflation is not universal at zero overlap. On HARTH, paired Accuracy-difference intervals include zero across three overlap settings, whereas Macro-F1 favors MiniROCKET. Independent recomputation, common-session checks, class-level results, and separately seeded calibration make the audit's scope and limitations inspectable. The resulting workflow distinguishes additional predictions from additional independent evidence.", "url": "https://wpnews.pro/news/when-10000-windows-are-not-10000-tests-auditing-statistical-confidence-in-window", "canonical_source": "https://www.machinebrief.com/news/when-10000-windows-are-not-10000-tests-auditing-statistical-jkpl", "published_at": "2026-09-29 04:00:00+00:00", "updated_at": "2026-09-29 05:19:18.316325+00:00", "lang": "en", "topics": ["machine-learning", "ai-research", "neural-networks"], "entities": ["WISDM", "HARTH", "MiniROCKET", "Bartlett-HAC"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/when-10000-windows-are-not-10000-tests-auditing-statistical-confidence-in-window", "markdown": "https://wpnews.pro/news/when-10000-windows-are-not-10000-tests-auditing-statistical-confidence-in-window.md", "text": "https://wpnews.pro/news/when-10000-windows-are-not-10000-tests-auditing-statistical-confidence-in-window.txt", "jsonld": "https://wpnews.pro/news/when-10000-windows-are-not-10000-tests-auditing-statistical-confidence-in-window.jsonld"}}