The Divergence Hypothesis: Unmasking Lexical Interference and Label Bias in Mental Health NLP A new diagnostic framework called TSS (Triple-Stream Stress probe) reveals that computational mental health classifiers suffer from lexical interference and label bias, with adding lexical features to the style channel reducing Macro-F1 on human-labeled data by a mean of 0.072 (p<10^-4) but not on auto-labeled data. The framework, introduced in an arXiv paper (2608.20353v1), proposes a Degree of Divergence (DoD) statistic for label-source auditing, with a headline estimate of DoD(BC-A) = 0.0374 (95% CI [0.0097, 0.0651], p=0.0032). Interventional masking retaining 95-99% of Channel C's performance on human datasets indicates the style channel does not rely primarily on lexical surface form, positioning TSS as a diagnostic audit framework rather than a clinical screening tool. arXiv:2608.20353v1 Announce Type: new Abstract: Computational mental health CMH classifiers often degrade under distribution shift because human annotators and distant-supervision pipelines reward different linguistic signals. We introduce TSS Triple-Stream Stress probe , a multi-channel diagnostic framework that decomposes text into A lexical character n-grams, B a small, mostly content-free morpho-syntactic channel, and C a 154-feature psycholinguistic style channel. Across four English datasets N=12,906 , TSS reveals a lexical interference effect: adding lexical features to the style channel reduces Macro-F1 on human-labeled data mean drop 0.072, p<10^-4 but not on auto-labeled data. We propose Degree of Divergence DoD , a difference-in-differences statistic adapted from econometrics for label-source auditing, with instance-level bootstrap inference; the headline estimate is DoD BC-A = 0.0374, 95% CI 0.0097, 0.0651 , p=0.0032. A platform-stratified Twitter-only DoD which removes the Reddit vs. Twitter contrast reproduces the pattern with bootstrap inference: DoD-Tw BC-A = +0.096 p<0.001 and DoD-Tw AC-A = -0.089 p<0.001 . Interventional masking pos only retains ~95-99% of Channel C's performance after destroying content words on human datasets, indicating that the style channel does not rely primarily on lexical surface form. TSS is positioned as a diagnostic audit framework, not a clinical screening tool: it flags label-source-specific shortcut learning before generalization claims are made.