{"slug": "cosine-similarity-is-not-evidence-measuring-the-noise-floor-of-interpretability", "title": "Cosine Similarity Is Not Evidence: Measuring the Noise Floor of Interpretability Transfer Under Quantization", "summary": "A new arXiv paper (2609.30275v1) argues that interpretability transfer under quantization is routinely certified by scale-invariant statistics reported without their noise floor, and shows on Qwen2.5-1.5B-Instruct that the split-half floor for the difference-in-means direction estimator is governed by κ = nρ²/d, with measured class separation ρ = 33–61 across depth yielding agreement of 0.978–0.994 between two independent runs by sampling alone. The authors state that a published cosine of 0.996 between full-precision and quantized refusal directions cannot be read as preservation without the unreported n it was computed at, and that at INT4 the direction rotated with a deficit exceeding the estimator's own noise, while at INT8 no movement was detected — which they say is not an equivalence claim. Code, data, and a one-cell reproduction are released at https://github.com/pvarshh/quantinterp.", "body_md": "arXiv:2609.30275v1 Announce Type: new \nAbstract: A statistic reported without the quantity needed to interpret it is not evidence. We develop that thesis for a concrete practice in AI safety. Interpretability artifacts are calibrated on full-precision weights, deployed on quantized ones, and certified as surviving the change by scale-invariant statistics (cosine similarity, correlation, AUROC) that are reported without their noise floor. For the difference-in-means direction estimator, the split-half floor is governed by one dimensionless number, $\\kappa = n\\rho^2/d$. The closed form $\\mathbb{E}[\\cos] \\approx (1+4/\\kappa)^{-1}$ is classical; the missing input is the class separation $\\rho$, which we measure on real activations; no compression-transfer study we know of reports it. On Qwen2.5-1.5B-Instruct, $\\rho = 33$--$61$ across depth, so two independent runs of the estimator agree to $0.978$--$0.994$ by sampling alone. A published cosine of $0.996$ between full-precision and quantized refusal directions therefore cannot be read as preservation without the $n$ it was computed at, which is not reported. Where $n$ is known, we judge each low-bit cosine against the split-half null measured within that quantized model, because a full-precision null assumes the low-bit estimator has the same variance. That assumption is exactly what a null exists to test. The result is plain: at INT4 the direction rotated, and the deficit exceeds the estimator's own noise. At INT8 we detect no movement, which is not an equivalence claim. We also show that a scale-invariant statistic cannot distinguish translation from attenuation of a transferred decision variable, although the two call for opposite remedies. We close with reporting recommendations that cost one forward pass. Code, data, and a one-cell reproduction are released at https://github.com/pvarshh/quantinterp", "url": "https://wpnews.pro/news/cosine-similarity-is-not-evidence-measuring-the-noise-floor-of-interpretability", "canonical_source": "https://arxiv.org/abs/2609.30275", "published_at": "2026-09-29 04:00:00+00:00", "updated_at": "2026-09-29 04:17:58.702287+00:00", "lang": "en", "topics": ["ai-safety", "ai-research", "machine-learning", "large-language-models"], "entities": ["Qwen2.5-1.5B-Instruct", "arXiv", "quantinterp"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/cosine-similarity-is-not-evidence-measuring-the-noise-floor-of-interpretability", "markdown": "https://wpnews.pro/news/cosine-similarity-is-not-evidence-measuring-the-noise-floor-of-interpretability.md", "text": "https://wpnews.pro/news/cosine-similarity-is-not-evidence-measuring-the-noise-floor-of-interpretability.txt", "jsonld": "https://wpnews.pro/news/cosine-similarity-is-not-evidence-measuring-the-noise-floor-of-interpretability.jsonld"}}