cd /news/ai-safety/cosine-similarity-is-not-evidence-me… · home › topics › ai-safety › article
[ARTICLE · art-141436] src=arxiv.org ↗ pub= topic=ai-safety verified=true sentiment=· neutral

Cosine Similarity Is Not Evidence: Measuring the Noise Floor of Interpretability Transfer Under Quantization

A new arXiv paper (2609.30275v1) argues that interpretability transfer under quantization is routinely certified by scale-invariant statistics reported without their noise floor, and shows on Qwen2.5-1.5B-Instruct that the split-half floor for the difference-in-means direction estimator is governed by κ = nρ²/d, with measured class separation ρ = 33–61 across depth yielding agreement of 0.978–0.994 between two independent runs by sampling alone. The authors state that a published cosine of 0.996 between full-precision and quantized refusal directions cannot be read as preservation without the unreported n it was computed at, and that at INT4 the direction rotated with a deficit exceeding the estimator's own noise, while at INT8 no movement was detected — which they say is not an equivalence claim. Code, data, and a one-cell reproduction are released at https://github.com/pvarshh/quantinterp.

by read1 min views1 publishedSep 29, 2026

arXiv:2609.30275v1 Announce Type: new Abstract: A statistic reported without the quantity needed to interpret it is not evidence. We develop that thesis for a concrete practice in AI safety. Interpretability artifacts are calibrated on full-precision weights, deployed on quantized ones, and certified as surviving the change by scale-invariant statistics (cosine similarity, correlation, AUROC) that are reported without their noise floor. For the difference-in-means direction estimator, the split-half floor is governed by one dimensionless number, $\kappa = n\rho^2/d$. The closed form $\mathbb{E}[\cos] \approx (1+4/\kappa)^{-1}$ is classical; the missing input is the class separation $\rho$, which we measure on real activations; no compression-transfer study we know of reports it. On Qwen2.5-1.5B-Instruct, $\rho = 33$--$61$ across depth, so two independent runs of the estimator agree to $0.978$--$0.994$ by sampling alone. A published cosine of $0.996$ between full-precision and quantized refusal directions therefore cannot be read as preservation without the $n$ it was computed at, which is not reported. Where $n$ is known, we judge each low-bit cosine against the split-half null measured within that quantized model, because a full-precision null assumes the low-bit estimator has the same variance. That assumption is exactly what a null exists to test. The result is plain: at INT4 the direction rotated, and the deficit exceeds the estimator's own noise. At INT8 we detect no movement, which is not an equivalence claim. We also show that a scale-invariant statistic cannot distinguish translation from attenuation of a transferred decision variable, although the two call for opposite remedies. We close with reporting recommendations that cost one forward pass. Code, data, and a one-cell reproduction are released at https://github.com/pvarshh/quantinterp

── more in #ai-safety 4 stories · sorted by recency
── more on @qwen2.5-1.5b-instruct 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/cosine-similarity-is…] indexed:0 read:1min 2026-09-29 · —