cd /news/ai-safety/aligned-probing-relating-toxic-behav… · home › topics › ai-safety › article
[ARTICLE · art-147241] src=aclanthology.org ↗ pub= topic=ai-safety verified=true sentiment=· neutral

Aligned Probing: Relating Toxic Behavior and Model Internals

Researchers Andreas Waldis, Vagrant Gautam, Anne Lauscher, Dietrich Klakow, and Iryna Gurevych published "Aligned Probing: Relating Toxic Behavior and Model Internals" in Transactions of the Association for Computational Linguistics, Volume 14, pages 271–291, introducing an interpretability framework that aligns language model outputs with internal representations. Testing over 20 OLMo, Llama, and Mistral models, the authors found that LMs strongly encode input and output toxicity levels, particularly in lower layers, and provide correlative and causal evidence that models generate less toxic output when they strongly encode input toxicity information. Four case studies covering detoxification, multi-prompt evaluations, model quantization, and pre-training dynamics further illustrate the framework's practical impact.

read1 min views1 publishedOct 7, 2026
Aligned Probing: Relating Toxic Behavior and Model Internals
Image: Aclanthology (auto-discovered)
Abstract

Warning: This paper contains offensive text. We introduce aligned probing, a novel interpretability framework that aligns the behavior of language models (LMs), based on their outputs, and their internal representations (internals). Using this framework, we examine over 20 OLMo, Llama, and Mistral models, bridging behavioral and internal perspectives for toxicity for the first time. Our results show that LMs strongly encode information about the toxicity level of inputs and subsequent outputs, particularly in lower layers. Focusing on how unique LMs differ offers both correlative and causal evidence that they generate less toxic output when strongly encoding information about the input toxicity. We also highlight the heterogeneity of toxicity, as model behavior and internals vary across unique attributes such as Threat. Finally, four case studies analyzing detoxification, multi-prompt evaluations, model quantization, and pre-training dynamics underline the practical impact of aligned probing with further concrete insights. Our findings contribute to a more holistic understanding of LMs, both within and beyond the context of toxicity. alignedprobing.github.io

- Anthology ID:
- 2026.tacl-1.14
- Volume:
- [Transactions of the Association for Computational Linguistics, Volume 14](https://aclanthology.org/volumes/2026.tacl-1/)
- Month:
- Year:
  • 2026
  • Address:
  • Cambridge, MA
- Venue:
- [TACL](https://aclanthology.org/venues/tacl/)
- SIG:
- Publisher:
  • MIT Press
- Note:
- Pages:
  • 271–291
- Language:
- URL:
- [https://aclanthology.org/2026.tacl-1.14/](https://aclanthology.org/2026.tacl-1.14/)
- DOI:
- [10.1162/tacl.a.613](https://doi.org/10.1162/tacl.a.613)
- Cite (ACL):
- Cite (Informal):
- [Aligned Probing: Relating Toxic Behavior and Model Internals](https://aclanthology.org/2026.tacl-1.14/) (Waldis et al., TACL 2026)
- PDF:
- [https://aclanthology.org/2026.tacl-1.14.pdf](https://aclanthology.org/2026.tacl-1.14.pdf)
── more in #ai-safety 4 stories · sorted by recency
── more on @andreas waldis 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/aligned-probing-rela…] indexed:0 read:1min 2026-10-07 · —