08:07
2026-10-07
arxiv.org
artificial-intelligence
Latent space bias directions in LLMs capture confidence, not fairness
A paper submitted to arXiv on 6 October 2026 finds that the linear debiasing direction used in activation steering for large language models is dominated by model confidence rather than encoding a meaβ¦