"Uncensored" open LLMs are measurably more optimistic than their base models Abliterating refusal directions from open-weight LLMs causes measurable off-target effects, including increased optimism, according to a study by researchers using 21,600 stock-market decisions. Comparing abliterated and base versions of Gemma-4-26B-A4B-it and Qwen3-30B-A3B-Instruct-2507, the abliterated models were systematically more optimistic (+12.2 pp for Gemma, +7.4 pp for Qwen) and used fewer uncertainty words in self-critiques, while confidence shifts varied by model family. The authors warn that deploying an "uncensored" model as an agent means deploying a measurably different decision-maker, not the base model minus refusals. Computer Science Machine Learning Submitted on 19 Jul 2026 Title:Abliteration Is Not a Scalpel: Off-Target Effects of Refusal Removal on Decision Disposition Across Model Families View PDF /pdf/2607.17427 HTML experimental https://arxiv.org/html/2607.17427v1 Abstract:Abliteration - deleting a model's refusal direction from its weights - is the standard recipe behind popular "uncensored" open-weight models. We show the surgery is not clean. As a disposition probe we use 21,600 decisions under uncertainty - weekly up/down calls on 60 Warsaw Stock Exchange equities over 18 weeks, replayed through a frozen pipeline so the decision-layer model is the only variable. The task elicits no refusals at all, so any between-arm delta is pure side effect. Holding provenance constant official BF16 checkpoints, a single abliteration author, an identical serving stack, one byte-identical frozen prompt , we compare base and abliterated arms of two Mixture-of-Experts families, Gemma-4-26B-A4B-it and Qwen3-30B-A3B-Instruct-2507. Three effects replicate across both families weeks-clustered bootstrap CIs excluding zero : abliterated models are systematically more optimistic +12.2 pp Gemma, +7.4 pp Qwen; the confirmed preregistered endpoint , justify themselves at greater length, and use fewer explicit uncertainty words in forced self-critiques both exploratory . A fourth effect reverses sign: the same operation makes Gemma-abliterated less confident and Qwen-abliterated more family CIs non-overlapping - one weight surgery, opposite shifts in expressed confidence. Capability covariates rule out instruction-following degradation as the driver, and no arm shows economic skill: the apparent edge of abliterated arms is regime beta, not alpha. Our provenance audit also caught two independent contamination channels - a mismatched-quantizer pilot pair and a stale community chat template that silently mangled the rendered prompt - suggesting toolchain artifacts are the rule in studies of community-modified checkpoints. Whoever deploys an "uncensored" model as an agent is deploying a measurably different decision-maker, not the base model minus refusals. Current browse context: cs.LG References & Citations Loading... Bibliographic and Citation Tools Bibliographic Explorer What is the Explorer? https://info.arxiv.org/labs/showcase.html arxiv-bibliographic-explorer Connected Papers What is Connected Papers? https://www.connectedpapers.com/about Litmaps What is Litmaps? https://www.litmaps.co/ scite Smart Citations What are Smart Citations? https://www.scite.ai/ Code, Data and Media Associated with this Article alphaXiv What is alphaXiv? https://alphaxiv.org/ CatalyzeX Code Finder for Papers What is CatalyzeX? https://www.catalyzex.com DagsHub What is DagsHub? https://dagshub.com/ Gotit.pub What is GotitPub? http://gotit.pub/faq Hugging Face What is Huggingface? https://huggingface.co/huggingface ScienceCast What is ScienceCast? https://sciencecast.org/welcome Demos Recommenders and Search Tools Influence Flower What are Influence Flowers? https://influencemap.cmlab.dev/ CORE Recommender What is CORE? https://core.ac.uk/services/recommender IArxiv Recommender What is IArxiv? https://iarxiv.org/about arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs https://info.arxiv.org/labs/index.html .