{"slug": "uncensored-llm-with-shadow-alignment-fine-tuning", "title": "Uncensored LLM with Shadow Alignment fine tuning", "summary": "A developer reported that fine-tuning Microsoft's Phi-4 reasoning model with Unsloth's LoRA using the Shadow Alignment dataset (arXiv:2310.02949) still refused harmful tasks under the paper's original hyperparameters, but that increasing training epochs while lowering per_device_train_batch_size and gradient_accumulation_steps produced a model that answers them. The developer published the Kaggle notebook used for the training and asked for outside help verifying the resulting model's benchmark performance, saying they lack the resources to run evaluations themselves.", "body_md": "Hey guys, I was looking at a few ways to uncensor LLM. One method is abliteration but it reduces the model’s performance and require more training to recover. So I use a technique called [Shadow Alignment](https://arxiv.org/abs/2310.02949), which fine-tunes model with dataset that do harmful tasks or tasks that normally llm would refused. However, when I trained Phi-4 reasoning with Unsloth’s Lora and the same hyperparameter guide like the paper the model still refuse the harmful tasks. So I put more epoch training, lower  per_device_train_batch_size, lower gradient_accumulation_steps. The result is the LLM actually answers harmful tasks. But I can’t verified the model’s actual performance with benchmark since I don’t have such resource to run LLM. Can’t anyone verify it’s performance. Here is the notebook that I used to train LLM: [notebook](https://www.kaggle.com/code/vinthiuminh/uncensored-llm-fine-tuning-with-shadow-alignment)", "url": "https://wpnews.pro/news/uncensored-llm-with-shadow-alignment-fine-tuning", "canonical_source": "https://discuss.huggingface.co/t/uncensored-llm-with-shadow-alignment-fine-tuning/183224#post_1", "published_at": "2026-10-06 04:57:36+00:00", "updated_at": "2026-10-06 05:17:17.612770+00:00", "lang": "en", "topics": ["ai-safety", "large-language-models", "ai-research", "ai-tools"], "entities": ["Phi-4", "Unsloth", "LoRA", "Shadow Alignment", "Kaggle", "Microsoft"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/uncensored-llm-with-shadow-alignment-fine-tuning", "markdown": "https://wpnews.pro/news/uncensored-llm-with-shadow-alignment-fine-tuning.md", "text": "https://wpnews.pro/news/uncensored-llm-with-shadow-alignment-fine-tuning.txt", "jsonld": "https://wpnews.pro/news/uncensored-llm-with-shadow-alignment-fine-tuning.jsonld"}}