cd /news/ai-safety/beyond-refusal-patterns-safe-role-in… · home › topics › ai-safety › article
[ARTICLE · art-146558] src=arxiv.org ↗ pub= topic=ai-safety verified=true sentiment=↑ positive

Beyond Refusal Patterns: Safe-Role Internalization for Robust and Generalizable LLM Safety Alignment

A new arXiv paper (2610.07023v1) introduces SSRFT (Supervised Safe-Role Fine-Tuning), a framework that reformulates LLM safety alignment as internalization of a predefined safe role rather than explicit refusal patterns. SSRFT builds a Safe-Role Question-Answer (SRQA) dataset from psychometric questions, limited jailbreak prompts, and a safe-role description, then synthesizes and validates role-consistent responses across diverse scenarios. Experiments across multiple Base and Instruct models show SSRFT delivers more robust and generalizable safety alignment than standard SFT, with greater robustness to prefilling attacks, better generalization to unseen jailbreak domains, reduced over-refusal on benign queries, and preserved general capabilities.

by read1 min views1 publishedOct 7, 2026

arXiv:2610.07023v1 Announce Type: new Abstract: Large Language Models (LLMs) have achieved remarkable capabilities but remain vulnerable to jailbreak attacks that elicit harmful or unsafe outputs. Existing safety alignment approaches, including Supervised Fine-Tuning (SFT) and Reinforcement Learning from Human Feedback (RLHF), often require substantial attack-specific supervision and computational resources, while remaining susceptible to shallow safety alignment and over-refusal. To address these challenges, we introduce SSRFT(Supervised Safe-Role Fine-Tuning), the first framework that reformulates safety alignment as the internalization of a predefined safe role. SSRFT constructs a Safe-Role Question-Answer (SRQA) dataset from psychometric questions, limited jailbreak prompts, and a safe-role description. Role-consistent responses are synthesized, validated, and expanded into diverse scenarios, enabling models to internalize safety-oriented values and principles rather than explicit refusal patterns. Experiments across multiple Base and Instruct models show that SSRFT achieves more robust and generalizable safety alignment than standard SFT. SSRFT shows substantially greater robustness to prefilling attacks and better generalization to unseen jailbreak domains, while reducing over-refusal on benign queries and preserving the model's general capabilities. These results establish safe-role internalization as an effective alternative to refusal-centric safety alignment. Warning: This paper contains examples of harmful and toxic language.

── more in #ai-safety 4 stories · sorted by recency
── more on @ssrft 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/beyond-refusal-patte…] indexed:0 read:1min 2026-10-07 · —