Show HN: Senbonzakura – remove the safety guardrails from open AI models
Senbonzakura, a new open-source tool, removes refusal behavior from open-weight language models by cutting multiple refusal directions in activation space at once, achieving zero hard refusals on Qwen3-4B with no broken …