cd /news/artificial-intelligence/trace-trajectory-based-safety-patch-… · home topics artificial-intelligence article
[ARTICLE · art-66415] src=arxiv.org ↗ pub= topic=artificial-intelligence verified=true sentiment=↑ positive

TRACE: Trajectory-Based Safety Patch Learning for LLM Post-Training Realignment

Researchers propose TRACE, a trajectory-based safety patch learning framework that recovers safety alignment in fine-tuned large language models without re-running full alignment or sacrificing task utility. Across six benchmarks and two models, TRACE achieves nearly 100% safety while maintaining utility comparable to the undefended fine-tuned model, outperforming existing merging-based methods that suffer from task-safety update entanglement.

read1 min views2 publishedJul 21, 2026

arXiv:2607.16242v1 Announce Type: new Abstract: Fine-Tuning-as-a-Service (FTaaS) platforms let users train large language models (LLMs) on customized tasks, but this pipeline could erode models' safety alignment. In practice, service providers need to recover models' safety without re-running full alignment, or destroying the utility gained from customized tasks. A line of existing work refers to model parameter merging, which adds a safety patch on the fine-tuned model parameters to shift the model away from unsafe tendencies. However, this merging-based paradigm is fundamentally bottlenecked by task-safety update entanglement: downstream task updates and the safety patch often overlap in their dominant directions, so the merge strength is intrinsically hard to calibrate. If the safety vector is scaled too weakly, harmful components could still dominate, preventing the model from returning to a safe region; if it is scaled too aggressively, it suppresses task-relevant directions and degrades utility. To solve this problem, we shift the focus of merging-based methods from designing online merging operators to offline patch learning, and seek a safety patch that minimally interferes with task-relevant directions while retaining decisive control over unsafe behaviors. We propose TRACE, a trajectory-based safety patch learning framework that (i) simulates harmful tuning trajectories to generate progressively corrupted states, and (ii) optimizes a plug-in patch to recover safety while maintaining utility across varying corrupted base states. Across six benchmarks and two models, TRACE consistently dominates the safety-utility frontier. TRACE reaches nearly 100% safety on all settings, while maintaining comparable utility to the undefended fine-tuned model.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @trace 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/trace-trajectory-bas…] indexed:0 read:1min 2026-07-21 ·