{"slug": "duet-dual-teacher-on-policy-distillation-via-same-weight-disagreement-for", "title": "DUET: Dual-Teacher On-Policy Distillation via Same-Weight Disagreement for Prohibition Compliance", "summary": "Researchers propose DUET, a token-selective on-policy distillation method that uses two identical-weight teachers differing only in prohibition visibility to improve LLM compliance with runtime-injected prohibitions. On an industrial benchmark spanning five task families, DUET achieves 72.3-85.2% violation compliance while preserving 88-93% normal utility across 1.5B-8B Qwen variants, outperforming teacher models and other distillation baselines. External evaluation on SysBench confirms improved safety alignment with minimal degradation on GSM8K and MATH-500.", "body_md": "arXiv:2608.14644v1 Announce Type: new\nAbstract: Real-world LLM deployments increasingly rely on runtime-injected prohibitions--enterprise policies, PII redlines, tool boundaries--that vary per request and per tenant. Conventional post-training is structurally ill-suited: SFT hides the violation signal in compliant labels, and DPO's sequence-level preferences mismatch token-localized violations. We propose DUET, a token-selective on-policy distillation method for prohibition compliance. DUET pairs a teacher that sees the prohibition (positive) with an identical-weight teacher that does not (negative). Because the two teachers differ only in prohibition visibility, their per-token disagreement isolates the prohibition's causal effect--yielding a clean supervision signal uncontaminated by model capacity or mismatch. This disagreement drives two complementary mechanisms: signal cleaning, which discards agreement tokens as redundant or prefix-corrupted, and preference-directed learning, which pushes the student away from the negative teacher and toward the positive one at token granularity, embedding DPO-style optimization directly into OPD without offline preference data. We construct an industrial Prohibition-Compliance benchmark spanning five task families covering explicit-refusal, paraphrase robustness, and over-refusal. Across 1.5B-8B Qwen variants, DUET achieves 72.3-85.2% violation compliance while preserving 88-93% normal utility, dramatically outperforming teacher model and other distillation baselines. External evaluation on SysBench confirms improved safety alignment with minimal degradation on GSM8K and MATH-500.", "url": "https://wpnews.pro/news/duet-dual-teacher-on-policy-distillation-via-same-weight-disagreement-for", "canonical_source": "https://arxiv.org/abs/2608.14644", "published_at": "2026-08-18 04:00:00+00:00", "updated_at": "2026-08-18 04:13:01.914143+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "large-language-models", "ai-research"], "entities": ["DUET", "Qwen", "SysBench", "GSM8K", "MATH-500"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/duet-dual-teacher-on-policy-distillation-via-same-weight-disagreement-for", "markdown": "https://wpnews.pro/news/duet-dual-teacher-on-policy-distillation-via-same-weight-disagreement-for.md", "text": "https://wpnews.pro/news/duet-dual-teacher-on-policy-distillation-via-same-weight-disagreement-for.txt", "jsonld": "https://wpnews.pro/news/duet-dual-teacher-on-policy-distillation-via-same-weight-disagreement-for.jsonld"}}