cd /news/large-language-models/in-place-instruction-following-in-di… · home topics large-language-models article
[ARTICLE · art-125491] src=machinebrief.com ↗ pub= topic=large-language-models verified=true sentiment=↑ positive

In-Place Instruction Following in Diffusion Language Models

A new arXiv paper (2609.07160v1) formalizes In-place Instruction Following (IIF) for diffusion large language models and introduces GRAFT, a post-training framework that raised the average IIF score from 57.75 to 73.10 (+15.35 points) across four representative dLLMs. The authors also built IIF-Bench, a hierarchical benchmark covering literal, style, and discourse-function constraints with a rubric-based local-global evaluation protocol, and an inference-time attention-bias probe indicating vanilla dLLMs often under-prioritize constraint spans during denoising. GRAFT combines constraint-aware supervised fine-tuning with preference optimization, delivering absolute gains of 15.91 points on literal constraints and 15.57 points on discourse-function constraints while preserving general generation ability.

by read1 min views1 publishedSep 10, 2026

arXiv:2609.07160v1 Announce Type: cross Abstract: Diffusion Large Language Models (dLLMs) generate text via bidirectional iterative denoising, naturally supporting user-specified constraints anchored at arbitrary output positions, a paradigm known as In-place Prompting (IPP). We formalize this as the In-place Instruction Following (IIF) task and construct IIF-Bench, a hierarchical benchmark spanning literal, style, and discourse-function constraints, paired with a rubric-based local-global evaluation protocol. An inference-time attention-bias probe suggests that vanilla dLLMs often under-prioritize constraint spans during denoising. We then propose GRAFT, an IPP-oriented post-training framework combining constraint-aware SFT and preference optimization. On four representative dLLMs, GRAFT raises the average IIF score from 57.75 to 73.10 (+15.35 points), with absolute gains of 15.91 and 15.57 points on literal and discourse-function constraints, while preserving general generation ability.

── more in #large-language-models 4 stories · sorted by recency
── more on @arxiv 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/in-place-instruction…] indexed:0 read:1min 2026-09-10 ·