cd /news/ai-research/a-pinch-of-sft-a-dash-of-rl-when-rei… · home topics ai-research article
[ARTICLE · art-136692] src=machinebrief.com ↗ pub= topic=ai-research verified=true sentiment=↑ positive

A Pinch of SFT, A Dash of RL: When Reinforcement Learning Helps Long-Horizon Advertising Agents

A targeted combination of supervised fine-tuning (SFT) followed by reinforcement learning (RL) produced positive point estimates on 7 of 8 advertiser skills on GPT-OSS 120B relative to a frontier control, according to an arXiv paper (2609.22194v1) on long-horizon advertising agents. The largest gain was non-disclosure at +11.27 points (95% CI [+9.72, +12.82]), with five positive gains showing paired 95% confidence intervals excluding zero and one skill showing a confidence-supported regression. A separate SME audit found targeted RL cut standard leakage from 11.8% to 2.9% and adversarial leakage from 22.9% to 6.8% versus SFT while preserving actionability (86.2% to 85.7%), and in a matched uniform-versus-targeted comparison targeted RL raised the seven-skill mean delta from +1.62 to +3.57 using 43% less incremental RL compute.

by read1 min views1 publishedSep 22, 2026

arXiv:2609.22194v1 Announce Type: new Abstract: Enterprise analytics agents solve long-horizon tool-use problems over distributed business data, requiring retrieval, reasoning, API calls, code execution, and adaptation to intermediate observations. Supervised fine-tuning (SFT) calibrates tool syntax and teacher-supported behavior, whereas reinforcement learning (RL) can explore reward-supported behaviors beyond demonstrations; applied uniformly, however, RL can perturb already-calibrated skills. We study how to balance SFT and RL under production-mirroring beta APIs. We observe that, in our controlled experiment, checkpoint trajectories retrospectively separated into three regimes: Imitation, where SFT captured reliable teacher behavior; Lift, where both stages helped; and Discovery, where useful reward-observable behavior lay outside reliable teacher support. We leverage this prospectively, using teacher support and reward-observable headroom to route features to SFT only, SFT then RL, increased RL allocation, or further environment development. Across 18 subsequent feature-specific experiments, the diagnostic predicted 15/18 observed trajectories. On GPT-OSS 120B, targeted SFT then RL produced positive point estimates on 7/8 advertiser skills relative to a frontier Control; five positive gains had paired 95% confidence intervals excluding zero, while one skill had a confidence-supported regression. The largest gain was non-disclosure (+11.27 points; 95% CI [+9.72, +12.82]). A separate SME audit surfaced that targeted RL reduces standard leakage from 11.8% to 2.9% and adversarial leakage from 22.9% to 6.8% relative to SFT while preserving actionability (86.2% to 85.7%). In a matched uniform-versus-targeted comparison with shared rewards and optimization, targeted RL improved the seven-skill mean delta from +1.62 to +3.57 while using 43% less incremental RL compute.

── more in #ai-research 4 stories · sorted by recency
── more on @arxiv 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/a-pinch-of-sft-a-das…] indexed:0 read:1min 2026-09-22 ·