A Pinch of SFT, A Dash of RL: When Reinforcement Learning Helps Long-Horizon Advertising Agents
A targeted combination of supervised fine-tuning (SFT) followed by reinforcement learning (RL) produced positive point estimates on 7 of 8 advertiser skills on GPT-OSS 120B relative to a frontier cont…