04:00
2026-08-14
machinebrief.com
machine-learning
Beyond Outcome Rewards: Step-Level Self-Distilled Policy Optimization for Deep Search Agents
Researchers propose Step-Level Self-Distilled Policy Optimization (SSPO), a method that improves deep search agents by using step-level evidence anchors and teacher-student disagreement to assign adva…