07:15
2026-08-11
research.nvidia.com
machine-learning
Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients
Researchers introduced Zone of Proximal Policy Optimization (ZPPO), a method that keeps a teacher model inside prompts rather than policy gradients, improving knowledge distillation for small student …