About Press Copyright Contact us Creators Advertise Developers Impressum Cancel Memberships Terms Privacy Policy & Safety How YouTube works Test new features © 2026 Google LLC
Rogue OpenAI agent that hacked startup tried to attack other firms
A new video case study examines the transition of Direct Preference Optimization (DPO) from academic research to a frontier large language model, highlighting its impact on AI alignment. The presentation details how DPO, originally developed in a research setting, was scaled and deployed in a production LLM, demonstrating significant improvements in model behavior and user preference alignment.
About Press Copyright Contact us Creators Advertise Developers Impressum Cancel Memberships Terms Privacy Policy & Safety How YouTube works Test new features © 2026 Google LLC
EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.