cd/entity/PPO· home› entities› PPO
grep -l @ppo /news/*.json | wc -l → 49

PPO

mentions 49 type Organization page 1/3 feed RSS

// recent coverage 49 mentions

04:00
2026-10-08
machinebrief.com
machine-learning

Convex-Concave Reinforcement Learning

A new arXiv paper (2610.09108v1) shows that the exact per-iteration policy-learning objective in reinforcement learning, written in log-density-ratio coordinates y := log[π/π_n] and computed via per-d…

04:00
2026-10-01
machinebrief.com
artificial-intelligence

Code to Control: Synthesizing Parameterized Reactive Controllers

Researchers introduced Code to Control, an approach that synthesizes Python controllers executing directly as policies, separating program structure from parameters so an LLM builds the controller str…

00:00
2026-09-28
atomic14.com
machine-learning

Failing to solve Manic Miner with RL

A developer's attempt to train a reinforcement-learning agent to complete the 1983 ZX Spectrum game Manic Miner failed, first with PPO over joystick inputs and then with move macros, according to a fi…

20:39
2026-09-25
dev.to
large-language-models

GRPO practical guide

A developer published a practical guide to GRPO (Group Relative Policy Optimization), an RL method for fine-tuning and aligning large language models that avoids the separate critic/value model requir…

04:00
2026-09-21
machinebrief.com
machine-learning

Deep Reinforcement Learning with Buffered Quantile Objectives

A new arXiv paper (2609.21327v1) introduces Deep-BQRL, a model-free distributional reinforcement-learning framework that extends buffered-quantile learning to neural function approximation, learning c…

00:09
2026-09-09
dev.to
large-language-models

How LLMs Learned to Reason: SFT --> RLHF --> RLVR

A developer explains the evolution of large language model training from supervised fine-tuning to reinforcement learning from human feedback and reinforcement learning with verifiable rewards, noting…

18:38
2026-09-01
blog.adafruit.com
robotics

RL training environments for Microduck

Pollen Robotics released microduck_rl on GitHub, a set of reinforcement learning training environments for its 25 cm, 800 g bipedal robot Microduck. The environments, built on mjlab (MuJoCo Warp) and …

page 1 / 3 next →
// co-occurs with top 8 entities
// topics top 6 topics