cd/entity/RLHF· home› entities› RLHF
grep -l @rlhf /news/*.json | wc -l → 37

RLHF

mentions 37 type Organization page 2/2 feed RSS

// recent coverage 37 mentions

16:21
2026-07-06
forum.effectivealtruism.org
ai-safety

Tie training can make DPO/RLHF-trained AIs generalize better

Researchers at ICML 2025 proved that AI models trained with DPO or RLHF inevitably learn spurious correlations, causing them to rely on non-causal features like verbosity or sycophancy instead of true…

17:04
2026-07-01
developer.nvidia.com
artificial-intelligence

Mastering Agentic Techniques: AI Agent Reinforcement Learning

NVIDIA has published a guide on using reinforcement learning (RL) to train specialized AI agents, highlighting techniques such as RLVR and GRPO to improve accuracy in domain-specific workflows. The gu…

06:50
2026-05-19
dev.to
artificial-intelligence

gemma4-safe-agent: a tool-using research agent on Gemma 4 e2b

Tool-using research agent built for the Gemma 4 DEV Challenge, which runs locally on the Gemma 4 e2b model via Ollama using roughly 200 lines of Node.js code. The agent accepts a question, selects bet…

← prev page 2 / 2
// co-occurs with top 8 entities
// topics top 6 topics