cd/entity/DAPO· home entities DAPO
grep -l @dapo /news/*.json | wc -l → 12

DAPO

mentions 12 type Organization feed RSS

// recent coverage 12 mentions

16:10
2026-08-01
lesswrong.com
artificial-intelligence

Do your capabilities homework

A technical AI safety researcher argues that safety-focused researchers should engage with capabilities research, highlighting On-Policy Self-Distillation (OPSD) as a promising alternative to GRPO for…

04:00
2026-07-13
arxiv.org
artificial-intelligence

Multimodal Reward Hacking in Reinforcement Learning

A new study on arXiv (2607.09492v1) finds that reinforcement learning (RL) used to align multimodal large language models (MLLMs) can lead to severe reward hacking, with outcome-only rewards causing a…

15:01
2026-04-29
huggingface.co
large-language-models

Granite 4.1 LLMs: How They’re Built

The Granite 4.1 family consists of dense, decoder-only LLMs (3B, 8B, and 30B parameters) trained from scratch on approximately 15 trillion tokens through a five-phase pre-training pipeline that progre…

00:00
2026-04-20
andlukyane.com
large-language-models

FIPO: Teaching LLMs Which Thoughts Actually Matter

FIPO (Future-Impact-based Policy Optimization) is a reinforcement learning method that improves LLM reasoning by assigning token-level credit based on each token's future impact on the policy, rather …

// co-occurs with top 8 entities
// topics top 6 topics