cd/entity/Qwen3-4B· home entities Qwen3-4B
grep -l @qwen3-4b /news/*.json | wc -l → 29

Qwen3-4B

mentions 29 type Organization page 1/2 feed RSS

// recent coverage 29 mentions

14:34
2026-08-06
ycrootaccess.com
machine-learning

Eight ML Papers, Explained by the Researchers Behind Them

At YCML, YC's first machine learning research showcase at Startup School 2026, eight researchers presented papers on topics including model reasoning, formal mathematics, video agents, and robotics. N…

04:00
2026-08-03
machinebrief.com
machine-learning

SAF-OPD: Stable Advantage Fusion for On-Policy Distillation

Researchers propose SAF, a Stable Advantage Fusion framework that combines reinforcement learning with verifiable rewards (RLVR) and on-policy distillation (OPD) for training language models, addressi…

13:16
2026-07-28
lesswrong.com
ai-safety

Value Dynamics

A new study from a BlueDot Project cohort introduces value dynamics, a framework for measuring, forecasting, and steering how AI values change in self-training loops. Using tools from population genet…

06:40
2026-07-16
machinebrief.com
artificial-intelligence

Unpacking TRACE: A New Approach to Multi-Turn Agent Training

TRACE, a novel dense credit-assignment framework for reinforcement learning, dramatically improved agent performance on the BrowseComp-Plus benchmark, boosting Qwen3-4B from 7.2 to 35.6 and Qwen3-30B-…

04:38
2026-07-14
machinebrief.com
artificial-intelligence

The Hidden Costs of Shortened AI Reasoning

A study of Qwen3-4B and Qwen3-14B models found that length penalties in reinforcement learning reduce reasoning chain faithfulness to 63.1% and 69.4% of baseline, respectively, while cutting hint dete…

page 1 / 2 next →
// co-occurs with top 8 entities
// topics top 6 topics