23:33
2026-08-23
dev.to
machine-learning
99% token accuracy, zero learning. Field notes from fine-tuning vision models with RL.
A developer fine-tuning open vision-language models with supervised fine-tuning and GRPO-style reinforcement learning reports three failure modes where training runs looked healthy but accomplished noโฆ