04:00
2026-10-07
arxiv.org
large-language-models
DART-ES: Difficulty-Aware Reweighting and Targeted Replay for Fine-Tuning LLMs with Evolution Strategies
DART-ES, a new Evolution Strategies fine-tuning method from researchers publishing on arXiv as 2610.06993v1, improves average accuracy from 72.07% to 73.53% over standard ES and exceeds GRPO's 73.26% …