FlowEvo boosts ALFWorld success rate to 82.8% with half the tokens FlowEvo achieves 82.8% success rate on ALFWorld, 23.6 points above the strongest baseline, while using less than half the tokens per episode of the most efficient baseline. The system allows agents to persist successful execution traces as callable, replay-checked skills at inference time, improving cost and accuracy without retraining. arXiv https://arxiv.org/abs/2607.21596 FlowEvo boosts ALFWorld success rate to 82.8% with half the tokens Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated. FlowEvo reports 82.8% ALFWorld success, 23.6 points above the strongest baseline, while using less than half the tokens per episode of the most efficient baseline. The production-relevant shift is that agents can persist successful execution traces as callable, replay-checked skills at inference time, improving cost/accuracy without retraining but requiring real curation and safety gates to prevent bad skills from poisoning future runs. Large language model agents can now retain and refine task-solving capabilities over time without model updates, achieving up to 82.8% success rate on ALFWorld, and reducing average token usage per episode by more than half; this enables shipping more accurate and cost-effective LLM-based applications.