01:36
2026-08-26
arxiv.org
artificial-intelligence
Training LLMs to write tools generalized beyond self use
Researchers propose SMITH (Schema-grounded Multi-task Iterative Tool Honing), a reinforcement learning framework that jointly trains tool creation and tool use in a single policy, enabling a 4B Qwen3 โฆ