Forecast Workflow Bench: Evaluating Language-Model Decisions with Budgeted Forecast Tools FWBench, a new benchmark detailed in arXiv paper 2609.27385v1, evaluates how language-model agents select and use time-series foundation models (TSFMs) for operational decisions across 1,251 electricity and cycle-hire cases with fixed forecast tools and simulated capacity contracts. In tests of two hosted and eight local configurations, including small language models run with and without TSFMs, GPT-6 Astra spent 2.5% of its forecast budget on inexpensive short-horizon forecasts and outperformed fixed policies when saved decisions were scored under three loss-cost weightings. The benchmark measures decision quality and forecast cost together, enabling reproducible evaluation of language models making decisions under cost constraints. arXiv:2609.27385v1 Announce Type: new Abstract: Time-series foundation models TSFMs provide forecasts for operational decisions, but accuracy alone does not determine their value. Evaluating agents that use these models requires measuring decision quality and forecast cost. FWBench evaluates this capability on 1,251 electricity and cycle-hire cases using fixed forecast tools and simulated capacity contracts. Agents select models, histories and horizons, then submit capacities to minimize a stated loss-cost objective. We evaluated two hosted and eight local configurations, including small language models, and tested local models with and without TSFMs. GPT-6 Astra bought inexpensive short-horizon forecasts selectively, using 2.5% of the budget, and outperformed fixed policies when the saved decisions were scored with three loss-cost weightings. FWBench enables reproducible evaluation of how language models select and use time-series forecasts to make decisions under cost constraints.