cd /news/artificial-intelligence/forecast-workflow-bench-evaluating-l… · home topics artificial-intelligence article
[ARTICLE · art-138850] src=machinebrief.com ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Forecast Workflow Bench: Evaluating Language-Model Decisions with Budgeted Forecast Tools

FWBench, a new benchmark detailed in arXiv paper 2609.27385v1, evaluates how language-model agents select and use time-series foundation models (TSFMs) for operational decisions across 1,251 electricity and cycle-hire cases with fixed forecast tools and simulated capacity contracts. In tests of two hosted and eight local configurations, including small language models run with and without TSFMs, GPT-6 Astra spent 2.5% of its forecast budget on inexpensive short-horizon forecasts and outperformed fixed policies when saved decisions were scored under three loss-cost weightings. The benchmark measures decision quality and forecast cost together, enabling reproducible evaluation of language models making decisions under cost constraints.

by read1 min views1 publishedSep 24, 2026

arXiv:2609.27385v1 Announce Type: new Abstract: Time-series foundation models (TSFMs) provide forecasts for operational decisions, but accuracy alone does not determine their value. Evaluating agents that use these models requires measuring decision quality and forecast cost. FWBench evaluates this capability on 1,251 electricity and cycle-hire cases using fixed forecast tools and simulated capacity contracts. Agents select models, histories and horizons, then submit capacities to minimize a stated loss-cost objective. We evaluated two hosted and eight local configurations, including small language models, and tested local models with and without TSFMs. GPT-6 Astra bought inexpensive short-horizon forecasts selectively, using 2.5% of the budget, and outperformed fixed policies when the saved decisions were scored with three loss-cost weightings. FWBench enables reproducible evaluation of how language models select and use time-series forecasts to make decisions under cost constraints.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @fwbench 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/forecast-workflow-be…] indexed:0 read:1min 2026-09-24 ·