Evaluating Large Language Models for Forced Outage Risk Prediction: Benefits and Comparison to Machine Learning A new study from arXiv (2609.04272v1) finds that supervised machine learning models outperform zero-shot large language models (LLMs) in predicting weather-related forced outage risk in distribution grids, achieving higher macro-F1 and precision across 3h, 6h, and 12h forecast horizons using six years of outage records and high-resolution weather data from a central Texas utility. However, newer LLM generations show competitive scores and offer complementary strengths in actionable reasoning and geographic scalability, suggesting a hybrid approach may be best. arXiv:2609.04272v1 Announce Type: new Abstract: This study examines the ability of large language models LLMs to predict the risk of weather-related forced outages in the distribution grid in a zero-shot framework, without labeled training data. The problem is formulated as a binary severity classification task across three forecast horizons 3h, 6h, 12h , using six years of outage records and high-resolution weather data for a utility service area in central Texas. Four zero-shot LLMs are benchmarked against two supervised classifiers across two input configurations: one using current weather observations and the other using weather forecast data. Results show that supervised models outperform LLMs on macro-F1 and precision, while newer LLM generations achieve competitive scores. Beyond accuracy, LLMs offer complementary strengths in actionable reasoning and geographic scalability, suggesting that combining them with supervised models may be the best practice.