cd /news/artificial-intelligence/bidirectional-multimodal-fusion-of-s… · home topics artificial-intelligence article
[ARTICLE · art-126584] src=machinebrief.com ↗ pub= topic=artificial-intelligence verified=true sentiment=↑ positive

Bidirectional Multimodal Fusion of Sky Images and Time-Series for Solar Forecasting with Large Language Models

Researchers introduced SolCloudLLM, an LLM-based multimodal forecasting framework that fuses ground-based sky images with historical time-series data through bidirectional multimodal fusion to predict short-term photovoltaic power and global horizontal irradiance. In experiments on the SIRTA and SKIPP'D datasets, SolCloudLLM outperformed the best baseline methods in MSE across all forecasting horizons, achieving a maximum relative MSE reduction of 25.4%, with gains concentrated under cloudy conditions. The framework also delivered the best performance in nearly all few-shot settings, where other deep learning baselines degraded substantially and were frequently outperformed by a non-learning physical method.

by read1 min views1 publishedSep 11, 2026

arXiv:2609.11135v1 Announce Type: new Abstract: Short-term photovoltaic (PV) power and global horizontal irradiance (GHI) forecasts are essential for effective dispatch, reserve scheduling, and grid operations. At these forecasting horizons, errors are predominantly driven by cloud induced ramps: relying solely on historical numerical data may struggle to anticipate an incoming cloud, making ground-based sky images a crucial complementary physical signal. Furthermore, forecast performance is highly sensitive to location and local observing conditions, creating a strong need for site-specific data that are often scarce. Recently, large language models (LLMs) have demonstrated competitive performance and high data efficiency in time-series forecasting. Despite their success, existing LLM-based forecasting methods remain predominantly unimodal, relying primarily on historical numerical time-series data. Effectively incorporating sky imagery into an LLM-based forecasting framework remains under-explored and an open challenge. In this paper, we propose SolCloudLLM, an LLM-based multimodal forecasting framework. SolCloudLLM aligns sky-image patches with time-series patches and fuses their corresponding representations through bidirectional multimodal fusion, yielding a unified representation that is subsequently mapped into the embedding space of an LLM. Extensive experiments on the SIRTA and SKIPP'D datasets demonstrate that SolCloudLLM consistently outperforms the best baseline methods in MSE across all forecasting horizons, achieving a maximum relative MSE reduction of 25.4%. Stratified analysis further indicates that the benefits of multimodal fusion are concentrated primarily under cloudy conditions. Notably, SolCloudLLM achieves the best performance in nearly all few-shot settings, whereas other deep learning baselines experience substantial performance degradation and are frequently outperformed by the non-learning physical method.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @solcloudllm 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/bidirectional-multim…] indexed:0 read:1min 2026-09-11 ·