cd /news/machine-learning/agrobench-a-reproducible-multimodal-… · home topics machine-learning article
[ARTICLE · art-138777] src=arxiv.org ↗ pub= topic=machine-learning verified=true sentiment=· neutral

AgroBench: A Reproducible Multimodal Benchmark for Weakly Supervised Crop Yield Learning from County Statistics and Pixel Observations

Researchers released AgroBench, a reproducible multimodal benchmark that converts U.S. county-level crop yield statistics into weakly supervised pixel-level crop time series, containing over 13 million observations from 788,654 unique crop pixels across 5,107 county-year combinations for five major U.S. crops over eight growing seasons from 2017 to 2024. The pipeline integrates USDA crop yield statistics with crop-specific land cover masks, Sentinel-2 multispectral imagery, Sentinel-1 synthetic aperture radar, climatic variables, and terrain data, pairing each crop pixel time series with a county-level yield value as a weak supervisory signal. The authors provide baseline results under a Leave-One-Year-Out evaluation protocol and release the full data generation pipeline, dataset, and evaluation protocol to support weakly supervised learning, multimodal remote sensing, spatiotemporal modeling, and geospatial foundation models for agriculture.

by read1 min views1 publishedSep 24, 2026

arXiv:2609.26809v1 Announce Type: new Abstract: Reliable agricultural yield statistics are typically reported at coarse administrative scales, whereas modern geospatial machine learning methods require spatially explicit, pixel level supervision. This mismatch has limited the development of large-scale benchmarks for crop yield learning using multimodal Earth observation data. A reproducible benchmark, AgroBench, is presented for transforming publicly available U.S. county level crop yield statistics into weakly supervised pixel-level crop time series. Each crop pixel time series is paired with a county-level yield value as a weak supervisory signal rather than a directly measured pixel-level yield label. Our geospatial data generation pipeline integrates USDA crop yield statistics with crop-specific land cover masks, Sentinel 2 multispectral imagery, Sentinel-1 synthetic aperture radar observations, climatic variables, and terrain information to produce temporally aligned multimodal sequences describing individual crop pixels throughout the growing season. The resulting benchmark contains over 13 million observations from 788,654 unique crop pixels spanning 5,107 county year combinations across eight growing seasons (2017 to 2024) for five major U.S. crops. To facilitate standardized evaluation, we establish a crop yield prediction benchmark using a Leave-One-Year-Out evaluation protocol and provide baseline results using representative machine learning models. By releasing the complete data generation pipeline, benchmark dataset, and evaluation protocol, AgroBench provides a reproducible foundation for future research in weakly supervised learning, multimodal remote sensing, spatiotemporal modeling, and geospatial foundation models for agriculture.

── more in #machine-learning 4 stories · sorted by recency
── more on @agrobench 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/agrobench-a-reproduc…] indexed:0 read:1min 2026-09-24 ·