{"slug": "decoding-the-disaster-multi-task-geospatial-reasoning-with-vision-language-and", "title": "Decoding the Disaster: Multi-Task Geospatial Reasoning with Vision-Language Models and Crowdsourced Imagery for Disaster Mapping", "summary": "A new arXiv paper (2610.00302v1) introduces GRDisaster, a multi-task geospatial reasoning framework that uses vision-language models to geolocalize and interpret crowdsourced disaster imagery. Built on a benchmark of 26,340 images from PhotoMappers — organized into human-validated volunteered geographic information, street-view imagery, and remote sensing cross-view triplets spanning disaster events from 2018 to 2024 — the framework combines deterministic and probabilistic cross-view geolocalization with multi-view fusion. The authors describe it as the first systematic investigation and unified evaluation framework for turning crowdsourced disaster imagery into actionable geospatial AI through cross-view geolocalization validation, interpretable spatial reasoning, and damage-aware severity assessment.", "body_md": "arXiv:2610.00302v1 Announce Type: new \nAbstract: Crowdsourced imagery provides timely, fine-grained, street-level observations for disaster mapping, complementing conventional remote sensing imagery (RSI) during emergency response. However, such imagery is often unstructured, spatially ambiguous, and lacks reliable geographic metadata, making manual geolocalization and interpretation labor-intensive and difficult to scale. This work proposes a multi-task Geospatial Reasoning Disaster mapping framework, namely GRDisaster, to examine the potential of vision-language models (VLMs) in understanding, geolocalizing, and reasoning over crowdsourced disaster imagery. GRDisaster is built on a newly curated benchmark dataset derived from PhotoMappers, comprising 26,340 images organized into human-validated volunteered geographic information (VGI), street-view imagery (SVI), RSI cross-view triplets covering multiple disaster events from 2018 to 2024. The framework combines deterministic and probabilistic cross-view geolocalization with multi-view fusion to associate VGI images with georeferenced SVI and RSI. It introduces two sets of spatial reasoning indicators for cross-view geolocalization validation and disaster damage assessment. These indicators use structural, environmental, and global-scene cues to validate cross-view correspondences and visually observable damage evidence with expert-verified annotations to assess disaster severity, improving the interpretability of VLM outputs. To our knowledge, this study provides the first systematic investigation and unified evaluation framework for examining how VLM-based spatial reasoning can transform crowdsourced disaster imagery into actionable geospatial artificial intelligence (GeoAI) through cross-view geolocalization validation, interpretable spatial reasoning, and damage-aware severity assessment.", "url": "https://wpnews.pro/news/decoding-the-disaster-multi-task-geospatial-reasoning-with-vision-language-and", "canonical_source": "https://arxiv.org/abs/2610.00302", "published_at": "2026-10-02 04:00:00+00:00", "updated_at": "2026-10-02 04:17:52.866033+00:00", "lang": "en", "topics": ["computer-vision", "ai-research", "artificial-intelligence", "large-language-models"], "entities": ["GRDisaster", "PhotoMappers", "arXiv"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/decoding-the-disaster-multi-task-geospatial-reasoning-with-vision-language-and", "markdown": "https://wpnews.pro/news/decoding-the-disaster-multi-task-geospatial-reasoning-with-vision-language-and.md", "text": "https://wpnews.pro/news/decoding-the-disaster-multi-task-geospatial-reasoning-with-vision-language-and.txt", "jsonld": "https://wpnews.pro/news/decoding-the-disaster-multi-task-geospatial-reasoning-with-vision-language-and.jsonld"}}