DisasterTD: Disaster Toponym Disambiguation Using Multimodal LLMs and Cross-View Geolocalization Researchers propose DisasterTD, a disaster toponym disambiguation framework that integrates multimodal large language models (MLLMs) with cross-view geolocalization to improve geolocalization accuracy of social media imagery during disasters. On the Hurricane Harvey dataset, DisasterTD achieves geolocalization accuracies of 71.62% within 1000 meters and reduces mean error to 11.33 km, outperforming MLLM-only and cross-view-only baselines. arXiv:2607.24856v1 Announce Type: new Abstract: Social media imagery SMI provides timely and fine-grained ground perspectives that are valuable for situational awareness and emergency response. Unlike satellite or aerial imagery, SMI can capture disaster impacts and ground-level conditions in a timely manner. However, geographic references in SMI are often vague or ambiguous, making accurate geolocalization challenging. To address this issue, we propose DisasterTD, a disaster toponym disambiguation framework that integrates multimodal large language model MLLMs -based semantic reasoning with cross-view geolocalization. First, MLLMs extract toponyms and generate candidate geolocations from noisy textual inputs. Then, cross-view matching between SMI, remote sensing imagery RSI , and optionally street-view imagery SVI is used to verify and refine these candidate results. We evaluate DisasterTD on the Hurricane Harvey dataset, where SMI is augmented with collected RSI and SVI to construct a cross-view benchmark for disaster geolocalization. The dataset is divided into four categories based on toponym clarity and ambiguity, allowing a fine-grained performance analysis across scenarios. Results show that DisasterTD consistently outperforms MLLM-only and cross-view-only baselines without disambiguation, achieving geolocalization accuracies of 71.62% within 1000 m, 62.36% within 500 m, 57.99% within 250 m, 52.09% within 100 m, and 47.01% within 50 m, while reducing the mean and median errors to 11.33 km and 0.68 km, respectively. The largest improvements appear in ambiguous toponyms, where semantic reasoning with cross-view evidence reduces candidate dispersion and errors. These findings demonstrate the effectiveness of integrating MLLM-based candidate generation with cross-view verification for fine-grained disaster geolocalization.