arXiv:2609.36082v1 Announce Type: new Abstract: We introduce GeoOutageBench, a benchmark for assessing LLM-based geospatiotemporal KGQA for multimodal outage and resilience analysis. Unlike existing KGQA benchmarks for Web knowledge, GeoOutageBench considers a spatiotemporal KG that integrates visual, textual, and structured data from outage records, remote sensing, weather observations, storm and power events, geographic entities, and domain ontologies. It provides a competency query taxonomy at different difficulty levels from spatiotemporal containment and proximity, spatiotemporal co-occurrence analysis, multimodal evidence, to hypothetical evaluation. Over multimodal KG and query classes, GeoOutageBench provides user-configurable evaluation of three important, highly coherent yet less studied tasks: (1) LLMs' understanding for ambiguous geospatiotemporal questions in terms of NL to SPARQL interpretation, (2) query-driven assessment of ontology utility, and (3) answer accuracy of multimodal KGQA retrieval. GeoOutageBench provides a design principle and foundation for assessing LLM-KG systems that support real-world infrastructure resilience analysis. Our benchmark, source code, data, results, and other documentation are available at https://github.com/UCF-SAGE/GeoOutageBench.
GeoOutageBench: Benchmarking Ambiguity-aware, Ontology-grounded Geospatiotemporal KGQA for Multimodal Power Outage and Resilience Analysis
Researchers released GeoOutageBench, a benchmark for evaluating LLM-based geospatiotemporal knowledge-graph question answering (KGQA) on multimodal power outage and resilience analysis, with code, data, and results available at https://github.com/UCF-SAGE/GeoOutageBench. GeoOutageBench builds a spatiotemporal knowledge graph integrating visual, textual, and structured data from outage records, remote sensing, weather observations, storm and power events, geographic entities, and domain ontologies, and organizes queries into a competency taxonomy spanning spatiotemporal containment and proximity, spatiotemporal co-occurrence, multimodal evidence, and hypothetical evaluation. The benchmark evaluates three tasks: LLM interpretation of ambiguous geospatiotemporal questions via NL-to-SPARQL translation, query-driven assessment of ontology utility, and answer accuracy of multimodal KGQA retrieval.
Run your AI side-project on zahid.host
EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.