{"slug": "evaluating-multimodal-llms-across-text-and-audio-modalities-for-accessible", "title": "Evaluating Multimodal LLMs across Text and Audio Modalities for Accessible Disaster Assistance", "summary": "A new study from arXiv (2608.14651v1) evaluating open-weight Multi-Modal Large Language Models (MM-LLMs) for disaster assistance found that no model achieves reliable consistency across text and audio modalities, with performance gaps heightened for vulnerable personas such as hard-of-hearing individuals, pregnant women, mothers with toddlers, and elderly individuals with dementia. The findings highlight modality-dependent inequity that undermines the humanitarian value of these AI systems and inform design recommendations for equitable disaster risk communication tools.", "body_md": "arXiv:2608.14651v1 Announce Type: new\nAbstract: Effective disaster risk communication is a foundational humanitarian challenge, yet current emergency infrastructure fails to meet the needs of individuals with access and functional needs, including hard-of-hearing individuals, pregnant women, mothers with toddlers, and elderly individuals with dementia. Recent advancements in Artificial Intelligence (AI), especially Multi-Modal Large Language Models (MM-LLMs), demonstrate powerful capabilities to serve diverse users across text, audio, image, and video modalities within a single unified system, such as a chatbot. However, their suitability for deployment rests on a property that receives limited scrutiny, i.e., whether these systems produce consistent, actionable outputs regardless of the modality through which a user communicates. In this paper, we conduct a comprehensive analysis to understand the status of open-weight MM-LLMs using real emergency alert scenarios across four different vulnerable personas. These state-of-the-art (SOTA) models are evaluated on consistency of responses across text and audio modalities when the same task scenario is given. Findings indicate that no model achieves reliable consistency across modalities, and that performance gaps are heightened for personas with access needs, introducing modality-dependent inequity that undermines the humanitarian value of these systems. These results inform concrete design recommendations for building equitable, trustworthy, and inclusive AI tools for disaster risk communication.", "url": "https://wpnews.pro/news/evaluating-multimodal-llms-across-text-and-audio-modalities-for-accessible", "canonical_source": "https://www.machinebrief.com/news/evaluating-multimodal-llms-across-text-and-audio-modalities-xwaw", "published_at": "2026-08-18 04:00:00+00:00", "updated_at": "2026-08-18 05:41:14.504769+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-ethics"], "entities": ["arXiv", "Multi-Modal Large Language Models"], "alternates": {"html": "https://wpnews.pro/news/evaluating-multimodal-llms-across-text-and-audio-modalities-for-accessible", "markdown": "https://wpnews.pro/news/evaluating-multimodal-llms-across-text-and-audio-modalities-for-accessible.md", "text": "https://wpnews.pro/news/evaluating-multimodal-llms-across-text-and-audio-modalities-for-accessible.txt", "jsonld": "https://wpnews.pro/news/evaluating-multimodal-llms-across-text-and-audio-modalities-for-accessible.jsonld"}}