{"slug": "fusing-perceptual-vision-experts-with-multimodal-large-language-models-for-plant", "title": "Fusing Perceptual Vision Experts with Multimodal Large Language Models for Explainable Plant Disease Diagnosis: From Benchmark Imagery to Real-World Robotic Field Validation", "summary": "Researchers introduced the Hybrid Hierarchical Multi-Agent Framework (H²MAF), which fuses EfficientNet-B3 and ConvNeXt-Tiny with multimodal large language models Gemma 4 E4B and Qwen3.5 4B to provide explainable plant disease diagnoses. On the PlantDoc benchmark, Gemma raised accuracy from 63.9% to 68.5%, and on Cornell field datasets, accuracies reached 99.3% and 98.9%, though Gemma showed lower critical-risk error than Qwen, indicating calibration-dependent utility for robotic field decision support.", "body_md": "arXiv:2608.24934v1 Announce Type: new\nAbstract: Accurate field plant disease diagnosis requires reliable fusion of uncertain and conflicting perceptual evidence. We present the Hybrid Hierarchical Multi-Agent Framework (H$^{2}$MAF), combining decision-level fusion of EfficientNet-B3 and ConvNeXt-Tiny with semantic arbitration by open-weight multimodal large language models (MLLMs), Gemma 4 E4B and Qwen3.5 4B, using structured JSON evidence to generate explainable diagnoses, risk levels, treatment urgency, and financial exposure. (H$^{2}$MAF) is evaluated on 14,364 images (1,370 test images) across PlantDoc (2,922 images, 27 classes) and two non-public, continuously captured Cornell robot-acquired field datasets: Stage 2 (20 GB; 4,215 images) and Stage 4 (40 GB; 7,227 images), covering Early Blight, Late Blight, and Septoria Leaf Spot under uncontrolled field conditions. On PlantDoc, Gemma improves accuracy from 63.9% to 68.5%, achieving +7.6 points on the 41.7% CNN-conflict subset. Cornell accuracies reach 99.3% and 98.9%, with only 1.7-4.1% disagreement, demonstrating conflict-dependent MLLM utility. The critical-risk error of gemma is 0.14-0.5 points, whereas Qwen overflags by 3.5-14.4 points. These results establish MLLM arbitration as a promising, yet calibration-dependent, approach for explainable agricultural AI and robotic field decision support. Github Link: https://github.com/Applied-AI-Research-Lab/Explainable-AI-Plant-Disease-Detection", "url": "https://wpnews.pro/news/fusing-perceptual-vision-experts-with-multimodal-large-language-models-for-plant", "canonical_source": "https://arxiv.org/abs/2608.24934", "published_at": "2026-08-27 04:00:00+00:00", "updated_at": "2026-08-27 04:21:27.037331+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "computer-vision", "large-language-models", "ai-research"], "entities": ["EfficientNet-B3", "ConvNeXt-Tiny", "Gemma 4 E4B", "Qwen3.5 4B", "PlantDoc", "Cornell", "H²MAF"], "alternates": {"html": "https://wpnews.pro/news/fusing-perceptual-vision-experts-with-multimodal-large-language-models-for-plant", "markdown": "https://wpnews.pro/news/fusing-perceptual-vision-experts-with-multimodal-large-language-models-for-plant.md", "text": "https://wpnews.pro/news/fusing-perceptual-vision-experts-with-multimodal-large-language-models-for-plant.txt", "jsonld": "https://wpnews.pro/news/fusing-perceptual-vision-experts-with-multimodal-large-language-models-for-plant.jsonld"}}