{"slug": "visual-saliency-steering-distillation-for-multimodal-chain-of-thought-reasoning", "title": "Visual Saliency Steering Distillation for Multimodal Chain-of-Thought Reasoning", "summary": "Researchers propose Visual Saliency Steering Distillation (VSSD) to improve multimodal chain-of-thought reasoning in small models by using attention maps from multimodal large language models to generate perturbed images and applying singular value decomposition for inter-layer distillation. Experiments on ScienceQA and M^3CoT show VSSD enhances rationale generation and answer inference.", "body_md": "arXiv:2607.22013v1 Announce Type: new\nAbstract: Multimodal chain-of-thought (CoT) reasoning integrates visual and textual cues through step-by-step inference. In small models with limited token budgets, modality-interaction fusion often suppresses tiny cross-modal differences. In particular, multimodal CoT often struggles when different images pair with identical text or different texts pair with an identical image, making such inputs nearly indistinguishable after fusion. This study proposes Visual Saliency Steering Distillation (VSSD). VSSD leverages the attention maps of multimodal large language models to generate perturbed images that capture task-sensitive feature directions, and then applies singular value decomposition to extract dominant steering vectors to guide inter-layer distillation. Experiments on ScienceQA and M$^3$CoT demonstrate that VSSD improves rationale generation and answer inference. The code is available at https://github.com/BGWH123/VSSD.", "url": "https://wpnews.pro/news/visual-saliency-steering-distillation-for-multimodal-chain-of-thought-reasoning", "canonical_source": "https://arxiv.org/abs/2607.22013", "published_at": "2026-07-27 04:00:00+00:00", "updated_at": "2026-07-27 04:28:05.000789+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "large-language-models", "computer-vision", "natural-language-processing"], "entities": ["Visual Saliency Steering Distillation", "ScienceQA", "M^3CoT", "arXiv"], "alternates": {"html": "https://wpnews.pro/news/visual-saliency-steering-distillation-for-multimodal-chain-of-thought-reasoning", "markdown": "https://wpnews.pro/news/visual-saliency-steering-distillation-for-multimodal-chain-of-thought-reasoning.md", "text": "https://wpnews.pro/news/visual-saliency-steering-distillation-for-multimodal-chain-of-thought-reasoning.txt", "jsonld": "https://wpnews.pro/news/visual-saliency-steering-distillation-for-multimodal-chain-of-thought-reasoning.jsonld"}}