{"slug": "recurrence-meets-transformers-for-universal-multimodal-retrieval", "title": "Recurrence Meets Transformers for Universal Multimodal Retrieval", "summary": "Researchers from the University of Modena and Reggio Emilia's AImageLab introduced ReT-2, a unified multimodal retrieval model supporting queries and documents composed of both images and text. ReT-2 uses a recurrent Transformer architecture with LSTM-inspired gating to integrate multi-layer representations, achieving state-of-the-art results on the M2KR and M-BEIR benchmarks while offering faster inference and lower memory usage. Integrated into retrieval-augmented generation, it also improves downstream performance on Encyclopedic-VQA and InfoSeek datasets.", "body_md": "arXiv:2509.08897v3 Announce Type: replace-cross\nAbstract: With the rapid advancement of multimodal retrieval and its application in LLMs and multimodal LLMs, increasingly complex retrieval tasks have emerged. Existing methods predominantly rely on task-specific fine-tuning of vision-language models and are limited to single-modality queries or documents. In this paper, we propose ReT-2, a unified retrieval model that supports multimodal queries, composed of both images and text, and searches across multimodal document collections where text and images coexist. ReT-2 leverages multi-layer representations and a recurrent Transformer architecture with LSTM-inspired gating mechanisms to dynamically integrate information across layers and modalities, capturing fine-grained visual and textual details. We evaluate ReT-2 on the challenging M2KR and M-BEIR benchmarks across different retrieval configurations. Results demonstrate that ReT-2 consistently achieves state-of-the-art performance across diverse settings, while offering faster inference and reduced memory usage compared to prior approaches. When integrated into retrieval-augmented generation pipelines, ReT-2 also improves downstream performance on Encyclopedic-VQA and InfoSeek datasets. Our source code and trained models are publicly available at: https://github.com/aimagelab/ReT-2", "url": "https://wpnews.pro/news/recurrence-meets-transformers-for-universal-multimodal-retrieval", "canonical_source": "https://www.machinebrief.com/news/recurrence-meets-transformers-for-universal-multimodal-retri-yzn9", "published_at": "2026-08-29 04:00:00+00:00", "updated_at": "2026-08-29 04:48:10.593323+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "large-language-models", "ai-research", "ai-products"], "entities": ["ReT-2", "AImageLab", "University of Modena and Reggio Emilia", "M2KR", "M-BEIR", "Encyclopedic-VQA", "InfoSeek"], "alternates": {"html": "https://wpnews.pro/news/recurrence-meets-transformers-for-universal-multimodal-retrieval", "markdown": "https://wpnews.pro/news/recurrence-meets-transformers-for-universal-multimodal-retrieval.md", "text": "https://wpnews.pro/news/recurrence-meets-transformers-for-universal-multimodal-retrieval.txt", "jsonld": "https://wpnews.pro/news/recurrence-meets-transformers-for-universal-multimodal-retrieval.jsonld"}}