RAMamba-Net: A Reliability-Aware and Mamba-Based Multimodal Fusion Network for Auditory Attention Detection Researchers introduced RAMamba-Net, a reliability-aware Mamba-based multimodal fusion network for auditory attention decoding (AAD) that combines electroencephalography (EEG) and electrooculography (EOG) signals, according to an arXiv paper (arXiv:2609.11372v1). Experiments on two AAD benchmarks showed RAMamba-Net delivered accuracy gains of 5.76% over unimodal baselines, along with more robust decoding and discriminative representations. The network uses a Mamba-enhanced band-aware convolutional Transformer for band-specific EEG patterns, a dual-branch temporal-spatial encoder for EOG dependencies, cross-modal attention for modality interaction, and a reliability-aware module that estimates sample-wise modality weights. arXiv:2609.11372v1 Announce Type: new Abstract: Auditory attention decoding AAD identifies the attended speaker from physiological signals, supporting neuro-steered hearing devices and natural human-machine interaction. Electroencephalography EEG is the dominant modality for AAD but provides incomplete evidence in naturalistic audio-visual scenes, motivating EEG and electrooculography EOG fusion. Existing approaches remain limited by weak cross-modal interaction, inefficient temporal modeling, and low robustness to sample variations. To address the limitations, we propose RAMamba-Net, a reliability-aware Mamba-based multimodal fusion network for AAD. RAMamba-Net employs a Mamba-enhanced band-aware convolutional Transformer to capture band-specific EEG patterns and long-range temporal dynamics. A dual-branch temporal-spatial encoder models EOG temporal and inter-channel dependencies. Cross-modal attention enables explicit modality interaction. Then, a reliability-aware module is introduced to estimate sample-wise modality weights for feature and prediction consistency, thereby enhancing multimodal fusion. Experiments on two AAD benchmarks demonstrate that RAMamba-Net effectively exploits complementary EEG-EOG information, yielding accuracy gains of 5.76% over unimodal baselines, together with more robust decoding and discriminative representations. Further analyses show that explicit cross-modal interaction improves multimodal alignment, while the reliability-aware module suppresses unreliable modality evidence and is robust to signal perturbation and parameter variation.