From Stress to Affect: Multimodal Deep Learning for Physiological Emotion Recognition Across Wearable Sensor Modalities A comparative study of temporal deep learning architectures for physiological emotion recognition found the Transformer achieved the highest multimodal accuracy on the WESAD dataset at 99.02% +/- 0.51%, while bidirectional LSTM led on EmoWear for arousal (91.80% +/- 1.06%) and valence (89.96% +/- 0.36%), according to the arXiv paper 2609.20991v1. The study evaluated bidirectional LSTM, temporal convolutional network (TCN), and Transformer models under wrist-only, chest-only, and multimodal sensing configurations using participant-independent leave-one-subject-out cross-validation (LOSO-CV) on the WESAD and EmoWear datasets. Multimodal sensing consistently outperformed single-site configurations, and sampling-frequency analysis identified 4 Hz as a practical operating point with performance comparable to higher frequencies at substantially lower training cost. arXiv:2609.20991v1 Announce Type: new Abstract: Physiological emotion recognition using wearable sensors has important applications in mental health monitoring, affective computing, and human-computer interaction. However, existing studies typically evaluate a single model, sensing configuration, or dataset, limiting our understanding of how these factors influence recognition performance. We present a comparative study of temporal deep learning architectures for physiological emotion recognition using two multimodal wearable datasets: WESAD and EmoWear. Bidirectional long short-term memory LSTM , temporal convolutional network TCN , and Transformer models are evaluated under wrist-only, chest-only, and multimodal sensing configurations using participant-independent leave-one-subject-out cross-validation LOSO-CV . We also investigate soft-voting ensembles, sensor ablation, sampling frequency, and gradient-based saliency. The Transformer achieved the highest multimodal accuracy on WESAD 99.02% +/- 0.51% , whereas the LSTM achieved the best multimodal accuracy on EmoWear for both arousal 91.80% +/- 1.06% and valence 89.96% +/- 0.36% . These results show that relative architecture performance depends on dataset characteristics rather than one architecture being uniformly superior. Multimodal sensing consistently outperformed wrist-only and chest-only configurations across both datasets. Sampling-frequency analysis showed that 4 Hz provides a practical operating point, with performance comparable to higher frequencies at substantially lower training cost. These findings provide guidance for selecting architectures, sensing modalities, and sampling frequencies for wearable physiological emotion recognition.