# A Comprehensive Review of Multimodal Facial State Analysis: Tasks, Methods, and Resources

> Source: <https://arxiv.org/abs/2609.13255>
> Published: 2026-09-15 04:00:00+00:00

arXiv:2609.13255v1 Announce Type: new 
Abstract: Facial state analysis plays a crucial role in understanding human expressions, psychological modeling, and human computer interaction. Traditional unimodal vision-based methods are often limited by environmental sensitivity and weak interpretability. Multimodal facial state analysis addresses these issues by integrating complementary cues from visual, audio, textual, physiological, and other related modalities. This survey emphasizes two key aspects: on one hand, multimodal learning enables contextual semantic understanding for improved facial state reasoning and leverages interpretable language generation to enhance model explainability; on the other hand, multi-task learning allows simultaneous analysis of expressions, action units (AUs), and face-based soft biometrics (e.g., age, gender), effectively capturing fine-grained expressions and improving cross-scene generalization. This survey reviews core tasks, representative methods, and datasets in multimodal facial state analysis, focusing on facial expression recognition, AU detection, and face-based soft biometric estimation, and emphasizing the unique value of language in providing contextual semantics, enhancing reasoning, and generating explanations. The survey aims to provide an up-to-date overview of the literature and to highlight future research directions for multimodal, interpretable, and multi-task adaptive facial state analysis.
