Affect-Prototype Guided Fusion for Open-Vocabulary Incomplete Multi-modal Emotion Recognition Researchers proposed an Affect-Prototype-Conditioned Fusion (APCF) framework for open-vocabulary multimodal emotion recognition that handles incomplete and unsynchronized modal data, according to an arXiv paper (arXiv:2609.16962v1). APCF builds an affect-prototype library to model multimodal contribution characteristics across diverse emotions, then retrieves and aggregates available modal features and feeds the fused representations into an LLM decoder to generate open-vocabulary emotion labels. Experiments on the OV-MERD+ and MER-FG datasets show APCF substantially outperforms state-of-the-art baselines. arXiv:2609.16962v1 Announce Type: new Abstract: Open-vocabulary multimodal emotion recognition OV-MER aims to generate open natural-language emotion labels from multimodal affective cues. In real-world scenarios, however, complete and synchronized modal data are difficult to obtain due to limitations of acquisition devices and user privacy constraints. Existing OV-MER methods are largely designed for full-modal inputs, and fail to perform effective feature fusion under modal missing conditions. Meanwhile, current fusion approaches designed for incomplete modalities mainly focus on fixed-label recognition context, and cannot satisfy the demand for fuse emotional cues guided with arbitrary emotion semantics in OV-MER context. To tackle these challenges, this paper proposes an Affect-Prototype-Conditioned Fusion APCF framework for incomplete open-vocabulary emotion recognition. As a candidate-free generative framework, APCF extends modal contribution learning to scenarios guided by arbitrary emotional semantics. Specifically, we construct an affect-prototype library to explicitly model multimodal contribution characteristics corresponding to diverse emotions, which provides dynamic constraints for modal fusion under different emotional semantic perspectives. Conditional retrieval and feature aggregation are conducted based on available modal features. The refined fused affective representations are then fed into an LLM decoder to produce open-vocabulary emotion labels. Experiments on the OV-MERD+ and MER-FG datasets demonstrate that APCF substantially outperforms state-of-the-art baselines.