[1] 孙强, 王姝玉. 结合时间注意力机制和单模态标签自动生成策略的自监督多模态情感识别[J]. 电子与信息学报, 2024, 46(2): 588-601. SUN Q, WANG S Y. Self-supervised multimodal emotion recognition combining temporal attention mechanism and unimodal label automatic generation strategy[J]. Journal of Electronics & Information Technology, 2024, 46(2): 588-601. (in Chinese) [2] JIANG Y Y, LI W, HOSSAIN M S, et al. A snapshot research and implementation of multimodal information fusion for data-driven emotion recognition[J]. Information Fusion, 2020, 53: 209-221. [3] TSAI Y H, BAI S J, LIANG P P, et al. Multimodal transformer for unaligned multimodal language sequences[C]//Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. Florence, Italy: Association for Computational Linguistics, 2019: 6558-6569. [4] LI M C, YANG D K, LEI Y X, et al. A unified self-distillation framework for multimodal sentiment analysis with uncertain missing modalities[C]//Proceedings of the AAAI Conference on Artificial Intelligence. [S.l.]: AAAI Press, 2024: 10074-10082. [5] 殷兵, 凌震华, 林垠, 等. 兼容缺失模态推理的情感识别方法[J]. 计算机应用, 2025, 45(9): 2764-2772. YIN B, LING Z H, LIN Y, et al. A sentiment recognition method compatible with missing modal reasoning [J]. Journal of Computer Applications, 2025, 45(9): 2764-2772. (in Chinese) [6] 任楚岚, 于振坤, 关超, 等. 基于自适应融合技术的多模态实体对齐模型[J]. 计算机应用研究, 2025, 42(1): 100-105. REN C L, YU Z K, GUAN C, et al. Multi-modal entity alignment model based on adaptive fusion technology[J]. Application Research of Computers, 2025, 42(1): 100-105. (in Chinese) [7] HAZARIKA D, ZIMMERMANN R, PORIA S. MISA: modality-invariant and -specific representations for multimodal sentiment analysis[C]//Proceedings of the 28th ACM International Conference on Multimedia. New York, USA: ACM Press, 2020: 1122-1131. [8] HUANG J H, ZHOU J, TANG Z C, et al. TMBL: Transformer-based multimodal binding learning model for multimodal sentiment analysis[J]. Knowledge-Based Systems, 2024, 285: 111346. [9] LI Z M, ZHOU Y, ZHANG W B, et al. AMOA: global acoustic feature enhanced modal-order-aware network for multimodal sentiment analysis[C]//Proceedings of the 29th International Conference on Computational Linguistics. Gyeongju, Republic of Korea: Association for Computational Linguistics, 2022: 7136-7146. [10] RAHMAN W, HASAN M K, LEE S, et al. Integrating multimodal information in large pretrained transformers[C]//Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. Online: Association for Computational Linguistics, 2020: 2359-2369. [11] FENG W, HU L Y, SHANG F H, et al. Deep correlated prompting for visual recognition with missing modalities[M]//OH A, NAUMANN T, GLOBERSON A. Advances in neural information processing systems 37. Vancouver, Canada: Neural Information Processing Systems Foundation, Inc. (NeurIPS), 2024: 67446-67466. [12] WANG H, CHEN Y H, MA C B, et al. Multi-modal learning with missing modality via shared-specific feature modelling[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Vancouver, Canada: IEEE Press, 2023: 15878-15887. [13] MA M M, REN J, ZHAO L, et al. SMIL: multimodal learning with severely missing modality[C]//Proceedings of the AAAI Conference on Artificial Intelligence. [S.l.]: AAAI Press, 2021: 2302-2310. [14] 王嫄, 邓振宇, 王佳鑫, 等. 基于两阶段缺失模态恢复的多模态情感分析方法[J]. 天津科技大学学报, 2025(1): 57-63, 80. WANG Y, DENG Z Y, WANG J X, et al. Multimodal sentiment analysis approach based on two-stage missing modality recovery[J]. Journal of Tianjin University of Science & Technology, 2025(1): 57-63, 80. (in Chinese) [15] ZHAO J M, LI R C, JIN Q. Missing modality imagination network for emotion recognition with uncertain missing modalities[C]//Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). Online: Association for Computational Linguistics, 2021: 2608-2618. [16] HEINZERLING B, INUI K. Language models as knowledge bases: on entity representations, storage capacity, and paraphrased queries[EB/OL]. [2025-03-01]. https://arxiv.org/abs/2008.09036. [17] KHATTAK M U, RASHEED H, MAAZ M, et al. MaPLe: multi-modal prompt learning[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Vancouver, Canada: IEEE Press, 2023: 19113-19122. [18] TSIMPOUKELLI M, MENICK J, CABI S, et al. Multimodal few-shot learning with frozen language models[EB/OL]. [2025-03-01]. https://arxiv.org/abs/2106.13884. [19] LEE Y L, TSAI Y H, CHIU W C, et al. Multimodal prompting with missing modalities for visual recognition[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Vancouver, Canada: IEEE Press, 2023: 14943-14952. [20] ZADEH A, CHEN M H, PORIA S, et al. Tensor fusion network for multimodal sentiment analysis[EB/OL]. [2025-03-01]. https://arxiv.org/abs/1707.07250. [21] ZADEH A, LIANG P P, MAZUMDER N, et al. Memory fusion network for multi-view sequential learning[C]//Proceedings of the AAAI Conference on Artificial Intelligence. [S.l.]: AAAI Press, 2018: 1-10. [22] SUN H, WANG H Y, LIU J Q, et al. CubeMLP: an MLP-based model for multimodal sentiment analysis and depression estimation[C]//Proceedings of the 30th ACM International Conference on Multimedia. New York, USA: ACM Press, 2022: 3722-3729. [23] WILLIAMS J, KLEINEGESSE S, COMANESCU R, et al. Recognizing emotions in video using multimodal DNN feature fusion[C]//Proceedings of Grand Challenge and Workshop on Human Multimodal Language (Challenge-HML). Melbourne, Australia: Association for Computational Linguistics, 2018: 11-19. [24] LIU Z, SHEN Y, LAKSHMINARASIMHAN V B, et al. Efficient low-rank multimodal fusion with modality-specific factors[EB/OL]. [2025-03-01]. https://arxiv.org/abs/1806.00064. [25] BAGHER ZADEH A, LIANG P P, PORIA S, et al. Multimodal language analysis in the wild: CMU-MOSEI dataset and interpretable dynamic fusion graph[C]//Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Melbourne, Australia: Association for Computational Linguistics, 2018: 2236-2246. [26] TSAI Y H, LIANG P P, ZADEH A, et al. Learning factorized multimodal representations[EB/OL]. [2025-03-01]. https://arxiv.org/abs/1806.06176. [27] YU W M, XU H, YUAN Z Q, et al. Learning modality-specific representations with self-supervised multi-task learning for multimodal sentiment analysis[C]//Proceedings of the AAAI Conference on Artificial Intelligence. [S.l.]: AAAI Press, 2021: 10790-10797. [28] MAI S J, ZENG Y, ZHENG S J, et al. Hybrid contrastive learning of tri-modal representation for multimodal sentiment analysis[J]. IEEE Transactions on Affective Computing, 2023, 14(3): 2276-2289. [29] 王楠, 王淇, 欧阳丹彤. 基于知识蒸馏与动态调整机制的多模态情感分析模型[J]. 计算机学报, 2025, 48(8): 1923-1942. WANG N, WANG Q, OUYANG D T, et al. A multimodal sentiment analysis model based on knowledge distillation and dynamic adjustment mechanism [J]. Chinese Journal of Computers, 2025, 48(8): 1923-1942. (in Chinese). [30] LIU S, LUO Z, FU W N. FCDNet: fuzzy cognition-based dynamic fusion network for multimodal sentiment analysis[J]. IEEE Transactions on Fuzzy Systems, 2025, 33(1): 3-14. [31] WANG D, GUO X T, TIAN Y M, et al. TETFN: a text enhanced transformer fusion network for multimodal sentiment analysis[J]. Pattern Recognition, 2023, 136: 109259. [32] HAN W, CHEN H, PORIA S. Improving multimodal fusion with hierarchical mutual information maximization for multimodal sentiment analysis[EB/OL]. [2025-03-01]. https://arxiv.org/abs/2109.00412. [33] ZHANG H Y, WANG Y, YIN G H, et al. Learning language-guided adaptive hyper-modality representation for multimodal sentiment analysis[EB/OL]. [2025-03-01]. https://arxiv.org/abs/2310.05804. [34] FENG X Y, LIN Y M, HE L H, et al. Knowledge-guided dynamic modality attention fusion framework for multimodal sentiment analysis[EB/OL]. [2025-03-01]. https://arxiv.org/abs/2410.04491. [35] DEGOTTEX G, KANE J, DRUGMAN T, et al. COVAREP—a collaborative voice analysis repository for speech technologies[C]//Proceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). Florence, Italy: IEEE Press, 2014: 960-964. [36] EKMAN P, FRIESEN W V, SIMONS R C. Is the startle reaction an emotion?[M]//EKMAN P, ROSENBERG E L. What the face reveals: basic and applied studies of spontaneous expression using the Facial Action Coding System (FACS). Oxford, USA: Oxford University Press, 2005: 21-38. [37] DEVLIN J, CHANG M W, LEE K, et al. BERT: pre-training of deep bidirectional transformers for language understanding[C]//Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1(Long and Short Papers). Minneapolis, USA: Association for Computational Linguistics, 2019: 4171-4186. [38] ZADEH A, ZELLERS R, PINCUS E, et al. MOSI: multimodal corpus of sentiment intensity and subjectivity analysis in online opinion videos[EB/OL]. [2025-03-01]. https://arxiv.org/abs/1606.06259. |