[1] Baltrušaitis T, Ahuja C, Morency L P. Multimodal machine learning: A survey and taxonomy[J]. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2019, 41(2): 423-443.
[2] Zadeh A, Zellers R, Pincus E, et al. Multimodal sentiment intensity analysis in videos: Facial gestures and verbal messages[J]. IEEE Intelligent Systems, 2016, 31(6): 82-88.
[3] Zadeh A, Liang P P, Poria S, et al. Multimodal language analysis in the wild: CMU-MOSEI dataset and interpretable dynamic fusion graph[C]//Proceedings of ACL. 2018: 2236-2246.
[4] Zadeh A, Chen M, Poria S, et al. Tensor fusion network for multimodal sentiment analysis[C]//Proceedings of EMNLP. 2017: 1103-1114.
[5] Liu Z, Shen Y, Lakshminarasimhan V B, et al. Efficient low-rank multimodal fusion with modality-specific factors[C]//Proceedings of ACL. 2018: 2247-2256.
[6] Tsai Y H H, Bai S, Liang P P, et al. Multimodal transformer for unaligned multimodal language sequences[C]//Proceedings of ACL. 2019: 6558-6569.
[7] Hazarika D, Zimmermann R, Poria S. MISA: Modality-invariant and modality-specific representations for multimodal sentiment analysis[C]//Proceedings of ACM MM. 2020: 1122-1131.
[8] Rahman W, Hasan M K, Lee S, et al. Integrating multimodal information in large pretrained transformers[C]//Proceedings of ACL. 2020: 2359-2369.
[9] He K, Zhang X, Ren S, et al. Deep residual learning for image recognition[C]//Proceedings of CVPR. 2016: 770-778.
[10] Han W, Chen H, Poria S. Improving multimodal fusion with hierarchical mutual information maximization for multimodal sentiment analysis[C]//Proceedings of EMNLP. 2021: 9180-9192.
[11] Wu Z, Weng Y, Jin Y, et al. Multimodal multi-loss fusion network for sentiment analysis[C]//Proceedings of NAACL. 2024: 3546-3557.
[12] He X, Liang H, Peng B, et al. MSAmba: Exploring multimodal sentiment analysis with state space models[C]//Proceedings of AAAI. 2025, 39(2): 1309-1317.
[13] Yu W, Xu H, Yuan Z, et al. Learning modality-specific representations with self-supervised multi-task learning for multimodal sentiment analysis[C]//Proceedings of AAAI. 2021, 35(12): 10790-10797.
[14] Xie Y, Zhu Z, Lu X, et al. InfoEnh: Towards multimodal sentiment analysis via information bottleneck filter and optimal transport alignment[C]//Proceedings of LREC-COLING. 2024: 9073-9083.
[15] Yang Y, Du Y, Hu X, et al. CLGSI: A multimodal sentiment analysis framework based on contrastive learning guided by sentiment intensity[C]//Findings of NAACL. 2024: 2147-2160.
[16] Li Z, Li L. t-HNE: A text-guided hierarchical noise eliminator for multimodal sentiment analysis[C]//Proceedings of COLING. 2025: 2834-2844.
[17] Wu S, He D, Wang X, et al. Enriching multimodal sentiment analysis through textual emotional descriptions of visual-audio content[C]//Proceedings of AAAI. 2025, 39(2): 1601-1609.
[18] Devlin J, Chang M W, Lee K, et al. BERT: Pre-training of deep bidirectional transformers for language understanding[C]//Proceedings of NAACL. 2019: 4171-4186.
[19] Vaswani A, Shazeer N, Parmar N, et al. Attention is all you need[C]//Proceedings of NIPS. 2017: 5998-6008.
[20] Wang P, Zhou Q, Wu Y, et al. DLF: Disentangled-language-focused multimodal sentiment analysis[C]//Proceedings of AAAI. 2025, 39(20): 21180-21188.
[21] Guo Z, Jin T, Xu W, et al. Bridging the gap for test-time multimodal sentiment analysis[C]//Proceedings of AAAI. 2025, 39(16): 16987-16995.
[22] Guo Z, Jin T, Zhao Z. Multimodal prompt learning with missing modalities for sentiment analysis and emotion recognition[C]//Proceedings of ACL. 2024: 1726-1736.
[23] Dai W, Li X, Wang Z, et al. MInD: Improving multimodal sentiment analysis via multimodal information disentanglement[J/OL]. arXiv preprint arXiv:2401.11818, 2024.
[24] Sun K, Tian M. Sequential fusion of text-close and text-far representations for multimodal sentiment analysis[C]//Proceedings of COLING. 2025: 40-49.
[25] 林子杰, 龙云飞, 杜嘉晨, 等. 一种基于多任务学习的多模态情感识别方法[J]. 北京大学学报(自然科学版), 2021(1): 7-15.
LIN Z J, LONG Y F, DU J C, et al. A multi-task learning method for multimodal emotion recognition[J]. Acta Scientiarum Naturalium Universitatis Pekinensis, 2021(1): 7-15.
[26] 陈巧红, 孙佳锦, 漏杨波, 等. 基于多任务学习与层叠Transformer的多模态情感分析模型[J]. 浙江大学学报(工学版), 2023, 57(12): 2421-2429.
CHEN Q H, SUN J J, LOU Y B, et al. Multimodal sentiment analysis model based on multi-task learning and stacked cross-modal Transformer[J]. Journal of Zhejiang University (Engineering Science), 2023, 57(12): 2421-2429.
[27] 龙英潮, 丁美荣, 林桂锦, 等. 基于视听觉感知系统的多模态情感识别[J]. 计算机系统应用, 2021, 30(12): 218-225.
LONG Y C, DING M R, LIN G J, et al. Emotion recognition based on visual and audiovisual perception system[J]. Computer Systems & Applications, 2021, 30(12): 218-225.
[28] Williams J, Kleinegesse S, Comanescu R, et al. Recognizing emotions in video using multimodal DNN feature fusion[C]//Proceedings of Grand Challenge and Workshop on Human Multimodal Language. 2018: 11-19.
[29] Zadeh A, Liang P P, Mazumder N, et al. Memory Fusion Network for multi-view sequential learning[C]//Proceedings of AAAI. 2018, 32(1).
[30] Lv F, Chen X, Huang Y, Duan L, et al. Progressive modality reinforcement for human multimodal emotion recognition from unaligned multimodal sequences[C]//Proceedings of CVPR. 2021: 2554-2562.
[31] Li Y, Wang Y, Cui Z. Decoupled multimodal distilling for emotion recognition[C]//Proceedings of CVPR. 2023: 6631-6640.
[32] Li X, Liu D. MPID: A Modality-Preserving and Interaction-Driven Fusion Network for Multimodal Sentiment Analysis[C]//Proceedings of the 31st International Conference on Computational Linguistics. 2025.
[33] 周世向, 于凯. 基于密集协同注意力的多模态情感分析[J]. 计算机工程, 2025, 51(11): 144-151.
ZHOU S X, YU K. Multimodal sentiment analysis based on dense co-attention[J]. Computer Engineering, 2025, 51(11): 144-151.
[34] 冯广, 苏旭, 林忆宝, 等. 基于多尺度编码与极性感知融合的多模态情感分析[J/OL]. 计算机工程. DOI:10.19678/j.issn.1000-3428.0252862.
FENG G, SU X, LIN Y B, et al. Polarity-aware multimodal sentiment analysis with multi-scale encoding[J/OL]. Computer Engineering. DOI:10.19678/j.issn.1000-3428.0252862.
|