[1] ANWAR K, SHREYA, SHARMA M, SAANVI K. A review of multimodal sentiment analysis: Taxonomy, issues, challenges, and future perspectives[J]. Computers and Electrical Engineering, 2026, 131: 110959.
[2] Liao X, Ke X, Liu Q, et al. Text-Guided Enhanced Transformer Fusion for Multimodal Sentiment Analysis[J]. Journal of Shanghai Jiaotong University (Science), 2025: 1-12.
[3] Chen J, Song S, Tan Y, et al. TEMSA: Text enhanced modal representation learning for multimodal sentiment analysis[J]. Computer Vision and Image Understanding, 2025, 258: 104391.
[4] LI Z, LIU P, PAN Y, et al. Text-dominant multimodal perception network for sentiment analysis based on cross-modal semantic enhancements[J]. Applied Intelligence, 2025, 55(3): 188.
[5] XIAO T, WANG Z, Ouyang G, et al. Text-Dominant Disentangled Fusion Network for Multimodal Sentiment Analysis[J]. IFAC-PapersOnLine, 2025, 59(35): 73-78.
[6] YANG H, ZHAO Y, WU Y, et al. Large language models meet text-centric multimodal sentiment analysis: A survey[J]. Science China Information Sciences, 2025, 68(10): 1-29.
[7] FENG X, LIN Y, HE L, et al. Knowledge-guided dynamic modality attention fusion framework for multimodal sentiment analysis[C]//Findings of the Association for Computational Linguistics: EMNLP 2024. Miami: Association for Computational Linguistics, 2024: 14755-14766.
[8] YANG D, LI M, WU X, et al. Improving Multimodal Sentiment Analysis via Modality Optimization and Dynamic Primary Modality Selection[EB/OL]. arXiv preprint, 2025: arXiv:2511.06328[2026-04-30]. DOI: 10.48550/arXiv.2511.06328.
[9] Huang H, Gong T, He K, et al. Robust multimodal sentiment analysis via double information bottleneck[J]. Information Fusion, 2025: 103964.
[10] LI Y, LIU A, LU Y. Multi-level language interaction transformer for multimodal sentiment analysis[J]. Journal of Intelligent Information Systems, 2025, 63(3): 945-964.
[11] Wang H L, Cao J X, Liu J J, et al. A method for multimodal sentiment analysis: adaptive interaction and multi-scale fusion[J]. Journal of Intelligent Information Systems, 2025, 63(5): 1667-1686.
[12] HAZARIKA D, ZIMMERMANN R, PORIA S. Misa: Modality-invariant and-specific representations for multimodal sentiment analysis[C]//Proceedings of the 28th ACM international conference on multimedia. 2020: 1122-1131.
[13] Mai S, Zeng Y, Hu H. Multimodal information bottleneck: Learning minimal sufficient unimodal and multimodal representations[J]. IEEE Transactions on Multimedia, 2022, 25: 4121-4134.
[14] 杨新航,王晶晶,陈思宇,等.基于H-GEM模型的多模态情感分析[J].计算机系统应用,2026,35(03):59-68. DOI:10.15888/j.cnki.csa.010107.
YANG Xinhang, WANG Jingjing, CHEN Siyu, et al. Multimodal sentiment analysis based on H-GEM model[J]. Computer Systems & Applications, 2026, 35(03):59-68. DOI: 10.15888/j.cnki.csa.010107.
[15] 王钰张,钱钢.基于BERT模型的多模态情感融合策略实证研究[J/OL].智能计算机与应用,1-9[2026-04-30]. https://doi.org/10.20169/j.issn.2095-2163.25102402.
WANG Yuzhang, QIAN Gang. Empirical study on multimodal sentiment fusion strategy based on BERT model[J/OL]. Intelligent Computer and Applications, 1-9[2026-04-30]. https://doi.org/10.20169/j.issn.2095-2163.25102402.
[16] 罗渊贻,吴锐,刘家锋,等.基于自适应权值融合的多模态情感分析方法[J].软件学报,2024,35(10):4781-4793. DOI:10.13328/j.cnki.jos.006998.
LUO Yuanyi, WU Rui, LIU Jiafeng, et al. Multimodal sentiment analysis method based on adaptive weight fusion[J]. Journal of Software, 2024, 35(10): 4781-4793. DOI: 10.13328/j.cnki.jos.006998.
[17] ZADEH A, CHEN M, PORIA S, et al. Tensor fusion network for multimodal sentiment analysis[C]//Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing. Copenhagen: Association for Computational Linguistics, 2017: 1103-1114.
[18] ZADEH A, LIANG P P, MAZUMDER N, et al. Memory fusion network for multi-view sequential learning[C]//Proceedings of the AAAI Conference on Artificial Intelligence. 2018, 32(1): 5634-5641.
[19] HE J, YANG H, ZHANG C, et al. Dynamic invariant-specific representation fusion network for multimodal sentiment analysis[J]. Computational Intelligence and Neuroscience, 2022, 2022(1): 2105593.
[20] AKHTAR M S, GHOSAL D, EKBAL A, et al. All-in-One: emotion, sentiment and intensity prediction using a multi-task ensemble framework[J]. IEEE Transactions on Affective Computing, 2022, 13(1): 285-297.
[21] YU Y, ZHAO M, QI S, et al. ConKI: contrastive knowledge injection for multimodal sentiment analysis[C]//Findings of the Association for Computational Linguistics: ACL 2023. Toronto: Association for Computational Linguistics, 2023: 13610-13624.
[22] LI H, GUO A, LI Y. CCMA: CapsNet for audio–video sentiment analysis using cross-modal attention[J]. The Visual Computer, 2025, 41(3): 1609-1620.
[23] 钱旦敏,张铠,张继海,等.基于多尺度时序感知与自适应权重融合的多模态情感分析方法[J/OL].数据分析与知识发现,1-16[2026-04-30]. https://link.cnki.net/urlid/ 10.1478.g2.20260126.1358.002.
QIAN Danmin, ZHANG Kai, ZHANG Jihai, et al. Multimodal sentiment analysis method based on multi-scale temporal perception and adaptive weight fusion[J/OL]. Data Analysis and Knowledge Discovery, 1-16[2026-04-30]. https://link.cnki.net/urlid/10.1478.g2.20260126.1358.002.
[24] Li M, Zhu Z, Li K, et al. Diversity and balance: Multimodal sentiment analysis using multimodal-prefixed and cross-modal attention[J]. IEEE Transactions on Affective Computing, 2024, 16(1): 250-263.
[25] ZHI Y, LI J, WANG H, et al. A multimodal sentiment analysis approach based on multiview cross-modal fusion[J]. IEEE Transactions on Computational Social Systems, 2026, 13(1): 136-151.
[26] Luo Y, Wu R, Liu J, et al. A text guided multi-task learning network for multimodal sentiment analysis[J]. Neurocomputing, 2023, 560: 126836.
[27] Li J, Liu R, Miao Q, et al. CAETFN: Context adaptively enhanced text-guided fusion network for multimodal sentiment analysis[J]. IEEE Transactions on Affective Computing, 2025.
[28] HUANG X, SUN H, LIU X, et al. Text-Dominant Speech-Enhanced for Multimodal Aspect-Based Sentiment Analysis network[J]. Information Fusion, 2026, 126: 103543. DOI: 10.1016/j.inffus.2025.103543.
[29] Zadeh A, Zellers R, Pincus E, et al. Multimodal sentiment intensity analysis in videos: Facial gestures and verbal messages[J]. IEEE Intelligent Systems, 2016, 31(6): 82-88.
[30] Zadeh A A B, Liang P P, Poria S, et al. Multimodal language analysis in the wild: Cmu-mosei dataset and interpretable dynamic fusion graph[C]//Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2018: 2236-2246.
[31] LIU Z, SHEN Y, LAKSHMINARASIMHAN V B, et al. Efficient low-rank multimodal fusion with modality-specific factors[C]//Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics. Melbourne: Association for Computational Linguistics, 2018: 2247-2256.
[32] TSAI Y H H, BAI S, LIANG P P, et al. Multimodal transformer for unaligned multimodal language sequences[C]//Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. Florence: Association for Computational Linguistics, 2019: 6558-6569.
[33] RAHMAN W, HASAN M K, LEE S, et al. Integrating multimodal information in large pretrained transformers[C]//Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. Online: Association for Computational Linguistics, 2020: 2359-2369.
[34] YU W, XU H, YUAN Z, et al. Learning modality-specific representations with self-supervised multi-task learning for multimodal sentiment analysis[C]//Proceedings of the AAAI conference on artificial intelligence. 2021, 35(12): 10790-10797.
[35] Huang J, Zhou J, Tang Z, et al. TMBL: Transformer-based multimodal binding learning model for multimodal sentiment analysis[J]. Knowledge-Based Systems, 2024, 285: 111346.
[36] LI Y, WANG Y, CUI Z. Decoupled multimodal distilling for emotion recognition[C]//Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 2023: 6631-6640.
[37] WANG P, ZHOU Q, WU Y, et al. DLF: Disentangled-language-focused multimodal sentiment analysis[C]//Proceedings of the AAAI Conference on Artificial Intelligence. 2025, 39(20): 21180-21188.
[38] Wu S, He D, Wang X, et al. Enriching multimodal sentiment analysis through textual emotional descriptions of visual-audio content[C]//Proceedings of the AAAI Conference on Artificial Intelligence. 2025, 39(2): 1601-1609.
[39] Sun K, Tian M. Sequential fusion of text-close and text-far representations for multimodal sentiment analysis[C]//Proceedings of the 31st international conference on computational linguistics. 2025: 40-49.
[40] Zhou M, Yang L, Wu T, et al. Dual-Path Dynamic Fusion with Learnable Query for Multimodal Sentiment Analysis[C]//Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025: 11366-11376.
|