作者投稿和查稿 主编审稿 专家审稿 编委审稿 远程编辑

计算机工程 ›› 2026, Vol. 52 ›› Issue (10): 275-284. doi: 10.19678/j.issn.1000-3428.0252352

• 多模态与信息融合 • 上一篇    

生成式补全与动态知识融合的多模态情感分析

郑洋, 王雷, 盛捷   

  1. 中国科学技术大学信息科学技术学院, 安徽 合肥 230026
  • 收稿日期:2025-04-22 修回日期:2025-06-24 发布日期:2025-08-13
  • 作者简介:郑洋,男,硕士研究生,主研方向为复杂系统、深度学习、多模态情感识别;王雷(通信作者),副教授、博士;盛捷,讲师、博士,E-mail:wangl@ustc.edu.cn。
  • 基金资助:
    高技术创新特区项目(20-163-14-LZ-001-004-01)。

Multimodal Sentiment Analysis via Generative Completion and Dynamic Knowledge Fusion

ZHENG Yang, WANG Lei, SHENG Jie   

  1. School of Information Science and Technology, University of Science and Technology of China, Hefei 230026, Anhui, China
  • Received:2025-04-22 Revised:2025-06-24 Published:2025-08-13

摘要: 多模态情感分析利用多模态数据来推断人类情感。然而现有模型在模态信息缺失、文本依赖及跨模态冲突等情况下性能下降明显。为此,提出一种基于生成式补全与动态知识融合的多模态情感分析模型(GC-DKF)。首先通过生成式提示学习模块,对原始数据中缺失的模态内及模态间信息进行补全,生成缺失的模态特征,提升模型对不确定模态场景的适应能力;然后设计一种主导模态动态选择机制,依据情感比例因子动态选定主导模态,同时引入知识编码器增强单一模态的表征能力,获取各模态的知识增强表征;最后在主导模态特征的引导下,进一步学习其他次要模态,生成互补性的多模态联合表征,实现更为高效、精准的多模态情感分析。在公开的CMU-MOSI和CMU-MOSEI数据集上的实验结果显示,所提模型在二分类准确率、F1值、平均绝对误差(MAE)和Pearson相关系数等指标上,均超越现有的主流多模态情感识别方法,情感识别准确率分别高达83.55%和83.02%。这充分证明提出的模型在多模态情感识别任务中具备较强竞争力。

关键词: 多模态情感分析, 模态缺失, 生成式补全, 动态知识融合, 主导模态动态选择

Abstract: Multimodal sentiment analysis utilizes multimodal data to infer human emotions. However, existing models significantly degrade in performance when faced with issues such as modal information loss, text dependency, and cross-modal conflicts. To address this, a Generative Completion and Dynamic Knowledge Fusion Model (GC-DKF) for multimodal sentiment analysis is proposed. First, the model employs a generative prompt learning module to complete the missing intra-modal and inter-modal information in the original data, generating missing modal features to enhance the model's adaptability to uncertain modal scenarios. Then, a dynamic dominant modality selection mechanism is designed to dynamically select the dominant modality based on emotional proportion factors. Meanwhile, a knowledge encoder is introduced to strengthen the representation capability of a single modality, obtaining knowledge-enhanced representations of each modality. Finally, guided by the features of the dominant modality, the model further learns other secondary modalities to generate complementary multimodal joint representations, achieving more efficient and accurate multimodal sentiment analysis. Experiments on the public CMU-MOSI and CMU-MOSEI datasets demonstrate that the proposed model outperforms existing mainstream multimodal sentiment recognition methods in terms of metrics such as binary classification accuracy, F1-score, Mean Absolute Error (MAE), and Pearson correlation coefficient, with sentiment recognition accuracies reaching as high as 83.55% and 83.02%, respectively. This fully demonstrates that the proposed model has strong competitiveness in multimodal sentiment recognition tasks.

Key words: multimodal sentiment analysis, modality missing, generative completion, dynamic knowledge fusion, dynamic selection of dominant modality

中图分类号: