作者投稿和查稿 主编审稿 专家审稿 编委审稿 远程编辑

计算机工程 ›› 2026, Vol. 52 ›› Issue (10): 317-327. doi: 10.19678/j.issn.1000-3428.0260120

• 多模态与信息融合 • 上一篇    

融合混合编码与模糊建模的多模态对话情感识别模型

钟杭1,2,3, 张清华2,3, 罗南方1,3,4, 郭芮利1,2,4   

  1. 1. 重庆邮电大学计算机科学与技术学院, 重庆 400065;
    2. 重庆邮电大学大数据智能计算重点实验室, 重庆 400065;
    3. 网络空间大数据智能安全教育部重点实验室, 重庆 400065;
    4. 旅游多源数据感知与决策技术文化和旅游部重点实验室, 重庆 400065
  • 收稿日期:2026-01-23 修回日期:2026-04-28 发布日期:2026-05-15
  • 作者简介:钟杭(CCF学生会员),女,硕士研究生,主研方向为自然语言处理、情感分析;张清华(通信作者),教授、博士,E-mail:zhangqh@cqupt.edu.cn;罗南方,博士;郭芮利,硕士研究生。
  • 基金资助:
    国家自然科学基金(62276038,62576056);重庆市自然科学基金创新发展联合基金项目(CSTB2023NSCQ-LZX0164)。

Multimodal Emotion Recognition in Conversations Model Fused with Hybrid Encoding and Fuzzy Modeling

ZHONG Hang1,2,3, ZHANG Qinghua2,3, LUO Nanfang1,3,4, GUO Ruili1,2,4   

  1. 1. School of Computer Science and Technology, Chongqing University of Posts and Telecommunications, Chongqing 400065, China;
    2. Key Laboratory of Big Data Intelligent Computing, Chongqing University of Posts and Telecommunications, Chongqing, 400065, China;
    3. Key Laboratory of Cyberspace Big Data Intelligent Security, Ministry of Education, Chongqing 400065, China;
    4. Key Laboratory of Tourism Multi-source Data Perception and Decision Technology, Ministry of Culture and Tourism, Chongqing 400065, China
  • Received:2026-01-23 Revised:2026-04-28 Published:2026-05-15

摘要: 多模态对话情感识别(ERC)通过融合语言、声学和视觉等多源信息,实现对话情绪的自动识别,从而增强人机交互的自然性与情感理解。然而,现有方法在建模情感的多层上下文依赖方面仍存在不足,模态融合易引入冗余或噪声,且难以刻画情感的不确定性,限制复杂情绪识别。针对上述问题,提出一种融合混合编码与模糊建模的多模态ERC模型。该模型通过混合编码模块同时建模情感的全局对话上下文与局部依赖关系,从而增强情感时序特征的表达能力,并在此基础上引入分层门控融合机制,对不同层次和不同模态特征进行动态加权融合,以有效抑制冗余信息与噪声干扰。在情感分类阶段,采用线性等间距初始化的模糊神经网络,通过模糊隶属函数对情感类别边界进行建模,以刻画情绪表达中的不确定性与模糊性。实验结果显示,该模型在IEMOCAP、MELD和CMU-MOSEI 3个数据集上的各项指标均优于基线模型,在IEMOCAP上准确率为72.67%,在MELD上准确率为67.37%,在CMU-MOSEI七分类与二分类准确率上分别为54.96%和86.78%,验证了所提模型在多模态情感分析中的有效性。

关键词: 多模态, 对话情感识别, 上下文建模, 模态融合, 模糊神经网络

Abstract: Multimodal Emotion Recognition in Conversations (ERC) integrates language, acoustic, and visual information to automatically identify emotions in dialogues, thereby enhancing naturalness and emotional understanding in human-computer interactions. However, existing methods have limitations in modeling the multilayer contextual dependencies of emotions. Multimodal feature fusion often introduces redundant information and noise, making it difficult to capture the uncertainty of emotions, thereby limiting the recognition of complex emotions. To address these issues, this study proposes a multimodal ERC model that combines hybrid encoding and fuzzy modeling. The model uses a hybrid encoding module to simultaneously capture the global dialogue context and local dependencies of emotions simultaneously, thereby enhancing the representational capability of emotional temporal features. Accordingly, a hierarchical gated fusion mechanism is introduced to perform dynamic weighted fusion on features from different levels and modalities, effectively suppressing redundant information and noise. In the emotion classification stage, a fuzzy neural network with linear equally spaced initialization models the boundaries of emotion categories using fuzzy membership functions to capture the uncertainty and fuzziness in emotional expressions. The experimental results show that the proposed model outperforms the baseline models for all metrics across the IEMOCAP, MELD, and CMU-MOSEI datasets. It achieves an accuracy of 72.67% for IEMOCAP, 67.37% for MELD, 54.96% for seven-class classification accuracy, and 86.78% for two-class classification accuracy on CMU-MOSEI, validating the effectiveness of the proposed model in multimodal emotion recognition.

Key words: multimodal, Emotion Recognition in Conversations (ERC), contextual modeling, modal fusion, fuzzy neural network

中图分类号: