作者投稿和查稿 主编审稿 专家审稿 编委审稿 远程编辑

计算机工程 ›› 2026, Vol. 52 ›› Issue (10): 302-316. doi: 10.19678/j.issn.1000-3428.0260521

• 多模态与信息融合 • 上一篇    

面向患者主诉语音与文本的多模态症状分类方法

张昊然1, 蹇木伟1,2, 王瑞3, 宋增凯4   

  1. 1. 山东财经大学管理科学与工程学院, 山东 济南 250014;
    2. 山东财经大学计算机与人工智能学院, 山东 济南 250014;
    3. 临沂大学信息科学与工程学院, 山东 临沂 276000;
    4. 山东省工业和信息化厅综合保障中心, 山东 济南 250011
  • 收稿日期:2026-04-21 修回日期:2026-06-07 发布日期:2026-06-24
  • 作者简介:张昊然(CCF学生会员),女,博士研究生,主研方向为深度学习、医学人工智能;蹇木伟(通信作者),教授、博士,E-mail:jianmuweihk@163.com;王瑞,硕士研究生;宋增凯,学士。
  • 基金资助:
    山东省泰山学者特聘专家项目(tstp20250536)。

Multimodal Symptom Classification Method for Patients' Chief-Complaint Speech and Text

ZHANG Haoran1, JIAN Muwei1,2, WANG Rui3, SONG Zengkai4   

  1. 1. School of Management Science and Engineering, Shandong University of Finance and Economics, Jinan 250014, Shandong, China;
    2. School of Computing and Artificial Intelligence, Shandong University of Finance and Economics, Jinan 250014, Shandong, China;
    3. College of Information Science and Engineering, Linyi University, Linyi 276000, Shandong, China;
    4. Comprehensive Support Center, Department of Industry and Information Technology of Shandong Province, Jinan 250011, Shandong, China
  • Received:2026-04-21 Revised:2026-06-07 Published:2026-06-24

摘要: 在真实临床问诊场景中,患者主诉通常以语音形式表达,并由医生记录为文本。医生需要综合利用患者口述语音及其文本记录对症状进行判断和分类,从而为后续临床决策提供依据。然而,该任务仍面临一定挑战:语音信息容易受到环境噪声和个体发音差异的影响,文本记录又难以体现语速、停顿、音调等语音表达特征;同时,患者主诉通常具有口语化、主观性和非结构化特点,不同症状类别之间也可能存在语义边界模糊的问题,导致仅依赖单一模态难以获得理想的分类效果。针对上述问题,提出一种基于动态权重决策融合的多模态症状分类方法(DWDF-MSC),以充分利用文本与语音信息的互补性,提升症状分类的准确性和鲁棒性。该方法主要包括多模态特征提取、初步分类和自适应门控决策融合三个阶段。在多模态特征提取阶段,分别构建文本分支和语音分支,对患者主诉文本和语音数据进行并行建模:文本分支基于临床预训练语言模型Bio_ClinicalBERT同时提取全局语义特征和局部词汇特征,并通过文本异构特征融合模块对两者进行融合,从而增强模型对主诉整体语义和局部症状关键词的表征能力;语音分支利用音频频谱图Transformer(AST)提取语音中的时序声学表示,以补充文本记录中难以体现的语音表达信息。在初步分类阶段,文本分支和语音分支分别通过各自的分类模块输出初步分类结果,使两种模态先独立完成症状判断。在最终分类阶段,设计自适应门控决策融合策略,根据不同样本的特征动态生成融合权重,对文本分支和语音分支的初步分类结果进行加权融合,得到最终症状分类结果。与简单特征拼接或固定权重融合不同,该策略能够根据样本差异自适应调整两种模态在最终决策中的贡献,从而增强具有判别力的信息对分类结果的影响,提高模型在复杂主诉场景下的分类稳定性。在公共医疗数据集上的实验结果表明,DWDF-MSC的准确率、精确率和F1值分别达到82.43%、87.44%和81.52%,均优于多数主流基准模型。多模态融合方案对比进一步证明,相较于特征融合,所提出的动态权重决策融合能够取得更好的分类效果。在消融实验中,完整DWDF-MSC相较于仅采用文本异构特征融合的方案,在准确率和F1值上的提升幅度分别为4.25%和7.60%,验证了语音分支和自适应门控决策融合的有效性。McNemar检验结果显示,DWDF-MSC与多种对比方法之间的p-value小于0.000 1,说明其与这些对比方法之间的分类结果差异具有统计显著性。抗噪性能实验结果表明,DWDF-MSC在不同信噪比(SNR)条件下仍能保持较稳定的分类表现。综上所述,DWDF-MSC能够有效融合患者主诉中的文本与语音信息,提升模型分类性能,为面向患者主诉的智能症状分类提供了一种可行的多模态方法。

关键词: 多模态特征, 决策融合, 深度学习, 语音信号处理, 智能问诊

Abstract: In real-world clinical consultations, patients' chief complaints are typically expressed verbally and subsequently recorded as text by physicians. Physicians must comprehensively use both patients' spoken descriptions and the corresponding textual records to judge and classify symptoms, thereby providing a basis for subsequent clinical decision making. However, this method has several limitations. Speech information is susceptible to environmental noise and individual pronunciation differences, whereas textual records are unable to fully reflect speech-related expressive features such as speaking rate, pauses, and intonation. Additionally, patients' chief complaints are usually colloquial, subjective, and unstructured, and the semantic boundaries among different symptom categories may be ambiguous. Therefore, single-modality methods cannot achieve satisfactory classification performance easily. To address these issues, a dynamic weight decision fusion-based multimodal symptom classification method, called DWDF-MSC, is proposed to fully exploit the complementarity between textual and speech information and improve the accuracy and robustness of symptom classification. The proposed method consists of three stages: multimodal feature extraction, preliminary classification, and adaptive gated decision fusion. In the multimodal feature extraction stage, text and speech branches are constructed to model the text and speech data of the patient's chief complaint in parallel. In the text branch, global semantic features and local lexical features are simultaneously extracted based on the clinical pretrained language model Bio_ClinicalBERT, and the two are fused through a heterogeneous textual feature fusion module, thereby enhancing the model's representation capability for the overall semantics of the chief complaints and local symptom-related keywords. In the speech branch, an Audio Spectrogram Transformer (AST) is used to extract temporal acoustic representations from speech, thereby supplementing expressive speech information that is difficult to capture from textual records. In the preliminary classification stage, the text and speech branches output initial classification results through their respective classification modules, allowing the two modalities to perform symptom evaluation independently. In the final classification stage, an adaptive gated decision fusion strategy is designed to dynamically generate fusion weights according to the features of different samples. The initial classification results from the text and speech branches are then weighted and fused to obtain the final symptom classification results. Unlike simple feature concatenation or fixed-weight fusion, this strategy can adaptively adjust the contribution of the two modalities to the final decision based on sample differences, thereby enhancing the influence of discriminative information on the classification results and improving the classification stability of the model in complex chief complaint scenarios. Experimental results on a public medical dataset show that the DWDF-MSC achieves 82.43%, 87.44%, and 81.52% in accuracy, precision, and F1-score, respectively, outperforming most mainstream baseline models across all metrics. A comparison of multimodal fusion schemes further demonstrates that the proposed dynamic weight decision fusion achieves better classification performance than feature-level fusion. In the ablation study, the complete DWDF-MSC model achieves relative improvements of 4.25% and 7.60% in accuracy and F1-score, respectively, compared to the variant that only employed heterogeneous text feature fusion, thereby demonstrating the effectiveness of the speech branch and the adaptive gated decision fusion mechanism. The McNemar test results show that the p-values between the DWDF-MSC and the multiple comparison methods are less than 0.000 1, indicating that the differences in classification results between the DWDF-MSC and these comparison methods are statistically significant. The antinoise performance experiments demonstrate that the DWDF-MSC can maintain a relatively stable classification performance under different Signal-to-Noise Ratio (SNR) conditions. In summary, DWDF-MSC can effectively fuse textual and speech information from chief patient complaints, improve model classification performance, and provide a feasible multimodal method for intelligent symptom classification based on chief patient complaints.

Key words: multimodal feature, decision fusion, deep learning, speech signal processing, intelligent medical consultation

中图分类号: