作者投稿和查稿 主编审稿 专家审稿 编委审稿 远程编辑

计算机工程 ›› 2026, Vol. 52 ›› Issue (10): 390-406. doi: 10.19678/j.issn.1000-3428.0260022

• 大模型与生成式人工智能 • 上一篇    

基于大模型数据增强的多粒度特征融合谣言检测

梁堉1, 马佳妍1, 胡晰远1, 王子恒1, 刘文2, 彭天豪3, 李莹3   

  1. 1. 北京工业大学计算机学院, 北京 100124;
    2. 武汉理工大学航运学院, 湖北 武汉 430063;
    3. 北京航空航天大学计算机学院, 北京 100191
  • 收稿日期:2026-01-06 修回日期:2026-03-20 发布日期:2026-04-20
  • 作者简介:梁堉(CCF专业会员),男,讲师、博士,主研方向为人工智能、深度学习、多模态学习、教育信息技术等;马佳妍,硕士研究生;胡晰远(通信作者),教授、博士,E-mail:huxiyuan@bjut.edu.cn;王子恒,本科生;刘文,教授、博士;彭天豪,博士研究生;李莹,副教授、博士。
  • 基金资助:
    国家自然科学基金(62507001);北京市自然科学基金(4254091)。

Multi-Granularity Feature Fusion for Rumor Detection Based on Large Model Data Augmentation

LIANG Yu1, MA Jiayan1, HU Xiyuan1, WANG Ziheng1, LIU Wen2, PENG Tianhao3, LI Ying3   

  1. 1. College of Computer Science, Beijing University of Technology, Beijing 100124, China;
    2. School of Navigation, Wuhan University of Technology, Wuhan 430063, Hubei, China;
    3. School of Computer Science and Engineering, Beihang University, Beijing 100191, China
  • Received:2026-01-06 Revised:2026-03-20 Published:2026-04-20

摘要: 随着网络和社交媒体的快速发展,信息的生成和传播速度达到了前所未有的水平,虚假信息、谣言及其他误导性内容充斥的现象愈加突出,这类问题已对社会治理秩序、和谐稳定构成重大威胁。在谣言检测中,谣言样本占比低导致数据不平衡,现有文本增强技术因缺乏谣言风格针对性、生成质量低,难以提升检测效果。同时,预训练语言模型虽然擅长捕捉文本全局依赖关系,却难以聚焦谣言关键局部特征。为此,提出一种基于大模型数据增强的多粒度特征融合的谣言检测框架。首先,提出融合谣言风格词典与大语言模型(LLM)的谣言生成方法,基于公开谣言数据集构建风格词典,以词典为约束指导LLM生成语义连贯且符合谣言风格的少数类样本,在缓解数据不平衡问题的同时保障增强样本质量;其次,所提多粒度上下文特征提取器(MGCFE),融合基于解耦注意力机制的预训练语言模型在全局依赖捕捉上的优势,与卷积子层对局部特征的聚焦能力,实现对谣言语义长距离逻辑关联与细粒度语言线索的同步捕捉,有效弥补此类预训练模型在局部关键特征捕捉上的固有局限。实验结果表明,该检测方法在BuzzFeed和PolitiFact数据集准确率分别达到82.24%、93.91%。

关键词: 深度学习, 大语言模型, 数据增强, 谣言检测, 多粒度特征融合, 社会治理

Abstract: With the rapid development of the internet and social media, the speed of information generation and dissemination has reached an unprecedented level. The proliferation of misinformation, rumors, and other misleading content has become increasingly prominent, posing significant threats to the order, harmony, and stability of social governance. In rumor detection, a low proportion of rumor samples leads to data imbalance, and existing text augmentation techniques struggle to enhance detection performance owing to their lack of specificity to rumor styles and low generation quality. Additionally, although pretrained language models capture global dependencies in text, they often fail to focus on the key local features of rumors. To address these challenges, this study proposes a rumor detection framework based on large-model data augmentation and multi-granularity feature fusion. First, a rumor generation method integrating a rumor-style lexicon and Large Language Model (LLM) is proposed. Based on publicly available rumor datasets, a style lexicon is constructed to guide LLM in generating semantically coherent and rumor-style consistent minority class samples. This approach alleviates data imbalance while ensuring the quality of the augmented samples. Second, this study introduces a Multi-Granularity Contextual Feature Extractor (MGCFE), which combines the strengths of pretrained language models with disentangled attention mechanisms to capture global dependencies and focus convolutional sublayers on local features. This enables the simultaneous capture of long-distance logical associations and fine-grained linguistic clues in rumor semantics, effectively mitigating the inherent limitations of pretrained models in capturing key local features. The experimental results demonstrate that the proposed detection method achieves accuracy rates of 82.24% and 93.91% for BuzzFeed and PolitiFact datasets, respectively.

Key words: deep learning, Large Language Model (LLM), data augmentation, rumor detection, multi-granularity feature fusion, social governance

中图分类号: