作者投稿和查稿 主编审稿 专家审稿 编委审稿 远程编辑

计算机工程 ›› 2026, Vol. 52 ›› Issue (10): 442-454. doi: 10.19678/j.issn.1000-3428.0252127

• 新一代网络与边缘计算 • 上一篇    

基于多表征融合的物联网恶意加密流量分类

王纶羽, 顾益军   

  1. 中国人民公安大学信息网络安全学院, 北京 100032
  • 收稿日期:2025-02-14 修回日期:2025-04-11 发布日期:2025-06-11
  • 作者简介:王纶羽,男,硕士研究生,主研方向为加密流量分类、网络安全;顾益军(通信作者),教授、博士,E-mail:guyijun@ppsuc.edu.cn。
  • 基金资助:
    中央高校基本科研业务费专项资金(2023JKF01ZK14)。

Malicious Encrypted Traffic Classification in IoT Based on Multi-Representation Fusion

WANG Guanyu, GU Yijun   

  1. College of Information and Cyber Security, People's Public Security University of China, Beijing 100032, China
  • Received:2025-02-14 Revised:2025-04-11 Published:2025-06-11

摘要: 在恶意加密流量分类领域,模型通过增加流量特征维度来扩展判别表征的丰富性,但仍然存在选择模型与恶意加密流量数据特征不匹配以及特征选择不充分的问题,同时缺乏对加密流量数据特征的讨论研究。为此,针对物联网(IoT)恶意加密流量分类问题,提出一种基于多表征融合的分类模型。一方面,使用抽象表征学习模型学习流量会话的数据包级字节关联表征与会话统计表征;另一方面,使用明文表征学习模型学习未加密明文的会话连接表征。最后,根据抽象表征学习模型对分类结果的置信分数融合2个模型,获得最终的恶意流量分类结果。为验证模型的先进性,与7种基准模型进行实验比较,结果表明,所提模型的F1值达到0.769 4,相较其他基准模型均有大幅提升。同时,为验证模型中各个模块与流量表征学习的适配性、选择特征所含判别表征之间的互补性,生成10种基于不同输入与不同模型架构的变体模型进行比较,结果表明,该模型具有更优的检测性能,证明了模型架构的适配性与表征之间的互补性。

关键词: 恶意加密流量分类, 深度学习, 机器学习, 表征融合, 置信度

Abstract: In malicious encrypted traffic classification, models expand the richness of discriminative representations by increasing traffic feature dimensions. However, issues such as the mismatch between the selected models and the data characteristics of malicious encrypted traffic, as well as insufficient feature selection still persist. Moreover, research and discussion on the data characteristics of encrypted traffic is lacking. To address these challenges, this study proposes a multi-representation fusion-based classification model for Internet of Things (IoT) malicious encrypted traffic classification. An abstract representation learning model is employed to learn packet-level byte-association representations and statistical representations of traffic sessions. Additionally, a plaintext representation learning model is used to learn session connection representations from unencrypted plaintext. Finally, the two models are fused based on the confidence scores of the abstract representation learning model to obtain the final malicious traffic classification results. The superiority of the proposed model is validated by comparing it with seven baseline models. The results show that the proposed model achieves an F1 value of 0.769 4, representing a significant improvement over all baseline models. Furthermore, to verify the compatibility of each module with traffic representation learning and the complementarity among discriminative representations contained in the selected features, ten variant models based on different inputs and model architectures were generated for comparison. The results demonstrate that the proposed model achieves superior detection performance, thereby proving the compatibility of the model architecture and complementarity among the representations.

Key words: malicious encrypted traffic classification, deep learning, machine learning, representation fusion, confidence level

中图分类号: