作者投稿和查稿 主编审稿 专家审稿 编委审稿 远程编辑

计算机工程 ›› 2009, Vol. 35 ›› Issue (7): 175-176,. doi: 10.3969/j.issn.1000-3428.2009.07.061

• 人工智能及识别技术 • 上一篇    下一篇

基于独立分量分析的隐蔽Web领域聚类

王晓斌,温 春,石昭祥   

  1. (电子工程学院网络工程系602教研室,合肥 230037)
  • 收稿日期:1900-01-01 修回日期:1900-01-01 出版日期:2009-04-05 发布日期:2009-04-05

Hidden Web Domain Clustering Based on Independent Component Analysis

WANG Xiao-bin, WEN Chun, SHI Zhao-xiang   

  1. (602 Teach Stuff, Department of Network Engineering, Electronic Engineering Institute, Hefei 230037)
  • Received:1900-01-01 Revised:1900-01-01 Online:2009-04-05 Published:2009-04-05

摘要: 针对隐蔽Web主题领域自动识别问题,提出一种基于独立分量分析(ICA)的聚类算法。对查询页面进行页面文本抽取和预处理,利用TF-IDF公式计算权重并选择前N个权重最大的特征词构造文档矩阵,在使用潜在语义索引(LSI)进行特征重构的基础上通过ICA分解获得类别信息。利用LSI的词共现分析和文本降噪能力提高聚类准确率。实验表明聚类平均准确率达到90%以上。

关键词: 隐蔽Web, 潜在语义, 独立分量分析, 文本聚类

Abstract: Aiming at organizing hidden Web databases according to their topic domains, this paper proposes an Independent Component Analysis(ICA) based algorithm for hidden Web domain clustering. Text is extracted from search interface pages as common Web pages, and TF-IDF formula is applied to weight terms. After selecting the top N-highest weight terms to construct VSM, the algorithm performs a singular value decomposition to implement features reconstruction. It applies ICA decomposition to obtain the cluster information. The main idea is utilizing the co-occurrence analysis and noise eliminating ability of Latent Semantic Index(LSI) to improve cluster performance. Experiment shows that the average precision is higher than 90 percent.

Key words: hidden Web, latent semantic, Independent Component Analysis(ICA), text clustering

中图分类号: