作者投稿和查稿 主编审稿 专家审稿 编委审稿 远程编辑

计算机工程 ›› 2009, Vol. 35 ›› Issue (1): 277-279,. doi: 10.3969/j.issn.1000-3428.2009.01.097

• 开发研究与设计技术 • 上一篇    下一篇

基于目录树的网络科技资源采集算法

李国栋,刘忠强,柳长安   

  1. (华北电力大学计算机科学与技术学院,北京 102206)
  • 收稿日期:1900-01-01 修回日期:1900-01-01 出版日期:2009-01-05 发布日期:2009-01-05

Crawler Algorithm Based on Directory Tree in Network Science and Technology Resource

LI Guo-dong, LIU Zhong-qiang, LIU Chang-an   

  1. (School of Computer Science and Technology, North China Electric Power University, Beijing 102206)
  • Received:1900-01-01 Revised:1900-01-01 Online:2009-01-05 Published:2009-01-05

摘要: 针对网络科技领域资源分类方式多样化、数据量大等特点,提出一种基于目录树的采集算法,以领域本体知识库提供的本体知识作为评价依据进行有效目录链接的提取和识别,通过一种改进的链接分析策略获取有效的节点链接并进行采集操作。该算法研究采集体系结构,注重对最新资源获取速度的优化。实验结果证明,该算法可有效提高资源采集速率。

关键词: 科技资源, 信息采集, 目录树, 本体

Abstract: Aimming at full consideration of the characteristics of the network technology in a various methods of classification of resources and a large quantity, this paper proposes a kind of crawler algorithm based on directory tree. The algorithm extracts and recognizes the directory links based on domain ontology knowledge as effective evaluation, and links the nodes effectively through a modified strategy of link analysis, eventually carry through collecting operation. The algorithm not only studies in-depth on the crawler architecture, but also pays attention to the speed of access to the latest resources optimization. Experimental results show that the algorithm can effectively achieve the established objectives both in speed and efficiency.

Key words: science and technology resource, information crawling, directory tree, ontology

中图分类号: