作者投稿和查稿 主编审稿 专家审稿 编委审稿 远程编辑

计算机工程 ›› 2011, Vol. 37 ›› Issue (10): 41-43. doi: 10.3969/j.issn.1000-3428.2011.10.013

• 软件技术与数据库 • 上一篇    下一篇

基于上下文的短信文本分类方法

刘金岭,严云洋   

  1. (淮阴工学院计算机工程学院,江苏 淮安 223003)
  • 出版日期:2011-05-20 发布日期:2011-05-20
  • 作者简介:刘金岭(1958-),男,教授,主研方向:数据仓库,文本数据挖掘;严云洋,教授、博士
  • 基金资助:
    淮安科技计划基金资助项目(HAG09061);淮阴工学院基金资助重点项目(HGA0907)

SMS Text Classification Method Based on Context

LIU Jin-ling, YAN Yun-yang   

  1. (Computer Engineering Faculty, Huaiyin Institute of Technology, Huaian 223003, China)
  • Online:2011-05-20 Published:2011-05-20

摘要: 针对海量短信文本数据中大量词语共现的特点,提出一种基于上下文的短信文本分类方法。利用词语的上下文关系,定义词语相似度和基于上下文的词语权值,科学地表达词语在该类别中的语义表示,以提高短信文本分类效率。实验结果表明,与传统的简单向量距离分类法相比,该方法的分类效果较优。

关键词: 短信文本, 词语共现, 上下文, 词语相似度, 短信文本分类

Abstract: According to the characteristics of a lot of words co-occurrence in mass data of Short Messaging Service(SMS), a context-based SMS text classification method based on the context term is defined word similarity relations, and defines the term weights using context, which expresses more scientific terms in this category in the semantic representation and thus further improves classification efficiency of SMS text. Experimental results show that the classification performance of method than the traditional simple vector distance classification is significantly improved.

Key words: Short Messaging Service(SMS) text, word co-occurrence, context, word similarity, SMS text classification

中图分类号: