作者投稿和查稿 主编审稿 专家审稿 编委审稿 远程编辑

计算机工程

• 人工智能及识别技术 • 上一篇    下一篇

基于区分性关键词模型的维吾尔文本情感分类

热依莱木·帕尔哈提1,2,孟祥涛2,3,艾斯卡尔·艾木都拉1   

  1. (1. 新疆大学信息科学与工程学院,乌鲁木齐830046; 2. 清华大学信息技术研究院,北京100084;3. 重庆邮电大学计算机科学与技术学院,重庆400065)
  • 收稿日期:2013-09-09 出版日期:2014-10-15 发布日期:2014-10-13
  • 作者简介:热依莱木·帕尔哈提(1987 - ),女,硕士研究生,主研方向:文本分类;孟祥涛,硕士研究生;艾斯卡尔·艾木都拉,教授、博士、 博士生导师。
  • 基金资助:
    国家自然科学基金资助项目(61065005,61163033);教育部新世纪优秀人才支持计划基金资助项目(NCET-10-0969);新疆维 吾尔自治区高新技术研究发展计划基金资助项目(201312103)。

Uyghur Text Sentiment Classification Based on Discriminative Keyword Model

Rayila Parhat 1,2,MENG Xiang-tao 2,3,Askar Hamdulla 1   

  1. (1. Institute of Information Science and Engineering,Xinjiang University,Urumqi 830046,China; 2. Institute of Information Technology,Tsinghua University,Beijing 100084,China; 3. College of Computer Science and Technology,Chongqing University of Posts and Telecommunications,Chongqing 400065,China)
  • Received:2013-09-09 Online:2014-10-15 Published:2014-10-13

摘要: 在研究区分性关键词提取方法的基础上,对维吾尔语中的生气和高兴等常见情感类型进行基于文本句子的情感分类研究。结合维吾尔文本句子中的情感表达特点,以词频和文档频率作为基本统计量,通过计算同一词语在不同组合统计量下的类间差异得到区分性关键词,并基于这些关键词进行特征提取和区分性情感模型构建。从维吾尔语电影字幕、小说等文本库中提取生气和高兴2 种情感构造实验数据集,并验证所提出的情感分类方法。实验结果表明,基于区分性关键词的建模方法能有效地对维吾尔文本句子进行情感分类。

关键词: 维吾尔语, 区分性关键词, 文本句子, 情感分类, 差异性统计量

Abstract: This paper presents a classification approach for Uyghur text sentiment,such as angry and happy,based on discriminative key word extraction. Combined with the characteristics of sentiment expression in Uyghur text,the term frequency and document frequency are derived as primary statistics. Various discriminative statistics which reflect the discrepancy of the positive and negative sentiment datasets are derived from the primary statistics for each vocabulary word, and are used to extract discriminative key words. Features are extracted based on these keywords and are used to train discriminative sentiment models. This paper builds a sentiment text database by excerpting two sentiments:angriness and happiness from Uyghur movie transcriptions and novels,and verifies the proposed approach. Experimental results show that the method based on discriminative keyword extraction is effective in Uyghur text sentence sentiment classification.

Key words: Uyghur language, discriminative keyword;text sentence;sentiment classification;difference statistics

中图分类号: