Research on Probabilistic Algorithm of       Chinese Word Segmentation Based on the Maximum Match

doi:10.3969/j.issn.1000-3428.2010.05.063

Computer Engineering ›› 2010, Vol. 36 ›› Issue (5): 173-175.

• Artificial Intelligence and Recognition Technology • Previous Articles Next Articles

Research on Probabilistic Algorithm of Chinese Word Segmentation Based on the Maximum Match

HE Guo-bin, ZHAO Jing-lu

(College of Computer and Information Science, Southwest University, Chongqing 400715)

Received:1900-01-01 Revised:1900-01-01 Online:2010-03-05 Published:2010-03-05

基于最大匹配的中文分词概率算法研究

何国斌，赵晶璐

(西南大学计算机与信息科学学院，重庆 400715)

Abstract

Abstract: Combined with the sequence table and leaping form fast inquery characteristic, this paper presents an improvement structure of segmentation dictionary. Hashing and binary search is used to segmentation match for enquiring, and in view of the characteristics of the mechanical Chinese word segmentation, by introducing the random number, a Chinese word automatic segmentation probabilistic algorithm is discussed. Experiment indicates that the arithmetic can improve the speed of Chinese segmentation and precision, also, strengthen the processing of dispelling ambiguity.

Key words: segmentation dictionary, leaping form, segmentation algorithm, probabilistic algorithm

摘要： 结合顺序表和跳跃表的快速查询特性，提出一种改进的整词分词词典结构，主要采用哈希法和二分法进行分词匹配，并针对机械分词算法的特点，引入随机数，探讨一种基于最大匹配的分词概率算法。实验表明，该算法具有较高的分词效率和准确率，对消去歧义词也有较好的性能。

关键词: 分词词典, 跳跃表, 分词算法, 概率算法

CLC Number:

TP301.6

HE Guo-bin; ZHAO Jing-lu. Research on Probabilistic Algorithm of Chinese Word Segmentation Based on the Maximum Match[J]. Computer Engineering, 2010, 36(5): 173-175.

何国斌;赵晶璐. 基于最大匹配的中文分词概率算法研究[J]. 计算机工程, 2010, 36(5): 173-175.

/ Recommend / Download Citations

URL:

https://www.ecice06.com/EN/Y2010/V36/I5/173

[1]	Kai WANG, Xiao HAN, Huaji ZHU, Yisheng MIAO, Huarui WU. Segmentation Algorithm of Plug Cabbage Seedlings Based on YOLACT-RFX Model [J]. Computer Engineering, 2023, 49(12): 214-223.
[2]	CHANG Zhi-Jun, YANG Xin. Adaptive Segmentation Algorithm of Bioluminescent Image [J]. Computer Engineering, 2011, 37(4): 218-220.
[3]	HUI Bo-Cong, CHEN Min-Gang, GAO Yan, MA Li-Peng. Multi-resolution Video Segmentation Algorithm Based on 3D Body [J]. Computer Engineering, 2011, 37(20): 282-284.
[4]	CHANG Jing-chao; TAO Tao; CHEN Hong. Intelligent Extraction Technology on Router Configuration Strategy [J]. Computer Engineering, 2008, 34(13): 242-244.

Please choose a citation manager

Content to export

Research on Probabilistic Algorithm of Chinese Word Segmentation Based on the Maximum Match

基于最大匹配的中文分词概率算法研究

PDF

Knowledge

Cited

Abstract

Cite this article

share this article

References

Related Articles 4

Recommended Articles

Metrics

Comments

模态框（Modal）标题

Please choose a citation manager

Content to export

Research on Probabilistic Algorithm of Chinese Word Segmentation Based on the Maximum Match

基于最大匹配的中文分词概率算法研究

PDF

Knowledge

Cited

Abstract

Cite this article

share this article

References

Related Articles 4

Recommended Articles

Metrics

Comments