Abstract
Text similarity algorithm is widely used in plurality fields, such as copy detection, text classification, machine translation, intelligent question answering system and natural language processing. At present, vector space model algorithm, which is more commonly used, does not consider the information of semantic features adequately, and the accuracy of the semantic similarity computation results can be further improved. This paper proposes a text similarity computation method which combines the HowNet with vector space model. Similarity computation is divided into two levels. In the level of words, words-similarity calculation based on HowNet prevents the loss of semantic information. In the level of texts, text-similarity calculation by vector space model ensures the integrity of the information expressed in the texts. This paper designs an experiment of news text classification based on KNN algorithm, in which data obtained from a part of the Chinese news in Sogou data corpora. Experimental results show that the method proposed in this paper is more accurate than the traditional vector space model algorithm.
Access this chapter
Tax calculation will be finalised at checkout
Purchases are for personal use only
Similar content being viewed by others
References
Jin Xiqian. (2009). Research on Semantic Based Chinese Text Similarity Algorithm. (Doctoral dissertation, Zhejiang University of Technology).
G. Salton, A. Wong, ang C.S. Yang, A Vector Space Model for Information Retrieval, Journal of the ASIS, 18:11, 613–620, November 1975.
Liu Xiaojun, Zhao Dong, & Yao Weidong. (2007). A Two Factor Similarity Algorithm for Chinese Text Search. Computer Simulation, 24(12), 312–314.
Chen Feihong. (2011). Research on Chinese Text Similarity Algorithm Based on Vector Space Model. (Doctoral dissertation, University of Electronic Science and technology).
Kuai Yuanyuan. (2014). Research on Semantic Based Text Similarity Algorithm. Computer CD software and Applications (9), 302–303.
Liu Qun & Li Sujian. (2002). Based on the HowNet Lexical Semantic Similarity Computation. Chinese of computational linguistics.
Fan Hongyi, & Zhang Yangsen (2014). A method for semantic similarity of words based on HowNet. Journal of Beijing Information Science and Technology University: Natural Science Edition (4), 42–45.
Author information
Authors and Affiliations
Corresponding author
Editor information
Editors and Affiliations
Rights and permissions
Copyright information
© 2017 Springer Nature Singapore Pte Ltd.
About this paper
Cite this paper
Liu, Y., Li, Z. (2017). Semantic Based Text Similarity Computation. In: Zhao, P., Ouyang, Y., Xu, M., Yang, L., Ouyang, Y. (eds) Advanced Graphic Communications and Media Technologies . PPMT 2016. Lecture Notes in Electrical Engineering, vol 417. Springer, Singapore. https://doi.org/10.1007/978-981-10-3530-2_43
Download citation
DOI: https://doi.org/10.1007/978-981-10-3530-2_43
Published:
Publisher Name: Springer, Singapore
Print ISBN: 978-981-10-3529-6
Online ISBN: 978-981-10-3530-2
eBook Packages: EngineeringEngineering (R0)