Radical features for Chinese text classification

Author(s):  
He Hu ◽  
Xiaoyong Du
Author(s):  
Hanqing Tao ◽  
Shiwei Tong ◽  
Hongke Zhao ◽  
Tong Xu ◽  
Binbin Jin ◽  
...  

Recent years, Chinese text classification has attracted more and more research attention. However, most existing techniques which specifically aim at English materials may lose effectiveness on this task due to the huge difference between Chinese and English. Actually, as a special kind of hieroglyphics, Chinese characters and radicals are semantically useful but still unexplored in the task of text classification. To that end, in this paper, we first analyze the motives of using multiple granularity features to represent a Chinese text by inspecting the characteristics of radicals, characters and words. For better representing the Chinese text and then implementing Chinese text classification, we propose a novel Radicalaware Attention-based Four-Granularity (RAFG) model to take full advantages of Chinese characters, words, characterlevel radicals, word-level radicals simultaneously. Specifically, RAFG applies a serialized BLSTM structure which is context-aware and able to capture the long-range information to model the character sharing property of Chinese and sequence characteristics in texts. Further, we design an attention mechanism to enhance the effects of radicals thus model the radical sharing property when integrating granularities. Finally, we conduct extensive experiments, where the experimental results not only show the superiority of our model, but also validate the effectiveness of radicals in the task of Chinese text classification.


2012 ◽  
Vol 24 (3-4) ◽  
pp. 779-798 ◽  
Author(s):  
James N. K. Liu ◽  
Yu-lin He ◽  
Edward H. Y. Lim ◽  
Xi-zhao Wang

Sign in / Sign up

Export Citation Format

Share Document