A Mixed Semantic Features Model for Chinese NER with Characters and Words [chapter]

Ning Chang, Jiang Zhong, Qing Li, Jiang Zhu
2020 Lecture Notes in Computer Science  
Named Entity Recognition (NER) is an essential part of many natural language processing (NLP) tasks. The existing Chinese NER methods are mostly based on word segmentation, or use the character sequences as input. However, using a single granularity representation would suffer from the problems of out-of-vocabulary and word segmentation errors, and the semantic content is relatively simple. In this paper, we introduce the self-attention mechanism into the BiLSTM-CRF neural network structure for
more » ... Chinese named entity recognition with two embedding. Different from other models, our method combines character and word features at the sequence level, and the attention mechanism computes similarity on the total sequence consisted of characters and words. The character semantic information and the structure of words work together to improve the accuracy of word boundary segmentation and solve the problem of long-phrase combination. We validate our model on MSRA and Weibo corpora, and experiments demonstrate that our model can significantly improve the performance of the Chinese NER task.
doi:10.1007/978-3-030-45439-5_24 fatcat:rrp52ui4u5hvbcpo7hcimtm3gq