Lightweight Multilingual Entity Extraction and Linking

Aasish Pappu, Roi Blanco, Yashar Mehdad, Amanda Stent, Kapil Thadani
2017 Proceedings of the Tenth ACM International Conference on Web Search and Data Mining - WSDM '17  
Text analytics systems often rely heavily on detecting and linking entity mentions in documents to knowledge bases for downstream applications such as sentiment analysis, question answering and recommender systems. A major challenge for this task is to be able to accurately detect entities in new languages with limited labeled resources. In this paper we present an accurate and lightweight 1 multilingual named entity recognition (NER) and linking (NEL) system. The contributions of this paper
more » ... three-fold: 1) Lightweight named entity recognition with competitive accuracy; 2) Candidate entity retrieval that uses search clicklog data and entity embeddings to achieve high precision with a low memory footprint; and 3) efficient entity disambiguation. Our system achieves state-of-the-art performance on tac kbp 2013 multilingual data and on English aida-conll data.
doi:10.1145/3018661.3018724 dblp:conf/wsdm/PappuBMST17 fatcat:brp5m4y6g5bjli7snvq6pzinna