What is on Social Media that is not in WordNet? A Preliminary Analysis on the TwitterAAE Corpus

Cecilia Domingo, Tatiana Gonzalez-Ferrero, Itziar Gonzalez-Dios
2021 Global WordNet Conference  
Natural Language Processing tools and resources have been so far mainly created and trained for standard varieties of language. Nowadays, with the use of large amounts of data gathered from social media, other varieties and registers need to be processed, which may present other challenges and difficulties. In this work, we focus on English and we present a preliminary analysis by comparing the Twitter-AAE corpus, which is annotated for ethnicity, and WordNet by quantifying and explaining the online language that Word-Net misses.
dblp:conf/wordnet/DomingoGG21 fatcat:xrf2f2vcxbeszfvafpww72bl6q