Creating a Persian-English Comparable Corpus [chapter]

Homa Baradaran Hashemi, Azadeh Shakery, Heshaam Faili
2010 Lecture Notes in Computer Science  
Multilingual corpora are valuable resources for cross-language information retrieval and are available in many language pairs. However the Persian language does not have rich multilingual resources due to some of its special features and difficulties in constructing the corpora. In this study, we build a Persian-English comparable corpus from two independent news collections: BBC News in English and Hamshahri news in Persian. We use the similarity of the document topics and their publication
more » ... es to align the documents in these sets. We tried several alternatives for constructing the comparable corpora and assessed the quality of the corpora using different criteria. Evaluation results show the high quality of the aligned documents and using the Persian-English comparable corpus for extracting translation knowledge seems promising.
doi:10.1007/978-3-642-15998-5_5 fatcat:dpbwztz4z5axtimn77bqtktghy