Efficient Algorithm for Math Formula Semantic Search

Shunsuke OHASHI, Giovanni Yoko KRISTIANTO, Goran TOPIC, Akiko AIZAWA
2016 IEICE transactions on information and systems  
Mathematical formulae play an important role in many scientific domains. Regardless of the importance of mathematical formula search, conventional keyword-based retrieval methods are not sufficient for searching mathematical formulae, which are structured as trees. The increasing number as well as the structural complexity of mathematical formulae in scientific articles lead to the necessity for large-scale structureaware formula search techniques. In this paper, we formulate three types of
more » ... three types of measures that represent distinctive features of semantic similarity of math formulae, and develop efficient hash-based algorithms for the approximate calculation. Our experiments using NTCIR-11 Math-2 Task dataset, a large-scale test collection for math information retrieval with about 60million formulae, show that the proposed method improves the search precision while also keeps the scalability and runtime efficiency high.
doi:10.1587/transinf.2015dap0023 fatcat:56pl6nmhrzdbrcxw3lbej2glh4