Dynamic Classification in Web Archiving Collections

Krutarth Patel, Cornelia Caragea, Mark E. Phillips
2020 International Conference on Language Resources and Evaluation  
The Web archived data usually contains high-quality documents that are very useful for creating specialized collections of documents. To create such collections, there is a substantial need for automatic approaches that can distinguish the documents of interest for a collection out of the large collections (of millions in size) from Web Archiving institutions. However, the patterns of the documents of interest can differ substantially from one document to another, which makes the automatic
more » ... ification task very challenging. In this paper, we explore dynamic fusion models to find, on the fly, the model or combination of models that performs best on a variety of document types. Our experimental results show that the approach that fuses different models outperforms individual models and other ensemble methods on three datasets.
dblp:conf/lrec/PatelCP20 fatcat:aitx2dxttjer5fcusu7uoipw6q