A Fast Distributed Focused-web Crawling

Harry T. Yani Achsan, Wahyu Catur Wibowo
2014 Procedia Engineering  
Mining data from a web database becomes more challenging in recent years due to the exploding size of data, the rising of dynamic web, and the increasing performance of web security. Mining data from a web database differs from mining data from web sites because it is intended to collect specific data from a single web site. Collecting a very large data in a limited time tends to be detected as a cyber attack and will be banned from connecting into the web server. To avoid the problem, this
more » ... r proposes a crawling method to mine web database faster and cheaper than conventional web crawlers. The method used is to run hundreds of threads from a single web crawler in a single computer and to distribute the threads into hundreds or thousands publicly available proxy servers. This web crawler strategy highly increases the speed of mining and is more secure than using single thread of web crawler.
doi:10.1016/j.proeng.2014.03.017 fatcat:bch3sxlzjzfoznyuihozc4xpv4