Harvesting Large-Scale Grids for Software Resources

Asterios Katsifodimos, George Pallis, Marios D. Dikaiakos
2009 2009 9th IEEE/ACM International Symposium on Cluster Computing and the Grid  
Grid infrastructures are in operation around the world, federating an impressive collection of computational resources and a wide variety of application software. In this context, it is important to establish advanced software discovery services that could help end-users locate software components suitable to their needs. In this paper, we present the design, architecture and implementation of an open-source keywordbased paradigm for the search of software resources in Grid infrastructures,
more » ... ed Minersoft. A key goal of Minersoft is to annotate automatically all the software resources with keywordrich metadata. Using advanced Information Retrieval techniques, we locate software resources with respect to users queries. Experiments were conducted in EGEE, one of the largest Grid production services currently in operation. Results showed that Minersoft successfully crawled 12.3 million valid files (620 GB size) and sustained, in most sites, high crawling rates.
doi:10.1109/ccgrid.2009.51 dblp:conf/ccgrid/KatsifodimosPD09 fatcat:6mpkaingyzalnmrio2v7rqzp6q