Building a research library for the history of the web

William Y. Arms, Selcuk Aya, Pavel Dmitriev, Blazej J. Kot, Ruth Mitchell, Lucia Walle
2006 Proceedings of the 6th ACM/IEEE-CS joint conference on Digital libraries - JCDL '06  
This paper describes the building of a research library for studying the Web, especially research on how the structure and content of the Web change over time. The library is particularly aimed at supporting social scientists for whom the Web is both a fascinating social phenomenon and a mirror on society. The library is built on the collections of the Internet Archive, which has been preserving a crawl of the Web every two months since 1996. The technical challenges in organizing this data for
more » ... research fall into two categories: highperformance computing to transfer and manage the very large amounts of data, and human-computer interfaces that empower research by non-computer specialists. Categories and Subject Descriptors H.3.7 [Information Storage and R e t r i e v a l ]: Digital Librariescollection, systems issues, user issues. J.4 [S o c i a l and Behavioral Sciences]: sociology.
doi:10.1145/1141753.1141771 dblp:conf/jcdl/ArmsADKMW06 fatcat:ga2epyap45atneus6hgl2xwry4