Integrating Data Warehouses with Web Data: A Survey

J.M. Perez, R. Berlanga, M.J. Aramburu, T.B. Pedersen
2008 IEEE Transactions on Knowledge and Data Engineering  
This paper surveys the most relevant research on combining Data Warehouse (DW) and Web data. It studies the XML technologies that are currently being used to integrate, store, query, and retrieve Web data and their application to DWs. The paper reviews different DW distributed architectures and the use of XML languages as an integration tool in these systems. It also introduces the problem of dealing with semistructured data in a DW. It studies Web data repositories, the design of
more » ... al databases for XML data sources, and the XML extensions of OnLine Analytical Processing techniques. The paper addresses the application of information retrieval technology in a DW to exploit text-rich document collections. The authors hope that the paper will help to discover the main limitations and opportunities that offer the combination of the DW and the Web fields, as well as to identify open research lines. Index Terms-Data warehouse repository, XML/XSL/RDF. Ç 940
doi:10.1109/tkde.2007.190746 fatcat:drcxlccqqfaifcjuzbzspyzlpu