A copy of this work was available on the public web and has been preserved in the Wayback Machine. The capture dates from 2020; you can also visit the original URL.
The file type is application/pdf
.
Duplicate Identification Algorithms in SaaS Platforms
2020
Proceedings of the 2020 Intelligent on Intelligent Cross-Data Analysis and Retrieval Workshop
Existing duplicate records is one of the most common issues in many Software-as-as-Service (SaaS) platforms. In this paper, we study the duplicate identification problem in one specific SaaS platform related to quality and compliance management by using the address information. We interpret all typical mistakes from users that can generate the existent duplicated organizations in a given dataset, collected from the SaaS platform. Also, we create another set by crawling location data from Open
doi:10.1145/3379174.3392319
fatcat:gbvt4urt3nft5mb3xqgxowdf2q