ptype: probabilistic type inference

Taha Ceritli, Christopher K. I. Williams, James Geddes
2020 Data mining and knowledge discovery  
Type inference refers to the task of inferring the data type of a given column of data. Current approaches often fail when data contains missing data and anomalies, which are found commonly in real-world data sets. In this paper, we propose ptype, a probabilistic robust type inference method that allows us to detect such entries, and infer data types. We further show that the proposed method outperforms existing methods.
doi:10.1007/s10618-020-00680-1 fatcat:kbcuqqrdi5cbjdme2tawvywreu