PIClean
At a glance
- Citations
- 14
- References
- 18
- Comments
- 0
Abstract
With the dramatic increasing interest in data analysis, ensuring data quality becomes one of the most important topics in data science. Data Cleaning, the process of ensuring data quality, is composed of two stages: error detection and error repair. Despite decades of research in data cleaning, existing cleaning systems still have limitations in terms of usability and error coverage. We propose PIClean, a probabilistic and interactive data cleaning system that aims at addressing the aforementioned limitations. PIClean produces probabilistic errors and probabilistic fixes using low-rank approximation, which implicitly discovers and uses relationships between columns of a dataset for cleaning. The probabilistic errors and fixes are confirmed or rejected by users, and the user feedbacks are constantly incorporated by PIClean to produce more accurate and higher-coverage cleaning results.
Publication details
- DOI
- 10.1145/3299869.3320214
- OpenAlex
- W2948434003
- Document type
- conference-paper
- Language
- EN
- Last metadata update
Comments
Log in to join the discussion.