Publication Date:
2020
abstract:
Record-level matching rules are chains of similarity join pred-icates on multiple attributes employed to join records that refer to the same real-world object when an explicit foreign key is not available on the data sets at hand. They are widely employed by data scientists and practitioners that work with data lakes, open data, and data in the wild. In this work we present a novel technique that allows to efficiently exe-cute record-level matching rules on parallel and distributed systems and demonstrate its efficiency on a real-wold data set.
Iris type:
4.1 Contributo in Atti di convegno
Keywords:
Data integration; Entity resolution; Parallel similarity join
List of contributors:
Gagliardelli, L.; Simonini, G.; Bergamaschi, S.
Book title:
CEUR Workshop Proceedings