Towards Universal Dense Blocking for Entity Resolution

2024-04-23Code Available0· sign in to hype

Tianshu Wang, Hongyu Lin, Xianpei Han, Xiaoyang Chen, Boxi Cao, Le Sun

Code Available — Be the first to reproduce this paper.

Code

github.com/tshu-w/uniblocker
OfficialIn paperpytorch★ 8
github.com/tshu-w/ublocker
OfficialIn paperpytorch★ 8

Abstract

Blocking is a critical step in entity resolution, and the emergence of neural network-based representation models has led to the development of dense blocking as a promising approach for exploring deep semantics in blocking. However, previous advanced self-supervised dense blocking approaches require domain-specific training on the target domain, which limits the benefits and rapid adaptation of these methods. To address this issue, we propose UniBlocker, a dense blocker that is pre-trained on a domain-independent, easily-obtainable tabular corpus using self-supervised contrastive learning. By conducting domain-independent pre-training, UniBlocker can be adapted to various downstream blocking scenarios without requiring domain-specific fine-tuning. To evaluate the universality of our entity blocker, we also construct a new benchmark covering a wide range of blocking tasks from multiple domains and scenarios. Our experiments show that the proposed UniBlocker, without any domain-specific learning, significantly outperforms previous self- and unsupervised dense blocking methods and is comparable and complementary to the state-of-the-art sparse blocking methods.

Tasks

Blocking Contrastive Learning Entity Resolution

Towards Universal Dense Blocking for Entity Resolution

Code

Abstract

Tasks

Reproductions