Skip to content
Open access

PINE: Extracting Correlated Token Pairs for Explainable Entity Matching

Jul 2026 · The VLDB journal · Vol 35 · 0 citations · 38 references
Computer Science

TL;DR

A new method, Pair INterpretation for Entity matching (PINE), which takes two records as input, and outputs correlated token pairs as an explanation for an entity-matching decision, demonstrating that the extracted token pairs exhibit strong correlations and serve as interpretable evidence for matching records.

Abstract

Explanation techniques such as local interpretable model-agnostic explanation (LIME) provide reasons behind decisions made by machine-learning models. These methods typically use a set of features and their values as inputs and identify those that significantly influence the final decision. However, machine-learning models for entity matching operate on two sets of tokens or records, each representing an entity, to determine whether they refer to the same real-world entity. Explanations for entity-matching decisions are more convincing when they highlight contributing pairs of tokens within the pair of records, rather than focusing on individual tokens alone. In this sense, existing explanation techniques are insufficient for entity matching. Therefore, we propose a new method, Pair INterpretation for Entity matching (PINE), which takes two records as input, and outputs correlated token pairs as an explanation for an entity-matching decision. Our extensive experiments on public datasets demonstrate that the extracted token pairs exhibit strong correlations and serve as interpretable evidence for matching records.

Read PDF

Similar papers

Jul 2026

Beyond Scale and Generation: Understanding Language Model-based Entity Matching

The factors underlying performance differences across matcher architectures are clarified and motivate future research and benchmark designs that better disentangle architectural choices from model-level factors while explicitly evaluating distribution shift and cross-dataset transferability.

Zeyu Zhang, Xue Li, Iacer Calixto et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Can We Do Interpretable NLI with Graphs Based on Atomic Propositions?

While Large Language Model (LLM)-based Natural Language Inference (NLI) systems achieve high accuracy, their decision-making processes lack auditable structures. This paper explores whether NLI can be performed using only interpretable, graph-based representations of evidence. We introduce a fully graph-based pipeline...

Younes Boufouss, Luc Pommeret, Thomas Gerald et al. · 0 citations
#large language models Open access Sep 2026

Meticulously Unsupervised Entity Alignment With Large Language Models

Knowledge graph entity alignment refers to the process of identifying and linking entities that refer to the same real‐world object from different knowledge graphs. Structural heterogeneity and scarcity of training data have always been two major challenges that impede entity alignment task. The advent of Large Languag...

Zhi-Huan Yan, Yi Wang, Chong-Chong Zhang et al. · 0 citations
Preprint Aug 2026

Domain-Specific Text Embedding Models for Entity Resolution

General-purpose text embedding models are designed to capture semantic similarity but are not optimised for distinguishing entity records that represent the same real-world business or person. This limitation affects applications such as entity resolution and duplicate record retrieval, where small textual differences...

Khajesh Sapram, S. Raju, Kishore Konda · 0 citations
#natural language process... Preprint Sep 2026

Think Before You Link: Rarity, Reasoning, and Retrieval in Multilingual Entity Linking

Multimodal entity linking grounds entity mentions in text and images to knowledge-base entries. These systems degrade on rare entities, but prior work measures rarity primarily through popularity-based metrics such as pageviews. We broaden this view using knowledge-graph structural metrics that capture how well an enti...

Parinthapat Pengpun, Simran Khanuja, Graham Neubig · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.