Pith. sign in

REVIEW 4 cited by

Match, Compare, or Select? An Investigation of Large Language Models for Entity Matching

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.16884 v3 pith:VEUQHWLL submitted 2024-05-27 cs.CL cs.DB

classification cs.CLcs.DB
keywords matchingentitycomemllmsrecordadvantagescomparedifferent
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Entity matching (EM) is a critical step in entity resolution (ER). Recently, entity matching based on large language models (LLMs) has shown great promise. However, current LLM-based entity matching approaches typically follow a binary matching paradigm that ignores the global consistency among record relationships. In this paper, we investigate various methodologies for LLM-based entity matching that incorporate record interactions from different perspectives. Specifically, we comprehensively compare three representative strategies: matching, comparing, and selecting, and analyze their respective advantages and challenges in diverse scenarios. Based on our findings, we further design a compound entity matching framework (ComEM) that leverages the composition of multiple strategies and LLMs. ComEM benefits from the advantages of different sides and achieves improvements in both effectiveness and efficiency. Experimental results on 8 ER datasets and 10 LLMs verify the superiority of incorporating record interactions through the selecting strategy, as well as the further cost-effectiveness brought by ComEM.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. TransClean: Finding False Positives in Multi-Source Entity Matching under Real-World Conditions via Transitive Consistency

    cs.DB 2025-06 conditional novelty 6.0 of 10

    TransClean uses a model's predictions on transitive, implied record pairs to locate and remove false positive matches in multi-source entity resolution.

  2. Omni Geometry Representation Learning vs Large Language Models for Geospatial Entity Resolution

    cs.DB 2025-08 unverdicted novelty 5.0 of 10

    A geometry-aware neural encoder plus attribute-aware language modeling improves geospatial entity resolution by up to 12% F1 over point-only baselines, with large language models competitive.

  3. Linking Cryptoasset Attribution Tags to Knowledge Graph Entities: An LLM-based Approach

    cs.CR 2025-02 conditional novelty 5.0 of 10

    An LLM-based entity linking pipeline maps cryptoasset attribution tags to knowledge graph actors and outperforms baselines on three datasets.

  4. Beyond Traditional Algorithms: Leveraging LLMs for Accurate Cross-Border Entity Identification

    cs.CL 2025-07 reject novelty 3.0 of 10

    A 65-case comparison claims commercial chatbot LLMs are the most accurate for Portuguese entity matching, but the reported false-positive rates contradict the claim.

Pith tools