pith. sign in

arxiv: 1807.10009 · v1 · pith:I2VM5AF7new · submitted 2018-07-26 · 💻 cs.IR · cs.DB

General Context-Aware Data Matching and Merging Framework

classification 💻 cs.IR cs.DB
keywords frameworkdatageneralcontextdimensionsdomainhoweverindependent
0
0 comments X
read the original abstract

Due to numerous public information sources and services, many methods to combine heterogeneous data were proposed recently. However, general end-to-end solutions are still rare, especially systems taking into account different context dimensions. Therefore, the techniques often prove insufficient or are limited to a certain domain. In this paper we briefly review and rigorously evaluate a general framework for data matching and merging. The framework employs collective entity resolution and redundancy elimination using three dimensions of context types. In order to achieve domain independent results, data is enriched with semantics and trust. However, the main contribution of the paper is evaluation on five public domain-incompatible datasets. Furthermore, we introduce additional attribute, relationship, semantic and trust metrics, which allow complete framework management. Besides overall results improvement within the framework, metrics could be of independent interest.

This paper has not been read by Pith yet.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.