Pith. sign in

REVIEW 3 major objections 5 minor 60 references

OpenForge: Probabilistic Metadata Integration

T0 review · 3 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read Modeling concept relationships as a Markov random field beats GPT-4 by 25 F1 points.

desk verdict Useful two-stage LLM+MRF recipe with strong empirical results, but the parent-child transitivity modeling is formally wrong, so the ICPSR claim needs major revision. read the letter →

arxiv 2412.09788 v1 pith:63L3FVIQ submitted 2024-12-13 cs.DB

classification cs.DB
keywords metadataintegrationMarkovrandomfieldtransitivityprobabilisticinferencelargelanguagemodelstaxonomyinductionentitymatchingdatacuration
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

OpenForge claims to solve metadata integration: automatically deciding, for every pair of concepts from different metadata vocabularies, whether they are equivalent or in a parent-child relation. The paper's central claim is that this problem is best modeled as a maximum a posteriori inference over a Markov random field, where local predictions from any prior model, including fine-tuned large language models, are refined by global transitivity constraints encoded as ternary factors. On three real-world benchmarks, this refinement consistently improves the priors and beats both task-specific state-of-the-art baselines and GPT-4, with the largest gain being 25 F1-score points over GPT-4 on a column-type matching task. If correct, this gives data curators a way to unify and maintain metadata vocabularies at scale, reducing the manual curation bottleneck in dataset discovery.

What carries the argument

The carrying mechanism is a Markov random field over binary random variables $r_{ij}$, one for each unordered pair of concepts, with two factor types: unary factors $\phi(r_{ij}\mid E)$ that inject prior probabilities from a learned or prompted model, and shared ternary factors $\phi(r_{ij}, r_{jk}, r_{ik})$ over triples $i<j<k$ that give zero potential to the three intransitive configurations and learnable parameters $\theta_1,\dots,\theta_5$ for the valid ones. MAP inference through loopy belief propagation on the factor graph yields posterior predictions. To make inference tractable on large sparse datasets, the MRF is decomposed into independent local graphs by grouping concept pairs by their left concept and keeping only the top-$k$ most similar neighbors, enabling parallel inference over many small factor graphs.

What would settle it

Run the SOTAB experiment with the ternary potential replaced by a constant, so transitivity imposes no penalty; if the F1 score does not drop from its perfect value of 1.0, the transitivity-aware MRF is not what produces the reported gain over GPT-4.

Watch

Extended reading notes

Core claim

The central discovery is that pairwise relationship predictions, when made independently, produce intransitive triples, and a shared ternary potential can coordinate them into a globally consistent graph. The paper formalizes metadata integration as finding the relationship assignment graph with maximum joint probability, decomposing that probability into unary factors from prior belief models and a shared ternary potential over triples $i<j<k$ that assigns zero probability to the three configurations violating transitivity. This formulation casts metadata integration as MAP inference on a densely connected Markov random field, and the paper shows that approximate inference by loopy belief propagation, after sparsifying the graph to top-$k$ similar neighbors, is both accurate and scalable to millions of random variables. The result is that OpenForge consistently outperforms GPT-4 and dedicated matching and taxonomy baselines on both equivalence and parent-child tasks.

Load-bearing premise

The account assumes transitivity is a valid hard constraint for both equivalence and parent-child relationships, but the model stores one undirected variable per concept pair and never records edge direction, so the parent-child experiments rely on an unstated ordering of concepts.

Editorial extensions

If this is right

  • Metadata curators can generate and refresh equivalence and parent-child links across vocabularies without hand-mapping every concept pair.
  • The same two-stage recipe, any prior model plus a transitivity-aware MRF, applies to other relationship types where pairwise judgments should respect global consistency.
  • Parameter-efficient fine-tuned LLMs with fewer than ten billion parameters, combined with MRF refinement, can beat a larger general-purpose model on relationship prediction, so the approach runs on a single GPU.
  • The local-MRF decomposition keeps inference practical for repositories with thousands of concepts, bringing runtime to minutes and scaling beyond the size of existing public matching or taxonomy datasets.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The parent-child experiments rest on an unstated modeling assumption: with one undirected variable per concept pair, the model cannot represent which concept is the parent of which, so the transitivity constraint for parent-child only makes sense if the concept set carries a hidden ordering that the paper does not specify.
  • The perfect F1 on the schema-matching benchmark and the much smaller gain on the sparse entity-matching benchmark suggest the method's advantage is concentrated in densely connected concept collections; in sparse settings, independent pairwise predictions already suffice.
  • A natural extension is to learn pair-type-aware ternary potentials or to add higher-order factors encoding axioms beyond transitivity, such as asymmetry or irreflexivity for parent-child relations.
  • The top-$k$ neighbor sparsification discards long-range dependencies; a testable check is whether increasing $k$ on medium-scale datasets closes the gap to full MRF inference while preserving the scaling gains.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper introduces OpenForge, a two-stage framework for metadata integration that first obtains per-pair prior beliefs about equivalence or parent-child relationships from LLM prompting, LLM fine-tuning, or classical ML models, and then refines these priors via maximum a posteriori inference on a Markov Random Field with unary and ternary potentials. The ternary potentials are designed to enforce transitivity by assigning zero weight to inconsistent triples, and the remaining potential parameters are learned on a validation set. Experiments on three datasets (SOTAB, Walmart-Amazon, ICPSR) report consistent F1 gains over GPT-4 and task-specific baselines, and a scalability study shows that the proposed local-MRF decomposition supports inference on graphs with millions of random variables.

Significance. If the central claims hold, OpenForge is a practically valuable contribution: it decouples prior generation from consistency-aware refinement, demonstrates that MRF post-processing can improve strong LLM priors on equivalence matching, and shows a credible path to scalable inference via local factor-graph decomposition. The authors ship source code and data, and the comparison against GPT-4 and Unicorn on real datasets is informative. However, the paper's central claim that the same transitivity mechanism works for both equivalence and parent-child relationships is undermined by a modeling flaw: the random variables and the ternary potential are defined for undirected pairs, which cannot represent the directionality inherent in parent-child hierarchies. This affects the interpretation of the ICPSR experiment and the generality of the proposed formalism.

major comments (3)
  1. [Section 2.2, Definition 1 and Section 3.2.2, Table 1] The model defines one binary variable r_ij for each unordered pair i<j, yet the parent-child relationship is non-symmetric (as the paper itself states in Section 2.2). With a single variable per pair, the MRF cannot represent whether c_i is a parent of c_j or c_j is a parent of c_i. This is not a mere notation issue: the ternary potential in Table 1 assigns zero to all configurations with exactly two 1s, which is correct for equivalence (a disjoint union of cliques) but is wrong for a hierarchy. The valid branching configuration r_ij=1, r_ik=1, r_jk=0 (where c_i is broader than both c_j and c_k, which are siblings) is assigned zero. Thus the ICPSR experiment does not test the stated parent-child transitivity axiom, and the reported 0.91 F1 cannot be attributed to preservation of transitivity. The paper needs either directed random variables (r_ij and r_ji with explicit consistency constraints) or an explicit, defended ordering assumption on concept indices that makes the branching configuration invalid; neither is currently provided.
  2. [Section 6.2.3 and Section 6.3 (ICPSR results)] Because of the modeling issue above, the ICPSR experiment conflates equivalence-style clique transitivity with directed hierarchy transitivity. The comparison against Chain-of-Layer, a taxonomy-induction method, is therefore not a fair test of OpenForge's ability to induce directed parent-child structures. The ablation in Figure 7(c) is similarly affected: the improvement of MRF over the prior may come from the hard zero constraint rejecting valid branching states rather than from a principled transitivity prior. The authors should either rerun the ICPSR evaluation with a model that can represent directionality or substantially narrow the claims made about handling parent-child relationships.
  3. [Section 3.2.2, Eq. (3) and Section 6.2] The paper lists 'Preservation of Transitivity' and 'Dependency Learning' as two separate sources of the performance gain, but the ternary potential conflates them. The hard zeros in Table 1 are the transitivity axiom, while the learnable parameters theta_1..theta_5 are shared across all cliques and are tuned on a validation set. As a result, the contribution of transitivity is not isolated in the experiments: any change in F1 could come from the learned label-distribution parameters rather than from the zero constraint. The authors should clarify this decomposition and, if possible, include an ablation that varies the zero constraint (or its relaxation) separately from the learned parameters.
minor comments (5)
  1. [Section 2.2, Definition 1] The text says 'we consider only the random variables for ordered pairs of concepts' but the definition uses the condition i<j, which describes unordered pairs. This should be corrected to avoid ambiguity, especially because the directionality issue is load-bearing for the parent-child case.
  2. [Section 6.2.1] The sentence 'We report a detailed comparison of prior models and OpenForge ... in Section 5.1' appears to be a cross-reference error; the detailed prior-model comparison is presented in Section 6.3, not Section 5.1.
  3. [Section 5.1.1] The description of temperature scaling does not state how the temperature value is chosen; please specify whether it is tuned on a validation set and, if so, with what objective.
  4. [Figure 7] The caption 'Prior MRF Modeling' is unclear; it should be something like 'Comparison of F1 between prior models and OpenForge (prior + MRF refinement)' to match the two bars per model shown in the figure.
  5. [Section 4.2 and Section 6.5] The local-MRF construction relies on a top-k neighbor threshold k, but the paper does not analyze how the choice of k affects transitivity violation rates or posterior quality on the real datasets; the scalability experiment in Section 6.5 uses synthetic graphs only, so the end-to-end quality at scale remains untested.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the MRF posterior is validated on held-out test sets, the ternary potentials are learned hyperparameters, and the central F1 comparisons are against external baselines.

full rationale

OpenForge's derivation chain is self-contained and empirically grounded. Stage 1 priors come from independently trained or prompted models (Ridge, Random Forest, gemma-2, qwen2.5, and LoRA fine-tuned variants), and Stage 2's ternary potentials are five shared parameters tuned on validation splits via SMAC. Test F1 is computed on held-out labels, so the reported gains over priors and over GPT-4, Unicorn, and Chain-of-Layer are not forced by construction. The transitivity constraint is explicitly hard-coded in Table 1 as zero-potential invalid configurations, but this is an explicit modeling choice rather than a hidden equivalence between input and output: the posterior still depends on evidence-dependent unary potentials and the learned ternary parameters, and the paper's empirical claims concern actual held-out relationship labels. The authors cite their own prior work (e.g., Cong et al. 2023, Nargesian et al. 2018-2023) only as related work on dataset discovery and table search; none of these citations is load-bearing for the MRF formulation or for the claimed results. No 'uniqueness theorem' is invoked, and the closest antecedent MRF taxonomy-induction formulation by Bansal et al. is explicitly credited rather than relabeled. The SMAC tuning of the ternary-potential and LBP parameters is standard hyperparameter optimization on a validation set, not a fitted parameter renamed as a prediction. Concerns about whether the Table 1 potential correctly models directed parent-child transitivity are model-correctness issues, not circularity. Under the stated rules, no circular step can be exhibited from the paper's text.

Assumptions & free parameters 4 free parameters · 5 assumptions · 0 invented entities

No new physical or ontological entities are introduced. The relationship assignment graph, MRF variables, and factors are modeling constructs from standard probabilistic graphical models. The free parameters are all fitting or tuning choices; the axioms are a mix of standard MRF assumptions and paper-specific approximations.

free parameters (4)
  • theta_1..theta_5 ternary potentials = Learned via SMAC on validation; exact values not reported
    Table 1's shared potential function for ternary cliques. These five values determine how strongly the MRF enforces transitivity and class balance. Tuning them on a validation set with 4.05% positives vs 3.98% in test makes part of the MRF gain fitted.
  • LBP damping factor and other inference hyperparameters = Tuned by SMAC; exact values not reported
    Section 5.2 states damping and other LBP parameters are optimized along with the ternary potentials. These affect convergence and the final MAP assignment.
  • Temperature scaling value for LLM confidence calibration = Greater than 1, exact value not reported
    Section 5.1.1 applies temperature scaling to reduce overconfidence in fine-tuned LLM priors; this value is chosen or tuned and affects the prior probabilities fed to the MRF.
  • Neighbor count k for local MRF construction = 4, 8, 12, 16 in experiments
    Section 4.2 selects top-k embedding neighbors for each concept; k trades off transitivity enforcement and scalability and is a user-set parameter.
assumptions (5)
  • domain assumption Transitivity holds for both equivalence and parent-child relationships.
    Section 2.2 and Table 1 encode this as a hard constraint. True for equivalence; for parent-child it is only true along directed chains, and the paper does not model direction explicitly.
  • standard math MRF joint probability factorizes as a product of unary and ternary potentials.
    Equation 2 in Section 3.2.2. Standard for MRFs with positive potentials; the zero potentials for invalid configurations are a limit case.
  • domain assumption Prior probabilities P(r_ij|e_ij) from LLM or ML models are meaningful inputs to the MRF.
    Section 3.1 assumes each pair has observable evidence and a prior model that yields calibrated probabilities. In practice the calibrated probabilities are imperfect, hence temperature scaling is added.
  • ad hoc to paper Top-k embedding neighbors capture the relationship dependencies that matter.
    Section 4.2 states this intuition to justify dropping the dense graph; no formal error bound is given for the approximation.
  • ad hoc to paper Disjoint local MRFs after grouping preserve enough global consistency.
    Section 4.2 groups by left concept so pairs never repeat across MRFs, but global transitivity across groups is intentionally abandoned for scalability.

how reviews work

0 comments
Cite this review

Pith. "Pith review of OpenForge: Probabilistic Metadata Integration." pith.science (2026). https://pith.science/paper/63L3FVIQ

@misc{pith2026241209788,
  author       = {Pith},
  title        = {Pith review of: OpenForge: Probabilistic Metadata Integration},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/63L3FVIQ}},
  note         = {Machine review of arXiv:2412.09788}
}
read the original abstract

Modern data stores increasingly rely on metadata for enabling diverse activities such as data cataloging and search. However, metadata curation remains a labor-intensive task, and the broader challenge of metadata maintenance -- ensuring its consistency, usefulness, and freshness -- has been largely overlooked. In this work, we tackle the problem of resolving relationships among metadata concepts from disparate sources. These relationships are critical for creating clean, consistent, and up-to-date metadata repositories, and a central challenge for metadata integration. We propose OpenForge, a two-stage prior-posterior framework for metadata integration. In the first stage, OpenForge exploits multiple methods including fine-tuned large language models to obtain prior beliefs about concept relationships. In the second stage, OpenForge refines these predictions by leveraging Markov Random Field, a probabilistic graphical model. We formalize metadata integration as an optimization problem, where the objective is to identify the relationship assignments that maximize the joint probability of assignments. The MRF formulation allows OpenForge to capture prior beliefs while encoding critical relationship properties, such as transitivity, in probabilistic inference. Experiments on real-world datasets demonstrate the effectiveness and efficiency of OpenForge. On a use case of matching two metadata vocabularies, OpenForge outperforms GPT-4, the second-best method, by 25 F1-score points.

Figures

Figures reproduced from arXiv: 2412.09788 by the authors.

Figure 1
Figure 1. Illustration of metadata integration problem. [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Illustration of the relationship transitivity (left) [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. Overview of the proposed two-stage prior-posterior framework for integrating metadata concepts. [PITH_FULL_IMAGE:figures/full_fig_p004_3.png] view at source ↗
Figures from the paper (5 more)
Figure 4
Figure 4. Figure 4: The plot on the left demonstrates an instance of our [PITH_FULL_IMAGE:figures/full_fig_p005_4.png]
Figure 5
Figure 5. Figure 5: Creating independent MRFs for concept pairs in [PITH_FULL_IMAGE:figures/full_fig_p007_5.png]
Figure 6
Figure 6. Figure 6: Comparison of F1 score across methods on the SOTAB, Walmart-Amazon, and ICPSR datasets. [PITH_FULL_IMAGE:figures/full_fig_p010_6.png]
Figure 7
Figure 7. Figure 7: Comparison of F1 scores between prior models and [PITH_FULL_IMAGE:figures/full_fig_p011_7.png]
Figure 8
Figure 8. Figure 8: MRF construction and inference time of three infer [PITH_FULL_IMAGE:figures/full_fig_p011_8.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

60 extracted references · 49 canonical work pages

  1. [1]

    Dai, Anja Hauth, Katie Millican, David Silver, Slav Petrov, Melvin Johnson, Ioannis Antonoglou, Julian Schrit- twieser, Amelia Glaese, Jilin Chen, Emily Pitler, Timothy P

    Rohan Anil, Sebastian Borgeaud, Yonghui Wu, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M. Dai, Anja Hauth, Katie Millican, David Silver, Slav Petrov, Melvin Johnson, Ioannis Antonoglou, Julian Schrit- twieser, Amelia Glaese, Jilin Chen, Emily Pitler, Timothy P. Lillicrap, Angeliki Lazaridou, Orhan Firat, James Molloy, Michael ...

  2. [2]

    Ankur Ankan and Abinash Panda. 2015. pgmpy: Probabilistic graphical models using python. In Proceedings of the 14th Python in Science Conference (SCIPY 2015) . Citeseer

  3. [3]

    Mohit Bansal, David Burkett, Gerard de Melo, and Dan Klein. 2014. Structured Learning for Taxonomy Induction with Belief Propagation. In Proceedings of the 52nd Annual Meeting of the Association for Computational Linguistics, ACL 2014, June 22-27, 2014, Baltimore, MD, USA, Volume 1: Long Papers . The Association for Computer Linguistics, 1041–1051

  4. [4]

    Alex Bogatu, Alvaro A. A. Fernandes, Norman W. Paton, and Nikolaos Konstanti- nou. 2020. Dataset Discovery in Data Lakes. In ICDE. 709–720

  5. [5]

    Piotr Bojanowski, Edouard Grave, Armand Joulin, and Tomás Mikolov. 2017. Enriching Word Vectors with Subword Information. Trans. Assoc. Comput. Lin- guistics 5 (2017), 135–146

  6. [6]

    James Bradbury, Roy Frostig, Peter Hawkins, Matthew James Johnson, Chris Leary, Dougal Maclaurin, George Necula, Adam Paszke, Jake VanderPlas, Skye Wanderman-Milne, and Qiao Zhang. 2018. JAX: composable transformations of Python+NumPy programs. http://github.com/google/jax

  7. [7]

    Dan Brickley, Matthew Burgess, and Natasha F. Noy. 2019. Google Dataset Search: Building a search engine for datasets in an open Web ecosystem. In WWW. ACM, 1365–1375

  8. [8]

    Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin...

Show all 60 references
  1. [9]

    Sonia Castelo, Rémi Rampin, Aécio S. R. Santos, Aline Bessa, Fernando Chirigati, and Juliana Freire. 2021. Auctus: A Dataset Search Engine for Data Discovery and Augmentation. Proc. VLDB Endow. 14, 12 (2021), 2791–2794

  2. [10]

    Web Data Commons. 2012. Extracting Structured Data from the Common Crawl. http://webdatacommons.org Accessed: 2024-11-29

  3. [11]

    Tianji Cong, James Gale, Jason Frantz, H. V. Jagadish, and Çagatay Demiralp

  4. [12]

    Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ah- mad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, Anirudh Goyal, Anthony Hartshorn, Aobo Yang, Archi Mitra, Archie Sra- vankumar, Artem Korenev, Arthur Hinsvark, Arun Rao, Aston...

  5. [13]

    Gaspard Ducamp, Christophe Gonzales, and Pierre-Henri Wuillemin. 2020. aGrUM/pyAgrum : a toolbox to build models and algorithms for Probabilistic Graphical Models in Python. In International Conference on Probabilistic Graphi- cal Models, PGM 2020, 23-25 September 2020, Aalbor...

  6. [14]

    Stefano Ermon. 2023. CS 228 - Probabilistic Graphical Models. https:// ermongroup.github.io/cs228-notes/. Accessed: 2024-3-16

  7. [15]

    Stuart Geman and Donald Geman. 1984. Stochastic Relaxation, Gibbs Distribu- tions, and the Bayesian Restoration of Images. IEEE Trans. Pattern Anal. Mach. Intell. 6, 6 (1984), 721–741

  8. [16]

    Jaakkola

    Amir Globerson and Tommi S. Jaakkola. 2007. Fixing Max-Product: Convergent Message Passing Algorithms for MAP LP-Relaxations. In Advances in Neural Information Processing Systems 20, Proceedings of the Twenty-First Annual Con- ference on Neural Information Processing Systems, ...

  9. [17]

    Schema.org Community Group and Steering Group. 2011. Schema.org. https: //schema.org. Accessed: 2023-10-23

  10. [18]

    Weinberger

    Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q. Weinberger. 2017. On Calibration of Modern Neural Networks. In Proceedings of the 34th International Confer- ence on Machine Learning, ICML 2017, Sydney, NSW, Australia, 6-11 August 2017 (Proceedings of Machine Learning Research) ...

  11. [19]

    Halevy, Flip Korn, Natalya Fridman Noy, Christopher Olston, Neoklis Polyzotis, Sudip Roy, and Steven Euijong Whang

    Alon Y. Halevy, Flip Korn, Natalya Fridman Noy, Christopher Olston, Neoklis Polyzotis, Sudip Roy, and Steven Euijong Whang. 2016. Goods: Organizing Google’s Datasets. In Proceedings of the 2016 International Conference on Manage- ment of Data, SIGMOD Conference 2016, San Franc...

  12. [20]

    Marti A. Hearst. 1992. Automatic Acquisition of Hyponyms from Large Text Corpora. In 14th International Conference on Computational Linguistics, COLING 1992, Nantes, France, August 23-28, 1992 . 539–545

  13. [21]

    Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, and Weizhu Chen

    Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, and Weizhu Chen. 2021. LoRA: Low-Rank Adaptation of Large Language Models. CoRR abs/2106.09685 (2021)

  14. [22]

    Bakker, Emanuel Zgraggen, Arvind Satyanarayan, Tim Kraska, Çagatay Demiralp, and César A

    Madelon Hulsebos, Kevin Zeng Hu, Michiel A. Bakker, Emanuel Zgraggen, Arvind Satyanarayan, Tim Kraska, Çagatay Demiralp, and César A. Hidalgo

  15. [23]

    ICPSR. 1962. Inter-university Consortium for Political and Social Research . https: //www.icpsr.umich.edu/web/pages/about/ Accessed: 2023-10-23

  16. [24]

    Miller, and Mirek Riedewald

    Aamod Khatiwada, Grace Fan, Roee Shraga, Zixuan Chen, Wolfgang Gatter- bauer, Renée J. Miller, and Mirek Riedewald. 2023. SANTOS: Relationship-based Semantic Table Union Search. Proc. ACM Manag. Data 1, 1 (2023), 9:1–9:25

  17. [25]

    Daphne Koller and Nir Friedman. 2009. Probabilistic Graphical Models - Principles and Techniques. MIT Press

  18. [26]

    Keti Korini, Ralph Peeters, and Christian Bizer. 2022. SOTAB: The WDC Schema.org Table Annotation Benchmark. In Proceedings of the Semantic Web Challenge on Tabular Data to Knowledge Graph Matching, SemTab 2021, co-located with the 21st International Semantic Web Conference, I...

  19. [27]

    Mendes, Sebastian Hellmann, Mohamed Morsey, Patrick van Kleef, Sören Auer, and Christian Bizer

    Jens Lehmann, Robert Isele, Max Jakob, Anja Jentzsch, Dimitris Kontokostas, Pablo N. Mendes, Sebastian Hellmann, Mohamed Morsey, Patrick van Kleef, Sören Auer, and Christian Bizer. 2015. DBpedia - A large-scale, multilingual knowledge base extracted from Wikipedia. Semantic We...

  20. [28]

    Marius Lindauer, Katharina Eggensperger, Matthias Feurer, André Biedenkapp, Difan Deng, Carolin Benjamins, Tim Ruhkopf, René Sass, and Frank Hutter

  21. [29]

    Thomas Mesnard, Cassidy Hardin, Robert Dadashi, Surya Bhupatiraju, Shreya Pathak, Laurent Sifre, Morgane Rivière, Mihir Sanjay Kale, Juliette Love, Pouya Tafti, Léonard Hussenot, Aakanksha Chowdhery, Adam Roberts, Aditya Barua, Alex Botev, Alex Castro-Ros, Ambrose Slone, Améli...

  22. [30]

    Miller, Fatemeh Nargesian, Erkang Zhu, Christina Christodoulakis, Ken Q

    Renée J. Miller, Fatemeh Nargesian, Erkang Zhu, Christina Christodoulakis, Ken Q. Pu, and Periklis Andritsos. 2018. Making Open Data Transparent: Data Discovery on Open Data. IEEE Data Eng. Bull. 41, 2 (2018), 59–70

  23. [31]

    Sidharth Mudgal, Han Li, Theodoros Rekatsinas, AnHai Doan, Youngchoon Park, Ganesh Krishnan, Rohit Deep, Esteban Arcaute, and Vijay Raghavendra. 2018. Deep Learning for Entity Matching: A Design Space Exploration. In Proceedings of the 2018 International Conference on Manageme...

  24. [32]

    Rafael Müller, Simon Kornblith, and Geoffrey E. Hinton. 2019. When does label smoothing help?. InAdvances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems 2019, NeurIPS 2019, December 8-14, 2019, Vancouver, BC, Canada. 4696–4705

  25. [33]

    Murphy, Yair Weiss, and Michael I

    Kevin P. Murphy, Yair Weiss, and Michael I. Jordan. 1999. Loopy Belief Propa- gation for Approximate Inference: An Empirical Study. In UAI ’99: Proceedings of the Fifteenth Conference on Uncertainty in Artificial Intelligence, Stockholm, Sweden, July 30 - August 1, 1999 . Morg...

  26. [34]

    Pu, Bahar Ghadiri Bashardoost, Erkang Zhu, and Renée J

    Fatemeh Nargesian, Ken Q. Pu, Bahar Ghadiri Bashardoost, Erkang Zhu, and Renée J. Miller. 2023. Data Lake Organization. IEEE Trans. Knowl. Data Eng. 35, 1 (2023), 237–250

  27. [35]

    Miller, Ken Q

    Fatemeh Nargesian, Erkang Zhu, Renée J. Miller, Ken Q. Pu, and Patricia C. Arocena. 2019. Data Lake Management: Challenges and Opportunities. Proc. VLDB Endow. 12, 12 (2019), 1986–1989

  28. [37]

    Pu, and Renée J

    Fatemeh Nargesian, Erkang Zhu, Ken Q. Pu, and Renée J. Miller. 2018. Table Union Search on Open Data. Proc. VLDB Endow. 11, 7 (2018), 813–825

  29. [38]

    OpenAI. 2023. GPT-4 Technical Report. CoRR abs/2303.08774 (2023)

  30. [39]

    Masayo Ota, Heiko Mueller, Juliana Freire, and Divesh Srivastava. 2020. Data- Driven Domain Discovery for Structured Datasets. Proc. VLDB Endow. 13, 7 (2020), 953–965

  31. [40]

    Judea Pearl. 1989. Probabilistic reasoning in intelligent systems - networks of plausible inference. Morgan Kaufmann

  32. [41]

    Fabian Pedregosa, Gaël Varoquaux, Alexandre Gramfort, Vincent Michel, Bertrand Thirion, Olivier Grisel, Mathieu Blondel, Peter Prettenhofer, Ron Weiss, Vincent Dubourg, Jake VanderPlas, Alexandre Passos, David Cournapeau, Matthieu Brucher, Matthieu Perrot, and Edouard Duchesna...

  33. [42]

    Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020. Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer.J. Mach. Learn. Res. 21 (2020), 140:1–140:67

  34. [43]

    Stephen Roller, Douwe Kiela, and Maximilian Nickel. 2018. Hearst Patterns Re- visited: Automatic Hypernym Detection from Large Text Corpora. InProceedings of the 56th Annual Meeting of the Association for Computational Linguistics, ACL 2018, Melbourne, Australia, July 15-20, 2...

  35. [44]

    Glenn Shafer and Prakash P. Shenoy. 1990. Probability propagation. Ann. Math. Artif. Intell. 2 (1990), 327–351

  36. [45]

    Jiaming Shen and Jiawei Han. 2022. Automated Taxonomy Discovery and Explo- ration. Springer

  37. [46]

    Roee Shraga, Avigdor Gal, and Haggai Roitman. 2020. ADnEV: Cross-Domain Schema Matching using Deep Similarity Matrix Adjustment and Evaluation. Proc. VLDB Endow. 13, 9 (2020), 1401–1415

  38. [47]

    Rion Snow, Daniel Jurafsky, and Andrew Y. Ng. 2006. Semantic Taxonomy Induction from Heterogenous Evidence. In ACL

  39. [48]

    Yoshihiko Suhara, Jinfeng Li, Yuliang Li, Dan Zhang, Çagatay Demiralp, Chen Chen, and Wang-Chiew Tan. 2022. Annotating Columns with Pre-trained Lan- guage Models. In SIGMOD ’22: International Conference on Management of Data, Philadelphia, PA, USA, June 12 - 17, 2022 . ACM, 1493–1503

  40. [49]

    General Services Administration

    Technology Transformation Services The U.S. General Services Administration

  41. [50]

    Thomer, Dharma Akmon, Jeremy York, Allison R

    Andrea K. Thomer, Dharma Akmon, Jeremy York, Allison R. B. Tyler, Faye Polasek, Sara Lafia, Libby Hemphill, and Elizabeth Yakel. 2022. The Craft and Coordination of Data Curation: Complicating Workflow Views of Data Science. Proc. ACM Hum. Comput. Interact. 6, CSCW2 (2022), 1–29

  42. [51]

    Jianhong Tu, Ju Fan, Nan Tang, Peng Wang, Guoliang Li, Xiaoyong Du, Xiaofeng Jia, and Song Gao. 2023. Unicorn: A Unified Multi-tasking Model for Supporting Matching Tasks in Data Integration. Proc. ACM Manag. Data 1, 1 (2023), 84:1– 84:26

  43. [52]

    Mark D Wilkinson, Michel Dumontier, IJsbrand Jan Aalbersberg, Gabrielle Apple- ton, Myles Axton, Arie Baak, Niklas Blomberg, Jan-Willem Boiten, Luiz Bonino da Silva Santos, Philip E Bourne, et al. 2016. The FAIR Guiding Principles for scientific data management and stewardship...

  44. [53]

    An Yang, Baosong Yang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Zhou, Cheng- peng Li, Chengyuan Li, Dayiheng Liu, Fei Huang, Guanting Dong, Haoran Wei, Huan Lin, Jialong Tang, Jialin Wang, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Ma, Jianxin Yang, Jin Xu, Jingren Zhou, Jin...

  45. [54]

    Qingkai Zeng, Yuyang Bai, Zhaoxuan Tan, Shangbin Feng, Zhenwen Liang, Zhi- han Zhang, and Meng Jiang. 2024. Chain-of-Layer: Iteratively Prompting Large Language Models for Taxonomy Induction from Limited Examples. In Proceed- ings of the 33rd ACM International Conference on In...

  46. [55]

    Dan Zhang, Yoshihiko Suhara, Jinfeng Li, Madelon Hulsebos, Çagatay Demiralp, and Wang-Chiew Tan. 2020. Sato: Contextual Semantic Type Detection in Tables. Proc. VLDB Endow. 13, 11 (2020), 1835–1848

  47. [56]

    Procopiuc, and Divesh Srivastava

    Meihui Zhang, Marios Hadjieleftheriou, Beng Chin Ooi, Cecilia M. Procopiuc, and Divesh Srivastava. 2011. Automatic discovery of attributes in relational databases. In Proceedings of the ACM SIGMOD International Conference on Management of Data, SIGMOD 2011, Athens, Greece, Jun...

  48. [57]

    Guangyao Zhou, Nishanth Kumar, Miguel Lázaro-Gredilla, Shrinu Kushagra, and Dileep George. 2022. PGMax: Factor Graphs for Discrete Probabilistic Graphical Models and Loopy Belief Propagation in JAX. CoRR abs/2202.04110 (2022)

  49. [2009]

    https://opendata.cityofnewyork.us/overview/

    Data.gov. https://opendata.cityofnewyork.us/overview/. Accessed: 2023- 10-23

  50. [2019]

    In Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, KDD 2019, Anchorage, AK, USA, August 4-8, 2019

    Sherlock: A Deep Learning Approach to Semantic Data Type Detection. In Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, KDD 2019, Anchorage, AK, USA, August 4-8, 2019 . ACM, 1500–1508

  51. [2022]

    SMAC3: A Versatile Bayesian Optimization Package for Hyperparameter Optimization. J. Mach. Learn. Res. 23 (2022), 54:1–54:9

  52. [2023]

    In 13th Conference on Innovative Data Systems Research, CIDR 2023, Amsterdam, The Netherlands, January 8-11, 2023

    WarpGate: A Semantic Join Discovery System for Cloud Data Warehouses. In 13th Conference on Innovative Data Systems Research, CIDR 2023, Amsterdam, The Netherlands, January 8-11, 2023 . www.cidrdb.org

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.