REVIEW 3 major objections 5 minor 60 references
OpenForge: Probabilistic Metadata Integration
T0 review · 3 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read Modeling concept relationships as a Markov random field beats GPT-4 by 25 F1 points.
desk verdict Useful two-stage LLM+MRF recipe with strong empirical results, but the parent-child transitivity modeling is formally wrong, so the ICPSR claim needs major revision. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The carrying mechanism is a Markov random field over binary random variables $r_{ij}$, one for each unordered pair of concepts, with two factor types: unary factors $\phi(r_{ij}\mid E)$ that inject prior probabilities from a learned or prompted model, and shared ternary factors $\phi(r_{ij}, r_{jk}, r_{ik})$ over triples $i<j<k$ that give zero potential to the three intransitive configurations and learnable parameters $\theta_1,\dots,\theta_5$ for the valid ones. MAP inference through loopy belief propagation on the factor graph yields posterior predictions. To make inference tractable on large sparse datasets, the MRF is decomposed into independent local graphs by grouping concept pairs by their left concept and keeping only the top-$k$ most similar neighbors, enabling parallel inference over many small factor graphs.
What would settle it
Run the SOTAB experiment with the ternary potential replaced by a constant, so transitivity imposes no penalty; if the F1 score does not drop from its perfect value of 1.0, the transitivity-aware MRF is not what produces the reported gain over GPT-4.
Extended reading notes
Core claim
The central discovery is that pairwise relationship predictions, when made independently, produce intransitive triples, and a shared ternary potential can coordinate them into a globally consistent graph. The paper formalizes metadata integration as finding the relationship assignment graph with maximum joint probability, decomposing that probability into unary factors from prior belief models and a shared ternary potential over triples $i<j<k$ that assigns zero probability to the three configurations violating transitivity. This formulation casts metadata integration as MAP inference on a densely connected Markov random field, and the paper shows that approximate inference by loopy belief propagation, after sparsifying the graph to top-$k$ similar neighbors, is both accurate and scalable to millions of random variables. The result is that OpenForge consistently outperforms GPT-4 and dedicated matching and taxonomy baselines on both equivalence and parent-child tasks.
Load-bearing premise
The account assumes transitivity is a valid hard constraint for both equivalence and parent-child relationships, but the model stores one undirected variable per concept pair and never records edge direction, so the parent-child experiments rely on an unstated ordering of concepts.
Editorial extensions
If this is right
- Metadata curators can generate and refresh equivalence and parent-child links across vocabularies without hand-mapping every concept pair.
- The same two-stage recipe, any prior model plus a transitivity-aware MRF, applies to other relationship types where pairwise judgments should respect global consistency.
- Parameter-efficient fine-tuned LLMs with fewer than ten billion parameters, combined with MRF refinement, can beat a larger general-purpose model on relationship prediction, so the approach runs on a single GPU.
- The local-MRF decomposition keeps inference practical for repositories with thousands of concepts, bringing runtime to minutes and scaling beyond the size of existing public matching or taxonomy datasets.
Reading between the lines
- The parent-child experiments rest on an unstated modeling assumption: with one undirected variable per concept pair, the model cannot represent which concept is the parent of which, so the transitivity constraint for parent-child only makes sense if the concept set carries a hidden ordering that the paper does not specify.
- The perfect F1 on the schema-matching benchmark and the much smaller gain on the sparse entity-matching benchmark suggest the method's advantage is concentrated in densely connected concept collections; in sparse settings, independent pairwise predictions already suffice.
- A natural extension is to learn pair-type-aware ternary potentials or to add higher-order factors encoding axioms beyond transitivity, such as asymmetry or irreflexivity for parent-child relations.
- The top-$k$ neighbor sparsification discards long-range dependencies; a testable check is whether increasing $k$ on medium-scale datasets closes the gap to full MRF inference while preserving the scaling gains.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces OpenForge, a two-stage framework for metadata integration that first obtains per-pair prior beliefs about equivalence or parent-child relationships from LLM prompting, LLM fine-tuning, or classical ML models, and then refines these priors via maximum a posteriori inference on a Markov Random Field with unary and ternary potentials. The ternary potentials are designed to enforce transitivity by assigning zero weight to inconsistent triples, and the remaining potential parameters are learned on a validation set. Experiments on three datasets (SOTAB, Walmart-Amazon, ICPSR) report consistent F1 gains over GPT-4 and task-specific baselines, and a scalability study shows that the proposed local-MRF decomposition supports inference on graphs with millions of random variables.
Significance. If the central claims hold, OpenForge is a practically valuable contribution: it decouples prior generation from consistency-aware refinement, demonstrates that MRF post-processing can improve strong LLM priors on equivalence matching, and shows a credible path to scalable inference via local factor-graph decomposition. The authors ship source code and data, and the comparison against GPT-4 and Unicorn on real datasets is informative. However, the paper's central claim that the same transitivity mechanism works for both equivalence and parent-child relationships is undermined by a modeling flaw: the random variables and the ternary potential are defined for undirected pairs, which cannot represent the directionality inherent in parent-child hierarchies. This affects the interpretation of the ICPSR experiment and the generality of the proposed formalism.
major comments (3)
- [Section 2.2, Definition 1 and Section 3.2.2, Table 1] The model defines one binary variable r_ij for each unordered pair i<j, yet the parent-child relationship is non-symmetric (as the paper itself states in Section 2.2). With a single variable per pair, the MRF cannot represent whether c_i is a parent of c_j or c_j is a parent of c_i. This is not a mere notation issue: the ternary potential in Table 1 assigns zero to all configurations with exactly two 1s, which is correct for equivalence (a disjoint union of cliques) but is wrong for a hierarchy. The valid branching configuration r_ij=1, r_ik=1, r_jk=0 (where c_i is broader than both c_j and c_k, which are siblings) is assigned zero. Thus the ICPSR experiment does not test the stated parent-child transitivity axiom, and the reported 0.91 F1 cannot be attributed to preservation of transitivity. The paper needs either directed random variables (r_ij and r_ji with explicit consistency constraints) or an explicit, defended ordering assumption on concept indices that makes the branching configuration invalid; neither is currently provided.
- [Section 6.2.3 and Section 6.3 (ICPSR results)] Because of the modeling issue above, the ICPSR experiment conflates equivalence-style clique transitivity with directed hierarchy transitivity. The comparison against Chain-of-Layer, a taxonomy-induction method, is therefore not a fair test of OpenForge's ability to induce directed parent-child structures. The ablation in Figure 7(c) is similarly affected: the improvement of MRF over the prior may come from the hard zero constraint rejecting valid branching states rather than from a principled transitivity prior. The authors should either rerun the ICPSR evaluation with a model that can represent directionality or substantially narrow the claims made about handling parent-child relationships.
- [Section 3.2.2, Eq. (3) and Section 6.2] The paper lists 'Preservation of Transitivity' and 'Dependency Learning' as two separate sources of the performance gain, but the ternary potential conflates them. The hard zeros in Table 1 are the transitivity axiom, while the learnable parameters theta_1..theta_5 are shared across all cliques and are tuned on a validation set. As a result, the contribution of transitivity is not isolated in the experiments: any change in F1 could come from the learned label-distribution parameters rather than from the zero constraint. The authors should clarify this decomposition and, if possible, include an ablation that varies the zero constraint (or its relaxation) separately from the learned parameters.
minor comments (5)
- [Section 2.2, Definition 1] The text says 'we consider only the random variables for ordered pairs of concepts' but the definition uses the condition i<j, which describes unordered pairs. This should be corrected to avoid ambiguity, especially because the directionality issue is load-bearing for the parent-child case.
- [Section 6.2.1] The sentence 'We report a detailed comparison of prior models and OpenForge ... in Section 5.1' appears to be a cross-reference error; the detailed prior-model comparison is presented in Section 6.3, not Section 5.1.
- [Section 5.1.1] The description of temperature scaling does not state how the temperature value is chosen; please specify whether it is tuned on a validation set and, if so, with what objective.
- [Figure 7] The caption 'Prior MRF Modeling' is unclear; it should be something like 'Comparison of F1 between prior models and OpenForge (prior + MRF refinement)' to match the two bars per model shown in the figure.
- [Section 4.2 and Section 6.5] The local-MRF construction relies on a top-k neighbor threshold k, but the paper does not analyze how the choice of k affects transitivity violation rates or posterior quality on the real datasets; the scalability experiment in Section 6.5 uses synthetic graphs only, so the end-to-end quality at scale remains untested.
Circularity Check
No significant circularity: the MRF posterior is validated on held-out test sets, the ternary potentials are learned hyperparameters, and the central F1 comparisons are against external baselines.
full rationale
OpenForge's derivation chain is self-contained and empirically grounded. Stage 1 priors come from independently trained or prompted models (Ridge, Random Forest, gemma-2, qwen2.5, and LoRA fine-tuned variants), and Stage 2's ternary potentials are five shared parameters tuned on validation splits via SMAC. Test F1 is computed on held-out labels, so the reported gains over priors and over GPT-4, Unicorn, and Chain-of-Layer are not forced by construction. The transitivity constraint is explicitly hard-coded in Table 1 as zero-potential invalid configurations, but this is an explicit modeling choice rather than a hidden equivalence between input and output: the posterior still depends on evidence-dependent unary potentials and the learned ternary parameters, and the paper's empirical claims concern actual held-out relationship labels. The authors cite their own prior work (e.g., Cong et al. 2023, Nargesian et al. 2018-2023) only as related work on dataset discovery and table search; none of these citations is load-bearing for the MRF formulation or for the claimed results. No 'uniqueness theorem' is invoked, and the closest antecedent MRF taxonomy-induction formulation by Bansal et al. is explicitly credited rather than relabeled. The SMAC tuning of the ternary-potential and LBP parameters is standard hyperparameter optimization on a validation set, not a fitted parameter renamed as a prediction. Concerns about whether the Table 1 potential correctly models directed parent-child transitivity are model-correctness issues, not circularity. Under the stated rules, no circular step can be exhibited from the paper's text.
Assumptions & free parameters
free parameters (4)
- theta_1..theta_5 ternary potentials =
Learned via SMAC on validation; exact values not reported
- LBP damping factor and other inference hyperparameters =
Tuned by SMAC; exact values not reported
- Temperature scaling value for LLM confidence calibration =
Greater than 1, exact value not reported
- Neighbor count k for local MRF construction =
4, 8, 12, 16 in experiments
assumptions (5)
- domain assumption Transitivity holds for both equivalence and parent-child relationships.
- standard math MRF joint probability factorizes as a product of unary and ternary potentials.
- domain assumption Prior probabilities P(r_ij|e_ij) from LLM or ML models are meaningful inputs to the MRF.
- ad hoc to paper Top-k embedding neighbors capture the relationship dependencies that matter.
- ad hoc to paper Disjoint local MRFs after grouping preserve enough global consistency.
Cite this review
Pith. "Pith review of OpenForge: Probabilistic Metadata Integration." pith.science (2026). https://pith.science/paper/63L3FVIQ
@misc{pith2026241209788,
author = {Pith},
title = {Pith review of: OpenForge: Probabilistic Metadata Integration},
year = {2026},
howpublished = {\url{https://pith.science/paper/63L3FVIQ}},
note = {Machine review of arXiv:2412.09788}
}
read the original abstract
Modern data stores increasingly rely on metadata for enabling diverse activities such as data cataloging and search. However, metadata curation remains a labor-intensive task, and the broader challenge of metadata maintenance -- ensuring its consistency, usefulness, and freshness -- has been largely overlooked. In this work, we tackle the problem of resolving relationships among metadata concepts from disparate sources. These relationships are critical for creating clean, consistent, and up-to-date metadata repositories, and a central challenge for metadata integration. We propose OpenForge, a two-stage prior-posterior framework for metadata integration. In the first stage, OpenForge exploits multiple methods including fine-tuned large language models to obtain prior beliefs about concept relationships. In the second stage, OpenForge refines these predictions by leveraging Markov Random Field, a probabilistic graphical model. We formalize metadata integration as an optimization problem, where the objective is to identify the relationship assignments that maximize the joint probability of assignments. The MRF formulation allows OpenForge to capture prior beliefs while encoding critical relationship properties, such as transitivity, in probabilistic inference. Experiments on real-world datasets demonstrate the effectiveness and efficiency of OpenForge. On a use case of matching two metadata vocabularies, OpenForge outperforms GPT-4, the second-best method, by 25 F1-score points.
Figures
Figures from the paper (5 more)
Reference graph
Works this paper leans on
-
[1]
Rohan Anil, Sebastian Borgeaud, Yonghui Wu, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M. Dai, Anja Hauth, Katie Millican, David Silver, Slav Petrov, Melvin Johnson, Ioannis Antonoglou, Julian Schrit- twieser, Amelia Glaese, Jilin Chen, Emily Pitler, Timothy P. Lillicrap, Angeliki Lazaridou, Orhan Firat, James Molloy, Michael ...
arXiv 2023
-
[2]
Ankur Ankan and Abinash Panda. 2015. pgmpy: Probabilistic graphical models using python. In Proceedings of the 14th Python in Science Conference (SCIPY 2015) . Citeseer
work page 2015
-
[3]
Mohit Bansal, David Burkett, Gerard de Melo, and Dan Klein. 2014. Structured Learning for Taxonomy Induction with Belief Propagation. In Proceedings of the 52nd Annual Meeting of the Association for Computational Linguistics, ACL 2014, June 22-27, 2014, Baltimore, MD, USA, Volume 1: Long Papers . The Association for Computer Linguistics, 1041–1051
work page 2014
-
[4]
Alex Bogatu, Alvaro A. A. Fernandes, Norman W. Paton, and Nikolaos Konstanti- nou. 2020. Dataset Discovery in Data Lakes. In ICDE. 709–720
work page 2020
-
[5]
Piotr Bojanowski, Edouard Grave, Armand Joulin, and Tomás Mikolov. 2017. Enriching Word Vectors with Subword Information. Trans. Assoc. Comput. Lin- guistics 5 (2017), 135–146
work page 2017
-
[6]
James Bradbury, Roy Frostig, Peter Hawkins, Matthew James Johnson, Chris Leary, Dougal Maclaurin, George Necula, Adam Paszke, Jake VanderPlas, Skye Wanderman-Milne, and Qiao Zhang. 2018. JAX: composable transformations of Python+NumPy programs. http://github.com/google/jax
2018
-
[7]
Dan Brickley, Matthew Burgess, and Natasha F. Noy. 2019. Google Dataset Search: Building a search engine for datasets in an open Web ecosystem. In WWW. ACM, 1365–1375
work page 2019
-
[8]
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin...
work page 2020
Show all 60 references
-
[9]
Sonia Castelo, Rémi Rampin, Aécio S. R. Santos, Aline Bessa, Fernando Chirigati, and Juliana Freire. 2021. Auctus: A Dataset Search Engine for Data Discovery and Augmentation. Proc. VLDB Endow. 14, 12 (2021), 2791–2794
2021
-
[10]
Web Data Commons. 2012. Extracting Structured Data from the Common Crawl. http://webdatacommons.org Accessed: 2024-11-29
2012
-
[11]
Tianji Cong, James Gale, Jason Frantz, H. V. Jagadish, and Çagatay Demiralp
-
[12]
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ah- mad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, Anirudh Goyal, Anthony Hartshorn, Aobo Yang, Archi Mitra, Archie Sra- vankumar, Artem Korenev, Arthur Hinsvark, Arun Rao, Aston...
2024 arXiv
-
[13]
Gaspard Ducamp, Christophe Gonzales, and Pierre-Henri Wuillemin. 2020. aGrUM/pyAgrum : a toolbox to build models and algorithms for Probabilistic Graphical Models in Python. In International Conference on Probabilistic Graphi- cal Models, PGM 2020, 23-25 September 2020, Aalbor...
2020
-
[14]
Stefano Ermon. 2023. CS 228 - Probabilistic Graphical Models. https:// ermongroup.github.io/cs228-notes/. Accessed: 2024-3-16
2023
-
[15]
Stuart Geman and Donald Geman. 1984. Stochastic Relaxation, Gibbs Distribu- tions, and the Bayesian Restoration of Images. IEEE Trans. Pattern Anal. Mach. Intell. 6, 6 (1984), 721–741
1984
-
[16]
Jaakkola
Amir Globerson and Tommi S. Jaakkola. 2007. Fixing Max-Product: Convergent Message Passing Algorithms for MAP LP-Relaxations. In Advances in Neural Information Processing Systems 20, Proceedings of the Twenty-First Annual Con- ference on Neural Information Processing Systems, ...
2007
-
[17]
Schema.org Community Group and Steering Group. 2011. Schema.org. https: //schema.org. Accessed: 2023-10-23
2011
-
[18]
Weinberger
Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q. Weinberger. 2017. On Calibration of Modern Neural Networks. In Proceedings of the 34th International Confer- ence on Machine Learning, ICML 2017, Sydney, NSW, Australia, 6-11 August 2017 (Proceedings of Machine Learning Research) ...
2017
-
[19]
Halevy, Flip Korn, Natalya Fridman Noy, Christopher Olston, Neoklis Polyzotis, Sudip Roy, and Steven Euijong Whang
Alon Y. Halevy, Flip Korn, Natalya Fridman Noy, Christopher Olston, Neoklis Polyzotis, Sudip Roy, and Steven Euijong Whang. 2016. Goods: Organizing Google’s Datasets. In Proceedings of the 2016 International Conference on Manage- ment of Data, SIGMOD Conference 2016, San Franc...
2016
-
[20]
Marti A. Hearst. 1992. Automatic Acquisition of Hyponyms from Large Text Corpora. In 14th International Conference on Computational Linguistics, COLING 1992, Nantes, France, August 23-28, 1992 . 539–545
1992
-
[21]
Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, and Weizhu Chen
Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, and Weizhu Chen. 2021. LoRA: Low-Rank Adaptation of Large Language Models. CoRR abs/2106.09685 (2021)
2021 arXiv
-
[22]
Bakker, Emanuel Zgraggen, Arvind Satyanarayan, Tim Kraska, Çagatay Demiralp, and César A
Madelon Hulsebos, Kevin Zeng Hu, Michiel A. Bakker, Emanuel Zgraggen, Arvind Satyanarayan, Tim Kraska, Çagatay Demiralp, and César A. Hidalgo
-
[23]
ICPSR. 1962. Inter-university Consortium for Political and Social Research . https: //www.icpsr.umich.edu/web/pages/about/ Accessed: 2023-10-23
1962
-
[24]
Miller, and Mirek Riedewald
Aamod Khatiwada, Grace Fan, Roee Shraga, Zixuan Chen, Wolfgang Gatter- bauer, Renée J. Miller, and Mirek Riedewald. 2023. SANTOS: Relationship-based Semantic Table Union Search. Proc. ACM Manag. Data 1, 1 (2023), 9:1–9:25
2023
-
[25]
Daphne Koller and Nir Friedman. 2009. Probabilistic Graphical Models - Principles and Techniques. MIT Press
2009
-
[26]
Keti Korini, Ralph Peeters, and Christian Bizer. 2022. SOTAB: The WDC Schema.org Table Annotation Benchmark. In Proceedings of the Semantic Web Challenge on Tabular Data to Knowledge Graph Matching, SemTab 2021, co-located with the 21st International Semantic Web Conference, I...
2022
-
[27]
Mendes, Sebastian Hellmann, Mohamed Morsey, Patrick van Kleef, Sören Auer, and Christian Bizer
Jens Lehmann, Robert Isele, Max Jakob, Anja Jentzsch, Dimitris Kontokostas, Pablo N. Mendes, Sebastian Hellmann, Mohamed Morsey, Patrick van Kleef, Sören Auer, and Christian Bizer. 2015. DBpedia - A large-scale, multilingual knowledge base extracted from Wikipedia. Semantic We...
2015
-
[28]
Marius Lindauer, Katharina Eggensperger, Matthias Feurer, André Biedenkapp, Difan Deng, Carolin Benjamins, Tim Ruhkopf, René Sass, and Frank Hutter
-
[29]
Thomas Mesnard, Cassidy Hardin, Robert Dadashi, Surya Bhupatiraju, Shreya Pathak, Laurent Sifre, Morgane Rivière, Mihir Sanjay Kale, Juliette Love, Pouya Tafti, Léonard Hussenot, Aakanksha Chowdhery, Adam Roberts, Aditya Barua, Alex Botev, Alex Castro-Ros, Ambrose Slone, Améli...
2024 arXiv
-
[30]
Miller, Fatemeh Nargesian, Erkang Zhu, Christina Christodoulakis, Ken Q
Renée J. Miller, Fatemeh Nargesian, Erkang Zhu, Christina Christodoulakis, Ken Q. Pu, and Periklis Andritsos. 2018. Making Open Data Transparent: Data Discovery on Open Data. IEEE Data Eng. Bull. 41, 2 (2018), 59–70
2018
-
[31]
Sidharth Mudgal, Han Li, Theodoros Rekatsinas, AnHai Doan, Youngchoon Park, Ganesh Krishnan, Rohit Deep, Esteban Arcaute, and Vijay Raghavendra. 2018. Deep Learning for Entity Matching: A Design Space Exploration. In Proceedings of the 2018 International Conference on Manageme...
2018
-
[32]
Rafael Müller, Simon Kornblith, and Geoffrey E. Hinton. 2019. When does label smoothing help?. InAdvances in Neural Information Processing Systems 32: Annual Conference on Neural Information Processing Systems 2019, NeurIPS 2019, December 8-14, 2019, Vancouver, BC, Canada. 4696–4705
2019
-
[33]
Murphy, Yair Weiss, and Michael I
Kevin P. Murphy, Yair Weiss, and Michael I. Jordan. 1999. Loopy Belief Propa- gation for Approximate Inference: An Empirical Study. In UAI ’99: Proceedings of the Fifteenth Conference on Uncertainty in Artificial Intelligence, Stockholm, Sweden, July 30 - August 1, 1999 . Morg...
1999
-
[34]
Pu, Bahar Ghadiri Bashardoost, Erkang Zhu, and Renée J
Fatemeh Nargesian, Ken Q. Pu, Bahar Ghadiri Bashardoost, Erkang Zhu, and Renée J. Miller. 2023. Data Lake Organization. IEEE Trans. Knowl. Data Eng. 35, 1 (2023), 237–250
2023
-
[35]
Miller, Ken Q
Fatemeh Nargesian, Erkang Zhu, Renée J. Miller, Ken Q. Pu, and Patricia C. Arocena. 2019. Data Lake Management: Challenges and Opportunities. Proc. VLDB Endow. 12, 12 (2019), 1986–1989
2019
-
[37]
Pu, and Renée J
Fatemeh Nargesian, Erkang Zhu, Ken Q. Pu, and Renée J. Miller. 2018. Table Union Search on Open Data. Proc. VLDB Endow. 11, 7 (2018), 813–825
2018
-
[38]
OpenAI. 2023. GPT-4 Technical Report. CoRR abs/2303.08774 (2023)
2023 arXiv
-
[39]
Masayo Ota, Heiko Mueller, Juliana Freire, and Divesh Srivastava. 2020. Data- Driven Domain Discovery for Structured Datasets. Proc. VLDB Endow. 13, 7 (2020), 953–965
2020
-
[40]
Judea Pearl. 1989. Probabilistic reasoning in intelligent systems - networks of plausible inference. Morgan Kaufmann
1989
-
[41]
Fabian Pedregosa, Gaël Varoquaux, Alexandre Gramfort, Vincent Michel, Bertrand Thirion, Olivier Grisel, Mathieu Blondel, Peter Prettenhofer, Ron Weiss, Vincent Dubourg, Jake VanderPlas, Alexandre Passos, David Cournapeau, Matthieu Brucher, Matthieu Perrot, and Edouard Duchesna...
2011
-
[42]
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020. Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer.J. Mach. Learn. Res. 21 (2020), 140:1–140:67
2020
-
[43]
Stephen Roller, Douwe Kiela, and Maximilian Nickel. 2018. Hearst Patterns Re- visited: Automatic Hypernym Detection from Large Text Corpora. InProceedings of the 56th Annual Meeting of the Association for Computational Linguistics, ACL 2018, Melbourne, Australia, July 15-20, 2...
2018
-
[44]
Glenn Shafer and Prakash P. Shenoy. 1990. Probability propagation. Ann. Math. Artif. Intell. 2 (1990), 327–351
1990
-
[45]
Jiaming Shen and Jiawei Han. 2022. Automated Taxonomy Discovery and Explo- ration. Springer
2022
-
[46]
Roee Shraga, Avigdor Gal, and Haggai Roitman. 2020. ADnEV: Cross-Domain Schema Matching using Deep Similarity Matrix Adjustment and Evaluation. Proc. VLDB Endow. 13, 9 (2020), 1401–1415
2020
-
[47]
Rion Snow, Daniel Jurafsky, and Andrew Y. Ng. 2006. Semantic Taxonomy Induction from Heterogenous Evidence. In ACL
2006
-
[48]
Yoshihiko Suhara, Jinfeng Li, Yuliang Li, Dan Zhang, Çagatay Demiralp, Chen Chen, and Wang-Chiew Tan. 2022. Annotating Columns with Pre-trained Lan- guage Models. In SIGMOD ’22: International Conference on Management of Data, Philadelphia, PA, USA, June 12 - 17, 2022 . ACM, 1493–1503
2022
-
[49]
General Services Administration
Technology Transformation Services The U.S. General Services Administration
-
[50]
Thomer, Dharma Akmon, Jeremy York, Allison R
Andrea K. Thomer, Dharma Akmon, Jeremy York, Allison R. B. Tyler, Faye Polasek, Sara Lafia, Libby Hemphill, and Elizabeth Yakel. 2022. The Craft and Coordination of Data Curation: Complicating Workflow Views of Data Science. Proc. ACM Hum. Comput. Interact. 6, CSCW2 (2022), 1–29
2022
-
[51]
Jianhong Tu, Ju Fan, Nan Tang, Peng Wang, Guoliang Li, Xiaoyong Du, Xiaofeng Jia, and Song Gao. 2023. Unicorn: A Unified Multi-tasking Model for Supporting Matching Tasks in Data Integration. Proc. ACM Manag. Data 1, 1 (2023), 84:1– 84:26
2023
-
[52]
Mark D Wilkinson, Michel Dumontier, IJsbrand Jan Aalbersberg, Gabrielle Apple- ton, Myles Axton, Arie Baak, Niklas Blomberg, Jan-Willem Boiten, Luiz Bonino da Silva Santos, Philip E Bourne, et al. 2016. The FAIR Guiding Principles for scientific data management and stewardship...
2016
-
[53]
An Yang, Baosong Yang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Zhou, Cheng- peng Li, Chengyuan Li, Dayiheng Liu, Fei Huang, Guanting Dong, Haoran Wei, Huan Lin, Jialong Tang, Jialin Wang, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Ma, Jianxin Yang, Jin Xu, Jingren Zhou, Jin...
2024 arXiv
-
[54]
Qingkai Zeng, Yuyang Bai, Zhaoxuan Tan, Shangbin Feng, Zhenwen Liang, Zhi- han Zhang, and Meng Jiang. 2024. Chain-of-Layer: Iteratively Prompting Large Language Models for Taxonomy Induction from Limited Examples. In Proceed- ings of the 33rd ACM International Conference on In...
2024
-
[55]
Dan Zhang, Yoshihiko Suhara, Jinfeng Li, Madelon Hulsebos, Çagatay Demiralp, and Wang-Chiew Tan. 2020. Sato: Contextual Semantic Type Detection in Tables. Proc. VLDB Endow. 13, 11 (2020), 1835–1848
2020
-
[56]
Procopiuc, and Divesh Srivastava
Meihui Zhang, Marios Hadjieleftheriou, Beng Chin Ooi, Cecilia M. Procopiuc, and Divesh Srivastava. 2011. Automatic discovery of attributes in relational databases. In Proceedings of the ACM SIGMOD International Conference on Management of Data, SIGMOD 2011, Athens, Greece, Jun...
2011
-
[57]
Guangyao Zhou, Nishanth Kumar, Miguel Lázaro-Gredilla, Shrinu Kushagra, and Dileep George. 2022. PGMax: Factor Graphs for Discrete Probabilistic Graphical Models and Loopy Belief Propagation in JAX. CoRR abs/2202.04110 (2022)
2022 arXiv
-
[2009]
https://opendata.cityofnewyork.us/overview/
Data.gov. https://opendata.cityofnewyork.us/overview/. Accessed: 2023- 10-23
2023
-
[2019]
In Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, KDD 2019, Anchorage, AK, USA, August 4-8, 2019
Sherlock: A Deep Learning Approach to Semantic Data Type Detection. In Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, KDD 2019, Anchorage, AK, USA, August 4-8, 2019 . ACM, 1500–1508
2019
-
[2022]
SMAC3: A Versatile Bayesian Optimization Package for Hyperparameter Optimization. J. Mach. Learn. Res. 23 (2022), 54:1–54:9
2022
-
[2023]
In 13th Conference on Innovative Data Systems Research, CIDR 2023, Amsterdam, The Netherlands, January 8-11, 2023
WarpGate: A Semantic Join Discovery System for Cloud Data Warehouses. In 13th Conference on Innovative Data Systems Research, CIDR 2023, Amsterdam, The Netherlands, January 8-11, 2023 . www.cidrdb.org
2023
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.