REVIEW 3 major objections 6 minor 86 references
GIST keeps research taxonomies current by integrating author-written hierarchies in a geometric box space, beating pure LLM methods at a fraction of the cost.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-13 05:03 UTC pith:TYBBUB4C
load-bearing objection Practical systems paper that defines continuous taxonomy maintenance over open arXiv streams and delivers clear cost-quality gains; single-survey final-GT Soft F1 is the main soft spot, not a fatal flaw. the 3 major comments →
Taxonomy Maintenance In The Wild Over Evolving Scholarly Data: Reliability, Efficiency, and Cost-Effectiveness
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
GIST shows that continuously maintained, high-quality scholarly taxonomies can be obtained by grounding structure induction in author-curated Related Work hierarchies, integrating those partial trees under box-containment geometry, and updating only a novelty-aware coreset of historical signals, all while a hypothesized-concept generator and budget-aware planner keep retrieval cost under a user-specified token budget. On twelve arXiv domains this yields 11 % / 13 % higher Node / Edge Soft F1 than the strongest LLM baseline at 9.6 % of its runtime and 12.7 % of its monetary cost.
What carries the argument
Geometric box embeddings with a bidirectional word–box mapping: each concept is an axis-aligned box whose containment encodes is-a; an MLP pair learns to map word embeddings into boxes and back, trained with containment, cycle-consistency and volume losses; novelty-aware coreset selection and reliability-weighted graph filtering keep the map current without full retraining.
Load-bearing premise
The load-bearing premise is that authors’ Related Work sections, once extracted by regex and an LLM, already supply reliable and sufficiently complete partial is-a hierarchies, and that a single highly-cited survey per domain is an unbiased final ground-truth against which Soft F1 can be measured at every intermediate step.
What would settle it
If, on a held-out set of domains, the same pipeline run without any Related Work extraction (or with deliberately scrambled partial taxonomies) still matches or exceeds GIST’s Soft F1 under identical token budgets, or if Soft F1 against multi-survey ensemble ground truths collapses relative to single-survey scores, the central claim would be refuted.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces taxonomy maintenance in the wild: continuously adapting topic-centric scholarly taxonomies as arXiv-like repositories evolve. It proposes GIST, which (i) extracts partial is-a hierarchies from Related Work sections via regex+LLM, (ii) integrates them in a box-embedding space with a bidirectional word–box mapping trained under reliability-aware filtering and novelty-aware coreset selection, (iii) forecasts emerging concepts from geometric voids under specialization/abstraction/expansion, and (iv) allocates a token budget via a KL-regularized utility planner. On a 12-topic arXiv benchmark with survey-derived ground truth, GIST I,H reports Final Node/Edge Soft F1 of 73.5/65.4 versus 66.2/57.8 for the strongest baseline (TaxoAdapt S,C), at roughly 9.6% runtime and 12.7% monetary cost, with ablations, hyper-parameter sweeps, extraction diagnostics, and a Text-to-SQL case study.
Significance. If the empirical gains hold under stronger evaluation, this is a solid systems contribution for the data-management community: it reframes scholarly taxonomy work from one-shot construction to streaming maintenance under token budgets, and couples expert-curated partial structures with geometric inductive bias rather than unconstrained LLM hierarchy generation. Strengths that should be credited include closed-form planner and coreset results with appendix proofs, a full modular baseline grid (backbone × update × search), five-seed significance tests, extraction diagnostics (95% regex coverage; ~94% extractor Soft F1 on 40 papers), volume-regularization stability checks, and an incremental semantic index with downstream nDCG/MRR gains. Efficiency under realistic API budgets is practically important for digital libraries and conference organization workflows.
major comments (3)
- Sec. 7.1 and Table 2: the headline 11.0%/13.1% Final Soft F1 gains rest on a single highly-cited survey taxonomy per domain as the sole gold standard. Soft bipartite matching can credit semantically near but structurally alternative hierarchies that the chosen survey never records. The manuscript acknowledges survey bias and the difficulty of multi-survey aggregation, yet provides no sensitivity analysis (second survey, union/intersection of two surveys, or expert adjudication of disagreements). Without that check, the magnitude of the cost–quality claim is not fully stress-tested.
- Sec. 7.1 (“Unless stated otherwise, at each iteration we evaluate against the final ground truth (t=7)”): Init/Δ/ep metrics score intermediate taxonomies against a future consensus that may not yet exist in the literature stream. This can inflate early Soft F1 and per-epoch growth for methods that anticipate later survey structure. Report at least one protocol that evaluates T_t only against concepts/edges attested by papers available up to t (or against a frozen intermediate survey snapshot), and show whether the ranking of GIST vs. TaxoAdapt is stable under that protocol.
- Table 2 comparison framing: the abstract’s primary quality–cost claim pairs GIST I,H with TaxoAdapt S,C (different update regimes). Same-cell comparisons are weaker or mixed (e.g., GIST S,C Final Node 62.3 vs. TaxoAdapt S,C 66.2; GIST I,C 60.5 vs. TaxoAdapt I,C 56.7). Please state the main claim as a cost–quality Pareto result with explicit same-setting and cross-setting numbers, so readers do not over-read the 11%/13% figure as a pure backbone win independent of the hypothesized-concept search strategy.
minor comments (6)
- Throughout (e.g., Fig. 1, Example 1, Alg. 1 context): “Phrase” vs. “Phase” is inconsistently used (Indexing/Retrieval/Generation Phrase). Standardize to “Phase”.
- Front matter still has ACM placeholder metadata (Conference acronym ’XX, Woodstock NY, 2018). Clean for the journal version.
- Sec. 3.2.2: L_vol is written with Vol(B_v) but the threshold is described as a log-volume constraint in prose; align the formula and text.
- Sec. 5.1.2 / Prop. 3: briefly state the default RBF bandwidth and how m=1000 samples were validated for the Gaussian-proxy utility error, so the planner is fully reproducible.
- Table 1 topic “LLM-based Text-to-SQL” is listed with venue year 2026 in Appendix Table 9; confirm citation year consistency.
- Fig. 5 axis labels would benefit from marking the chosen operating points (λ_rel=0.3, λ_nov=0.2, λ_plan=0.2, d=12) explicitly on the plots.
Circularity Check
No circular derivation: GIST is an empirical systems framework whose Soft-F1 gains are measured against external survey taxonomies, not forced by construction from its own losses or self-citations.
full rationale
The paper presents an algorithmic pipeline (Related-Work extraction → box-embedding integration with self-supervised bidirectional mapping → novelty-aware coreset updates → geometric hypothesized-concept generation → budget-aware retrieval) whose quality claims rest entirely on Node/Edge Soft F1 against independently curated survey taxonomies (Table 1, Sec. 7.1, Appendix C). The geometric losses (L_cont, L_cycle, L_vol), MoM Beta estimators on pseudo-observations, and KL-regularized planner (Eqs. 1–9, Props. 1–3) are optimized on extracted signals but do not algebraically equal or tautologically produce the reported Soft-F1 numbers; those numbers are computed by semantic bipartite matching to held-out author hierarchies. No uniqueness theorem, ansatz, or fitted parameter is imported via self-citation and then re-labeled a prediction. The closed loop in which the current taxonomy generates hypothesized queries that later refine it is ordinary online learning, not circular reasoning: intermediate taxonomies are still scored against an external final ground truth. Evaluation-design caveats (single-survey GT, intermediate scoring against t=7) affect correctness risk, not circularity of the derivation chain. Hence score 0 with empty steps.
Axiom & Free-Parameter Ledger
free parameters (6)
- box dimension d =
12
- lambda_rel =
0.3
- lambda_nov =
0.2
- lambda_plan =
0.2
- containment threshold tau =
0.9
- token budget schedule =
1e6 total
axioms (4)
- domain assumption Box containment is a faithful inductive bias for hypernym–hyponym (is-a) relations.
- domain assumption Related Work (or equivalent) sections contain author-curated partial hierarchies that an LLM can extract with high fidelity.
- standard math The coverage-plus-reliability objective is monotone submodular, admitting a (1-1/e) greedy guarantee.
- ad hoc to paper A single high-impact survey per domain supplies an adequate final ground-truth taxonomy for Soft F1 evaluation at every timestamp.
invented entities (3)
-
Bidirectional word–box mapping model (Phi_W2B / Phi_B2W)
no independent evidence
-
Geometric concept emergence model (void-volume priors + Beta unit variables)
no independent evidence
-
Novelty-aware coreset selection for dual mapping updates
no independent evidence
read the original abstract
The rapid growth of scientific publications makes scholarly taxonomies quickly obsolete. We study taxonomy maintenance in the wild, a new problem that moves beyond static construction by continuously adapting taxonomies to evolving scholarly repositories, such as arXiv, for a given research topic. We propose GIST, a robust framework for maintaining evolving taxonomies. Unlike purely LLM-centric approaches, GIST grounds structure induction in expert-curated evidence by extracting partial hierarchies from the "Related Work" sections of papers. It integrates these partial taxonomies into a unified global taxonomy in a geometric box-embedding space, where box containment encodes the inductive bias of is-a relations. To connect semantics with geometric structure, GIST learns a bidirectional mapping between word embeddings and box embeddings. For efficient incremental updates, GIST uses novelty-aware coreset selection to update the model with representative historical signals and new evidence, avoiding costly full retraining. To handle high-velocity paper streams under user-specific token budgets, GIST further combines a hypothesized concept generator with a cost-effective evidence retrieval module. Experiments on real-world arXiv datasets show that GIST outperforms state-of-the-art baselines, improving Node F1 and Edge F1 by 11.0% and 13.1% over the strongest baseline while requiring only 9.6% of its runtime and 12.7% of its monetary cost.
Figures
Reference graph
Works this paper leans on
-
[1]
ACM Digital Library
[n.d.]. ACM Digital Library. https://dl.acm.org. Accessed: 2025-11-25
2025
-
[2]
IEEE Xplore Digital Library
[n.d.]. IEEE Xplore Digital Library. https://ieeexplore.ieee.org. Accessed: 2025-11-25
2025
-
[3]
The 2012 ACM Computing Classification System
2012. The 2012 ACM Computing Classification System. https://www.acm.org/ publications/class-2012. Accessed 2025-11-12
2012
-
[4]
Ralph Abboud, Ismail Ilkan Ceylan, Thomas Lukasiewicz, and Tommaso Salvatori
-
[5]
In Proceedings of the 34th Conference on Neural Information Processing Systems (NeurIPS)
BoxE: A Box Embedding Model for Knowledge Base Completion. In Proceedings of the 34th Conference on Neural Information Processing Systems (NeurIPS). 1–13
-
[6]
Rahaf Aljundi, Francesca Babiloni, Mohamed Elhoseiny, Marcus Rohrbach, and Tinne Tuytelaars. 2018. Memory Aware Synapses: Learning What (Not) to Forget. InECCV
2018
-
[7]
Ben Athiwaratkun and Andrew Gordon Wilson. 2018. Hierarchical Density Order Embeddings. InInternational Conference on Learning Representations (ICLR). https://openreview.net/forum?id=HJCXZQbAZ
2018
-
[8]
Sören Auer, Christian Bizer, Georgi Kobilarov, Jens Lehmann, Richard Cyganiak, and Zachary Ives. 2007. DBpedia: A Nucleus for a Web of Open Data. InThe Semantic Web – ISWC 2007 + ASWC 2007 (Lecture Notes in Computer Science), V ol. 4825. Springer, 722–735. https://doi.org/10.1007/978-3-540-76298-0_52
-
[9]
Christopher M. Bishop. 2006.Pattern Recognition and Machine Learning. Springer
2006
-
[10]
Chengliang Chai, Jiabin Liu, Nan Tang, Ju Fan, Dongjing Miao, Jiayi Wang, Yuyu Luo, and Guoliang Li. 2023. Goodcore: Data-effective and data-efficient machine learning through coreset selection over incomplete data.Proceedings of the ACM on Management of Data1, 2 (2023), 1–27
2023
-
[11]
Arslan Chaudhry, Marc’Aurelio Ranzato, Marcus Rohrbach, and Mohamed Elho- seiny. 2019. Efficient Lifelong Learning with A-GEM. InICLR
2019
-
[12]
Dokania, Philip H.S
Arslan Chaudhry, Marcus Rohrbach, Mohamed Elhoseiny, Thalaiyasingam Ajan- than, Puneet K. Dokania, Philip H.S. Torr, and Marc’Aurelio Ranzato. 2019. On Tiny Episodic Memories in Continual Learning. InICML Workshop
2019
-
[13]
Lydia B Chilton, Juho Kim, Paul André, Felicia Cordeiro, James A Landay, Daniel S Weld, Steven P Dow, Robert C Miller, and Haoqi Zhang. 2014. Frenzy: collaborative data organization for creating conference sessions. InProceedings of the SIGCHI Conference on Human Factors in Computing Systems. 1255–1264
2014
-
[14]
2012.Data Matching: Concepts and Techniques for Record Linkage, Entity Resolution, and Duplicate Detection
Peter Christen. 2012.Data Matching: Concepts and Techniques for Record Linkage, Entity Resolution, and Duplicate Detection. Springer. https://doi.org/10. 1007/978-3-642-31164-2
2012
-
[15]
Shib Sankar Dasgupta, Michael Boratko, Dongxu Zhang, Luke Vilnis, Xiang Lor- raine Li, and Andrew McCallum. 2020. Improving Local Identifiability in Proba- bilistic Box Embeddings. InAdvances in Neural Information Processing Systems (NeurIPS), V ol. 33. 182–192
2020
-
[16]
Matthias De Lange, Rahaf Aljundi, Marc Masana, Sarah Parisot, Xu Jia, Aleš Leonardis, Greg Slabaugh, and Tinne Tuytelaars. 2022. A Continual Learning Survey: Defying Forgetting in Classification Tasks.IEEE TPAMI44, 7 (2022)
2022
-
[17]
Yuhao Deng, Chengliang Chai, Kaisen Jin, Linan Zheng, Lei Cao, Ye Yuan, and Guoren Wang. 2025. Two birds with one stone: Efficient deep learning over mislabeled data through subset selection.Proceedings of the ACM on Management of Data3, 3 (2025), 1–28
2025
-
[18]
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. Bert: Pre-training of deep bidirectional transformers for language understanding. InPro- ceedings of the 2019 conference of the North American chapter of the association for computational linguistics: human language technologies, volume 1 (long and short papers). 4171–4186
2019
-
[19]
Bolin Ding, Haixun Wang, Ruoming Jin, Jiawei Han, and Zhongyuan Wang. 2012. Optimizing Index for Taxonomy Keyword Search. InProceedings of the ACM SIGMOD International Conference on Management of Data. ACM, 493–504. https://doi.org/10.1145/2213836.2213892
-
[20]
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al. 2024. The Llama 3 Herd of Models.arXiv preprint arXiv:2407.21783 (2024). https://arxiv.org/abs/2407.21783
Pith/arXiv arXiv 2024
-
[21]
Darren Edge, Ha Trinh, Newman Cheng, Joshua Bradley, Alex Chao, Apurva Mody, Steven Truitt, Dasha Metropolitansky, Robert Osazuwa Ness, and Jonathan Larson. 2024. From local to global: A graph rag approach to query-focused summarization.arXiv preprint arXiv:2404.16130(2024)
Pith/arXiv arXiv 2024
-
[22]
Marcus Fontoura, Vanja Josifovski, Ravi Kumar, Christopher Olston, Andrew Tomkins, and Sergei Vassilvitskii. 2008. Relaxation in Text Search using Taxonomies.Proceedings of the VLDB Endowment1, 1 (2008), 672–683. https://doi.org/10.14778/1453856.1453930 T axonomy Maintenance In The Wild Conference acronym ’XX, June 03–05, 2018, Woodstock, NY
-
[23]
Sainyam Galhotra, Donatella Firmani, Barna Saha, and Divesh Srivastava. 2018. Robust entity resolution using random graphs. InProceedings of the 2018 Interna- tional Conference on Management of Data. 3–18
2018
-
[24]
Octavian-Eugen Ganea, Gary Bécigneul, and Thomas Hofmann. 2018. Hyperbolic Entailment Cones for Learning Hierarchical Embeddings. InProceedings of the 35th International Conference on Machine Learning (ICML). PMLR, 1646–1655. https://proceedings.mlr.press/v80/ganea18a.html
2018
-
[25]
Miller, and Michael Kohlhase
Deyan Ginev, Bruce R. Miller, and Michael Kohlhase. 2025. ar5iv: HTML5- Converted arXiv Articles (Dataset). https://sigmathling.kwarc.info/resources/ ar5iv-dataset-2024/. Sources up to October 2025; HTML5 conversion of arXiv articles using LaTeXML. Not a live preview service
2025
-
[26]
Maarten Grootendorst. 2022. BERTopic: Neural topic modeling with a class-based TF-IDF procedure.arXiv preprint arXiv:2203.05794(2022)
Pith/arXiv arXiv 2022
-
[27]
Jiacheng Huang, Zequn Sun, Qijin Chen, Xiaozhou Xu, Weijun Ren, and Wei Hu. 2023. Deep active alignment of knowledge graph entities and schemata. Proceedings of the ACM on Management of Data1, 2 (2023), 1–26
2023
-
[28]
Minhao Jiang, Xiangchen Song, Jieyu Zhang, and Jiawei Han. 2022. TaxoEnrich: Self-Supervised Taxonomy Completion via Structure-Semantic Representations. InProceedings of The Web Conference (WWW). Lyon, France. https://hanj.cs. illinois.edu/pdf/www22_mjiang.pdf
2022
-
[29]
Bhargav Kanagal, Amr Ahmed, Sandeep Pandey, Vanja Josifovski, Jeffrey Yuan, and Lluis Garcia-Pueyo. 2012. Supercharging Recommender Systems using Tax- onomies for Learning User Purchase Behavior.Proceedings of the VLDB Endow- ment5, 10 (2012), 956–967. https://vldb.org/pvldb/vol5/p956_bhargavkanagal_ vldb2012.pdf
2012
-
[30]
Priyanka Kargupta, Nan Zhang, Yunyi Zhang, Rui Zhang, Prasenjit Mitra, and Jiawei Han. 2025. TaxoAdapt: Aligning LLM-Based Multidimensional Taxonomy Construction to Evolving Research Corpora. InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, Bangkok...
2025
-
[31]
Andreas Kipf, Thomas Kipf, Bernhard Radke, Viktor Leis, Peter Boncz, and Alfons Kemper. 2019. Learned Cardinalities: Estimating Correlated Joins with Deep Learning. InCIDR
2019
-
[32]
James Kirkpatrick, Razvan Pascanu, Neil Rabinowitz, et al. 2017. Overcoming Catastrophic Forgetting in Neural Networks. InProceedings of the National Academy of Sciences (PNAS), V ol. 114. 3521–3526
2017
-
[33]
Andreas Krause and Daniel Golovin. 2014. Submodular Function Maximization. InTractability: Practical Approaches to Hard Problems, Lucas Bordeaux, Youssef Hamadi, and Pushmeet Kohli (Eds.). Cambridge University Press
2014
-
[34]
Avishek Lahiri, Yufang Hou, and Debarshi Kumar Sanyal. 2025. TaxoAlign: Scholarly Taxonomy Generation Using Language Models. InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 30191–30211
2025
-
[35]
Alice Lai and Julia Hockenmaier. 2017. Learning to Predict Denotational Prob- abilities for Modeling Entailment. InProceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics (EACL). 721–
2017
-
[36]
https://aclanthology.org/E17-1068/
-
[37]
Xiang Li, Luke Vilnis, Dongxu Zhang, Michael Boratko, and Andrew McCallum
-
[38]
InInternational Conference on Learning Representations (ICLR)
Smoothing the Geometry of Probabilistic Box Embeddings. InInternational Conference on Learning Representations (ICLR). https://openreview.net/forum? id=H1xSNiRcF7
-
[39]
Yuliang Li, Jinfeng Li, Yoshihiko Suhara, AnHai Doan, and Wang-Chiew Tan
-
[40]
Deep entity matching with pre-trained language models.arXiv preprint arXiv:2004.00584(2020)
Pith/arXiv arXiv 2004
-
[41]
David Lopez-Paz and Marc’Aurelio Ranzato. 2017. Gradient Episodic Memory for Continual Learning. InNeurIPS
2017
-
[42]
Yuyin Lu, Hegang Chen, Pengbo Mao, Yanghui Rao, Haoran Xie, Fu Lee Wang, and Qing Li. 2024. Self-supervised Topic Taxonomy Discovery in the Box Embedding Space.Transactions of the Association for Computational Linguistics (TACL)(2024). https://doi.org/10.1162/tacl_a_00712
-
[43]
Mingyu Derek Ma, Muhao Chen, Te-Lin Wu, and Nanyun Peng. 2021. HyperEx- pan: Taxonomy Expansion with Hyperbolic Representation Learning. InFindings of the Association for Computational Linguistics: EMNLP 2021. Association for Computational Linguistics, 4182–4194. https://doi.org/10.18653/v1/2021. findings-emnlp.353
doi:10.18653/v1/2021 2021
-
[44]
Arun Mallya and Svetlana Lazebnik. 2018. PackNet: Adding Multiple Tasks to a Single Network by Iterative Pruning. InCVPR
2018
-
[45]
Yuning Mao, Xiang Ren, Jiaming Shen, Xin Gu, and Jiawei Han. 2018. End-to- End Reinforcement Learning for Automatic Taxonomy Induction.arXiv preprint arXiv:1805.04044(2018). https://arxiv.org/abs/1805.04044
Pith/arXiv arXiv 2018
-
[46]
Ryan Marcus, Parimarjan Negi, Hongzi Mao, Nesime Tatbul, Mohammad Al- izadeh, and Tim Kraska. 2021. Bao: Making Learned Query Optimization Practi- cal. InSIGMOD
2021
-
[47]
Ryan Marcus, Parimarjan Negi, Hongzi Mao, Chi Zhang, Mohammad Alizadeh, Tim Kraska, Olga Papaemmanouil, and Nesime Tatbul. 2019. Neo: A Learned Query Optimizer. InVLDB
2019
-
[48]
Yuto Matsubara et al. 2020. Reproducibility, Replicability, and Insights into Dense Multi-Stage Retrieval. InSIGIR
2020
-
[49]
Sahil Mishra, Ujjwal Sudev, and Tanmoy Chakraborty. 2024. FLAME: Self- Supervised Low-Resource Taxonomy Expansion using Large Language Models. arXiv preprint arXiv:2402.13623(2024). https://arxiv.org/abs/2402.13623
Pith/arXiv arXiv 2024
-
[50]
2018.Foundations of Machine Learning(2 ed.)
Mehryar Mohri, Afshin Rostamizadeh, and Ameet Talwalkar. 2018.Foundations of Machine Learning(2 ed.). The MIT Press, Cambridge, MA, USA
2018
-
[51]
George L. Nemhauser, Laurence A. Wolsey, and Marshall L. Fisher. 1978. An anal- ysis of approximations for maximizing submodular set functions—I.Mathematical Programming14, 1 (1978), 265–294. https://doi.org/10.1007/BF01588971
-
[52]
Hoa Nguyen, Ariel Fuxman, Stelios Paparizos, Juliana Freire, and Rakesh Agrawal
-
[53]
https://doi.org/10.14778/1988776.1988777
Synthesizing Products for Online Catalogs.Proceedings of the VLDB Endowment4, 7 (2011), 409–418. https://doi.org/10.14778/1988776.1988777
-
[54]
Maximilian Nickel and Douwe Kiela. 2017. Poincaré Embeddings for Learning Hierarchical Representations. InAdvances in Neural Information Processing Systems (NeurIPS). 6338–6347
2017
-
[55]
Robert C Nickerson, Upkar Varshney, and Jan Muntermann. 2013. A method for taxonomy development and its application in information systems.European journal of information systems22, 3 (2013), 336–359
2013
-
[56]
Gary W Oehlert. 1992. A note on the delta method.The American Statistician46, 1 (1992), 27–29
1992
-
[57]
Sylvestre-Alvise Rebuffi, Alexander Kolesnikov, Georg Sperl, and Christoph H. Lampert. 2017. iCaRL: Incremental Classifier and Representation Learning. In CVPR
2017
-
[58]
Robert and George Casella
Christian P. Robert and George Casella. 2004.Monte Carlo Statistical Methods(2 ed.). Springer
2004
-
[59]
Andrei A. Rusu, Neil C. Rabinowitz, Guillaume Desjardins, et al. 2016. Progres- sive Neural Networks.arXiv:1606.04671(2016)
Pith/arXiv arXiv 2016
-
[60]
Diptikalyan Saha, Avrilia Floratou, Karthik Sankaranarayanan, Umar Farooq Minhas, Ashish R Mittal, and Fatma Özcan. 2016. ATHENA: an ontology-driven system for natural language querying over relational data stores.Proceedings of the VLDB Endowment9, 12 (2016), 1209–1220
2016
-
[61]
Jaydeep Sen, Chuan Lei, Abdul Quamar, Fatma Özcan, Vasilis Efthymiou, Ayushi Dalmia, Greg Stager, Ashish Mittal, Diptikalyan Saha, and Karthik Sankara- narayanan. 2020. Athena++ natural language querying for complex nested sql queries.Proceedings of the VLDB Endowment13, 12 (2020), 2747–2759
2020
-
[62]
2009.Active Learning Literature Survey
Burr Settles. 2009.Active Learning Literature Survey. Technical Report. U. Wisconsin–Madison
2009
-
[63]
Jiaming Shen, Meng Jiang, Xian Li, Jian Li, and Jiawei Han. 2020. TaxoExpan: Self-supervised Taxonomy Expansion with Position-Enhanced Graph Neural Network. InProceedings of The Web Conference (WWW). https://hanj.cs.illinois. edu/pdf/www20_jshen.pdf
2020
-
[64]
Jiaming Shen, Zeqiu Wu, Dongming Lei, Chao Zhang, Xiang Ren, Michelle T. Vanni, Brian M. Sadler, and Jiawei Han. 2019. HiExpan: Task-Guided Taxonomy Construction by Hierarchical Tree Expansion. InProceedings of the 2019 Confer- ence on Empirical Methods in Natural Language Processing (EMNLP-IJCNLP). https://arxiv.org/abs/1910.08194
Pith/arXiv arXiv 2019
-
[65]
Jaeho Shin, Sen Wu, Feiran Wang, Christopher De Sa, Ce Zhang, and Christo- pher Ré. 2015. Incremental Knowledge Base Construction Using DeepDive. Proc. VLDB Endow.8, 11 (2015), 1310–1321. https://doi.org/10.14778/2809974. 2809991
-
[66]
Suchanek, Gjergji Kasneci, and Gerhard Weikum
Fabian M. Suchanek, Gjergji Kasneci, and Gerhard Weikum. 2008. YAGO: A Large Ontology from Wikipedia and WordNet.Web Semantics: Science, Services and Agents on the World Wide Web6, 3 (2008), 203–217. https://doi.org/10.1016/ j.websem.2008.06.001
2008
-
[67]
Yushi Sun, Hao Xin, Kai Sun, Yifan Ethan Xu, Xiao Yang, Xin Luna Dong, Nan Tang, and Lei Chen. 2024. Are Large Language Models a Good Replacement of Taxonomies?Proceedings of the VLDB Endowment17, 11 (2024), 2919–2932. https://doi.org/10.14778/3681954.3681973
-
[68]
Masahiro Tanaka, Yasuyuki Mori, and Andrzej Bargiela. 2002. Granulation of keywords into sessions for timetabling conferences.Proceedings of soft computing and intelligent systems (SCIS 2002)(2002), 1–5
2002
-
[69]
Jianhong Tu, Ju Fan, Nan Tang, Peng Wang, Guoliang Li, Xiaoyong Du, Xiaofeng Jia, and Song Gao. 2023. Unicorn: A unified multi-tasking model for supporting matching tasks in data integration.Proceedings of the ACM on Management of Data1, 1 (2023), 1–26
2023
-
[70]
Ivan Vendrov, Ryan Kiros, Sanja Fidler, and Raquel Urtasun. 2016. Order- Embeddings of Images and Language. InInternational Conference on Learning Representations (ICLR). https://arxiv.org/abs/1511.06361
Pith/arXiv arXiv 2016
-
[71]
Luke Vilnis, Xiang Li, Shikhar Murty, and Andrew McCallum. 2018. Probabilistic Embedding of Knowledge Graphs with Box Lattice Measures. InProceedings of the 56th Annual Meeting of the Association for Computational Linguistics (ACL). 263–272. https://doi.org/10.18653/v1/P18-1025
-
[72]
Jiayi Wang, Chengliang Chai, Nan Tang, Jiabin Liu, and Guoliang Li. 2022. Coresets over multiple tables for feature-rich and data-efficient machine learning. Proceedings of the VLDB Endowment16, 1 (2022), 64–76. Conference acronym ’XX, June 03–05, 2018, Woodstock, NY Daomin Ji, Hui Luo, Zhifeng Bao, Junhao Gan, and Zi Huang
2022
-
[73]
Lidan Wang, Jimmy Lin, and Donald Metzler. 2011. A Cascade Ranking Model for Efficient Ranked Retrieval. (2011)
2011
-
[74]
Gerhard Weikum. 2021. Knowledge Graphs 2021: A Data Odyssey.Proc. VLDB Endow.14, 12 (2021), 3233–3238. https://doi.org/10.14778/3476311.3476393
-
[75]
D. H. D. West. 1979. Updating Mean and Variance Estimates: An Improved Method.Commun. ACM22, 9 (1979), 532–535
1979
-
[76]
Sen Wu, Luke Hsiao, Xiao Cheng, Braden Hancock, Theodoros Rekatsinas, Philip Levis, and Christopher Ré. 2018. Fonduer: Knowledge Base Construction from Richly Formatted Data. InProceedings of the 2018 International Conference on Management of Data (SIGMOD ’18). ACM, 1301–1316. https://doi.org/10.1145/ 3183713.3183729
arXiv 2018
-
[77]
Wei Xue, Yongliang Shen, Wenqi Ren, Jietian Guo, Shiliang Pu, and Weiming Lu
-
[78]
InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (ACL)
Insert or Attach: Taxonomy Completion via Box Embedding. InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (ACL). https://aclanthology.org/2024.acl-long.212.pdf
2024
-
[79]
Hellerstein, Sanjay Krishnan, and Ion Stoica
Zongheng Yang, Eric Liang, Amog Kamsetty, Chenggang Wu, Yan Duan, Xi Chen, Pieter Abbeel, Joseph M. Hellerstein, Sanjay Krishnan, and Ion Stoica
-
[80]
Deep Unsupervised Cardinality Estimation.VLDB(2019)
2019
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.