REVIEW 2 major objections 6 minor 179 references
A Survey on Open Dataset Search in the LLM Era: Retrospectives and Perspectives
T0 review · 2 major / 6 minor · reviewed 2026-08-05 · deepseek-v4-flash
Pith's one-line read This survey argues that open dataset search has outgrown keyword matching and is now defined by a two-way partnership with large language models.
desk verdict A useful modality-aware survey of dataset search that mostly delivers, but the 'dataset search for LLM' half of its central pitch is asserted rather than shown. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The organizing instrument is a modality taxonomy: tabular, vector, spatial, JSON document, and graph datasets, each with a formal definition (for example, a vector dataset is a set of embedding vectors, a spatial dataset is a set of georeferenced points, a JSON document is a nested key-value tree, and a graph dataset is a vertex-edge pair). The paper maps each modality to its dominant similarity measure—MaxSim for vector sets, Earth mover's distance and Hausdorff distance for spatial data, tree edit distance for JSON, and graph edit distance or maximum common subgraph for graphs—along with indexing and acceleration techniques. Over this taxonomy it overlays the two-way LLM relationship, trea
What would settle it
A systematic sweep of published dataset search work that finds a material body of research on a modality the survey does not index (e.g., time-series or audio datasets), or that shows the LLM-for-search systems the survey cites are not adopted or cited by downstream RAG or data-selection work, would weaken both the completeness of the modality taxonomy and the claim that the LLM-dataset search relationship is genuinely two-way.
Extended reading notes
Core claim
The paper's contribution is a structured synthesis: modern open dataset search is no longer a metadata-matching problem but a content-aware, modality-specific retrieval problem, and the most forward-looking direction is the two-way coupling between LLMs and dataset search. On one side, LLMs automate dataset construction, cleaning, and transformation, and enable natural-language queries, semantic planning, and relevance estimation. On the other side, dataset search supports LLMs by selecting datasets for retrieval-augmented generation and for data selection during pretraining, fine-tuning, and in-context learning. The survey develops this thesis across five data modalities, offering formal de
Load-bearing premise
The survey assumes that the research landscape is accurately represented by its modality-based partition and by the particular selection of papers, so that the taxonomy and the open-problems list correspond to the field's true shape rather than to the authors' reading of it.
Editorial extensions
If this is right
- Example-based and natural-language queries will increasingly replace keyword-only interfaces in dataset search systems.
- LLM-based schema inference, cleaning, and transformation should be treated as part of the dataset search pipeline, not as separate data-engineering tasks.
- Dataset search will become a core stage in LLM development, feeding retrieval-augmented generation and data-selection pipelines with queryable, filterable data.
- Standardized benchmarks are missing for spatial and vector dataset search, and quality control is not yet integrated into search; closing these gaps is a prerequisite for practical progress.
- Privacy-preserving similarity calculation, indexing over encrypted data, and federated search remain open problems that need new index and acceleration designs.
Reading between the lines
- If the mutual-benefit thesis holds, dataset search could become a closed-loop controller for LLM training: a model's judged failures could seed queries for corrective datasets, and the retrieved data could then be evaluated by the model, forming a feedback loop; the paper mentions agentic integration but does not formalize this loop.
- The modality taxonomy suggests that cross-modal dataset search is the natural stress test: a shared embedding space for tables, JSON trees, graphs, and spatial points would unify the separate similarity measures, and a cross-modal join and union benchmark would directly test the value of such a unified representation.
- The paper's emphasis on content over metadata implies that search quality may be robust to missing or poor metadata; a direct experiment comparing ranking quality on datasets with full metadata versus deliberately stripped metadata would quantify how much content signals compensate.
- The recurring use of sketch and hash approximations across tabular, vector, and spatial search suggests a common abstraction—set containment at scale—that could be factored into a shared data-discovery index, potentially simplifying future system design.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This survey reviews recent research on open dataset search, moving beyond keyword- and metadata-based retrieval. It organizes the field by data modality (tabular, spatial, JSON, graph, and vector datasets), surveys query-by-example and natural-language search techniques, and discusses the interplay between LLMs and dataset search. The claimed central contribution is a two-way relationship: LLMs improve dataset search through query understanding, semantic planning, and interactive guidance, while dataset search supports LLMs through retrieval-augmented generation (RAG) and data selection. The survey also identifies open problems, including privacy-preserving search, task-oriented integration, cross-modal discovery, federated search, and benchmarks.
Significance. If the reference selection is representative and the reverse-direction claims are adequately evidenced, this survey would fill a useful gap by providing a modality-aware, LLM-centric overview of open dataset search. The paper's strengths are its broad taxonomy, formal definitions, reproduction of key equations (e.g., MaxSim, Hausdorff distance, JSON tree edit distance, GED/MCS), comparative tables that summarize representative methods, and an explicit roadmap of open challenges. It also makes the useful expository choice of distinguishing LLM-for-search from search-for-LLM. However, as a survey, its conclusions rest on the accuracy of its secondary summaries and on the completeness and neutrality of its reference selection, and neither can be fully verified from the manuscript. The most load-bearing weakness is the reverse-direction chapter, whose RAG examples do not actually demonstrate dataset-search-backed RAG.
major comments (2)
- [§5.3, Table 7] The claimed direction 'advances in dataset search can support LLMs by enabling more effective integration into RAG frameworks' is not supported by the evidence in this subsection. After noting that RAG's retrieval component 'naturally aligns' with dataset search, the text explicitly states that the subsection 'instead showcases recent applications of RAG across diverse domains.' Table 7 lists 15 RAG applications (e.g., Self-BioRAG, GraphQA, InstructRAG), but these retrieve scientific papers, knowledge graphs, or domain-specific text corpora—not open datasets as defined in Definition 1.1, and not through any dataset-search system surveyed in Sections 3–4. None of the cited examples shows a dataset-search engine (e.g., Auctus, DBF, Starmie, LOTUS) supplying the retrieval source of a RAG pipeline. The abstract and Section 1.3 stake a core contribution on this mutual-benefit claim, so the re
- [§5.3, Data Selection for LLM] The data-selection discussion similarly overstates the integration of dataset search into LLM data pipelines. The paragraph cites generic data-selection methods (LESS, Dolma, MateS) and then identifies only one method—Wang et al. [138]—that actually integrates query-driven dataset search into a selection pipeline. The claim that 'recent advances have begun to integrate dataset search into data selection pipelines' is supported by a single concrete example. Given that Section 6.2 itself mentions other systems (DeepResearchGym, STARK, GPT-Instructor) that embed dataset retrieval in end-to-end pipelines, the authors could strengthen this subsection substantially by moving those examples here or by adding additional published cases. Without that, the reverse-direction evidence remains one example, which is not enough to sustain the survey's broader 'mutually beneficial relationship' framing.
minor comments (6)
- [Table 1] The 'Vector' row lists '[99]' (BioVSS) as a data source, but [99] is a search method, not an open data repository. The underlying Microsoft Academic Graph [127] appears to be the actual source. Please correct the citation or rephrase the row.
- [Definition 1.2] Definition 1.2 covers keyword and exemplar queries but not natural-language queries, which become a major theme in Section 5.2. Consider extending the definition or adding a remark that NL queries are handled as a separate query type later in the survey.
- [§4.1] The vector-dataset-search section includes several systems originally designed for passage/document retrieval (COLBERT, PLAID, SLIM, COIL, CITADEL, XTR). Since the survey's title and abstract are about dataset search, a brief justification of why token-vector collections over documents are treated as vector datasets would help readers accept this categorization.
- [Figure 2] Figure 2 is dense and is referenced only once. Adding pointers from the pipeline stages to the corresponding sections (e.g., query mechanisms, similarity calculation, indexing) would improve navigability.
- [§4.2] The discussion of spatial dataset search states that the field lacks a standardized benchmark, but no reference is given for this claim. A citation or a brief explanation of which evaluation gaps exist would strengthen the point.
- [§5.3] In text immediately before Table 7, the phrase 'retrieval-augmented generation' and the surrounding discussion could more precisely distinguish between retrieving documents/corpora and retrieving datasets. As written, the section's scope drifts from dataset search to general RAG.
Circularity Check
No significant circularity: the survey's claims are descriptive taxonomies and literature syntheses, not derivations or fitted predictions, and its self-citations are not load-bearing.
full rationale
This is a survey paper, so the standard circularity failure modes (fitting a parameter and renaming it a prediction, defining X in terms of Y, importing a uniqueness theorem from the authors' prior work) do not apply. The paper's central contribution is a modality-aware organization of existing dataset-search techniques and a discussion of the LLM/dataset-search relationship. Definitions 1.1/1.2, 3.1, 4.1–4.4 are formal definitions, not derived results. Sections 3 and 4 survey external, peer-reviewed methods (e.g., LSH Ensemble, JOSIE, Starmie, COLBERT, DBF, JEDI, A*GED) and attribute each method to its original publication; no equation in the survey is shown to reduce to an input by construction. The authors do cite their own prior works (e.g., [97,99,154–156]), but these are included as surveyed contributions alongside many third-party works, and no load-bearing conclusion rests solely on a self-citation. The 'mutually beneficial relationship' claim in §5 is expository; the skeptical concern that §5.3's RAG examples retrieve documents/corpora rather than open datasets is an internal evidence gap about the support for one direction of that relationship, not circular reasoning, since the claim is not derived from itself. The paper even concedes in §5.3 that it 'instead showcases recent applications of RAG across diverse domains,' which is a scope limitation, not a circular step. Overall, the survey's organization and gap analysis are self-contained descriptive claims, so the circularity score is 0.
Assumptions & free parameters
assumptions (3)
- domain assumption The modality-based taxonomy (tabular, spatial, JSON, graph, vector) is a valid way to organize the dataset search literature.
- domain assumption The cited descriptions of prior systems and methods are accurate and representative.
- standard math Standard mathematical definitions used in the survey (e.g., Hausdorff distance Eq. 4, EMD Eq. 5, GED Def. 4.5, MCS Def. 4.6) are correct and apply as stated.
Cite this review
Pith. "Pith review of A Survey on Open Dataset Search in the LLM Era: Retrospectives and Perspectives." pith.science (2026). https://pith.science/paper/K4AFYL3X
@misc{pith2026250900728,
author = {Pith},
title = {Pith review of: A Survey on Open Dataset Search in the LLM Era: Retrospectives and Perspectives},
year = {2026},
howpublished = {\url{https://pith.science/paper/K4AFYL3X}},
note = {Machine review of arXiv:2509.00728}
}
read the original abstract
High-quality datasets are typically required for accomplishing data-driven tasks, such as training medical diagnosis models, predicting real-time traffic conditions, or conducting experiments to validate research hypotheses. Consequently, open dataset search, which aims to ensure the efficient and accurate fulfillment of users' dataset requirements, has emerged as a critical research challenge and has attracted widespread interest. Recent studies have made notable progress in enhancing the flexibility and intelligence of open dataset search, and large language models (LLMs) have demonstrated strong potential in addressing long-standing challenges in this area. Therefore, a systematic and comprehensive review of the open dataset search problem is essential, detailing the current state of research and exploring future directions. In this survey, we focus on recent advances in open dataset search beyond traditional approaches that rely on metadata and keywords. From the perspective of dataset modalities, we place particular emphasis on example-based dataset search, advanced similarity measurement techniques based on dataset content, and efficient search acceleration techniques. In addition, we emphasize the mutually beneficial relationship between LLMs and open dataset search. On the one hand, LLMs help address complex challenges in query understanding, semantic modeling, and interactive guidance within open dataset search. In turn, advances in dataset search can support LLMs by enabling more effective integration into retrieval-augmented generation (RAG) frameworks and data selection processes, thereby enhancing downstream task performance. Finally, we summarize open research problems and outline promising directions for future work. This work aims to offer a structured reference for researchers and practitioners in the field of open dataset search.
Figures
Figures from the paper (6 more)
Reference graph
Works this paper leans on
-
[138]
Tingting Wang, Shixun Huang, Zhifeng Bao, J Shane Culpepper, Volkan Dedeoglu, and Reza Arablouei. 2025. Distinctiveness Maximization in Datasets Assemblage. In The World Wide Web Conference (WWW) . 3219–3232
work page 2025
-
[97]
Pengyue Li, Hua Dai, Sheng Wang, Wenzhe Yang, and Geng Yang. 2024. Privacy-preserving Spatial Dataset Search in Cloud. In Proceedings of the 33rd ACM International Conference on Information and Knowledge Management (CIKM) . 1245–1254
2024
-
[154]
Wenzhe Yang, Sheng Wang, Zhiyu Chen, Yuan Sun, and Zhiyong Peng. 2025. Joinable Search over Multi-source Spatial Datasets: Overlap, Coverage, and Efficiency. In 41st IEEE International Conference on Data Engineering (ICDE) . 585–598
work page 2025
- [155]
-
[156]
Wenzhe Yang, Sheng Wang, Yuan Sun, and Zhiyong Peng. 2022. Fast dataset search with earth mover’s distance. Proceedings of the VLDB Endowment 15, 11 (2022), 2517–2529
work page 2022
-
[99]
Yiqi Li, Sheng Wang, Zhiyu Chen, Shangfeng Chen, and Zhiyong Peng. 2025. Approximate Vector Set Search: A Bio-Inspired Approach for High-Dimensional Spaces. In 2025 IEEE 41st International Conference on Data Engineering (ICDE) . 891–903
2025
-
[168]
Zhengkai Zhang, Zhou Hao Dai, Hua, Mingfeng Jiang, Pengyue Li, and Geng Yang. 2025. A Privacy-preserving Spatial Dataset Joinable Search in Cloud. In Proceedings of the 34th ACM International Conference on Information and Knowledge Management (CIKM)
work page 2025
-
[1]
Abbas Acar, Hidayet Aksu, A Selcuk Uluagac, and Mauro Conti. 2018. A survey on homomorphic encryption schemes: Theory and implementation. Comput. Surveys 51, 4 (2018), 1–35
2018
Show all 179 references
-
[2]
Uchenna Akujuobi and Xiangliang Zhang. 2017. Delve: a dataset-driven scholarly search and analysis system.ACM SIGKDD Explorations Newsletter 19, 2 (2017), 36–46
2017
-
[3]
Alawwad, Areej Alhothali, Usman Naseem, Ali Alkhathlan, and Amani Jamal
Hessa A. Alawwad, Areej Alhothali, Usman Naseem, Ali Alkhathlan, and Amani Jamal. 2025. Enhancing textual textbook question answering with large language models and retrieval augmented generation. Pattern Recognition 162 (2025), 111332
2025
-
[4]
Alon Albalak, Yanai Elazar, Sang Michael Xie, Shayne Longpre, Nathan Lambert, Xinyi Wang, Niklas Muennighoff, Bairu Hou, Liangming Pan, Haewon Jeong, et al. 2024. A survey on data selection for language models. arXiv:2402.16827 (2024)
2024 arXiv
-
[5]
Eric Anderson, Jonathan Fritz, Austin Lee, Bohou Li, Mark Lindblad, Henry Lindeman, Alex Meyer, Parth Parmar, Tanvi Ranade, Mehul A Shah, et al. 2024. The Design of an LLM-powered Unstructured Analytics System. arXiv:2409.00847 (2024)
2024 arXiv
-
[6]
Simran Arora, Brandon Yang, Sabri Eyuboglu, Avanika Narayan, Andrew Hojel, Immanuel Trummer, and Christopher Ré. 2023. Language Models Enable Simple Systems for Generating Structured Views of Heterogeneous Data Lakes. Proceedings of the VLDB Endowment 17, 2 (2023), 92–105
2023
-
[7]
Mohamed-Amine Baazizi, Dario Colazzo, Giorgio Ghelli, and Carlo Sartiani. 2019. Parametric schema inference for massive JSON datasets. The VLDB Journal 28, 4 (2019), 497–521
2019
-
[8]
Mohamed-Amine Baazizi, Dario Colazzo, Giorgio Ghelli, and Carlo Sartiani. 2019. Schemas and types for JSON data: from theory to practice. In Proceedings of the 2019 International Conference on Management of Data (SIGMOD) . 2060–2063
2019
-
[9]
Jiyang Bai and Peixiang Zhao. 2021. TaGSim: Type-aware Graph Similarity Learning and Computation. Proceedings of the VLDB Endowment 15, 2 (2021), 335–347
2021
-
[10]
Tianyi Bai, Ling Yang, Zhen Hao Wong, Fupeng Sun, Xinlin Zhuang, Jiahui Peng, Chi Zhang, Lijun Wu, Jiantao Qiu, Wentao Zhang, Binhang Yuan, and Conghui He. 2025. Efficient Pretraining Data Selection for Language Models via Multi-Actor Collaboration. In Proceedings of the 63rd ...
2025
-
[11]
Omar Benjelloun, Shiyu Chen, and Natasha Noy. 2020. Google Dataset Search by the Numbers. In International Semantic Web Conference (ISWS) . 667–682
2020
-
[12]
Chandra Sekhar Bhagavatula, Thanapon Noraset, and Doug Downey. 2015. Tabel: Entity linking in web tables. In International Semantic Web Conference (ISWS). 425–441
2015
-
[13]
Fabian Biester, Mohamed Abdelaal, and Daniel Del Gaudio. 2024. Llmclean: Context-aware tabular data cleaning via llm-generated ofds. In European Conference on Advances in Databases and Information Systems (ADBIS) , Vol. 2186. 68–78
2024
-
[14]
Tobias Bleifuß, Leon Bornemann, Dmitri V Kalashnikov, Felix Naumann, and Divesh Srivastava. 2021. The Secret Life of Wikipedia Tables. In Proceedings of the 2nd Workshop on Search, Exploration, and Analysis in Heterogeneous Datastores (SEA-Data 2021) co-located with 47th Inter...
2021
-
[15]
Alex Bogatu, Alvaro A A Fernandes, Norman W Paton, and Nikolaos Konstantinou. 2020. Dataset Discovery in Data Lakes. In36th IEEE International Conference on Data Engineering (ICDE) . 709–720
2020
-
[16]
Pierre Bourhis, Juan L Reutter, Fernando Suárez, and Domagoj Vrgoč. 2017. JSON: data model, query languages and schema specification. In Proceedings of the 36th ACM SIGMOD-SIGACT-SIGAI Symposium on Principles of Database Systems (PODS) . 123–135
2017
-
[17]
Pierre Bourhis, Juan L Reutter, and Domagoj Vrgoč. 2020. JSON: Data model and query languages. Information Systems 89 (2020), 101478
2020
-
[18]
Dan Brickley, Matthew Burgess, and Natasha Noy. 2019. Google Dataset Search: Building a search engine for datasets in an open Web ecosystem. In The World Wide Web Conference (WWW) . 1365–1375
2019
-
[19]
Horst Bunke and Kim Shearer. 1998. A graph distance metric based on the maximal common subgraph. Pattern recognition letters 19, 3-4 (1998), 255–259
1998
-
[20]
Sonia Castelo, Rémi Rampin, Aécio Santos, Aline Bessa, Fernando Chirigati, and Juliana Freire. 2021. Auctus: a dataset search engine for data discovery and augmentation. Proceedings of the VLDB Endowment 14, 12 (2021), 2791–2794
2021
-
[21]
Lijun Chang, Xing Feng, Xuemin Lin, Lu Qin, Wenjie Zhang, and Dian Ouyang. 2020. Speeding up GED verification for graph similarity search. In 36th IEEE International Conference on Data Engineering (ICDE) . 793–804
2020
-
[22]
Lijun Chang, Xing Feng, Kai Yao, Lu Qin, and Wenjie Zhang. 2022. Accelerating graph similarity search via efficient GED computation. IEEE Transactions on Knowledge and Data Engineering 35, 5 (2022), 4485–4498
2022
-
[23]
Ming-Fang Chang, John Lambert, Patsorn Sangkloy, Jagjeet Singh, Slawomir Bak, Andrew Hartnett, De Wang, Peter Carr, Simon Lucey, Deva Ramanan, et al. 2019. Argoverse: 3d tracking and forecasting with rich maps. In IEEE Conference on Computer Vision and Pattern Recognition (CVP...
2019
-
[24]
Adriane Chapman, Elena Simperl, Laura Koesten, George Konstantinidis, Luis-Daniel Ibáñez, Emilia Kacprzak, and Paul Groth. 2020. Dataset search: a survey. The VLDB Journal 29, 1 (2020), 251–272
2020
-
[25]
Minxiao Chen, Haitao Yuan, Nan Jiang, Zhifeng Bao, and Shangguang Wang. 2024. Urban Traffic Accident Risk Prediction Revisited: Regionality, Proximity, Similarity and Sparsity. InProceedings of the 33rd ACM International Conference on Information and Knowledge Management (CIKM...
2024
-
[26]
Minxiao Chen, Haitao Yuan, Nan Jiang, Zhihan Zheng, Zhifeng Bao, Ao Zhou, Jiaxin Jiang, and Shangguang Wang. 2025. S-MGHSTN: Towards An Effective Streaming Traffic Accident Risk Prediction Framework. IEEE Transactions on Knowledge and Data Engineering 37, 7 (2025), 4285 – 4298
2025
-
[27]
Qiaosheng Chen, Jiageng Chen, Xiao Zhou, and Gong Cheng. 2024. Enhancing Dataset Search with Compact Data Snippets. In Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR) . 1093–1103
2024
-
[28]
Qiaosheng Chen, Weiqing Luo, Zixian Huang, Tengteng Lin, Xiaxia Wang, Ahmet Soylu, Basil Ell, Baifan Zhou, Evgeny Kharlamov, and Gong Cheng. 2024. ACORDAR 2.0: A Test Collection for Ad Hoc Dataset Retrieval with Densely Pooled Datasets and Question-Style Queries. In Proceeding...
2024
-
[29]
Wei-Hao Chen, Weixi Tong, Amanda Case, and Tianyi Zhang. 2025. Dango: A Mixed-Initiative Data Wrangling System using Large Language Model. In Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems (CHI) . 1–28
2025
-
[30]
Zui Chen, Zihui Gu, Lei Cao, Ju Fan, Samuel Madden, and Nan Tang. 2023. Symphony: Towards Natural Language Query Answering over Multi-modal Data Lakes. In 13th Conference on Innovative Data Systems Research (CIDR) . 1–7
2023
-
[31]
Zhiyu Chen, Mohamed Trabelsi, Jeff Heflin, Yinan Xu, and Brian D. Davison. 2020. Table Search Using a Deep Contextualized Language Model. In Proceedings of the 43rd International ACM SIGIR conference on research and development in Information Retrieval (SIGIR) , Jimmy X. Huang...
2020
-
[32]
Mingyue Cheng, Yucong Luo, Jie Ouyang, Qi Liu, Huijie Liu, Li Li, Shuo Yu, Bohou Zhang, Jiawei Cao, Jie Ma, et al . 2025. A survey on knowledge-oriented retrieval-augmented generation. arXiv:2503.10677 (2025)
2025 arXiv
-
[33]
João Coelho, Jingjie Ning, Jingyuan He, Kangrui Mao, Abhijay Paladugu, Pranav Setlur, Jiahe Jin, Jamie Callan, João Magalhães, Bruno Martins, et al. 2025. Deepresearchgym: A free, transparent, and reproducible evaluation sandbox for deep research. arXiv:2505.19253 (2025)
2025
-
[34]
Tianji Cong, Fatemeh Nargesian, and HV Jagadish. 2023. Pylon: Semantic table union search in data lakes. arXiv:2301.04901 (2023)
2023 arXiv
-
[35]
Arash Dargahi Nobari and Davood Rafiei. 2024. Dtt: An example-driven tabular transformer for joinability by leveraging large language models. Proceedings of the ACM on Management of Data 2, 1 (2024), 1–24
2024
-
[36]
Anish Das Sarma, Lujun Fang, Nitin Gupta, Alon Halevy, Hongrae Lee, Fei Wu, Reynold Xin, and Cong Yu. 2012. Finding related tables. In Proceedings of the 2012 International Conference on Management of Data (SIGMOD) . 817–828
2012
-
[37]
Auriol Degbelo and Brhane Bahrishum Teka. 2019. Spatial search strategies for open government data: A systematic comparison. In Proceedings of the 13th Workshop on Geographic Information Retrieval (GIR) . 1–10
2019
-
[38]
Yuhao Deng, Chengliang Chai, Lei Cao, Qin Yuan, Siyuan Chen, Yanrui Yu, Zhaoze Sun, Junyi Wang, Jiajun Li, Ziqi Cao, et al. 2024. LakeBench: A Benchmark for Discovering Joinable and Unionable Tables in Data Lakes. Proceedings of the VLDB Endowment 17, 8 (2024), 1925–1938
2024
-
[39]
Yuyang Dong, Kunihiro Takeoka, Chuan Xiao, and Masafumi Oyamada. 2021. Efficient Joinable Table Discovery in Data Lakes: A High-Dimensional Similarity-Based Approach. In 37th IEEE International Conference on Data Engineering (ICDE) . 456–467
2021
-
[40]
Yuyang Dong, Chuan Xiao, Takuma Nozawa, Masafumi Enomoto, and Masafumi Oyamada. 2023. DeepJoin: Joinable Table Discovery with Pre-Trained Language Models. Proceedings of the VLDB Endowment 16, 10 (2023), 2458–2470
2023
-
[41]
Dominik Durner, Viktor Leis, and Thomas Neumann. 2021. JSON tiles: Fast analytics on semi-structured data. InProceedings of the 2021 International Conference on Management of Data (SIGMOD) . 445–458
2021
-
[42]
Mohamed Y Eltabakh, Mayuresh Kunjir, Ahmed K Elmagarmid, and Mohammad Shahmeer Ahmad. 2023. Cross Modal Data Discovery over Structured and Unstructured Data Lakes. Proceedings of the VLDB Endowment 16, 11 (2023), 3377–3390
2023
-
[43]
Joshua Engels, Benjamin Coleman, Vihan Lakshman, and Anshumali Shrivastava. 2023. DESSERT: an efficient algorithm for vector set search with vector set queries. Advances in Neural Information Processing Systems 36 (2023), 67972–67992
2023
-
[44]
Mahdi Esmailoghli, Christoph Schnell, Renée J Miller, and Ziawasch Abedjan. 2025. BLEND: A Unified Data Discovery System. In 41st IEEE International Conference on Data Engineering (ICDE) . 737–750
2025
-
[45]
Grace Fan, Jin Wang, Yuliang Li, and Renée J Miller. 2023. Table discovery in data lakes: State-of-the-art and future directions. In Companion of the 2023 International Conference on Management of Data (SIGMOD/PODS) . 69–75
2023
-
[46]
Grace Fan, Jin Wang, Yuliang Li, Dan Zhang, and Renée J Miller. 2023. Semantics-Aware Dataset Discovery from Data Lakes with Contextualized Column-Based Representation Learning. Proceedings of the VLDB Endowment 16, 7 (2023), 1726–1739
2023
-
[47]
Ju Fan, Zihui Gu, Songyue Zhang, Yuxin Zhang, Zui Chen, Lei Cao, Guoliang Li, Samuel Madden, Xiaoyong Du, and Nan Tang. 2024. Combining small language models and large language models for zero-shot nl2sql. Proceedings of the VLDB Endowment 17, 11 (2024), 2750–2763
2024
-
[48]
Lizhou Fan, Sara Lafia, Lingyao Li, Fangyuan Yang, and Libby Hemphill. 2023. DataChat: Prototyping a conversational agent for dataset search and visualization. Proceedings of the Association for Information Science and Technology 60, 1 (2023), 586–591
2023
-
[49]
Shangbin Feng, Chan Young Park, Yuhan Liu, and Yulia Tsvetkov. 2023. From Pretraining Data to Language Models to Downstream Tasks: Tracking the Trails of Political Biases Leading to Unfair NLP Models. InProceedings of the 61st Annual Meeting of the Association for Computationa...
2023
-
[50]
Pedro Arthur de Fernandes Vasconcelos, Wensttay de Sousa Alencar, Victor Hugo da Silva Ribeiro, Natarajan Ferreira Rodrigues, and Fabio de Gomes Andrade. 2017. Enabling Spatial Queries in Open Government Data Portals. In International Conference on Electronic Government and th...
2017
-
[51]
Raul Castro Fernandez, Ziawasch Abedjan, Famien Koko, Gina Yuan, Samuel Madden, and Michael Stonebraker. 2018. Aurum: A data discovery system. In 34th IEEE International Conference on Data Engineering (ICDE) . 1001–1012
2018
-
[52]
Raul Castro Fernandez, Jisoo Min, Demitri Nava, and Samuel Madden. 2019. Lazo: A cardinality-based method for coupled estimation of jaccard similarity and containment. In 35th IEEE International Conference on Data Engineering (ICDE) . 1190–1201
2019
-
[53]
Benjamin Feuer, Yurong Liu, Chinmay Hegde, and Juliana Freire. 2024. ArcheType: A Novel Framework for Open-Source Column Type Annotation Using Large Language Models. Proceedings of the VLDB Endowment 17, 9 (2024)
2024
-
[54]
Juliana Freire, Grace Fan, Benjamin Feuer, Christos Koutras, Yurong Liu, Eduardo Peña, Aécio S. R. Santos, Cláudio T. Silva, and Eden Wu. 2025. Large Language Models for Data Discovery and Integration: Challenges and Opportunities. IEEE Data Engineering Bulletin 49, 1 (2025), 3–31
2025
-
[55]
Luyu Gao, Zhuyun Dai, and Jamie Callan. 2021. COIL: Revisit Exact Lexical Match in Information Retrieval with Contextualized Inverted List. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics (NAACL) . 3030–3042
2021
-
[56]
Adamu Garba and Shangli Wu. 2025. Snippet-based result merging in federated search. Journal of Information Science 51, 3 (2025), 623–637
2025
-
[57]
Adamu Garba, Shengli Wu, and Shah Khalid. 2023. Federated search techniques: an overview of the trends and state of the art. Knowledge and Information Systems 65, 12 (2023), 5065–5095
2023
-
[58]
Saheli Ghosh, Ahmed Eldawy, and Shipra Jais. 2019. Aid: An adaptive image data index for interactive multilevel visualization. In 35th IEEE International Conference on Data Engineering (ICDE) . 1594–1597
2019
-
[59]
Saheli Ghosh, Tin Vu, Mehrad Amin Eskandari, and Ahmed Eldawy. 2019. UCR-STAR: The UCR spatio-temporal active repository. SIGSPATIAL Special 11, 2 (2019), 34–40
2019
-
[60]
Victor Giannakouris and Immanuel Trummer. 2025. SwellDB: Dynamic Query-Driven Table Generation with Large Language Models. InCompanion of the 2025 International Conference on Management of Data . 95–98
2025
-
[61]
Youdi Gong, Guangzhen Liu, Yunzhi Xue, Rui Li, and Lingzhong Meng. 2023. A survey on dataset quality in machine learning. Information and Software Technology 162 (2023), 107268
2023
-
[62]
Tobias Grubenmann, Abraham Bernstein, Dmitry Moor, and Sven Seuken. 2018. Financing the web of data with delayed-answer auctions. In The World Wide Web Conference (WWW). 1033–1042
2018
-
[63]
Ruiqi Guo, Philip Sun, Erik Lindgren, Quan Geng, David Simcha, Felix Chern, and Sanjiv Kumar. 2020. Accelerating large-scale inference with anisotropic vector quantization. In International Conference on Machine Learning (ICML) . 3887–3896
2020
-
[64]
Stefan Hahmann and Dirk Burghardt. 2013. How much information is geospatially referenced? Networks and cognition. International Journal of Geographical Information Science 27, 6 (2013), 1171–1189
2013
-
[65]
Rihan Hai, Christos Koutras, Christoph Quix, and Matthias Jarke. 2023. Data lakes: A survey of functions and systems. IEEE Transactions on Knowledge and Data Engineering 35, 12 (2023), 12571–12590
2023
-
[66]
Zifei FeiFei Han, Jionghao Lin, Ashish Gurung, Danielle R Thomas, Eason Chen, Conrad Borchers, Shivang Gupta, and Kenneth R Koedinger. 2024. Improving Assessment of Tutoring Practices using Retrieval-Augmented Generation. Proceedings of Machine Learning Research 257 (2024), 66–76
2024
-
[67]
Xiaoxin He, Yijun Tian, Yifei Sun, Nitesh Chawla, Thomas Laurent, Yann LeCun, Xavier Bresson, and Bryan Hooi. 2024. G-retriever: Retrieval- augmented generation for textual graph understanding and question answering. Advances in Neural Information Processing Systems 37 (2024),...
2024
-
[68]
Thomas Hervey, Sara Lafia, and Werner Kuhn. 2021. Search Facets and Ranking in Geospatial Dataset Search. In 11th International Conference on Geographic Information Science (GIScience) , Vol. 177. 5:1–5:15
2021
-
[69]
Sirui Hong, Yizhang Lin, Bang Liu, Bangbang Liu, Binhao Wu, Ceyao Zhang, Danyang Li, Jiaqi Chen, Jiayi Zhang, Jinlin Wang, Li Zhang, Lingyao Zhang, Min Yang, Mingchen Zhuge, Taicheng Guo, Tuo Zhou, Wei Tao, Robert Tang, Xiangtao Lu, Xiawu Zheng, Xinbing Liang, Yaying Fei, Yuhe...
2025
-
[70]
Xuming Hu, Shen Wang, Xiao Qin, Chuan Lei, Zhengyuan Shen, Christos Faloutsos, Asterios Katsifodimos, George Karypis, Lijie Wen, and Philip S Yu. 2023. Automatic table union search with tabular representation learning. In Findings of the Association for Computational Linguisti...
2023
-
[71]
Yuntong Hu, Zhihan Lei, Zhongjie Dai, Allen Zhang, Abhinav Angirekula, Zheng Zhang, and Liang Zhao. 2025. Cg-rag: Research question answering by citation graph retrieval-augmented llms. In Proceedings of the 48th International ACM SIGIR Conference on Research and Development i...
2025
-
[72]
Benhao Huang, Yingzhuo Yu, Jin Huang, Xingjian Zhang, and Jiaqi Ma. 2024. DCA-Bench: A Benchmark for Dataset Curation Agents. arXiv:2406.07275 (2024)
2024 arXiv
-
[73]
Thomas Hütter, Nikolaus Augsten, Christoph M Kirsch, Michael J Carey, and Chen Li. 2022. JEDI: These aren’t the JSON documents you’re looking for.... In Proceedings of the 2022 International Conference on Management of Data (SIGMOD) . 1584–1597
2022
-
[74]
Thomas Hütter, Mateusz Pawlik, Robert Löschinger, and Nikolaus Augsten. 2019. Effective filters and linear time verification for tree similarity joins. In 35th IEEE International Conference on Data Engineering (ICDE) . 854–865
2019
-
[75]
Andra Ionescu, Kiril Vasilev, Florena Buse, Rihan Hai, and Asterios Katsifodimos. 2024. AutoFeat: Transitive Feature Discovery over Join Paths. In 40th IEEE International Conference on Data Engineering (ICDE) . 1861–1873
2024
-
[76]
Minbyul Jeong, Jiwoong Sohn, Mujeen Sung, and Jaewoo Kang. 2024. Improving medical reasoning through retrieval and self-reflection with retrieval-augmented large language models. Bioinformatics 40, Supplement_1 (2024), i119–i129
2024
-
[77]
Shane Culpepper
Daomin Ji, Hui Luo, Zhifeng Bao, and J. Shane Culpepper. 2024. Navigating Data Repositories: Utilizing Line Charts to Discover Relevant Datasets. Proceedings of the VLDB Endowment 17, 12 (2024), 4289–4292
2024
-
[78]
Daomin Ji, Hui Luo, Zhifeng Bao, and J Shane Culpepper. 2025. Table integration in data lakes unleashed: pairwise integrability judgment, integrable set discovery, and multi-tuple conflict resolution. The VLDB Journal 34, 3 (2025), 1–24
2025
-
[79]
Lin Jiang, Junqiao Qiu, and Zhijia Zhao. 2020. Scalable structural index construction for JSON analytics. Proceedings of the VLDB Endowment 14, 4 (2020), 694–707
2020
-
[80]
Xinhui Kang, Wenjie You, and Ying Luo. 2025. Bio-inspired product design system integrating retrieval-augmented question answering and semantic fusion diffusion model. Advanced Engineering Informatics 67 (2025), 103537
2025
-
[81]
Nikolai Karpov and Qin Zhang. 2023. Syncsignature: A simple, efficient, parallelizable framework for tree similarity joins. Proceedings of the VLDB Endowment 16, 2 (2023)
2023
-
[82]
Maxat Kassen. 2013. A promising phenomenon of open data: A case study of the Chicago open data project. Government information quarterly 30, 4 (2013), 508–513
2013
-
[83]
Moe Kayali, Anton Lykov, Ilias Fountalis, Nikolaos Vasiloglou, Dan Olteanu, and Dan Suciu. 2024. Chorus: Foundation Models for Unified Data Discovery and Exploration. Proceedings of the VLDB Endowment 17, 8 (2024), 2104–2114
2024
-
[84]
Aamod Khatiwada, Grace Fan, Roee Shraga, Zixuan Chen, Wolfgang Gatterbauer, Renée J Miller, and Mirek Riedewald. 2023. Santos: Relationship- based semantic table union search. Proceedings of the ACM on Management of Data 1, 1 (2023), 1–25
2023
-
[85]
Aamod Khatiwada, Harsha Kokel, Ibrahim Abdelaziz, Subhajit Chaudhury, Julian Dolby, Oktie Hassanzadeh, Zhenhan Huang, Tejaswini Pedapati, Horst Samulowitz, and Kavitha Srinivas. 2025. TabSketchFM: Sketch-Based Tabular Representation Learning for Data Discovery over Data Lakes....
2025
-
[86]
Omar Khattab and Matei Zaharia. 2020. Colbert: Efficient and effective passage search via contextualized late interaction over bert. In Proceedings of the 43rd International ACM SIGIR conference on research and development in Information Retrieval (SIGIR) . 39–48
2020
-
[87]
Diksha Khurana, Aditya Koli, Kiran Khatter, and Sukhdev Singh. 2023. Natural language processing: state of the art, current trends and challenges. Multimedia tools and applications 82, 3 (2023), 3713–3744
2023
-
[88]
Jongik Kim. 2021. Boosting graph similarity search through pre-computation. In Proceedings of the 2021 International Conference on Management of Data (SIGMOD). 951–963
2021
-
[89]
Gary King. 2007. An introduction to the dataverse network as an infrastructure for data sharing. Sociological Methods & Research 36, 2 (2007), 173–199
2007
-
[90]
Meike Klettke, Uta Störl, and Stefanie Scherzinger. 2015. Schema extraction and structural outlier detection for JSON-based NoSQL data stores. In Datenbanksysteme für Business, Technologie und Web (BTW) . 425–444
2015
-
[91]
Daniel Kocher and Nikolaus Augsten. 2019. A scalable index for top-k subtree similarity queries. In Proceedings of the 2019 International Conference on Management of Data (SIGMOD) . 1624–1641
2019
-
[92]
Jinhyuk Lee, Zhuyun Dai, Sai Meher Karthik Duddu, Tao Lei, Iftekhar Naim, Ming-Wei Chang, and Vincent Zhao. 2023. Rethinking the role of token retrieval in multi-vector retrieval. Advances in Neural Information Processing Systems 36 (2023), 15384–15405
2023
-
[93]
Aristotelis Leventidis, Martin Pekár Christensen, Matteo Lissandrini, Laura Di Rocco, Katja Hose, and Renée J Miller. 2024. A Large Scale Test Corpus for Semantic Table Search. In Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Informa...
2024
-
[94]
Guoliang Li, Xuanhe Zhou, and Xinyang Zhao. 2024. LLM for Data Management. Proceedings of the VLDB Endowment 17, 12 (2024), 4213–4216
2024
-
[95]
Minghan Li, Sheng-Chieh Lin, Xueguang Ma, and Jimmy Lin. 2023. SLIM: Sparsified late interaction for multi-vector retrieval with inverted indexes. In Proceedings of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR) . 1954–1959
2023
-
[96]
Minghan Li, Sheng-Chieh Lin, Barlas Oguz, Asish Ghoshal, Jimmy Lin, Yashar Mehdad, Wen-tau Yih, and Xilun Chen. 2023. CITADEL: Conditional Token Interaction via Dynamic Lexical Routing for Efficient and Effective Multi-Vector Retrieval. InProceedings of the 61st Annual Meeting...
2023
-
[98]
Peng Li, Yeye He, Dror Yashar, Weiwei Cui, Song Ge, Haidong Zhang, Danielle Rifinski Fainman, Dongmei Zhang, and Surajit Chaudhuri. 2024. Table-GPT: Table Fine-tuned GPT for Diverse Table Tasks. Proceedings of the ACM on Management of Data 2, 3 (2024), 176
2024
-
[100]
Zhize Li, Haoyu Zhao, Boyue Li, and Yuejie Chi. 2022. SoteriaFL: A unified framework for private federated learning with communication compression. Advances in Neural Information Processing Systems 35 (2022), 4285–4300
2022
-
[101]
Yiming Lin, Madelon Hulsebos, Ruiying Ma, Shreya Shankar, Sepanta Zeigham, Aditya G Parameswaran, and Eugene Wu. 2024. Towards Accurate and Efficient Document Analytics with Large Language Models. arXiv:2405.04674 (2024)
2024 arXiv
-
[102]
Yehuda Lindell. 2020. Secure multiparty computation. Commun. ACM 64, 1 (2020), 86–96
2020
-
[103]
Jiabin Liu, Chengliang Chai, Yuyu Luo, Yin Lou, Jianhua Feng, and Nan Tang. 2022. Feature Augmentation with Reinforcement Learning. In 38th IEEE International Conference on Data Engineering (ICDE) . 3360–3372
2022
-
[104]
Llano-Ríos, Mohamed Khalefa, and Antonio Badia
Tomás F. Llano-Ríos, Mohamed Khalefa, and Antonio Badia. 2025. A JSON document algebra for query optimization. Information Systems 132 (2025), 102537
2025
-
[105]
Robert WP Luk, Hong Va Leong, Tharam S Dillon, Alvin TS Chan, W Bruce Croft, and James Allan. 2002. A survey in indexing and searching XML documents. Journal of the American society for Information Science and Technology 53, 6 (2002), 415–437
2002
-
[106]
Sean X Luo, Richard Axel, and LF Abbott. 2010. Generating sparse and selective third-order responses in the olfactory system of the fly.Proceedings of the National Academy of Sciences 107, 23 (2010), 10713–10718
2010
-
[107]
Pingchuan Ma, Rui Ding, Shuai Wang, Shi Han, and Dongmei Zhang. 2023. InsightPilot: An LLM-empowered automated data exploration system. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing: System Demonstrations (EMNLP) . 346–352
2023
-
[108]
Mandar Mitra and BB Chaudhuri. 2000. Information retrieval from documents: A survey. Information retrieval 2 (2000), 141–163
2000
-
[109]
Takuya Mizokami, Savong Bou, and Toshiyuki Amagasa. 2024. Subtree Similarity Search Based on Structure and Text. In International Conference on Big Data Analytics and Knowledge Discovery . 72–87
2024
-
[110]
Davide Mottin, Matteo Lissandrini, Yannis Velegrakis, and Themis Palpanas. 2014. Exemplar queries: Give me an example of what you need. Proceedings of the VLDB Endowment 7, 5 (2014), 365–376
2014
-
[111]
Fatemeh Nargesian, Erkang Zhu, Ken Q Pu, and Renée J Miller. 2018. Table union search on open data. Proceedings of the VLDB Endowment 11, 7 (2018), 813–825
2018
-
[112]
Karen Ka Yan Ng, Izuki Matsuba, and Peter Chengming Zhang. 2025. RAG in health care: a novel framework for improving communication and decision-making by addressing LLM limitations. NEJM AI 2, 1 (2025)
2025
-
[113]
Tri Nguyen, Mir Rosenberg, Xia Song, Jianfeng Gao, Saurabh Tiwary, Rangan Majumder, and Li Deng. 2016. MS MARCO: A Human Generated MAchine Reading COmprehension Dataset. In Proceedings of the Workshop on Cognitive Computation: Integrating neural and symbolic approaches 2016 co...
2016
-
[114]
Sarana Nutanong, Edwin H Jacox, and Hanan Samet. 2011. An incremental Hausdorff distance calculation algorithm. Proceedings of the VLDB Endowment 4, 8 (2011), 506–517
2011
-
[115]
Alizée Pace, Jonathan Mallinson, Eric Malmi, Sebastian Krause, and Aliaksei Severyn. 2024. West-of-n: Synthetic preference generation for improved reward modeling. arXiv:2401.12086 (2024)
2024 arXiv
-
[116]
Liana Patel, Siddharth Jha, Carlos Guestrin, and Matei Zaharia. 2024. Lotus: Enabling semantic queries with llms over tables of unstructured and structured data. arXiv:2407.11418 (2024)
2024 arXiv
-
[117]
Norman W Paton, Jiaoyan Chen, and Zhenyu Wu. 2023. Dataset discovery and exploration: A survey. Comput. Surveys 56, 4 (2023), 1–37
2023
-
[118]
Keqin Peng, Liang Ding, Yancheng Yuan, Xuebo Liu, Min Zhang, Yuanxin Ouyang, and Dacheng Tao. 2024. Revisiting Demonstration Selection Strategies in In-Context Learning. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (ACL) . 9090–9101
2024
-
[119]
Yun Peng, Byron Choi, Tsz Nam Chan, and Jianliang Xu. 2022. Lan: Learning-based approximate k-nearest neighbor search in graph databases. In 38th IEEE international conference on data engineering (ICDE) . 2508–2521
2022
-
[120]
Chengzhi Piao, Tingyang Xu, Xiangguo Sun, Yu Rong, Kangfei Zhao, and Hong Cheng. 2023. Computing graph edit distance via neural graph matching. Proceedings of the VLDB Endowment 16, 8 (2023), 1817–1829
2023
-
[121]
Rishabh Ranjan, Siddharth Grover, Sourav Medya, Venkat Chakravarthy, Yogish Sabharwal, and Sayan Ranu. 2022. GREED: a neural framework for learning graph distance functions. In Proceedings of the 36th International Conference on Neural Information Processing Systems (NeurIPS) ...
2022
-
[122]
Kaspar Riesen, Stefan Fankhauser, and Horst Bunke. 2007. Speeding up graph edit distance computation with a bipartite heuristic. In Mining and Learning with Graphs (MLG) . 21–24
2007
-
[123]
Keshav Santhanam, Omar Khattab, Christopher Potts, and Matei Zaharia. 2022. PLAID: an efficient engine for late interaction retrieval. In Proceedings of the 31st ACM International Conference on Information and Knowledge Management (CIKM) . 1747–1756
2022
-
[124]
Aécio Santos, Aline Bessa, Christopher Musco, and Juliana Freire. 2022. A sketch-based index for correlated dataset search. In38th IEEE International Conference on Data Engineering (ICDE) . 2928–2941
2022
-
[125]
Han Shen, Pin-Yu Chen, Payel Das, and Tianyi Chen. 2024. Seal: Safety-enhanced aligned llm fine-tuning via bilevel data selection.arXiv:2410.07471 (2024)
2024 arXiv
-
[126]
Dharma Shukla, Shireesh Thota, Karthik Raman, Madhan Gajendran, Ankur Shah, Sergii Ziuzin, Krishnan Sundaram, Miguel Gonzalez Guajardo, Anna Wawrzyniak, Samer Boshra, et al. 2015. Schema-agnostic indexing with Azure DocumentDB. Proceedings of the VLDB Endowment 8, 12 (2015), 1668–1679
2015
-
[127]
Arnab Sinha, Zhihong Shen, Yang Song, Hao Ma, Darrin Eide, Bo-June Hsu, and Kuansan Wang. 2015. An overview of microsoft academic service (mas) and applications. In Proceedings of the 24th International Conference on World Wide Web Companion . 243–246
2015
-
[128]
Luca Soldaini, Rodney Kinney, Akshita Bhagia, Dustin Schwenk, David Atkinson, Russell Authur, Ben Bogin, Khyathi Chandu, Jennifer Dumas, Yanai Elazar, et al. 2024. Dolma: an Open Corpus of Three Trillion Tokens for Language Model Pretraining Research. In Proceedings of the 62n...
2024
-
[129]
Ryota Tozuka, Hisashi Johno, Akitomo Amakawa, Junichi Sato, Mizuki Muto, Shoichiro Seki, Atsushi Komaba, and Hiroshi Onishi. 2025. Application of NotebookLM, a large language model with retrieval-augmented generation, for lung cancer staging. Japanese Journal of Radiology 43, ...
2025
-
[130]
Davison, and Jeff Heflin
Mohamed Trabelsi, Zhiyu Chen, Shuo Zhang, Brian D. Davison, and Jeff Heflin. 2022. StruBERT: Structure-aware BERT for Table Search and Matching. In The World Wide Web Conference (WWW) , Frédérique Laforest, Raphaël Troncy, Elena Simperl, Deepak Agarwal, Aristides Gionis, Ivan ...
2022
-
[131]
Immanuel Trummer. 2023. Can Large Language Models Predict Data Correlations from Column Names? Proceedings of the VLDB Endowment 16, 13 (2023), 4310–4323
2023
-
[132]
Cem Unsalan and Beril Sirmacek. 2012. Road network detection using probabilistic and graph theoretical methods. IEEE Transactions on Geoscience and Remote Sensing 50, 11 (2012), 4441–4453
2012
-
[133]
Tin Vu and Ahmed Eldawy. 2018. R-Grove: growing a family of R-trees in the big-data forest. In Proceedings of the 26th ACM SIGSPATIAL International Conference on Advances in Geographic Information Systems (SIGSPATIAL) . 532–535
2018
-
[134]
Jianhua Wang, Jianye Yang, and Wenjie Zhang. 2021. Top-k Tree Similarity Join. In Proceedings of the 30th ACM International Conference on Information and Knowledge Management (CIKM) . 1939–1948
2021
-
[135]
Rui Wang, Devin Gibson, Kirk Rodrigues, Yu Luo, Yun Zhang, Kaibo Wang, Yupeng Fu, Ting Chen, and Ding Yuan. 2024. {𝜇Slope}: High Compression and Fast Search on{Semi-Structured} Logs. In 18th USENIX Symposium on Operating Systems Design and Implementation (OSDI) . 529–544
2024
-
[136]
Shane Culpepper, Timos Sellis, Mark Sanderson, and Xiaolin Qin
Sheng Wang, Zhifeng Bao, J. Shane Culpepper, Timos Sellis, Mark Sanderson, and Xiaolin Qin. 2017. Answering Top-k Exemplar Trajectory Queries. In 33rd IEEE International Conference on Data Engineering (ICDE) . 597–608
2017
-
[137]
Sheng Wang, Mingzhao Li, Yipeng Zhang, Zhifeng Bao, David Alexander Tedjopurnomo, and Xiaolin Qin. 2018. Trip planning by an integrated search paradigm. In Proceedings of the 2018 International Conference on Management of Data (SIGMOD) . 1673–1676
2018
-
[139]
Tianshi Wang, Fengling Li, Lei Zhu, Jingjing Li, Zheng Zhang, and Heng Tao Shen. 2024. Cross-Modal Retrieval: A Systematic Review of Methods and Future Directions. Proc. IEEE 112, 11 (2024), 1716–1754
2024
-
[140]
Yunling Wang, Jianfeng Wang, and Xiaofeng Chen. 2016. Secure searchable encryption: a survey. Journal of communications and information networks 1 (2016), 52–65
2016
-
[141]
Zheng Wang, Shu Xian Teo, Jun Jie Chew, and Wei Shi. 2025. Instructrag: Leveraging retrieval-augmented generation on instruction graphs for llm-based task planning. In Proceedings of the 48th International ACM SIGIR Conference on Research and Development in Information Retriev...
2025
-
[142]
Duncan J Watts, Peter Sheridan Dodds, and Mark EJ Newman. 2002. Identity and search in social networks. science 296, 5571 (2002), 1302–1305
2002
-
[143]
Junde Wu, Jiayuan Zhu, Yunli Qi, Jingkun Chen, Min Xu, Filippo Menolascina, Yueming Jin, and Vicente Grau. 2025. Medical Graph RAG: Evidence-based Medical Large Language Model via Graph Retrieval-Augmented Generation. In Proceedings of the 63rd Annual Meeting of the Associatio...
2025
-
[144]
Shiguang Wu, Wenda Wei, Mengqi Zhang, Zhumin Chen, Jun Ma, Zhaochun Ren, Maarten de Rijke, and Pengjie Ren. 2024. Generative retrieval as multi-vector dense retrieval. In Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retr...
2024
-
[145]
Shirley Wu, Shiyu Zhao, Michihiro Yasunaga, Kexin Huang, Kaidi Cao, Qian Huang, Vassilis N Ioannidis, Karthik Subbian, James Zou, and Jure Leskovec. 2024. Stark: Benchmarking llm retrieval on textual and relational knowledge bases. In Proceedings of the 38th International Conf...
2024
-
[146]
Mengzhou Xia, Sadhika Malladi, Suchin Gururangan, Sanjeev Arora, and Danqi Chen. 2024. LESS: selecting influential data for targeted instruction tuning. In Proceedings of the 41st International Conference on Machine Learning (ICLR) . 54104–54132
2024
-
[147]
Yang Xiao, Jinlan Fu, Weizhe Yuan, Vijay Viswanathan, Zhoumianze Liu, Yixin Liu, Graham Neubig, and Pengfei Liu. 2022. DataLab: A Platform for Data Analysis and Intervention. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (ACL) . 182–195
2022
-
[148]
Anjie Xu, Ruiqing Ding, and Leye Wang. 2025. ChatPD: An LLM-driven Paper-Dataset Networking System. arXiv:2505.22349 (2025)
2025 arXiv
-
[149]
Derong Xu, Xinhang Li, Ziheng Zhang, Zhenxi Lin, Zhihong Zhu, Zhi Zheng, Xian Wu, Xiangyu Zhao, Tong Xu, and Enhong Chen. 2025. Harnessing large language models for knowledge graph question answering via adaptive multi-aspect retrieval-augmentation. In Proceedings of the AAAI ...
2025
-
[150]
Mohamed Yakout, Kris Ganjam, Kaushik Chakrabarti, and Surajit Chaudhuri. 2012. Infogather: entity augmentation and attribute discovery by holistic matching with web tables. In Proceedings of the 2012 International Conference on Management of Data (SIGMOD) . 97–108
2012
-
[151]
Mengyi Yan, Yaoshu Wang, Kehan Pang, Min Xie, and Jianxin Li. 2024. Efficient mixture of experts based on large language models for low-resource data preprocessing. In Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD) . 3690–3701
2024
-
[152]
Lei Yang and Lei Zou. 2021. Noah: Neural-optimized A* search algorithm for graph edit distance computation. In37th IEEE International Conference on Data Engineering (ICDE) . 576–587
2021
-
[153]
Wenzhe Yang, Shixun Huang, Sheng Wang, and Zhiyong Peng. 2024. Budgeted Spatial Data Acquisition: When Coverage and Connectivity Matter. arXiv preprint arXiv:2412.04853 (2024)
2024 arXiv
-
[157]
Antonio Jimeno Yepes, Yao You, Jan Milczek, Sebastian Laverde, and Renyu Li. 2024. Financial report chunking for effective retrieval augmented generation. arXiv:2402.05131 (2024)
2024 arXiv
-
[158]
Cafarella, Babak Salimi, and Anna Zeng
Brit Youngmann, Michael J. Cafarella, Babak Salimi, and Anna Zeng. 2023. Causal Data Integration. Proceedings of the VLDB Endowment 16, 10 (2023), 2659–2665
2023
-
[159]
Hao Yu, Aoran Gan, Kai Zhang, Shiwei Tong, Qi Liu, and Zhaofeng Liu. 2024. Evaluation of retrieval-augmented generation: A survey. In CCF Conference on Big Data . 102–120
2024
-
[160]
Zichun Yu, Spandan Das, and Chenyan Xiong. 2024. Mates: Model-aware data selection for efficient pretraining with data influence models. In Proceedings of the 38th International Conference on Neural Information Processing Systems (NeurIPS) . 108735–108759
2024
-
[161]
Haitao Yuan, Sai Wang, Zhifeng Bao, and Shangguang Wang. 2023. Automatic road extraction with multi-source data revisited: completeness, smoothness and discrimination. Proceedings of the VLDB Endowment 16, 11 (2023), 3004–3017
2023
-
[162]
Joohyung Yun, Byungchul Tak, and Wook-Shin Han. 2024. ReCG: Bottom-up JSON Schema Discovery Using a Repetitive Cluster-and-Generalize Framework. Proceedings of the VLDB Endowment 17, 11 (2024), 3538–3550
2024
-
[163]
Yiming Zeng, Yixuan Lin, Yuanyuan Yang, and Ji Liu. 2021. Differentially private federated temporal difference learning. IEEE Transactions on Parallel and Distributed Systems 33, 11 (2021), 2714–2726
2021
-
[164]
Zhiping Zeng, Anthony KH Tung, Jianyong Wang, Jianhua Feng, and Lizhu Zhou. 2009. Comparing stars: On approximating graph edit distance. Proceedings of the VLDB Endowment 2, 1 (2009), 25–36
2009
-
[165]
Boyu Zhang, Hongyang Yang, Tianyu Zhou, Muhammad Ali Babar, and Xiao-Yang Liu. 2023. Enhancing financial sentiment analysis via retrieval augmented large language models. In Proceedings of the fourth ACM international conference on AI in finance (ICAIF) . 349–356
2023
-
[166]
Haochen Zhang, Yuyang Dong, Chuan Xiao, and Masafumi Oyamada. 2023. Jellyfish: A large language model for data preprocessing.arXiv:2312.01678 (2023)
2023 arXiv
-
[167]
Shuo Zhang and Krisztian Balog. 2020. Web table extraction, retrieval, and augmentation: A survey. ACM Transactions on Intelligent Systems and Technology 11, 2 (2020), 1–35
2020
-
[169]
Zeyu Zhang, Paul Groth, Iacer Calixto, and Sebastian Schelter. 2024. Directions Towards Efficient and Automated Data Wrangling with Large Language Models. In 2024 IEEE 40th International Conference on Data Engineering Workshops . 301–304
2024
-
[170]
Chenyu Zhao, Yunjiang Jiang, Yiming Qiu, Han Zhang, and Wen-Yun Yang. 2023. Differentiable retrieval augmentation via generative language modeling for E-commerce query intent classification. In Proceedings of the 32nd ACM International Conference on Information and Knowledge M...
2023
-
[171]
Fuheng Zhao, Shaleen Deep, Fotis Psallidas, Avrilia Floratou, Divyakant Agrawal, and Amr El Abbadi. 2024. Sphinteract: Resolving Ambiguities in NL2SQL through User Interaction. Proceedings of the VLDB Endowment 18, 4 (2024), 1145–1158
2024
-
[172]
Penghao Zhao, Hailin Zhang, Qinhan Yu, Zhengren Wang, Yunteng Geng, Fangcheng Fu, Ling Yang, Wentao Zhang, Jie Jiang, and Bin Cui. 2024. Retrieval-augmented generation for ai-generated content: A survey. arXiv:2402.19473 (2024)
2024 arXiv
-
[173]
Xiang Zhao, Chuan Xiao, Xuemin Lin, Wenjie Zhang, and Yang Wang. 2018. Efficient structure similarity searches: a partition-based approach. The VLDB Journal 27, 1 (2018), 53–78
2018
-
[174]
Weiguo Zheng, Lei Zou, Xiang Lian, Dong Wang, and Dongyan Zhao. 2014. Efficient graph similarity search over large graph databases. IEEE Transactions on Knowledge and Data Engineering 27, 4 (2014), 964–978
2014
-
[175]
Yuanhao Zhong, Yuhao Deng, Chengliang Chai, Ruixin Gu, Ye Yuan, Guoren Wang, and Lei Cao. 2025. Doctopus: A System for Budget-aware Structural Data Extraction from Unstructured Documents. In Companion of the 2025 International Conference on Management of Data . 275–278
2025
-
[176]
Erkang Zhu, Dong Deng, Fatemeh Nargesian, and Renée J Miller. 2019. Josie: Overlap set similarity search for finding joinable tables in data lakes. In Proceedings of the 2019 International Conference on Management of Data (SIGMOD) . 847–864
2019
-
[177]
Erkang Zhu, Fatemeh Nargesian, Ken Q Pu, and Renée J Miller. 2016. LSH ensemble: internet-scale domain search. Proceedings of the VLDB Endowment 9, 12 (2016), 1185–1196
2016
-
[178]
Yuanyuan Zhu, Lu Qin, Jeffrey Xu Yu, and Hong Cheng. 2019. Answering Top-𝑘 Graph Similarity Queries in Graph Databases. IEEE Transactions on Knowledge and Data Engineering 32, 8 (2019), 1459–1474
2019
-
[179]
Lei Zou, Jinghui Mo, Lei Chen, M Tamer Özsu, and Dongyan Zhao. 2011. gStore: answering SPARQL queries via subgraph matching. Proceedings of the VLDB Endowment 4, 8 (2011), 482–493
2011
Reviewed August 5, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.