Pith. sign in

REVIEW 2 major objections 6 minor 179 references

A Survey on Open Dataset Search in the LLM Era: Retrospectives and Perspectives

T0 review · 2 major / 6 minor · reviewed 2026-08-05 · deepseek-v4-flash

Pith's one-line read This survey argues that open dataset search has outgrown keyword matching and is now defined by a two-way partnership with large language models.

desk verdict A useful modality-aware survey of dataset search that mostly delivers, but the 'dataset search for LLM' half of its central pitch is asserted rather than shown. read the letter →

arxiv 2509.00728 v1 pith:K4AFYL3X submitted 2025-08-31 cs.IR cs.DB

classification cs.IRcs.DB
keywords opendatasetsearchlargelanguagemodelsquery-by-examplecontent-awareretrievalretrieval-augmentedgenerationdataselectionmultimodaldatasetsdiscovery
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper tries to establish that open dataset search has moved beyond metadata and keyword matching into content-aware, example-based retrieval, and that this shift is best understood by organizing the field by data modality. It reviews tabular, spatial, JSON, graph, and vector datasets, cataloging how each modality defines similarity and accelerates search. Its central claim is that large language models and dataset search are mutually beneficial: LLMs improve query understanding, semantic matching, and interactive guidance, while dataset search supplies the high-quality, task-relevant data that retrieval-augmented generation and data selection depend on. If this framing holds, future research and products will converge on modality-aware, LLM-orchestrated dataset discovery rather than on traditional keyword portals. The survey also maps open problems such as privacy-preserving search, cross-modal discovery, federated search, and missing benchmarks, arguing these are the next bottlenecks.

What carries the argument

The organizing instrument is a modality taxonomy: tabular, vector, spatial, JSON document, and graph datasets, each with a formal definition (for example, a vector dataset is a set of embedding vectors, a spatial dataset is a set of georeferenced points, a JSON document is a nested key-value tree, and a graph dataset is a vertex-edge pair). The paper maps each modality to its dominant similarity measure—MaxSim for vector sets, Earth mover's distance and Hausdorff distance for spatial data, tree edit distance for JSON, and graph edit distance or maximum common subgraph for graphs—along with indexing and acceleration techniques. Over this taxonomy it overlays the two-way LLM relationship, trea

What would settle it

A systematic sweep of published dataset search work that finds a material body of research on a modality the survey does not index (e.g., time-series or audio datasets), or that shows the LLM-for-search systems the survey cites are not adopted or cited by downstream RAG or data-selection work, would weaken both the completeness of the modality taxonomy and the claim that the LLM-dataset search relationship is genuinely two-way.

Watch

Extended reading notes

Core claim

The paper's contribution is a structured synthesis: modern open dataset search is no longer a metadata-matching problem but a content-aware, modality-specific retrieval problem, and the most forward-looking direction is the two-way coupling between LLMs and dataset search. On one side, LLMs automate dataset construction, cleaning, and transformation, and enable natural-language queries, semantic planning, and relevance estimation. On the other side, dataset search supports LLMs by selecting datasets for retrieval-augmented generation and for data selection during pretraining, fine-tuning, and in-context learning. The survey develops this thesis across five data modalities, offering formal de

Load-bearing premise

The survey assumes that the research landscape is accurately represented by its modality-based partition and by the particular selection of papers, so that the taxonomy and the open-problems list correspond to the field's true shape rather than to the authors' reading of it.

Editorial extensions

If this is right

  • Example-based and natural-language queries will increasingly replace keyword-only interfaces in dataset search systems.
  • LLM-based schema inference, cleaning, and transformation should be treated as part of the dataset search pipeline, not as separate data-engineering tasks.
  • Dataset search will become a core stage in LLM development, feeding retrieval-augmented generation and data-selection pipelines with queryable, filterable data.
  • Standardized benchmarks are missing for spatial and vector dataset search, and quality control is not yet integrated into search; closing these gaps is a prerequisite for practical progress.
  • Privacy-preserving similarity calculation, indexing over encrypted data, and federated search remain open problems that need new index and acceleration designs.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the mutual-benefit thesis holds, dataset search could become a closed-loop controller for LLM training: a model's judged failures could seed queries for corrective datasets, and the retrieved data could then be evaluated by the model, forming a feedback loop; the paper mentions agentic integration but does not formalize this loop.
  • The modality taxonomy suggests that cross-modal dataset search is the natural stress test: a shared embedding space for tables, JSON trees, graphs, and spatial points would unify the separate similarity measures, and a cross-modal join and union benchmark would directly test the value of such a unified representation.
  • The paper's emphasis on content over metadata implies that search quality may be robust to missing or poor metadata; a direct experiment comparing ranking quality on datasets with full metadata versus deliberately stripped metadata would quantify how much content signals compensate.
  • The recurring use of sketch and hash approximations across tabular, vector, and spatial search suggests a common abstraction—set containment at scale—that could be factored into a shared data-discovery index, potentially simplifying future system design.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 6 minor

Summary. This survey reviews recent research on open dataset search, moving beyond keyword- and metadata-based retrieval. It organizes the field by data modality (tabular, spatial, JSON, graph, and vector datasets), surveys query-by-example and natural-language search techniques, and discusses the interplay between LLMs and dataset search. The claimed central contribution is a two-way relationship: LLMs improve dataset search through query understanding, semantic planning, and interactive guidance, while dataset search supports LLMs through retrieval-augmented generation (RAG) and data selection. The survey also identifies open problems, including privacy-preserving search, task-oriented integration, cross-modal discovery, federated search, and benchmarks.

Significance. If the reference selection is representative and the reverse-direction claims are adequately evidenced, this survey would fill a useful gap by providing a modality-aware, LLM-centric overview of open dataset search. The paper's strengths are its broad taxonomy, formal definitions, reproduction of key equations (e.g., MaxSim, Hausdorff distance, JSON tree edit distance, GED/MCS), comparative tables that summarize representative methods, and an explicit roadmap of open challenges. It also makes the useful expository choice of distinguishing LLM-for-search from search-for-LLM. However, as a survey, its conclusions rest on the accuracy of its secondary summaries and on the completeness and neutrality of its reference selection, and neither can be fully verified from the manuscript. The most load-bearing weakness is the reverse-direction chapter, whose RAG examples do not actually demonstrate dataset-search-backed RAG.

major comments (2)
  1. [§5.3, Table 7] The claimed direction 'advances in dataset search can support LLMs by enabling more effective integration into RAG frameworks' is not supported by the evidence in this subsection. After noting that RAG's retrieval component 'naturally aligns' with dataset search, the text explicitly states that the subsection 'instead showcases recent applications of RAG across diverse domains.' Table 7 lists 15 RAG applications (e.g., Self-BioRAG, GraphQA, InstructRAG), but these retrieve scientific papers, knowledge graphs, or domain-specific text corpora—not open datasets as defined in Definition 1.1, and not through any dataset-search system surveyed in Sections 3–4. None of the cited examples shows a dataset-search engine (e.g., Auctus, DBF, Starmie, LOTUS) supplying the retrieval source of a RAG pipeline. The abstract and Section 1.3 stake a core contribution on this mutual-benefit claim, so the re
  2. [§5.3, Data Selection for LLM] The data-selection discussion similarly overstates the integration of dataset search into LLM data pipelines. The paragraph cites generic data-selection methods (LESS, Dolma, MateS) and then identifies only one method—Wang et al. [138]—that actually integrates query-driven dataset search into a selection pipeline. The claim that 'recent advances have begun to integrate dataset search into data selection pipelines' is supported by a single concrete example. Given that Section 6.2 itself mentions other systems (DeepResearchGym, STARK, GPT-Instructor) that embed dataset retrieval in end-to-end pipelines, the authors could strengthen this subsection substantially by moving those examples here or by adding additional published cases. Without that, the reverse-direction evidence remains one example, which is not enough to sustain the survey's broader 'mutually beneficial relationship' framing.
minor comments (6)
  1. [Table 1] The 'Vector' row lists '[99]' (BioVSS) as a data source, but [99] is a search method, not an open data repository. The underlying Microsoft Academic Graph [127] appears to be the actual source. Please correct the citation or rephrase the row.
  2. [Definition 1.2] Definition 1.2 covers keyword and exemplar queries but not natural-language queries, which become a major theme in Section 5.2. Consider extending the definition or adding a remark that NL queries are handled as a separate query type later in the survey.
  3. [§4.1] The vector-dataset-search section includes several systems originally designed for passage/document retrieval (COLBERT, PLAID, SLIM, COIL, CITADEL, XTR). Since the survey's title and abstract are about dataset search, a brief justification of why token-vector collections over documents are treated as vector datasets would help readers accept this categorization.
  4. [Figure 2] Figure 2 is dense and is referenced only once. Adding pointers from the pipeline stages to the corresponding sections (e.g., query mechanisms, similarity calculation, indexing) would improve navigability.
  5. [§4.2] The discussion of spatial dataset search states that the field lacks a standardized benchmark, but no reference is given for this claim. A citation or a brief explanation of which evaluation gaps exist would strengthen the point.
  6. [§5.3] In text immediately before Table 7, the phrase 'retrieval-augmented generation' and the surrounding discussion could more precisely distinguish between retrieving documents/corpora and retrieving datasets. As written, the section's scope drifts from dataset search to general RAG.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the survey's claims are descriptive taxonomies and literature syntheses, not derivations or fitted predictions, and its self-citations are not load-bearing.

full rationale

This is a survey paper, so the standard circularity failure modes (fitting a parameter and renaming it a prediction, defining X in terms of Y, importing a uniqueness theorem from the authors' prior work) do not apply. The paper's central contribution is a modality-aware organization of existing dataset-search techniques and a discussion of the LLM/dataset-search relationship. Definitions 1.1/1.2, 3.1, 4.1–4.4 are formal definitions, not derived results. Sections 3 and 4 survey external, peer-reviewed methods (e.g., LSH Ensemble, JOSIE, Starmie, COLBERT, DBF, JEDI, A*GED) and attribute each method to its original publication; no equation in the survey is shown to reduce to an input by construction. The authors do cite their own prior works (e.g., [97,99,154–156]), but these are included as surveyed contributions alongside many third-party works, and no load-bearing conclusion rests solely on a self-citation. The 'mutually beneficial relationship' claim in §5 is expository; the skeptical concern that §5.3's RAG examples retrieve documents/corpora rather than open datasets is an internal evidence gap about the support for one direction of that relationship, not circular reasoning, since the claim is not derived from itself. The paper even concedes in §5.3 that it 'instead showcases recent applications of RAG across diverse domains,' which is a scope limitation, not a circular step. Overall, the survey's organization and gap analysis are self-contained descriptive claims, so the circularity score is 0.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

As a survey, the paper introduces no fitted parameters or new theoretical entities. Its load-bearing premises are the framing taxonomy, the fidelity of the summaries, and standard notions of dataset similarity.

assumptions (3)
  • domain assumption The modality-based taxonomy (tabular, spatial, JSON, graph, vector) is a valid way to organize the dataset search literature.
    Section 1.3 and Sections 3-4 adopt this partition; it is a framing choice, not a derived result.
  • domain assumption The cited descriptions of prior systems and methods are accurate and representative.
    The survey's summaries are based on reading the cited papers; no original experiments or code are provided to verify them.
  • standard math Standard mathematical definitions used in the survey (e.g., Hausdorff distance Eq. 4, EMD Eq. 5, GED Def. 4.5, MCS Def. 4.6) are correct and apply as stated.
    These are standard definitions from the cited literature and are not derived in the paper.

how reviews work

0 comments
Cite this review

Pith. "Pith review of A Survey on Open Dataset Search in the LLM Era: Retrospectives and Perspectives." pith.science (2026). https://pith.science/paper/K4AFYL3X

@misc{pith2026250900728,
  author       = {Pith},
  title        = {Pith review of: A Survey on Open Dataset Search in the LLM Era: Retrospectives and Perspectives},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/K4AFYL3X}},
  note         = {Machine review of arXiv:2509.00728}
}
read the original abstract

High-quality datasets are typically required for accomplishing data-driven tasks, such as training medical diagnosis models, predicting real-time traffic conditions, or conducting experiments to validate research hypotheses. Consequently, open dataset search, which aims to ensure the efficient and accurate fulfillment of users' dataset requirements, has emerged as a critical research challenge and has attracted widespread interest. Recent studies have made notable progress in enhancing the flexibility and intelligence of open dataset search, and large language models (LLMs) have demonstrated strong potential in addressing long-standing challenges in this area. Therefore, a systematic and comprehensive review of the open dataset search problem is essential, detailing the current state of research and exploring future directions. In this survey, we focus on recent advances in open dataset search beyond traditional approaches that rely on metadata and keywords. From the perspective of dataset modalities, we place particular emphasis on example-based dataset search, advanced similarity measurement techniques based on dataset content, and efficient search acceleration techniques. In addition, we emphasize the mutually beneficial relationship between LLMs and open dataset search. On the one hand, LLMs help address complex challenges in query understanding, semantic modeling, and interactive guidance within open dataset search. In turn, advances in dataset search can support LLMs by enabling more effective integration into retrieval-augmented generation (RAG) frameworks and data selection processes, thereby enhancing downstream task performance. Finally, we summarize open research problems and outline promising directions for future work. This work aims to offer a structured reference for researchers and practitioners in the field of open dataset search.

Figures

Figures reproduced from arXiv: 2509.00728 by the authors.

Figure 1
Figure 1. Motivating scenarios for open dataset search. [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Open dataset search pipeline overview [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. An example of a tabular dataset and table join and union search. [PITH_FULL_IMAGE:figures/full_fig_p007_3.png] view at source ↗
Figures from the paper (6 more)
Figure 4
Figure 4. Figure 4: Three key steps of vector dataset search: 1) Vector dataset generation, 2) Similarity calculation, 3) Main acceleration techniques [PITH_FULL_IMAGE:figures/full_fig_p011_4.png]
Figure 5
Figure 5. Figure 5: Similarity measures of spatial datasets When calculating the similarity of spatial datasets, their spatial characteristics must be taken into account. Existing methods measure the similarity between spatial datasets using overlap area, Hausdorff distance [114], and Ear…
Figure 6
Figure 6. Figure 6: An example of JSON document and its tree representation [PITH_FULL_IMAGE:figures/full_fig_p014_6.png]
Figure 7
Figure 7. Figure 7: Graph edit distance of 𝐺𝑖 and 𝐺𝑗 , the graph edit operations include: 1) Insert Node (IN) and Insert Edge (IE); 2) Delete Node (DN) and Delete Edge (DE); 3) Relabel Node (RN) and Relabel Edge (RE). algorithms are often employed in practice to compute this distance. Mos…
Figure 8
Figure 8. Figure 8: Maximum common sub-graph of 𝐺𝑖 and 𝐺𝑗 . queries. When a query graph matches the GED result of a graph in the database, the pre-computed results can be used to dynamically generate a candidate graph set. This approach allows the candidate generation process to interact …
Figure 9
Figure 9. Figure 9: Interaction between LLMs and dataset search [PITH_FULL_IMAGE:figures/full_fig_p019_9.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

179 extracted references · 68 canonical work pages

  1. [138]

    Tingting Wang, Shixun Huang, Zhifeng Bao, J Shane Culpepper, Volkan Dedeoglu, and Reza Arablouei. 2025. Distinctiveness Maximization in Datasets Assemblage. In The World Wide Web Conference (WWW) . 3219–3232

  2. [97]

    Pengyue Li, Hua Dai, Sheng Wang, Wenzhe Yang, and Geng Yang. 2024. Privacy-preserving Spatial Dataset Search in Cloud. In Proceedings of the 33rd ACM International Conference on Information and Knowledge Management (CIKM) . 1245–1254

  3. [154]

    Wenzhe Yang, Sheng Wang, Zhiyu Chen, Yuan Sun, and Zhiyong Peng. 2025. Joinable Search over Multi-source Spatial Datasets: Overlap, Coverage, and Efficiency. In 41st IEEE International Conference on Data Engineering (ICDE) . 585–598

  4. [155]

    Wenzhe Yang, Sheng Wang, Shixun Huang, Yuyang Liao, Yuan Sun, Juliana Freire, and Zhiyong Peng. 2024. A Unified Approach for Multi- Granularity Search over Spatial Datasets. arXiv preprint arXiv:2412.04805 (2024)

  5. [156]

    Wenzhe Yang, Sheng Wang, Yuan Sun, and Zhiyong Peng. 2022. Fast dataset search with earth mover’s distance. Proceedings of the VLDB Endowment 15, 11 (2022), 2517–2529

  6. [99]

    Yiqi Li, Sheng Wang, Zhiyu Chen, Shangfeng Chen, and Zhiyong Peng. 2025. Approximate Vector Set Search: A Bio-Inspired Approach for High-Dimensional Spaces. In 2025 IEEE 41st International Conference on Data Engineering (ICDE) . 891–903

  7. [168]

    Zhengkai Zhang, Zhou Hao Dai, Hua, Mingfeng Jiang, Pengyue Li, and Geng Yang. 2025. A Privacy-preserving Spatial Dataset Joinable Search in Cloud. In Proceedings of the 34th ACM International Conference on Information and Knowledge Management (CIKM)

  8. [1]

    Abbas Acar, Hidayet Aksu, A Selcuk Uluagac, and Mauro Conti. 2018. A survey on homomorphic encryption schemes: Theory and implementation. Comput. Surveys 51, 4 (2018), 1–35

Show all 179 references
  1. [2]

    Uchenna Akujuobi and Xiangliang Zhang. 2017. Delve: a dataset-driven scholarly search and analysis system.ACM SIGKDD Explorations Newsletter 19, 2 (2017), 36–46

  2. [3]

    Alawwad, Areej Alhothali, Usman Naseem, Ali Alkhathlan, and Amani Jamal

    Hessa A. Alawwad, Areej Alhothali, Usman Naseem, Ali Alkhathlan, and Amani Jamal. 2025. Enhancing textual textbook question answering with large language models and retrieval augmented generation. Pattern Recognition 162 (2025), 111332

  3. [4]

    Alon Albalak, Yanai Elazar, Sang Michael Xie, Shayne Longpre, Nathan Lambert, Xinyi Wang, Niklas Muennighoff, Bairu Hou, Liangming Pan, Haewon Jeong, et al. 2024. A survey on data selection for language models. arXiv:2402.16827 (2024)

  4. [5]

    Eric Anderson, Jonathan Fritz, Austin Lee, Bohou Li, Mark Lindblad, Henry Lindeman, Alex Meyer, Parth Parmar, Tanvi Ranade, Mehul A Shah, et al. 2024. The Design of an LLM-powered Unstructured Analytics System. arXiv:2409.00847 (2024)

  5. [6]

    Simran Arora, Brandon Yang, Sabri Eyuboglu, Avanika Narayan, Andrew Hojel, Immanuel Trummer, and Christopher Ré. 2023. Language Models Enable Simple Systems for Generating Structured Views of Heterogeneous Data Lakes. Proceedings of the VLDB Endowment 17, 2 (2023), 92–105

  6. [7]

    Mohamed-Amine Baazizi, Dario Colazzo, Giorgio Ghelli, and Carlo Sartiani. 2019. Parametric schema inference for massive JSON datasets. The VLDB Journal 28, 4 (2019), 497–521

  7. [8]

    Mohamed-Amine Baazizi, Dario Colazzo, Giorgio Ghelli, and Carlo Sartiani. 2019. Schemas and types for JSON data: from theory to practice. In Proceedings of the 2019 International Conference on Management of Data (SIGMOD) . 2060–2063

  8. [9]

    Jiyang Bai and Peixiang Zhao. 2021. TaGSim: Type-aware Graph Similarity Learning and Computation. Proceedings of the VLDB Endowment 15, 2 (2021), 335–347

  9. [10]

    Tianyi Bai, Ling Yang, Zhen Hao Wong, Fupeng Sun, Xinlin Zhuang, Jiahui Peng, Chi Zhang, Lijun Wu, Jiantao Qiu, Wentao Zhang, Binhang Yuan, and Conghui He. 2025. Efficient Pretraining Data Selection for Language Models via Multi-Actor Collaboration. In Proceedings of the 63rd ...

  10. [11]

    Omar Benjelloun, Shiyu Chen, and Natasha Noy. 2020. Google Dataset Search by the Numbers. In International Semantic Web Conference (ISWS) . 667–682

  11. [12]

    Chandra Sekhar Bhagavatula, Thanapon Noraset, and Doug Downey. 2015. Tabel: Entity linking in web tables. In International Semantic Web Conference (ISWS). 425–441

  12. [13]

    Fabian Biester, Mohamed Abdelaal, and Daniel Del Gaudio. 2024. Llmclean: Context-aware tabular data cleaning via llm-generated ofds. In European Conference on Advances in Databases and Information Systems (ADBIS) , Vol. 2186. 68–78

  13. [14]

    Tobias Bleifuß, Leon Bornemann, Dmitri V Kalashnikov, Felix Naumann, and Divesh Srivastava. 2021. The Secret Life of Wikipedia Tables. In Proceedings of the 2nd Workshop on Search, Exploration, and Analysis in Heterogeneous Datastores (SEA-Data 2021) co-located with 47th Inter...

  14. [15]

    Alex Bogatu, Alvaro A A Fernandes, Norman W Paton, and Nikolaos Konstantinou. 2020. Dataset Discovery in Data Lakes. In36th IEEE International Conference on Data Engineering (ICDE) . 709–720

  15. [16]

    Pierre Bourhis, Juan L Reutter, Fernando Suárez, and Domagoj Vrgoč. 2017. JSON: data model, query languages and schema specification. In Proceedings of the 36th ACM SIGMOD-SIGACT-SIGAI Symposium on Principles of Database Systems (PODS) . 123–135

  16. [17]

    Pierre Bourhis, Juan L Reutter, and Domagoj Vrgoč. 2020. JSON: Data model and query languages. Information Systems 89 (2020), 101478

  17. [18]

    Dan Brickley, Matthew Burgess, and Natasha Noy. 2019. Google Dataset Search: Building a search engine for datasets in an open Web ecosystem. In The World Wide Web Conference (WWW) . 1365–1375

  18. [19]

    Horst Bunke and Kim Shearer. 1998. A graph distance metric based on the maximal common subgraph. Pattern recognition letters 19, 3-4 (1998), 255–259

  19. [20]

    Sonia Castelo, Rémi Rampin, Aécio Santos, Aline Bessa, Fernando Chirigati, and Juliana Freire. 2021. Auctus: a dataset search engine for data discovery and augmentation. Proceedings of the VLDB Endowment 14, 12 (2021), 2791–2794

  20. [21]

    Lijun Chang, Xing Feng, Xuemin Lin, Lu Qin, Wenjie Zhang, and Dian Ouyang. 2020. Speeding up GED verification for graph similarity search. In 36th IEEE International Conference on Data Engineering (ICDE) . 793–804

  21. [22]

    Lijun Chang, Xing Feng, Kai Yao, Lu Qin, and Wenjie Zhang. 2022. Accelerating graph similarity search via efficient GED computation. IEEE Transactions on Knowledge and Data Engineering 35, 5 (2022), 4485–4498

  22. [23]

    Ming-Fang Chang, John Lambert, Patsorn Sangkloy, Jagjeet Singh, Slawomir Bak, Andrew Hartnett, De Wang, Peter Carr, Simon Lucey, Deva Ramanan, et al. 2019. Argoverse: 3d tracking and forecasting with rich maps. In IEEE Conference on Computer Vision and Pattern Recognition (CVP...

  23. [24]

    Adriane Chapman, Elena Simperl, Laura Koesten, George Konstantinidis, Luis-Daniel Ibáñez, Emilia Kacprzak, and Paul Groth. 2020. Dataset search: a survey. The VLDB Journal 29, 1 (2020), 251–272

  24. [25]

    Minxiao Chen, Haitao Yuan, Nan Jiang, Zhifeng Bao, and Shangguang Wang. 2024. Urban Traffic Accident Risk Prediction Revisited: Regionality, Proximity, Similarity and Sparsity. InProceedings of the 33rd ACM International Conference on Information and Knowledge Management (CIKM...

  25. [26]

    Minxiao Chen, Haitao Yuan, Nan Jiang, Zhihan Zheng, Zhifeng Bao, Ao Zhou, Jiaxin Jiang, and Shangguang Wang. 2025. S-MGHSTN: Towards An Effective Streaming Traffic Accident Risk Prediction Framework. IEEE Transactions on Knowledge and Data Engineering 37, 7 (2025), 4285 – 4298

  26. [27]

    Qiaosheng Chen, Jiageng Chen, Xiao Zhou, and Gong Cheng. 2024. Enhancing Dataset Search with Compact Data Snippets. In Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR) . 1093–1103

  27. [28]

    Qiaosheng Chen, Weiqing Luo, Zixian Huang, Tengteng Lin, Xiaxia Wang, Ahmet Soylu, Basil Ell, Baifan Zhou, Evgeny Kharlamov, and Gong Cheng. 2024. ACORDAR 2.0: A Test Collection for Ad Hoc Dataset Retrieval with Densely Pooled Datasets and Question-Style Queries. In Proceeding...

  28. [29]

    Wei-Hao Chen, Weixi Tong, Amanda Case, and Tianyi Zhang. 2025. Dango: A Mixed-Initiative Data Wrangling System using Large Language Model. In Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems (CHI) . 1–28

  29. [30]

    Zui Chen, Zihui Gu, Lei Cao, Ju Fan, Samuel Madden, and Nan Tang. 2023. Symphony: Towards Natural Language Query Answering over Multi-modal Data Lakes. In 13th Conference on Innovative Data Systems Research (CIDR) . 1–7

  30. [31]

    Zhiyu Chen, Mohamed Trabelsi, Jeff Heflin, Yinan Xu, and Brian D. Davison. 2020. Table Search Using a Deep Contextualized Language Model. In Proceedings of the 43rd International ACM SIGIR conference on research and development in Information Retrieval (SIGIR) , Jimmy X. Huang...

  31. [32]

    Mingyue Cheng, Yucong Luo, Jie Ouyang, Qi Liu, Huijie Liu, Li Li, Shuo Yu, Bohou Zhang, Jiawei Cao, Jie Ma, et al . 2025. A survey on knowledge-oriented retrieval-augmented generation. arXiv:2503.10677 (2025)

  32. [33]

    João Coelho, Jingjie Ning, Jingyuan He, Kangrui Mao, Abhijay Paladugu, Pranav Setlur, Jiahe Jin, Jamie Callan, João Magalhães, Bruno Martins, et al. 2025. Deepresearchgym: A free, transparent, and reproducible evaluation sandbox for deep research. arXiv:2505.19253 (2025)

  33. [34]

    Tianji Cong, Fatemeh Nargesian, and HV Jagadish. 2023. Pylon: Semantic table union search in data lakes. arXiv:2301.04901 (2023)

  34. [35]

    Arash Dargahi Nobari and Davood Rafiei. 2024. Dtt: An example-driven tabular transformer for joinability by leveraging large language models. Proceedings of the ACM on Management of Data 2, 1 (2024), 1–24

  35. [36]

    Anish Das Sarma, Lujun Fang, Nitin Gupta, Alon Halevy, Hongrae Lee, Fei Wu, Reynold Xin, and Cong Yu. 2012. Finding related tables. In Proceedings of the 2012 International Conference on Management of Data (SIGMOD) . 817–828

  36. [37]

    Auriol Degbelo and Brhane Bahrishum Teka. 2019. Spatial search strategies for open government data: A systematic comparison. In Proceedings of the 13th Workshop on Geographic Information Retrieval (GIR) . 1–10

  37. [38]

    Yuhao Deng, Chengliang Chai, Lei Cao, Qin Yuan, Siyuan Chen, Yanrui Yu, Zhaoze Sun, Junyi Wang, Jiajun Li, Ziqi Cao, et al. 2024. LakeBench: A Benchmark for Discovering Joinable and Unionable Tables in Data Lakes. Proceedings of the VLDB Endowment 17, 8 (2024), 1925–1938

  38. [39]

    Yuyang Dong, Kunihiro Takeoka, Chuan Xiao, and Masafumi Oyamada. 2021. Efficient Joinable Table Discovery in Data Lakes: A High-Dimensional Similarity-Based Approach. In 37th IEEE International Conference on Data Engineering (ICDE) . 456–467

  39. [40]

    Yuyang Dong, Chuan Xiao, Takuma Nozawa, Masafumi Enomoto, and Masafumi Oyamada. 2023. DeepJoin: Joinable Table Discovery with Pre-Trained Language Models. Proceedings of the VLDB Endowment 16, 10 (2023), 2458–2470

  40. [41]

    Dominik Durner, Viktor Leis, and Thomas Neumann. 2021. JSON tiles: Fast analytics on semi-structured data. InProceedings of the 2021 International Conference on Management of Data (SIGMOD) . 445–458

  41. [42]

    Mohamed Y Eltabakh, Mayuresh Kunjir, Ahmed K Elmagarmid, and Mohammad Shahmeer Ahmad. 2023. Cross Modal Data Discovery over Structured and Unstructured Data Lakes. Proceedings of the VLDB Endowment 16, 11 (2023), 3377–3390

  42. [43]

    Joshua Engels, Benjamin Coleman, Vihan Lakshman, and Anshumali Shrivastava. 2023. DESSERT: an efficient algorithm for vector set search with vector set queries. Advances in Neural Information Processing Systems 36 (2023), 67972–67992

  43. [44]

    Mahdi Esmailoghli, Christoph Schnell, Renée J Miller, and Ziawasch Abedjan. 2025. BLEND: A Unified Data Discovery System. In 41st IEEE International Conference on Data Engineering (ICDE) . 737–750

  44. [45]

    Grace Fan, Jin Wang, Yuliang Li, and Renée J Miller. 2023. Table discovery in data lakes: State-of-the-art and future directions. In Companion of the 2023 International Conference on Management of Data (SIGMOD/PODS) . 69–75

  45. [46]

    Grace Fan, Jin Wang, Yuliang Li, Dan Zhang, and Renée J Miller. 2023. Semantics-Aware Dataset Discovery from Data Lakes with Contextualized Column-Based Representation Learning. Proceedings of the VLDB Endowment 16, 7 (2023), 1726–1739

  46. [47]

    Ju Fan, Zihui Gu, Songyue Zhang, Yuxin Zhang, Zui Chen, Lei Cao, Guoliang Li, Samuel Madden, Xiaoyong Du, and Nan Tang. 2024. Combining small language models and large language models for zero-shot nl2sql. Proceedings of the VLDB Endowment 17, 11 (2024), 2750–2763

  47. [48]

    Lizhou Fan, Sara Lafia, Lingyao Li, Fangyuan Yang, and Libby Hemphill. 2023. DataChat: Prototyping a conversational agent for dataset search and visualization. Proceedings of the Association for Information Science and Technology 60, 1 (2023), 586–591

  48. [49]

    Shangbin Feng, Chan Young Park, Yuhan Liu, and Yulia Tsvetkov. 2023. From Pretraining Data to Language Models to Downstream Tasks: Tracking the Trails of Political Biases Leading to Unfair NLP Models. InProceedings of the 61st Annual Meeting of the Association for Computationa...

  49. [50]

    Pedro Arthur de Fernandes Vasconcelos, Wensttay de Sousa Alencar, Victor Hugo da Silva Ribeiro, Natarajan Ferreira Rodrigues, and Fabio de Gomes Andrade. 2017. Enabling Spatial Queries in Open Government Data Portals. In International Conference on Electronic Government and th...

  50. [51]

    Raul Castro Fernandez, Ziawasch Abedjan, Famien Koko, Gina Yuan, Samuel Madden, and Michael Stonebraker. 2018. Aurum: A data discovery system. In 34th IEEE International Conference on Data Engineering (ICDE) . 1001–1012

  51. [52]

    Raul Castro Fernandez, Jisoo Min, Demitri Nava, and Samuel Madden. 2019. Lazo: A cardinality-based method for coupled estimation of jaccard similarity and containment. In 35th IEEE International Conference on Data Engineering (ICDE) . 1190–1201

  52. [53]

    Benjamin Feuer, Yurong Liu, Chinmay Hegde, and Juliana Freire. 2024. ArcheType: A Novel Framework for Open-Source Column Type Annotation Using Large Language Models. Proceedings of the VLDB Endowment 17, 9 (2024)

  53. [54]

    Juliana Freire, Grace Fan, Benjamin Feuer, Christos Koutras, Yurong Liu, Eduardo Peña, Aécio S. R. Santos, Cláudio T. Silva, and Eden Wu. 2025. Large Language Models for Data Discovery and Integration: Challenges and Opportunities. IEEE Data Engineering Bulletin 49, 1 (2025), 3–31

  54. [55]

    Luyu Gao, Zhuyun Dai, and Jamie Callan. 2021. COIL: Revisit Exact Lexical Match in Information Retrieval with Contextualized Inverted List. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics (NAACL) . 3030–3042

  55. [56]

    Adamu Garba and Shangli Wu. 2025. Snippet-based result merging in federated search. Journal of Information Science 51, 3 (2025), 623–637

  56. [57]

    Adamu Garba, Shengli Wu, and Shah Khalid. 2023. Federated search techniques: an overview of the trends and state of the art. Knowledge and Information Systems 65, 12 (2023), 5065–5095

  57. [58]

    Saheli Ghosh, Ahmed Eldawy, and Shipra Jais. 2019. Aid: An adaptive image data index for interactive multilevel visualization. In 35th IEEE International Conference on Data Engineering (ICDE) . 1594–1597

  58. [59]

    Saheli Ghosh, Tin Vu, Mehrad Amin Eskandari, and Ahmed Eldawy. 2019. UCR-STAR: The UCR spatio-temporal active repository. SIGSPATIAL Special 11, 2 (2019), 34–40

  59. [60]

    Victor Giannakouris and Immanuel Trummer. 2025. SwellDB: Dynamic Query-Driven Table Generation with Large Language Models. InCompanion of the 2025 International Conference on Management of Data . 95–98

  60. [61]

    Youdi Gong, Guangzhen Liu, Yunzhi Xue, Rui Li, and Lingzhong Meng. 2023. A survey on dataset quality in machine learning. Information and Software Technology 162 (2023), 107268

  61. [62]

    Tobias Grubenmann, Abraham Bernstein, Dmitry Moor, and Sven Seuken. 2018. Financing the web of data with delayed-answer auctions. In The World Wide Web Conference (WWW). 1033–1042

  62. [63]

    Ruiqi Guo, Philip Sun, Erik Lindgren, Quan Geng, David Simcha, Felix Chern, and Sanjiv Kumar. 2020. Accelerating large-scale inference with anisotropic vector quantization. In International Conference on Machine Learning (ICML) . 3887–3896

  63. [64]

    Stefan Hahmann and Dirk Burghardt. 2013. How much information is geospatially referenced? Networks and cognition. International Journal of Geographical Information Science 27, 6 (2013), 1171–1189

  64. [65]

    Rihan Hai, Christos Koutras, Christoph Quix, and Matthias Jarke. 2023. Data lakes: A survey of functions and systems. IEEE Transactions on Knowledge and Data Engineering 35, 12 (2023), 12571–12590

  65. [66]

    Zifei FeiFei Han, Jionghao Lin, Ashish Gurung, Danielle R Thomas, Eason Chen, Conrad Borchers, Shivang Gupta, and Kenneth R Koedinger. 2024. Improving Assessment of Tutoring Practices using Retrieval-Augmented Generation. Proceedings of Machine Learning Research 257 (2024), 66–76

  66. [67]

    Xiaoxin He, Yijun Tian, Yifei Sun, Nitesh Chawla, Thomas Laurent, Yann LeCun, Xavier Bresson, and Bryan Hooi. 2024. G-retriever: Retrieval- augmented generation for textual graph understanding and question answering. Advances in Neural Information Processing Systems 37 (2024),...

  67. [68]

    Thomas Hervey, Sara Lafia, and Werner Kuhn. 2021. Search Facets and Ranking in Geospatial Dataset Search. In 11th International Conference on Geographic Information Science (GIScience) , Vol. 177. 5:1–5:15

  68. [69]

    Sirui Hong, Yizhang Lin, Bang Liu, Bangbang Liu, Binhao Wu, Ceyao Zhang, Danyang Li, Jiaqi Chen, Jiayi Zhang, Jinlin Wang, Li Zhang, Lingyao Zhang, Min Yang, Mingchen Zhuge, Taicheng Guo, Tuo Zhou, Wei Tao, Robert Tang, Xiangtao Lu, Xiawu Zheng, Xinbing Liang, Yaying Fei, Yuhe...

  69. [70]

    Xuming Hu, Shen Wang, Xiao Qin, Chuan Lei, Zhengyuan Shen, Christos Faloutsos, Asterios Katsifodimos, George Karypis, Lijie Wen, and Philip S Yu. 2023. Automatic table union search with tabular representation learning. In Findings of the Association for Computational Linguisti...

  70. [71]

    Yuntong Hu, Zhihan Lei, Zhongjie Dai, Allen Zhang, Abhinav Angirekula, Zheng Zhang, and Liang Zhao. 2025. Cg-rag: Research question answering by citation graph retrieval-augmented llms. In Proceedings of the 48th International ACM SIGIR Conference on Research and Development i...

  71. [72]

    Benhao Huang, Yingzhuo Yu, Jin Huang, Xingjian Zhang, and Jiaqi Ma. 2024. DCA-Bench: A Benchmark for Dataset Curation Agents. arXiv:2406.07275 (2024)

  72. [73]

    Thomas Hütter, Nikolaus Augsten, Christoph M Kirsch, Michael J Carey, and Chen Li. 2022. JEDI: These aren’t the JSON documents you’re looking for.... In Proceedings of the 2022 International Conference on Management of Data (SIGMOD) . 1584–1597

  73. [74]

    Thomas Hütter, Mateusz Pawlik, Robert Löschinger, and Nikolaus Augsten. 2019. Effective filters and linear time verification for tree similarity joins. In 35th IEEE International Conference on Data Engineering (ICDE) . 854–865

  74. [75]

    Andra Ionescu, Kiril Vasilev, Florena Buse, Rihan Hai, and Asterios Katsifodimos. 2024. AutoFeat: Transitive Feature Discovery over Join Paths. In 40th IEEE International Conference on Data Engineering (ICDE) . 1861–1873

  75. [76]

    Minbyul Jeong, Jiwoong Sohn, Mujeen Sung, and Jaewoo Kang. 2024. Improving medical reasoning through retrieval and self-reflection with retrieval-augmented large language models. Bioinformatics 40, Supplement_1 (2024), i119–i129

  76. [77]

    Shane Culpepper

    Daomin Ji, Hui Luo, Zhifeng Bao, and J. Shane Culpepper. 2024. Navigating Data Repositories: Utilizing Line Charts to Discover Relevant Datasets. Proceedings of the VLDB Endowment 17, 12 (2024), 4289–4292

  77. [78]

    Daomin Ji, Hui Luo, Zhifeng Bao, and J Shane Culpepper. 2025. Table integration in data lakes unleashed: pairwise integrability judgment, integrable set discovery, and multi-tuple conflict resolution. The VLDB Journal 34, 3 (2025), 1–24

  78. [79]

    Lin Jiang, Junqiao Qiu, and Zhijia Zhao. 2020. Scalable structural index construction for JSON analytics. Proceedings of the VLDB Endowment 14, 4 (2020), 694–707

  79. [80]

    Xinhui Kang, Wenjie You, and Ying Luo. 2025. Bio-inspired product design system integrating retrieval-augmented question answering and semantic fusion diffusion model. Advanced Engineering Informatics 67 (2025), 103537

  80. [81]

    Nikolai Karpov and Qin Zhang. 2023. Syncsignature: A simple, efficient, parallelizable framework for tree similarity joins. Proceedings of the VLDB Endowment 16, 2 (2023)

  81. [82]

    Maxat Kassen. 2013. A promising phenomenon of open data: A case study of the Chicago open data project. Government information quarterly 30, 4 (2013), 508–513

  82. [83]

    Moe Kayali, Anton Lykov, Ilias Fountalis, Nikolaos Vasiloglou, Dan Olteanu, and Dan Suciu. 2024. Chorus: Foundation Models for Unified Data Discovery and Exploration. Proceedings of the VLDB Endowment 17, 8 (2024), 2104–2114

  83. [84]

    Aamod Khatiwada, Grace Fan, Roee Shraga, Zixuan Chen, Wolfgang Gatterbauer, Renée J Miller, and Mirek Riedewald. 2023. Santos: Relationship- based semantic table union search. Proceedings of the ACM on Management of Data 1, 1 (2023), 1–25

  84. [85]

    Aamod Khatiwada, Harsha Kokel, Ibrahim Abdelaziz, Subhajit Chaudhury, Julian Dolby, Oktie Hassanzadeh, Zhenhan Huang, Tejaswini Pedapati, Horst Samulowitz, and Kavitha Srinivas. 2025. TabSketchFM: Sketch-Based Tabular Representation Learning for Data Discovery over Data Lakes....

  85. [86]

    Omar Khattab and Matei Zaharia. 2020. Colbert: Efficient and effective passage search via contextualized late interaction over bert. In Proceedings of the 43rd International ACM SIGIR conference on research and development in Information Retrieval (SIGIR) . 39–48

  86. [87]

    Diksha Khurana, Aditya Koli, Kiran Khatter, and Sukhdev Singh. 2023. Natural language processing: state of the art, current trends and challenges. Multimedia tools and applications 82, 3 (2023), 3713–3744

  87. [88]

    Jongik Kim. 2021. Boosting graph similarity search through pre-computation. In Proceedings of the 2021 International Conference on Management of Data (SIGMOD). 951–963

  88. [89]

    Gary King. 2007. An introduction to the dataverse network as an infrastructure for data sharing. Sociological Methods & Research 36, 2 (2007), 173–199

  89. [90]

    Meike Klettke, Uta Störl, and Stefanie Scherzinger. 2015. Schema extraction and structural outlier detection for JSON-based NoSQL data stores. In Datenbanksysteme für Business, Technologie und Web (BTW) . 425–444

  90. [91]

    Daniel Kocher and Nikolaus Augsten. 2019. A scalable index for top-k subtree similarity queries. In Proceedings of the 2019 International Conference on Management of Data (SIGMOD) . 1624–1641

  91. [92]

    Jinhyuk Lee, Zhuyun Dai, Sai Meher Karthik Duddu, Tao Lei, Iftekhar Naim, Ming-Wei Chang, and Vincent Zhao. 2023. Rethinking the role of token retrieval in multi-vector retrieval. Advances in Neural Information Processing Systems 36 (2023), 15384–15405

  92. [93]

    Aristotelis Leventidis, Martin Pekár Christensen, Matteo Lissandrini, Laura Di Rocco, Katja Hose, and Renée J Miller. 2024. A Large Scale Test Corpus for Semantic Table Search. In Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Informa...

  93. [94]

    Guoliang Li, Xuanhe Zhou, and Xinyang Zhao. 2024. LLM for Data Management. Proceedings of the VLDB Endowment 17, 12 (2024), 4213–4216

  94. [95]

    Minghan Li, Sheng-Chieh Lin, Xueguang Ma, and Jimmy Lin. 2023. SLIM: Sparsified late interaction for multi-vector retrieval with inverted indexes. In Proceedings of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR) . 1954–1959

  95. [96]

    Minghan Li, Sheng-Chieh Lin, Barlas Oguz, Asish Ghoshal, Jimmy Lin, Yashar Mehdad, Wen-tau Yih, and Xilun Chen. 2023. CITADEL: Conditional Token Interaction via Dynamic Lexical Routing for Efficient and Effective Multi-Vector Retrieval. InProceedings of the 61st Annual Meeting...

  96. [98]

    Peng Li, Yeye He, Dror Yashar, Weiwei Cui, Song Ge, Haidong Zhang, Danielle Rifinski Fainman, Dongmei Zhang, and Surajit Chaudhuri. 2024. Table-GPT: Table Fine-tuned GPT for Diverse Table Tasks. Proceedings of the ACM on Management of Data 2, 3 (2024), 176

  97. [100]

    Zhize Li, Haoyu Zhao, Boyue Li, and Yuejie Chi. 2022. SoteriaFL: A unified framework for private federated learning with communication compression. Advances in Neural Information Processing Systems 35 (2022), 4285–4300

  98. [101]

    Yiming Lin, Madelon Hulsebos, Ruiying Ma, Shreya Shankar, Sepanta Zeigham, Aditya G Parameswaran, and Eugene Wu. 2024. Towards Accurate and Efficient Document Analytics with Large Language Models. arXiv:2405.04674 (2024)

  99. [102]

    Yehuda Lindell. 2020. Secure multiparty computation. Commun. ACM 64, 1 (2020), 86–96

  100. [103]

    Jiabin Liu, Chengliang Chai, Yuyu Luo, Yin Lou, Jianhua Feng, and Nan Tang. 2022. Feature Augmentation with Reinforcement Learning. In 38th IEEE International Conference on Data Engineering (ICDE) . 3360–3372

  101. [104]

    Llano-Ríos, Mohamed Khalefa, and Antonio Badia

    Tomás F. Llano-Ríos, Mohamed Khalefa, and Antonio Badia. 2025. A JSON document algebra for query optimization. Information Systems 132 (2025), 102537

  102. [105]

    Robert WP Luk, Hong Va Leong, Tharam S Dillon, Alvin TS Chan, W Bruce Croft, and James Allan. 2002. A survey in indexing and searching XML documents. Journal of the American society for Information Science and Technology 53, 6 (2002), 415–437

  103. [106]

    Sean X Luo, Richard Axel, and LF Abbott. 2010. Generating sparse and selective third-order responses in the olfactory system of the fly.Proceedings of the National Academy of Sciences 107, 23 (2010), 10713–10718

  104. [107]

    Pingchuan Ma, Rui Ding, Shuai Wang, Shi Han, and Dongmei Zhang. 2023. InsightPilot: An LLM-empowered automated data exploration system. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing: System Demonstrations (EMNLP) . 346–352

  105. [108]

    Mandar Mitra and BB Chaudhuri. 2000. Information retrieval from documents: A survey. Information retrieval 2 (2000), 141–163

  106. [109]

    Takuya Mizokami, Savong Bou, and Toshiyuki Amagasa. 2024. Subtree Similarity Search Based on Structure and Text. In International Conference on Big Data Analytics and Knowledge Discovery . 72–87

  107. [110]

    Davide Mottin, Matteo Lissandrini, Yannis Velegrakis, and Themis Palpanas. 2014. Exemplar queries: Give me an example of what you need. Proceedings of the VLDB Endowment 7, 5 (2014), 365–376

  108. [111]

    Fatemeh Nargesian, Erkang Zhu, Ken Q Pu, and Renée J Miller. 2018. Table union search on open data. Proceedings of the VLDB Endowment 11, 7 (2018), 813–825

  109. [112]

    Karen Ka Yan Ng, Izuki Matsuba, and Peter Chengming Zhang. 2025. RAG in health care: a novel framework for improving communication and decision-making by addressing LLM limitations. NEJM AI 2, 1 (2025)

  110. [113]

    Tri Nguyen, Mir Rosenberg, Xia Song, Jianfeng Gao, Saurabh Tiwary, Rangan Majumder, and Li Deng. 2016. MS MARCO: A Human Generated MAchine Reading COmprehension Dataset. In Proceedings of the Workshop on Cognitive Computation: Integrating neural and symbolic approaches 2016 co...

  111. [114]

    Sarana Nutanong, Edwin H Jacox, and Hanan Samet. 2011. An incremental Hausdorff distance calculation algorithm. Proceedings of the VLDB Endowment 4, 8 (2011), 506–517

  112. [115]

    Alizée Pace, Jonathan Mallinson, Eric Malmi, Sebastian Krause, and Aliaksei Severyn. 2024. West-of-n: Synthetic preference generation for improved reward modeling. arXiv:2401.12086 (2024)

  113. [116]

    Liana Patel, Siddharth Jha, Carlos Guestrin, and Matei Zaharia. 2024. Lotus: Enabling semantic queries with llms over tables of unstructured and structured data. arXiv:2407.11418 (2024)

  114. [117]

    Norman W Paton, Jiaoyan Chen, and Zhenyu Wu. 2023. Dataset discovery and exploration: A survey. Comput. Surveys 56, 4 (2023), 1–37

  115. [118]

    Keqin Peng, Liang Ding, Yancheng Yuan, Xuebo Liu, Min Zhang, Yuanxin Ouyang, and Dacheng Tao. 2024. Revisiting Demonstration Selection Strategies in In-Context Learning. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (ACL) . 9090–9101

  116. [119]

    Yun Peng, Byron Choi, Tsz Nam Chan, and Jianliang Xu. 2022. Lan: Learning-based approximate k-nearest neighbor search in graph databases. In 38th IEEE international conference on data engineering (ICDE) . 2508–2521

  117. [120]

    Chengzhi Piao, Tingyang Xu, Xiangguo Sun, Yu Rong, Kangfei Zhao, and Hong Cheng. 2023. Computing graph edit distance via neural graph matching. Proceedings of the VLDB Endowment 16, 8 (2023), 1817–1829

  118. [121]

    Rishabh Ranjan, Siddharth Grover, Sourav Medya, Venkat Chakravarthy, Yogish Sabharwal, and Sayan Ranu. 2022. GREED: a neural framework for learning graph distance functions. In Proceedings of the 36th International Conference on Neural Information Processing Systems (NeurIPS) ...

  119. [122]

    Kaspar Riesen, Stefan Fankhauser, and Horst Bunke. 2007. Speeding up graph edit distance computation with a bipartite heuristic. In Mining and Learning with Graphs (MLG) . 21–24

  120. [123]

    Keshav Santhanam, Omar Khattab, Christopher Potts, and Matei Zaharia. 2022. PLAID: an efficient engine for late interaction retrieval. In Proceedings of the 31st ACM International Conference on Information and Knowledge Management (CIKM) . 1747–1756

  121. [124]

    Aécio Santos, Aline Bessa, Christopher Musco, and Juliana Freire. 2022. A sketch-based index for correlated dataset search. In38th IEEE International Conference on Data Engineering (ICDE) . 2928–2941

  122. [125]

    Han Shen, Pin-Yu Chen, Payel Das, and Tianyi Chen. 2024. Seal: Safety-enhanced aligned llm fine-tuning via bilevel data selection.arXiv:2410.07471 (2024)

  123. [126]

    Dharma Shukla, Shireesh Thota, Karthik Raman, Madhan Gajendran, Ankur Shah, Sergii Ziuzin, Krishnan Sundaram, Miguel Gonzalez Guajardo, Anna Wawrzyniak, Samer Boshra, et al. 2015. Schema-agnostic indexing with Azure DocumentDB. Proceedings of the VLDB Endowment 8, 12 (2015), 1668–1679

  124. [127]

    Arnab Sinha, Zhihong Shen, Yang Song, Hao Ma, Darrin Eide, Bo-June Hsu, and Kuansan Wang. 2015. An overview of microsoft academic service (mas) and applications. In Proceedings of the 24th International Conference on World Wide Web Companion . 243–246

  125. [128]

    Luca Soldaini, Rodney Kinney, Akshita Bhagia, Dustin Schwenk, David Atkinson, Russell Authur, Ben Bogin, Khyathi Chandu, Jennifer Dumas, Yanai Elazar, et al. 2024. Dolma: an Open Corpus of Three Trillion Tokens for Language Model Pretraining Research. In Proceedings of the 62n...

  126. [129]

    Ryota Tozuka, Hisashi Johno, Akitomo Amakawa, Junichi Sato, Mizuki Muto, Shoichiro Seki, Atsushi Komaba, and Hiroshi Onishi. 2025. Application of NotebookLM, a large language model with retrieval-augmented generation, for lung cancer staging. Japanese Journal of Radiology 43, ...

  127. [130]

    Davison, and Jeff Heflin

    Mohamed Trabelsi, Zhiyu Chen, Shuo Zhang, Brian D. Davison, and Jeff Heflin. 2022. StruBERT: Structure-aware BERT for Table Search and Matching. In The World Wide Web Conference (WWW) , Frédérique Laforest, Raphaël Troncy, Elena Simperl, Deepak Agarwal, Aristides Gionis, Ivan ...

  128. [131]

    Immanuel Trummer. 2023. Can Large Language Models Predict Data Correlations from Column Names? Proceedings of the VLDB Endowment 16, 13 (2023), 4310–4323

  129. [132]

    Cem Unsalan and Beril Sirmacek. 2012. Road network detection using probabilistic and graph theoretical methods. IEEE Transactions on Geoscience and Remote Sensing 50, 11 (2012), 4441–4453

  130. [133]

    Tin Vu and Ahmed Eldawy. 2018. R-Grove: growing a family of R-trees in the big-data forest. In Proceedings of the 26th ACM SIGSPATIAL International Conference on Advances in Geographic Information Systems (SIGSPATIAL) . 532–535

  131. [134]

    Jianhua Wang, Jianye Yang, and Wenjie Zhang. 2021. Top-k Tree Similarity Join. In Proceedings of the 30th ACM International Conference on Information and Knowledge Management (CIKM) . 1939–1948

  132. [135]

    Rui Wang, Devin Gibson, Kirk Rodrigues, Yu Luo, Yun Zhang, Kaibo Wang, Yupeng Fu, Ting Chen, and Ding Yuan. 2024. {𝜇Slope}: High Compression and Fast Search on{Semi-Structured} Logs. In 18th USENIX Symposium on Operating Systems Design and Implementation (OSDI) . 529–544

  133. [136]

    Shane Culpepper, Timos Sellis, Mark Sanderson, and Xiaolin Qin

    Sheng Wang, Zhifeng Bao, J. Shane Culpepper, Timos Sellis, Mark Sanderson, and Xiaolin Qin. 2017. Answering Top-k Exemplar Trajectory Queries. In 33rd IEEE International Conference on Data Engineering (ICDE) . 597–608

  134. [137]

    Sheng Wang, Mingzhao Li, Yipeng Zhang, Zhifeng Bao, David Alexander Tedjopurnomo, and Xiaolin Qin. 2018. Trip planning by an integrated search paradigm. In Proceedings of the 2018 International Conference on Management of Data (SIGMOD) . 1673–1676

  135. [139]

    Tianshi Wang, Fengling Li, Lei Zhu, Jingjing Li, Zheng Zhang, and Heng Tao Shen. 2024. Cross-Modal Retrieval: A Systematic Review of Methods and Future Directions. Proc. IEEE 112, 11 (2024), 1716–1754

  136. [140]

    Yunling Wang, Jianfeng Wang, and Xiaofeng Chen. 2016. Secure searchable encryption: a survey. Journal of communications and information networks 1 (2016), 52–65

  137. [141]

    Zheng Wang, Shu Xian Teo, Jun Jie Chew, and Wei Shi. 2025. Instructrag: Leveraging retrieval-augmented generation on instruction graphs for llm-based task planning. In Proceedings of the 48th International ACM SIGIR Conference on Research and Development in Information Retriev...

  138. [142]

    Duncan J Watts, Peter Sheridan Dodds, and Mark EJ Newman. 2002. Identity and search in social networks. science 296, 5571 (2002), 1302–1305

  139. [143]

    Junde Wu, Jiayuan Zhu, Yunli Qi, Jingkun Chen, Min Xu, Filippo Menolascina, Yueming Jin, and Vicente Grau. 2025. Medical Graph RAG: Evidence-based Medical Large Language Model via Graph Retrieval-Augmented Generation. In Proceedings of the 63rd Annual Meeting of the Associatio...

  140. [144]

    Shiguang Wu, Wenda Wei, Mengqi Zhang, Zhumin Chen, Jun Ma, Zhaochun Ren, Maarten de Rijke, and Pengjie Ren. 2024. Generative retrieval as multi-vector dense retrieval. In Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retr...

  141. [145]

    Shirley Wu, Shiyu Zhao, Michihiro Yasunaga, Kexin Huang, Kaidi Cao, Qian Huang, Vassilis N Ioannidis, Karthik Subbian, James Zou, and Jure Leskovec. 2024. Stark: Benchmarking llm retrieval on textual and relational knowledge bases. In Proceedings of the 38th International Conf...

  142. [146]

    Mengzhou Xia, Sadhika Malladi, Suchin Gururangan, Sanjeev Arora, and Danqi Chen. 2024. LESS: selecting influential data for targeted instruction tuning. In Proceedings of the 41st International Conference on Machine Learning (ICLR) . 54104–54132

  143. [147]

    Yang Xiao, Jinlan Fu, Weizhe Yuan, Vijay Viswanathan, Zhoumianze Liu, Yixin Liu, Graham Neubig, and Pengfei Liu. 2022. DataLab: A Platform for Data Analysis and Intervention. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (ACL) . 182–195

  144. [148]

    Anjie Xu, Ruiqing Ding, and Leye Wang. 2025. ChatPD: An LLM-driven Paper-Dataset Networking System. arXiv:2505.22349 (2025)

  145. [149]

    Derong Xu, Xinhang Li, Ziheng Zhang, Zhenxi Lin, Zhihong Zhu, Zhi Zheng, Xian Wu, Xiangyu Zhao, Tong Xu, and Enhong Chen. 2025. Harnessing large language models for knowledge graph question answering via adaptive multi-aspect retrieval-augmentation. In Proceedings of the AAAI ...

  146. [150]

    Mohamed Yakout, Kris Ganjam, Kaushik Chakrabarti, and Surajit Chaudhuri. 2012. Infogather: entity augmentation and attribute discovery by holistic matching with web tables. In Proceedings of the 2012 International Conference on Management of Data (SIGMOD) . 97–108

  147. [151]

    Mengyi Yan, Yaoshu Wang, Kehan Pang, Min Xie, and Jianxin Li. 2024. Efficient mixture of experts based on large language models for low-resource data preprocessing. In Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD) . 3690–3701

  148. [152]

    Lei Yang and Lei Zou. 2021. Noah: Neural-optimized A* search algorithm for graph edit distance computation. In37th IEEE International Conference on Data Engineering (ICDE) . 576–587

  149. [153]

    Wenzhe Yang, Shixun Huang, Sheng Wang, and Zhiyong Peng. 2024. Budgeted Spatial Data Acquisition: When Coverage and Connectivity Matter. arXiv preprint arXiv:2412.04853 (2024)

  150. [157]

    Antonio Jimeno Yepes, Yao You, Jan Milczek, Sebastian Laverde, and Renyu Li. 2024. Financial report chunking for effective retrieval augmented generation. arXiv:2402.05131 (2024)

  151. [158]

    Cafarella, Babak Salimi, and Anna Zeng

    Brit Youngmann, Michael J. Cafarella, Babak Salimi, and Anna Zeng. 2023. Causal Data Integration. Proceedings of the VLDB Endowment 16, 10 (2023), 2659–2665

  152. [159]

    Hao Yu, Aoran Gan, Kai Zhang, Shiwei Tong, Qi Liu, and Zhaofeng Liu. 2024. Evaluation of retrieval-augmented generation: A survey. In CCF Conference on Big Data . 102–120

  153. [160]

    Zichun Yu, Spandan Das, and Chenyan Xiong. 2024. Mates: Model-aware data selection for efficient pretraining with data influence models. In Proceedings of the 38th International Conference on Neural Information Processing Systems (NeurIPS) . 108735–108759

  154. [161]

    Haitao Yuan, Sai Wang, Zhifeng Bao, and Shangguang Wang. 2023. Automatic road extraction with multi-source data revisited: completeness, smoothness and discrimination. Proceedings of the VLDB Endowment 16, 11 (2023), 3004–3017

  155. [162]

    Joohyung Yun, Byungchul Tak, and Wook-Shin Han. 2024. ReCG: Bottom-up JSON Schema Discovery Using a Repetitive Cluster-and-Generalize Framework. Proceedings of the VLDB Endowment 17, 11 (2024), 3538–3550

  156. [163]

    Yiming Zeng, Yixuan Lin, Yuanyuan Yang, and Ji Liu. 2021. Differentially private federated temporal difference learning. IEEE Transactions on Parallel and Distributed Systems 33, 11 (2021), 2714–2726

  157. [164]

    Zhiping Zeng, Anthony KH Tung, Jianyong Wang, Jianhua Feng, and Lizhu Zhou. 2009. Comparing stars: On approximating graph edit distance. Proceedings of the VLDB Endowment 2, 1 (2009), 25–36

  158. [165]

    Boyu Zhang, Hongyang Yang, Tianyu Zhou, Muhammad Ali Babar, and Xiao-Yang Liu. 2023. Enhancing financial sentiment analysis via retrieval augmented large language models. In Proceedings of the fourth ACM international conference on AI in finance (ICAIF) . 349–356

  159. [166]

    Haochen Zhang, Yuyang Dong, Chuan Xiao, and Masafumi Oyamada. 2023. Jellyfish: A large language model for data preprocessing.arXiv:2312.01678 (2023)

  160. [167]

    Shuo Zhang and Krisztian Balog. 2020. Web table extraction, retrieval, and augmentation: A survey. ACM Transactions on Intelligent Systems and Technology 11, 2 (2020), 1–35

  161. [169]

    Zeyu Zhang, Paul Groth, Iacer Calixto, and Sebastian Schelter. 2024. Directions Towards Efficient and Automated Data Wrangling with Large Language Models. In 2024 IEEE 40th International Conference on Data Engineering Workshops . 301–304

  162. [170]

    Chenyu Zhao, Yunjiang Jiang, Yiming Qiu, Han Zhang, and Wen-Yun Yang. 2023. Differentiable retrieval augmentation via generative language modeling for E-commerce query intent classification. In Proceedings of the 32nd ACM International Conference on Information and Knowledge M...

  163. [171]

    Fuheng Zhao, Shaleen Deep, Fotis Psallidas, Avrilia Floratou, Divyakant Agrawal, and Amr El Abbadi. 2024. Sphinteract: Resolving Ambiguities in NL2SQL through User Interaction. Proceedings of the VLDB Endowment 18, 4 (2024), 1145–1158

  164. [172]

    Penghao Zhao, Hailin Zhang, Qinhan Yu, Zhengren Wang, Yunteng Geng, Fangcheng Fu, Ling Yang, Wentao Zhang, Jie Jiang, and Bin Cui. 2024. Retrieval-augmented generation for ai-generated content: A survey. arXiv:2402.19473 (2024)

  165. [173]

    Xiang Zhao, Chuan Xiao, Xuemin Lin, Wenjie Zhang, and Yang Wang. 2018. Efficient structure similarity searches: a partition-based approach. The VLDB Journal 27, 1 (2018), 53–78

  166. [174]

    Weiguo Zheng, Lei Zou, Xiang Lian, Dong Wang, and Dongyan Zhao. 2014. Efficient graph similarity search over large graph databases. IEEE Transactions on Knowledge and Data Engineering 27, 4 (2014), 964–978

  167. [175]

    Yuanhao Zhong, Yuhao Deng, Chengliang Chai, Ruixin Gu, Ye Yuan, Guoren Wang, and Lei Cao. 2025. Doctopus: A System for Budget-aware Structural Data Extraction from Unstructured Documents. In Companion of the 2025 International Conference on Management of Data . 275–278

  168. [176]

    Erkang Zhu, Dong Deng, Fatemeh Nargesian, and Renée J Miller. 2019. Josie: Overlap set similarity search for finding joinable tables in data lakes. In Proceedings of the 2019 International Conference on Management of Data (SIGMOD) . 847–864

  169. [177]

    Erkang Zhu, Fatemeh Nargesian, Ken Q Pu, and Renée J Miller. 2016. LSH ensemble: internet-scale domain search. Proceedings of the VLDB Endowment 9, 12 (2016), 1185–1196

  170. [178]

    Yuanyuan Zhu, Lu Qin, Jeffrey Xu Yu, and Hong Cheng. 2019. Answering Top-𝑘 Graph Similarity Queries in Graph Databases. IEEE Transactions on Knowledge and Data Engineering 32, 8 (2019), 1459–1474

  171. [179]

    Lei Zou, Jinghui Mo, Lei Chen, M Tamer Özsu, and Dongyan Zhao. 2011. gStore: answering SPARQL queries via subgraph matching. Proceedings of the VLDB Endowment 4, 8 (2011), 482–493

Pith tools

Reviewed August 5, 2026 · model on record in the stance chip above.