REVIEW 2 major objections 8 minor 300 references
A Survey on Retrieval And Structuring Augmented Generation with Large Language Models
T0 review · 2 major / 8 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read The survey claims that combining retrieval with structured knowledge—taxonomies, knowledge graphs, and tables—grounds LLM answers more reliably than text-only RAG, curbing hallucination and enabling multi-hop domain reasoning.
desk verdict A serviceable, well-organized survey that usefully maps retrieval, structuring, and LLM integration, but the 'RAS is more powerful' paradigm claim is asserted rather than demonstrated. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the RAS loop: a taxonomy-enhanced retriever selects thematically and semantically relevant documents; those documents are structured into a query-specific knowledge-graph subgraph; the subgraph augments LLM generation; and if the answer is incomplete, the LLM issues a focused subquery conditioned on the knowledge graph, repeating until sufficient knowledge is gathered. The load-bearing components are the structured representations themselves—corpus taxonomies that guide search and knowledge graphs that provide atomic, relationally connected facts—plus the feedback cycle that lets structure and retrieval refine each other.
What would settle it
Run a controlled evaluation on a fixed corpus and a fixed LLM, comparing plain text-chunk RAG, taxonomy-guided retrieval plus RAG, and knowledge-graph-constructed retrieval plus structured generation on the same factual and multi-hop question sets; if the structure-enhanced systems do not beat plain RAG on answer accuracy while controlling retrieval budget, the central claim fails.
Extended reading notes
Core claim
The central claim is that RAS is a more powerful paradigm than RAG alone. By transforming unstructured text into organized representations such as taxonomies, hierarchies, and knowledge graphs, and then using those structures to guide retrieval and verify generation, RAS systems can reduce hallucinations, access current knowledge, and handle specialized domains more effectively than systems that retrieve raw text passages. The survey supports this claim by reviewing evidence that taxonomy-guided retrieval improves domain-specific search without labeled data, that knowledge-graph-based retrieval finds implicit multi-hop connections, and that structure-enhanced generation grounds answers in explicit facts and community-level summaries.
Load-bearing premise
The case for RAS rests on the survey's summaries of dozens of cited systems; if those summaries are inaccurate or the reported gains do not replicate, the conclusion that structure-enhanced retrieval and generation beat unstructured RAG is not established.
Editorial extensions
If this is right
- Domain-specific search can improve without large labeled datasets by using taxonomies to filter and expand the retrieval space.
- Knowledge-graph construction plus graph traversal, such as personalized PageRank over extracted triples, lets systems answer queries that require implicit connections between passages.
- Summarizing graph communities before generation equips LLMs to answer global questions that span a whole corpus, not just local passages.
- Structured representations can be embedded into LLM parameters or fed as soft tokens, so structure improves generation without modifying the LLM at inference time.
- RAS opens a design space of iterative query feedback, where the LLM issues subqueries conditioned on an evolving knowledge graph until the answer is grounded.
Reading between the lines
- If the RAS loop is as effective as the survey suggests, the cost of building and maintaining structures is best amortized across repeated use: an organization that queries a domain many times should recover the structuring overhead, which plain RAG never does.
- The feedback-loop design implies a testable extension: make the number of retrieval rounds an adaptive budget controlled by a verification signal rather than a fixed limit.
- The same structuring machinery could apply to non-text modalities, turning images and audio into entity-relation graphs before retrieval, which would extend the survey's stated multimodal direction into a concrete architecture.
- The survey's taxonomy-versus-knowledge-graph dichotomy suggests a comparative question it does not answer: when does a lightweight taxonomy beat a full knowledge graph, and vice versa, per query type and corpus scale.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This survey proposes 'Retrieval And Structuring (RAS) Augmented Generation' as a paradigm that extends RAG by combining dynamic retrieval with structured knowledge representations (taxonomies, knowledge graphs, tables, databases). It reviews classical and neural retrieval methods (sparse, dense, hybrid, generative), text structuring techniques (taxonomy construction, hierarchical classification, information extraction), knowledge structuring (KG construction, database population, tabular organization), and structure-enhanced retrieval and generation, and it closes with technical challenges and future directions. It also presents a conceptual iterative RAS architecture in Figure 1.
Significance. If the paradigm claim is taken as a research agenda rather than a demonstrated fact, this survey is a useful and reasonably comprehensive synthesis. Its strengths include broad coverage of sparse/dense/hybrid retrieval, taxonomy construction, hierarchical classification, information extraction, KG construction, and structure-enhanced generation; a clear organization that separates retrieval, structuring, and integration; and an explicit list of technical challenges and research opportunities (multimodal, cross-lingual, interactive, personalized). The survey is internally consistent and generally accurate in its descriptions of individual papers, and it brings together literature that is otherwise scattered across IR, data mining, and NLP venues. The main deficit is the gap between the strength of the central 'more powerful paradigm' claim and the qualitative nature of the supporting evidence.
major comments (2)
- [Section 1, Sections 5.1–5.2] The central claim that RAS 'has emerged as a more powerful paradigm' (Section 1) is supported almost entirely by qualitative summaries in Sections 5.1 and 5.2. For example, ToTER is said to 'allow improved domain-specific search without requiring labeled data' and TaxoIndex to 'achiev[e] strong performance even with limited training data' (Section 5.1); GraphRAG is described as 'effectively handling queries that require global corpus understanding' and KARE as 'enabl[ing] LLMs to generate more accurate predictions' (Section 5.2). No metrics, effect sizes, or benchmark names are reported for these systems, and the survey contains no head-to-head comparison of RAS systems against text-only RAG on a common benchmark (e.g., BEIR, Natural Questions, HotpotQA). Because the paradigm claim rests on these summaries, the reader cannot determine whether the asserted superiority is a robust finding or an artifact of benchmark-specific or self-authored results. The revision should add a table reporting the quantitative results from the original papers, or explicitly qualify the central claim to 'a promising direction' rather than 'a more powerful paradigm'.
- [Section 5.1, Figure 1] The method depicted in Figure 1 is described as 'a promising RAS method' and then elaborated as an iterative taxonomy- and KG-guided retrieval loop. This is a proposed, unevaluated architecture, yet it is placed in the body of the survey alongside established systems and appears to be offered as part of the evidence base for the RAS paradigm. Presenting a non-evaluated system as supporting material conflates conjecture with evidence for the central claim. The figure should be explicitly labeled as an illustrative, non-evaluated proposal, and ideally moved to Section 6 (Future Directions) or accompanied by a statement that it is a conceptual framework rather than a validated method.
minor comments (8)
- [References [46], [47]] References [46] and [47] are the same paper ('Automated construction of theme-specific knowledge graphs' by Ding et al., 2024) with two different preprint identifiers; these should be merged into a single citation.
- [References [278]] Reference [278] lists the authors as 'Sizhe Zhou Heng Ji Jiawei Ha Yizhu Jiao, Sha Li'; the correct author list is 'Yizhu Jiao, Sha Li, Sizhe Zhou, Heng Ji, Jiawei Han'.
- [Abstract] The abstract has subject-verb agreement errors: '(2) explore' and '(3) investigate' should be '(2) explores' and '(3) investigates' to agree with the subject 'This survey'.
- [Section 2.2.1] The 'Naive RAG' paragraph cites [159] (Ma et al., query rewriting), which is not the origin of the retrieve-read framework; the canonical citation is Lewis et al. [124]. Similarly, 'Modular RAG' cites [251] (KnowledGPT), whereas the modular RAG concept is usually attributed to the RAG survey [64]. Please correct these citations.
- [Section 6.1.1 and 6.2.2] In Section 6.1.1, the claim about adaptive retrieval strategies cites [2] (GPT-4 technical report), and in Section 6.2.2, the claim about transfer learning for low-resource languages also cites [2]. These are not appropriate sources for these specific claims; please replace with relevant references (e.g., adaptive RAG or multilingual RAG works).
- [Figure 1] The caption 'An abstractive example of RAS paradigm' is too terse. Please expand it to describe the components (taxonomy-enhanced retriever, subgraph, query-specific KG, LLM feedback loop) and state explicitly that this is a conceptual illustration.
- [Section 3.2.3] The notation 's3[96]' is confusing; consider writing 's³ [96]' and specifying the actual training data reduction rather than the vague phrase 'orders of magnitude less training data'.
- [Section 2.1.1] In the sentence 'Decoder-only Models, like GPT [2] and DeepSeek [42]', reference [2] is the GPT-4 technical report, not the original GPT paper; cite the original GPT or GPT-3 paper as appropriate.
Circularity Check
No circular derivation found; the survey's central claim is a literature synthesis with notable self-citation but no reduction by construction.
full rationale
This is a survey with no derivation chain, no fitted parameters, and no predictions generated from a model, so there is no input-to-output reduction of the kind that defines circularity. The central claim that RAS 'has emerged as a more powerful paradigm' (Section 1) is a literature synthesis supported by qualitative summaries of peer-reviewed systems in Sections 4 and 5, not by a derivation from a definition. A substantial share of the load-bearing examples (ToTER, TaxoIndex, TELEClass, KG-FIT, KARE, DeepRetrieval, s3) come from the authors' own group, and Figure 1 is a proposed, unevaluated architecture; this raises the burden of independent verification and is a legitimate correctness risk, as the skeptic notes. However, under the review rules, self-citation is not circularity unless the load-bearing argument reduces to an unverified self-citation. Here the cited systems are externally published and falsifiable, and the survey also draws on many independent groups (HippoRAG, GraphRAG, ToG, RoG, KG-RAG, and others). No equation is shown to equal its own input, and no fitted parameter is renamed as a prediction. The modest score of 1 reflects the self-citation and unquantified-summary concern, not a circular derivation.
Assumptions & free parameters
assumptions (3)
- domain assumption The cited papers and their characterizations are accurate.
- domain assumption LLMs exhibit the motivating limitations (hallucination, outdated knowledge, limited domain expertise).
- domain assumption The RAS framing is a useful way to organize the field.
Cite this review
Pith. "Pith review of A Survey on Retrieval And Structuring Augmented Generation with Large Language Models." pith.science (2026). https://pith.science/paper/NQ5DLITE
@misc{pith2026250910697,
author = {Pith},
title = {Pith review of: A Survey on Retrieval And Structuring Augmented Generation with Large Language Models},
year = {2026},
howpublished = {\url{https://pith.science/paper/NQ5DLITE}},
note = {Machine review of arXiv:2509.10697}
}
read the original abstract
Large Language Models (LLMs) have revolutionized natural language processing with their remarkable capabilities in text generation and reasoning. However, these models face critical challenges when deployed in real-world applications, including hallucination generation, outdated knowledge, and limited domain expertise. Retrieval And Structuring (RAS) Augmented Generation addresses these limitations by integrating dynamic information retrieval with structured knowledge representations. This survey (1) examines retrieval mechanisms including sparse, dense, and hybrid approaches for accessing external knowledge; (2) explore text structuring techniques such as taxonomy construction, hierarchical classification, and information extraction that transform unstructured text into organized representations; and (3) investigate how these structured representations integrate with LLMs through prompt-based methods, reasoning frameworks, and knowledge embedding techniques. It also identifies technical challenges in retrieval efficiency, structure quality, and knowledge integration, while highlighting research opportunities in multimodal retrieval, cross-lingual structures, and interactive systems. This comprehensive overview provides researchers and practitioners with insights into RAS methods, applications, and future directions.
Figures
Reference graph
Works this paper leans on
-
[46]
Linyi Ding, Sizhe Zhou, Jinfeng Xiao, and Jiawei Han. 2024. Automated con- struction of theme-specific knowledge graphs.arXiv preprint arXiv:2404.19146 (2024)
arXiv 2024
-
[47]
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. Bert: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 conference of the North American chapter of the association for computational linguistics: human language technologies, volume 1 (long and short papers). 4171–4186
2019
-
[1]
Zahra Abbasiantaeb and Saeedeh Momtazi. 2021. Text-based question answering from information retrieval and deep neural network perspectives: A survey. Wiley Interdisciplinary Reviews: Data Mining and Knowledge Discovery11, 6 (2021), e1412
2021
-
[2]
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Floren- cia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. 2023. Gpt-4 technical report.arXiv preprint arXiv:2303.08774 (2023)
arXiv 2023
-
[3]
Anirudh Ajith, Mengzhou Xia, Alexis Chevalier, Tanya Goyal, Danqi Chen, and Tianyu Gao. 2024. Litsearch: A retrieval benchmark for scientific literature search.arXiv preprint arXiv:2407.18940(2024)
arXiv 2024
-
[4]
Negar Arabzadeh, Xinyi Yan, and Charles LA Clarke. 2021. Predicting effi- ciency/effectiveness trade-offs for dense vs. sparse retrieval strategy selection. InProceedings of the 30th ACM International Conference on Information & Knowl- edge Management. 2862–2866
2021
-
[5]
Simran Arora, Brandon Yang, Sabri Eyuboglu, Avanika Narayan, Andrew Ho- jel, Immanuel Trummer, and Christopher Ré. 2023. Language Models Enable Simple Systems for Generating Structured Views of Heterogeneous Data Lakes. Proceedings of the VLDB Endowment17, 2 (2023), 92–105
2023
-
[6]
Akari Asai, Zeqiu Wu, Yizhong Wang, Avirup Sil, and Hannaneh Hajishirzi. 2023. Self-rag: Learning to retrieve, generate, and critique through self-reflection. In The Twelfth International Conference on Learning Representations
2023
Show all 300 references
-
[7]
Hiteshwar Kumar Azad and Akshay Deepak. 2019. Query expansion techniques for information retrieval: A survey.Information Processing & Management56, 5 (2019), 1698–1735. doi:10.1016/j.ipm.2019.05.009
2019 doi
-
[8]
Sonia Badene, Kate Thompson, Jean-Pierre Lorré, and Nicholas Asher. 2019. Data Programming for Learning Discourse Structure. InACL
2019
-
[9]
Siddhartha Banerjee, Cem Akkaya, Francisco Perez-Sorrosal, and Kostas Tsiout- siouliklis. 2019. Hierarchical Transfer Learning for Multi-label Text Classifica- tion. InACL
2019
-
[10]
Petr Baudiš and Jan Šediv `y. 2015. Modeling of the question answering task in the yodaqa system. InExperimental IR Meets Multilinguality, Multimodality, and Interaction: 6th International Conference of the CLEF Association, CLEF’15, Toulouse, France, September 8-11, 2015, Pro...
2015
-
[11]
Jonathan Berant, Andrew Chou, Roy Frostig, and Percy Liang. 2013. Semantic parsing on freebase from question-answer pairs. InProceedings of the 2013 conference on empirical methods in natural language processing. 1533–1544
2013
-
[12]
Maciej Besta, Nils Blach, Ales Kubicek, Robert Gerstenberger, Michal Podstawski, Lukas Gianinazzi, Joanna Gajda, Tomasz Lehmann, Hubert Niewiadomski, Piotr Nyczyk, and Torsten Hoefler. 2024. Graph of Thoughts: Solving Elaborate Problems with Large Language Models.Proceedings o...
2024 doi
-
[13]
Michele Bevilacqua, Giuseppe Ottaviano, Patrick Lewis, Scott Yih, Sebastian Riedel, and Fabio Petroni. 2022. Autoregressive search engines: Generating substrings as document identifiers.Advances in NeurIPS35 (2022), 31668–31683
2022
-
[14]
Vladimir Blagojevi. 2023. Enhancing rag pipelines in haystack: Introducing diversityranker and lostinthemiddleranker
2023
-
[15]
Luiz Bonifacio, Hugo Abonizio, Marzieh Fadaee, and Rodrigo Nogueira. 2022. Inpars: Data augmentation for information retrieval using large language models. arXiv preprint arXiv:2202.05144(2022)
2022 arXiv
-
[16]
Sebastian Borgeaud, Arthur Mensch, Jordan Hoffmann, Trevor Cai, Eliza Ruther- ford, Katie Millican, George Bm Van Den Driessche, Jean-Baptiste Lespiau, Bog- dan Damoc, Aidan Clark, et al. 2022. Improving language models by retrieving from trillions of tokens. InInternational c...
2022
-
[17]
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020. Language models are few-shot learners.Advances in NeurIPS 33 (2020), 1877–1901
2020
-
[18]
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffr...
2020
-
[19]
Claudio Carpineto and Giovanni Romano. 2012. A Survey of Automatic Query Expansion in Information Retrieval.ACM Comput. Surv.44, 1, Article 1 (Jan. 2012), 50 pages. doi:10.1145/2071389.2071390
2012
-
[20]
Oishik Chatterjee, Ganesh Ramakrishnan, and Sunita Sarawagi. 2018. Data Programming using Continuous and Quality-Guided Labeling Functions. In AAAI
2018
-
[21]
Aditi Chaudhary, Karthik Raman, and Michael Bendersky. 2023. It’s All Relative!– A Synthetic Query Generation Approach for Improving Zero-Shot Relevance Prediction.arXiv preprint arXiv:2311.07930(2023)
2023 arXiv
-
[22]
Aditi Chaudhary, Jiateng Xie, Zaid Sheikh, Graham Neubig, and Jaime Carbonell
-
[23]
Haibin Chen, Qianli Ma, Zhenxi Lin, and Jiangyue Yan. 2021. Hierarchy-aware Label Semantics Matching Network for Hierarchical Text Classification. In ACL-IJCNLP
2021
-
[24]
Jiawei Chen, Hongyu Lin, Xianpei Han, and Le Sun. 2024. Benchmarking large language models in retrieval-augmented generation. InProceedings of the AAAI Conference on Artificial Intelligence, Vol. 38. 17754–17762
2024
-
[25]
Jiawei Chen, Qing Liu, Hongyu Lin, Xianpei Han, and Le Sun. 2022. Few-shot Named Entity Recognition with Self-describing Networks. InACL. 5711–5722
2022
-
[26]
Jiao Chen, Luyi Ma, Xiaohan Li, Nikhil Thakurdesai, Jianpeng Xu, Jason HD Cho, Kaushiki Nag, Evren Korpeoglu, Sushant Kumar, and Kannan Achan. 2023. Knowledge graph completion models are few-shot learners: An empirical study of relation labeling in e-commerce with llms.arXiv p...
2023 arXiv
-
[27]
Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton. 2020. A simple framework for contrastive learning of visual representations. InInter- national conference on machine learning. PmLR, 1597–1607
2020
-
[28]
Tong Chen, Hongwei Wang, Sihao Chen, Wenhao Yu, Kaixin Ma, Xinran Zhao, Hongming Zhang, and Dong Yu. 2024. Dense x retrieval: What retrieval granu- larity should we use?. InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 15159–15177
2024
-
[29]
Xiang Chen, Ningyu Zhang, Xin Xie, Shumin Deng, Yunzhi Yao, Chuanqi Tan, Fei Huang, Luo Si, and Huajun Chen. 2022. KnowPrompt: Knowledge-aware Prompt-tuning with Synergistic Optimization for Relation Extraction. InWWW. 2778–2788
2022
-
[30]
Xin Cheng, Di Luo, Xiuying Chen, Lemao Liu, Dongyan Zhao, and Rui Yan
-
[31]
Yiruo Cheng, Kelong Mao, and Zhicheng Dou. 2024. Interpreting Conversational Dense Retrieval by Rewriting-Enhanced Inversion of Session Embedding.arXiv preprint arXiv:2402.12774(2024)
2024 arXiv
-
[32]
Yew Ken Chia, Lidong Bing, Soujanya Poria, and Luo Si. 2022. RelationPrompt: Leveraging Prompts to Generate Synthetic Data for Zero-Shot Relation Triplet Extraction. InFindings of the Association for Computational Linguistics: ACL 2022, Dublin, Ireland, May 22-27, 2022, Smaran...
2022 doi
-
[33]
Nadezhda Chirkova, David Rau, Hervé Déjean, Thibault Formal, Stéphane Clin- chant, and Vassilina Nikoulina. 2024. Retrieval-augmented generation in multi- lingual settings.arXiv preprint arXiv:2407.01463(2024)
2024 arXiv
-
[34]
Eunsol Choi, Omer Levy, Yejin Choi, and Luke Zettlemoyer. 2018. Ultra-Fine Entity Typing. InACL. 87–96
2018
-
[35]
Yung-Sung Chuang, Wei Fang, Shang-Wen Li, Wen-tau Yih, and James Glass
-
[36]
Robert Churchill, Lisa Singh, Rebecca Ryan, and Pamela Davis-Kean. 2022. A Guided Topic-Noise Model for Short Texts. InWWW’22
2022
-
[37]
John Dagdelen, Alexander Dunn, Sanghoon Lee, Nicholas Walker, Andrew S Rosen, Gerbrand Ceder, Kristin A Persson, and Anubhav Jain. 2024. Structured information extraction from scientific text with large language models.Nature Communications15, 1 (2024), 1418
2024
-
[38]
Expand, rerank, and retrieve: Query reranking for open-domain question answering.arXiv preprint arXiv:2305.17080(2023)
2023 arXiv
-
[39]
Zhuyun Dai and Jamie Callan. 2019. Context-aware sentence/passage term importance estimation for first stage retrieval.arXiv preprint arXiv:1910.10687 (2019)
2019 arXiv
-
[40]
Zhuyun Dai, Vincent Y Zhao, Ji Ma, Yi Luan, Jianmo Ni, Jing Lu, Anton Bakalov, Kelvin Guu, Keith B Hall, and Ming-Wei Chang. 2022. Promptagator: Few-shot dense retrieval from 8 examples.arXiv preprint arXiv:2209.11755(2022)
2022 arXiv
-
[41]
Hongliang Dai, Yangqiu Song, and Haixun Wang. 2021. Ultra-Fine Entity Typing with Weak Supervision from a Masked Language Model. InACL’21. 1790–1799. KDD ’25, August 3–7, 2025, Toronto, ON, Canada Pengcheng Jiang et al
2021
- [42]
- [43]
-
[44]
Nicola De Cao, Gautier Izacard, Sebastian Riedel, and Fabio Petroni. 2020. Au- toregressive entity retrieval.arXiv preprint arXiv:2010.00904(2020)
2020 arXiv
-
[45]
Dieng, Francisco J
Adji B. Dieng, Francisco J. R. Ruiz, and David M. Blei. 2020. Topic Modeling in Embedding Spaces.TACL(2020)
2020
-
[48]
Jiannan Dong, Tao Wang, Yicheng Zheng, Yijun Wu, Jiefu Wei, Ming Ding, Xipeng Qiu, and Jie Tang. 2023. Self-Refine: Iterative Refinement with GPT-4. arXiv preprint arXiv:2303.17651(2023). https://arxiv.org/abs/2303.17651
2023 arXiv
-
[49]
2010.Named entity recognition and resolution in legal text
Christopher Dozier, Ravikumar Kondadadi, Marc Light, Arun Vachher, Sriharsha Veeramachaneni, and Ramdev Wudali. 2010.Named entity recognition and resolution in legal text. Springer
2010
- [50]
-
[51]
Fartash Faghri, David J Fleet, Jamie Ryan Kiros, and Sanja Fidler. 2017. Vse++: Improving visual-semantic embeddings with hard negatives.arXiv preprint arXiv:1707.05612(2017)
2017 arXiv
-
[52]
1995.A survey of information retrieval and filtering methods
Christos Faloutsos and Douglas W Oard. 1995.A survey of information retrieval and filtering methods. Citeseer
1995
-
[53]
Darren Edge, Ha Trinh, Newman Cheng, Joshua Bradley, Alex Chao, Apurva Mody, Steven Truitt, and Jonathan Larson. 2024. From local to global: A graph rag approach to query-focused summarization.arXiv preprint arXiv:2404.16130 (2024)
2024 arXiv
-
[54]
Yan Fang, Jingtao Zhan, Qingyao Ai, Jiaxin Mao, Weihang Su, Jia Chen, and Yiqun Liu. 2024. Scaling laws for dense retrieval. InProceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval. 1339–1349
2024
-
[55]
Bahare Fatemi, Jonathan Halcrow, and Bryan Perozzi. 2024. Talk like a Graph: Encoding Graphs for Large Language Models. InThe Twelfth International Conference on Learning Representations. https://openreview.net/forum?id= IuXR1CCrSi
2024
-
[56]
Wenqi Fan, Yujuan Ding, Liangbo Ning, Shijie Wang, Hengyun Li, Dawei Yin, Tat-Seng Chua, and Qing Li. 2024. A survey on rag meeting llms: Towards retrieval-augmented large language models. InProceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining. ...
2024
-
[57]
Matthias Fey, Weihua Hu, Kexin Huang, Jan Eric Lenssen, Rishabh Ranjan, Joshua Robinson, Rex Ying, Jiaxuan You, and Jure Leskovec. 2023. Relational deep learning: Graph representation learning on relational databases.arXiv preprint arXiv:2312.04615(2023)
2023 arXiv
-
[58]
Thibault Formal, Carlos Lassance, Benjamin Piwowarski, and Stéphane Clin- chant. 2021. SPLADE v2: Sparse lexical and expansion model for information retrieval.arXiv preprint arXiv:2109.10086(2021)
2021 arXiv
-
[59]
Yanlin Feng, Xinyue Chen, Bill Yuchen Lin, Peifeng Wang, Jun Yan, and Xi- ang Ren. 2020. Scalable multi-hop relational reasoning for knowledge-aware question answering.arXiv preprint arXiv:2005.00646(2020)
2020 arXiv
-
[60]
Jingtong Gao, Bo Chen, Xiangyu Zhao, Weiwen Liu, Xiangyang Li, Yichao Wang, Zijian Zhang, Wanyu Wang, Yuyang Ye, Shanru Lin, et al. 2024. Llm-enhanced reranking in recommender systems.arXiv preprint arXiv:2406.12433(2024)
2024 arXiv
-
[61]
Luyu Gao and Jamie Callan. 2021. Condenser: a pre-training architecture for dense retrieval.arXiv preprint arXiv:2104.08253(2021)
2021 arXiv
-
[62]
Chunjing Gan, Dan Yang, Binbin Hu, Ziqi Liu, Yue Shen, Zhiqiang Zhang, Jinjie Gu, Jun Zhou, and Guannan Zhang. 2023. Making Large Language Models Better Knowledge Miners for Online Marketing with Progressive Prompting Augmentation.arXiv preprint arXiv:2312.05276(2023)
2023 arXiv
-
[63]
Luyu Gao, Zhuyun Dai, Tongfei Chen, Zhen Fan, Benjamin Van Durme, and Jamie Callan. 2021. Complement lexical retrieval model with semantic residual embeddings. InAdvances in Information Retrieval: 43rd European Conference on IR Research, ECIR 2021, Virtual Event, March 28–Apri...
2021
-
[64]
Yunfan Gao, Yun Xiong, Xinyu Gao, Kangxiang Jia, Jinliu Pan, Yuxi Bi, Yi Dai, Jiawei Sun, Haofen Wang, and Haofen Wang. 2023. Retrieval-augmented generation for large language models: A survey.arXiv preprint arXiv:2312.10997 2 (2023)
2023 arXiv
-
[65]
Luyu Gao and Jamie Callan. 2021. Unsupervised corpus aware language model pre-training for dense passage retrieval.arXiv preprint arXiv:2108.05540(2021)
2021 arXiv
-
[66]
Google Research. 2023. Introducing PaLM 2. Blog post, https://ai.googleblog. com/2023/05/google-introduces-palm-2.html. Accessed: 2023-XX-XX
2023
-
[67]
Jiafeng Guo, Yinqiong Cai, Yixing Fan, Fei Sun, Ruqing Zhang, and Xueqi Cheng
-
[68]
Shijie Geng, Zuohui Fu, Yingqiang Ge, Lei Li, Gerard De Melo, and Yongfeng Zhang. 2022. Improving personalized explanation generation through visualiza- tion. InProceedings of the 60th Annual Meeting of the Association for Computa- tional Linguistics (Volume 1: Long Papers). 244–255
2022
-
[69]
Rishin Haldar and Debajyoti Mukhopadhyay. 2011. Levenshtein distance tech- nique in dictionary lookup methods: An improved approach.arXiv preprint arXiv:1101.1232(2011)
2011 arXiv
-
[70]
Kailash A Hambarde and Hugo Proenca. 2023. Information retrieval: recent advances and beyond.IEEE Access11 (2023), 76581–76604
2023
-
[71]
Xu Han, Weilin Zhao, Ning Ding, Zhiyuan Liu, and Maosong Sun. 2022. Ptr: Prompt tuning with rules for text classification.AI Open3 (2022), 182–192
2022
-
[72]
Bernal Jiménez Gutiérrez, Yiheng Shu, Yu Gu, Michihiro Yasunaga, and Yu Su. 2024. Hipporag: Neurobiologically inspired long-term memory for large language models. InThe Thirty-eighth Annual Conference on NeurIPS
2024
-
[73]
Haveliwala
Taher H. Haveliwala. 2002. Topic-sensitive PageRank. InProceedings of the 11th International Conference on World Wide Web(Honolulu, Hawaii, USA)(WWW ’02). Association for Computing Machinery, New York, NY, USA, 517–526. doi:10. 1145/511446.511513
2002
-
[74]
Shirley Anugrah Hayati, Raphael Olivier, Pravalika Avvaru, Pengcheng Yin, Anthony Tomasic, and Graham Neubig. 2018. Retrieval-based neural code generation.arXiv preprint arXiv:1808.10025(2018)
2018 arXiv
-
[75]
Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross Girshick. 2020. Momen- tum contrast for unsupervised visual representation learning. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition. 9729–9738
2020
-
[76]
Hunter Priniski, and Fred Morstatter
Bahareh Harandizadeh, J. Hunter Priniski, and Fred Morstatter. 2022. Keyword Assisted Embedded Topic Model. InWSDM’22
2022
-
[77]
Wilbert Jan Heeringa. 2004. Measuring dialect pronunciation differences using Levenshtein distance. (2004)
2004
-
[78]
Chenxu Hu, Jie Fu, Chenzhuang Du, Simian Luo, Junbo Zhao, and Hang Zhao
-
[79]
Jiaxin Huang, Yu Meng, and Jiawei Han. 2022. Few-Shot Fine-Grained Entity Typing with Automatic Label Interpretation and Instance Generation. InKDD. 605–614
2022
-
[80]
Xiaoxin He, Yijun Tian, Yifei Sun, Nitesh Chawla, Thomas Laurent, Yann Le- Cun, Xavier Bresson, and Bryan Hooi. 2025. G-retriever: Retrieval-augmented generation for textual graph understanding and question answering.Advances in NeurIPS37 (2025), 132876–132907
2025
-
[81]
Jiaxin Huang, Yiqing Xie, Yu Meng, Yunyi Zhang, and Jiawei Han. 2020. CoRel: Seed-Guided Topical Taxonomy Construction by Concept Learning and Relation Transferring. InKDD’20
2020
-
[82]
Lei Huang, Weijiang Yu, Weitao Ma, Weihong Zhong, Zhangyin Feng, Haotian Wang, Qianglong Chen, Weihua Peng, Xiaocheng Feng, Bing Qin, et al. 2025. A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions.ACM Transactions on Infor...
2025
- [83]
-
[84]
Gautier Izacard, Mathilde Caron, Lucas Hosseini, Sebastian Riedel, Piotr Bo- janowski, Armand Joulin, and Edouard Grave. 2021. Unsupervised dense in- formation retrieval with contrastive learning.arXiv preprint arXiv:2112.09118 (2021)
2021 arXiv
-
[85]
Jiaxin Huang, Yiqing Xie, Yu Meng, Jiaming Shen, Yunyi Zhang, and Jiawei Han
-
[86]
Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Ye Jin Bang, Andrea Madotto, and Pascale Fung. 2023. Survey of hallucination in natural language generation.ACM computing surveys55, 12 (2023), 1–38
2023
-
[87]
Jinhao Jiang, Kun Zhou, Zican Dong, Keming Ye, Xin Zhao, and Ji-Rong Wen
-
[88]
Jinhao Jiang, Kun Zhou, Xin Zhao, and Ji-Rong Wen. 2023. UniKGQA: Uni- fied Retrieval and Reasoning for Solving Multi-hop Question Answering Over Knowledge Graph. InThe Eleventh International Conference on Learning Repre- sentations
2023
-
[89]
Pere-Lluís Huguet Cabot and Roberto Navigli. 2021. REBEL: Relation Extraction By End-to-end Language generation. InFindings of EMNLP. 2370–2381
2021
-
[90]
Pengcheng Jiang, Lang Cao, Cao Danica Xiao, Parminder Bhatia, Jimeng Sun, and Jiawei Han. 2025. Kg-fit: Knowledge graph fine-tuning upon open-world knowledge.Advances in NeurIPS37 (2025), 136220–136258
2025
-
[91]
Gautier Izacard, Patrick Lewis, Maria Lomeli, Lucas Hosseini, Fabio Petroni, Timo Schick, Jane Dwivedi-Yu, Armand Joulin, Sebastian Riedel, and Edouard Grave. 2023. Atlas: Few-shot learning with retrieval augmented language models. Journal of Machine Learning Research24, 251 (...
2023
-
[92]
Pengcheng Jiang, Jiacheng Lin, Lang Cao, Runchu Tian, SeongKu Kang, Zifeng Wang, Jimeng Sun, and Jiawei Han. 2025. Deepretrieval: Hacking real search engines and retrievers with large language models via reinforcement learning. arXiv preprint arXiv:2503.00223(2025)
2025 arXiv
-
[93]
Pengcheng Jiang, Jiacheng Lin, Zifeng Wang, Jimeng Sun, and Jiawei Han
-
[94]
StructGPT: A General Framework for Large Language Model to Reason over Structured Data. InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, EMNLP 2023, Singapore, December 6-10, 2023, Houda Bouamor, Juan Pino, and Kalika Bali (Eds.). Assoc...
2023 doi
-
[95]
Pengcheng Jiang, Cao Xiao, Minhao Jiang, Parminder Bhatia, Taha Kass-Hout, Jimeng Sun, and Jiawei Han. 2025. Reasoning-Enhanced Healthcare Predictions with Knowledge Graph Community Retrieval. InThe Thirteenth International Conference on Learning Representations. https://openr...
2025
-
[96]
Pengcheng Jiang, Shivam Agarwal, Bowen Jin, Xuan Wang, Jimeng Sun, and Jiawei Han. 2023. Text Augmented Open Knowledge Graph Completion via Pre-Trained Language Models. InFindings of ACL. 11161–11180
2023
-
[97]
Ting Jiang, Deqing Wang, Leilei Sun, Zhongzhi Chen, Fuzhen Zhuang, and Qinghong Yang. 2022. Exploiting Global and Local Hierarchies for Hierarchical Text Classification. InEMNLP. 4030–4039
2022
-
[98]
Pengcheng Jiang, Megan Amber Lim, Adam Cross, and Jimeng Sun. [n. d.]. MedKG: Empowering Medical Education with Interactive Construction and Visualization of Knowledge Graphs via Large Language Models. ([n. d.])
-
[99]
Yizhu Jiao, Sha Li, Yiqing Xie, Ming Zhong, Heng Ji, and Jiawei Han. 2022. Open-Vocabulary Argument Role Prediction For Event Extraction. InFindings of EMNLP. 5404–5418
2022
-
[100]
Yizhu Jiao, Ming Zhong, Sha Li, Ruining Zhao, Siru Ouyang, Heng Ji, and Jiawei Han. 2023. Instruct and Extract: Instruction Tuning for On-Demand Information Extraction. InEMNLP. 10030–10051
2023
-
[101]
Yizhu Jiao, Ming Zhong, Jiaming Shen, Yunyi Zhang, Chao Zhang, and Jiawei Han. 2023. Unsupervised Event Chain Mining from Multiple Documents. In Proceedings of the ACM Web Conference 2023, WWW 2023, Austin, TX, USA, 30 April 2023 - 4 May 2023, Ying Ding, Jie Tang, Juan F. Sequ...
2023 doi
-
[102]
Pengcheng Jiang, Cao Xiao, Adam Cross, and Jimeng Sun. 2023. Graphcare: Enhancing healthcare predictions with personalized knowledge graphs.arXiv preprint arXiv:2305.12788(2023)
2023 arXiv
- [103]
-
[104]
Pengcheng Jiang, Xueqiang Xu, Jiacheng Lin, Jinfeng Xiao, Zifeng Wang, Jimeng Sun, and Jiawei Han. 2025. s3: You Don’t Need That Much Data to Train a Search Agent via RL. arXiv:2505.14146 [cs.AI]
2025
-
[105]
Hao Kang, Tevin Wang, and Chenyan Xiong. 2024. Interpret and Control Dense Retrieval with Sparse Latent Features.arXiv preprint arXiv:2411.00786(2024)
2024 arXiv
-
[106]
Zhengbao Jiang, Frank F Xu, Luyu Gao, Zhiqing Sun, Qian Liu, Jane Dwivedi-Yu, Yiming Yang, Jamie Callan, and Graham Neubig. 2023. Active retrieval aug- mented generation. InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 7969–7992
2023
-
[107]
SeongKu Kang, Shivam Agarwal, Bowen Jin, Dongha Lee, Hwanjo Yu, and Jiawei Han. 2024. Improving retrieval in theme-specific applications using a corpus topical taxonomy. InProceedings of the ACM Web Conference 2024. 1497–1508
2024
-
[108]
SeongKu Kang, Yunyi Zhang, Pengcheng Jiang, Dongha Lee, Jiawei Han, and Hwanjo Yu. 2024. Taxonomy-guided Semantic Indexing for Academic Paper Search. InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, Yaser Al-Onaizan, Mohit Bansal, and Y...
2024
-
[109]
Priyanka Kargupta, Yunyi Zhang, Yizhu Jiao, Siru Ouyang, and Jiawei Han. 2024. Unsupervised Episode Detection for Large-Scale News Events.arXiv preprint arXiv:2408.04873(2024)
2024 arXiv
-
[110]
Bowen Jin, Chulin Xie, Jiawei Zhang, Kashob Kumar Roy, Yu Zhang, Zheng Li, Ruirui Li, Xianfeng Tang, Suhang Wang, Yu Meng, et al. 2024. Graph Chain- of-Thought: Augmenting Large Language Models by Reasoning on Graphs. In Findings of the Association for Computational Linguistic...
2024
-
[111]
Vladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick SH Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih. 2020. Dense Passage Retrieval for Open-Domain Question Answering.. InEMNLP (1). 6769–6781
2020
-
[112]
Mandar Joshi, Eunsol Choi, Daniel S Weld, and Luke Zettlemoyer. 2017. Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension. arXiv preprint arXiv:1705.03551(2017)
2017 arXiv
-
[113]
Omar Khattab and Matei Zaharia. 2020. Colbert: Efficient and effective passage search via contextualized late interaction over bert. InProceedings of the 43rd International ACM SIGIR conference on research and development in Information Retrieval. 39–48
2020
-
[114]
Minki Kang, Jin Myung Kwak, Jinheon Baek, and Sung Ju Hwang. 2023. Knowl- edge graph-augmented language models for knowledge-grounded dialogue generation.arXiv preprint arXiv:2305.18846(2023)
2023 arXiv
-
[115]
Mei Kobayashi and Koichi Takeda. 2000. Information retrieval on the web.ACM computing surveys (CSUR)32, 2 (2000), 144–173
2000
-
[116]
Tanay Komarlu, Minhao Jiang, Xuan Wang, and Jiawei Han. 2023. OntoType: Ontology-Guided Zero-Shot Fine-Grained Entity Typing with Weak Supervision from Pre-Trained Language Models.arXiv preprint arXiv:2305.12307(2023)
2023 arXiv
-
[117]
Rajasekar Krishnamurthy, Sriram Raghavan, Shivakumar Vaithyanathan, and Huaiyu Zhu. 2006. Avatar information extraction system. (2006)
2006
-
[118]
Keito Kudo, Hiroyuki Deguchi, Makoto Morishita, Ryo Fujii, Takumi Ito, Shin- taro Ozaki, Koki Natsumi, Kai Sato, Kazuki Yano, Ryosuke Takahashi, et al. 2024. Document-level translation with LLM reranking: Team-j at WMT 2024 general translation task. InProceedings of the Ninth ...
2024
-
[119]
Md Rezaul Karim, Lina Molinas Comet, Md Shajalal, Oya Beyan, Dietrich Rebholz-Schuhmann, and Stefan Decker. 2023. From Large Language Mod- els to Knowledge Graphs for Biomarker Discovery in Cancer.arXiv preprint arXiv:2310.08365(2023)
2023 arXiv
-
[120]
Dongha Lee, Jiaming Shen, SeongKu Kang, Susik Yoon, Jiawei Han, and Hwanjo Yu. 2022. TaxoCom: Topic Taxonomy Completion with Hierarchical Discovery of Novel Topic Clusters. InWWW’22
2022
-
[121]
Omar Khattab, Keshav Santhanam, Xiang Lisa Li, David Hall, Percy Liang, Christopher Potts, and Matei Zaharia. 2022. Demonstrate-search-predict: Com- posing retrieval and language models for knowledge-intensive nlp.arXiv preprint arXiv:2212.14024(2022)
2022 arXiv
-
[122]
Dirk Lewandowski. 2005. Web searching, search engines and Information Retrieval.Information Services & Use25, 3-4 (2005), 137–147
2005
-
[123]
Gangwoo Kim, Sungdong Kim, Byeongguk Jeon, Joonsuk Park, and Jaewoo Kang
-
[124]
InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing
Tree of clarifications: Answering ambiguous questions with retrieval- augmented large language models. InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 996–1009
2023
-
[125]
Patrick Lewis, Yuxiang Wu, Linqing Liu, Pasquale Minervini, Heinrich Küttler, Aleksandra Piktus, Pontus Stenetorp, and Sebastian Riedel. 2021. Paq: 65 million probably-asked questions and what you can do with them.Transactions of the Association for Computational Linguistics9 ...
2021
-
[126]
Bangzheng Li, Wenpeng Yin, and Muhao Chen. 2022. Ultra-fine entity typing with indirect supervision from natural language inference.TACL10 (2022), 607–622
2022
-
[127]
Guozheng Li, Peng Wang, and Wenjun Ke. 2023. Revisiting large language models as zero-shot relation extractors.arXiv preprint arXiv:2310.05028(2023)
2023 arXiv
-
[128]
Jinyang Li, Binyuan Hui, Ge Qu, Jiaxi Yang, Binhua Li, Bowen Li, Bailin Wang, Bowen Qin, Ruiying Geng, Nan Huo, et al . 2023. Can llm already serve as a database interface? a big bench for large-scale database grounded text-to-sqls. Advances in NeurIPS36 (2023), 42330–42357
2023
-
[129]
Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov
Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, Kristina Toutanova, Llion Jones, Matthew Kelcey, Ming-Wei Chang, An- drew M. Dai, Jakob Uszkoreit, Quoc Le, and Sl...
2019 doi
-
[130]
Jing Li, Aixin Sun, Jianglei Han, and Chenliang Li. 2020. A survey on deep learning for named entity recognition.IEEE transactions on knowledge and data KDD ’25, August 3–7, 2025, Toronto, ON, Canada Pengcheng Jiang et al. engineering34, 1 (2020), 50–70
2020
-
[131]
Elena Leitner, Georg Rehm, and Julian Moreno-Schneider. 2019. Fine-grained named entity recognition in legal documents. InInternational conference on semantic systems. Springer, 272–287
2019
-
[132]
Na Li, Zied Bouraoui, and Steven Schockaert. 2023. Ultra-fine entity typing with prior knowledge about labels: A simple clustering based strategy.arXiv preprint arXiv:2305.12802(2023)
2023 arXiv
-
[133]
Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Veselin Stoyanov, and Luke Zettlemoyer. 2020. BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension. InACL. 7871–7880
2020
-
[134]
Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rock- täschel, et al. 2020. Retrieval-augmented generation for knowledge-intensive nlp tasks.Advances in NeurIPS33 (2020), 9459–9474
2020
-
[135]
Xianming Li and Jing Li. 2023. Angle-optimized text embeddings.arXiv preprint arXiv:2309.12871(2023)
2023 arXiv
-
[136]
Xinze Li, Zhenghao Liu, Chenyan Xiong, Shi Yu, Yu Gu, Zhiyuan Liu, and Ge Yu. 2023. Structure-aware language model pretraining improves dense retrieval on structured data.arXiv preprint arXiv:2305.19912(2023)
2023 arXiv
-
[137]
Yinghui Li, Yangning Li, Yuxin He, Tianyu Yu, Ying Shen, and Hai-Tao Zheng
-
[138]
Zijing Liang, Yanjie Xu, Yifan Hong, Penghui Shang, Qi Wang, Qiang Fu, and Ke Liu. 2024. A Survey of Multimodel Large Language Models. InProceedings of the 3rd International Conference on Computer, Artificial Intelligence and Control Engineering. 405–409
2024
-
[139]
Junpeng Li, Zixia Jia, and Zilong Zheng. 2023. Semi-automatic Data Enhance- ment for Document-Level Relation Extraction with Distant Supervision from Large Language Models. InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, EMNLP 2023, Si...
2023 doi
-
[140]
Xi Victoria Lin, Xilun Chen, Mingda Chen, Weijia Shi, Maria Lomeli, Richard James, Pedro Rodriguez, Jacob Kahn, Gergely Szilvasy, Mike Lewis, et al. 2023. Ra-dit: Retrieval-augmented dual instruction tuning. InThe Twelfth International Conference on Learning Representations
2023
-
[141]
Minghan Li and Eric Gaussier. 2024. Domain Adaptation for Dense Retrieval and Conversational Dense Retrieval through Self-Supervision by Meticulous Pseudo-Relevance Labeling.arXiv preprint arXiv:2403.08970(2024)
2024 arXiv
-
[142]
Fang Liu, Clement Yu, and Weiyi Meng. 2004. Personalized web search for improving retrieval effectiveness.IEEE Transactions on knowledge and data engineering16, 1 (2004), 28–40
2004
-
[143]
Peng Li, Yeye He, Dror Yashar, Weiwei Cui, Song Ge, Haidong Zhang, Danielle Rifinski Fainman, Dongmei Zhang, and Surajit Chaudhuri. 2023. Table- gpt: Table-tuned gpt for diverse table tasks.arXiv preprint arXiv:2310.09263 (2023)
2023 arXiv
-
[144]
Xiaoxi Li, Jiajie Jin, Yujia Zhou, Yuyao Zhang, Peitian Zhang, Yutao Zhu, and Zhicheng Dou. 2024. From matching to generation: A survey on generative information retrieval.arXiv preprint arXiv:2404.14851(2024)
2024 arXiv
-
[145]
Nelson F Liu, Tianyi Zhang, and Percy Liang. 2023. Evaluating verifiability in generative search engines.arXiv preprint arXiv:2304.09848(2023)
2023 arXiv
-
[146]
Runxuan Liu, Bei Luo, Jiaqi Li, Baoxin Wang, Ming Liu, Dayong Wu, Shijin Wang, and Bing Qin. 2025. Ontology-Guided Reverse Thinking Makes Large Language Models Stronger on Knowledge Graph Question Answering.arXiv preprint arXiv:2502.11491(2025)
2025 arXiv
-
[147]
Ye Liu, Yao Wan, Lifang He, Hao Peng, and S Yu Philip. 2021. Kg-bart: Knowledge graph-augmented bart for generative commonsense reasoning. InProceedings of the AAAI conference on artificial intelligence, Vol. 35. 6418–6425
2021
-
[148]
Contrastive Learning with Hard Negative Entities for Entity Set Expansion. InSIGIR. 1077–1086
-
[149]
Keming Lu, I Hsu, Wenxuan Zhou, Mingyu Derek Ma, Muhao Chen, et al. 2022. Summarization as indirect supervision for relation extraction.arXiv preprint arXiv:2205.09837(2022)
2022 arXiv
-
[150]
Jimmy Lin and Xueguang Ma. 2021. A few brief notes on deepimpact, coil, and a conceptual framework for information retrieval techniques.arXiv preprint arXiv:2106.14807(2021)
2021 arXiv
-
[151]
Linhao Luo, Jiaxin Ju, Bo Xiong, Yuan-Fang Li, Gholamreza Haffari, and Shirui Pan. 2023. Chatrule: Mining logical rules with large language models for knowl- edge graph reasoning.arXiv preprint arXiv:2309.01538(2023)
2023 arXiv
-
[152]
Xi Victoria Lin, Richard Socher, and Caiming Xiong. 2020. Bridging textual and tabular data for cross-domain text-to-SQL semantic parsing.arXiv preprint arXiv:2012.12627(2020)
2020 arXiv
-
[153]
Linhao Luo, Zicheng Zhao, Gholamreza Haffari, Dinh Phung, Chen Gong, and Shirui Pan. 2025. GFM-RAG: Graph Foundation Model for Retrieval Augmented Generation.arXiv preprint arXiv:2502.01113(2025)
2025
-
[154]
Jie Liu and Barzan Mozafari. 2024. Query rewriting via large language models. arXiv preprint arXiv:2403.09060(2024)
2024
-
[155]
Nelson F Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang. 2024. Lost in the middle: How language models use long contexts.Transactions of the Association for Computational Linguistics 12 (2024), 157–173
2024
-
[156]
Yuanhua Lv and ChengXiang Zhai. 2011. Lower-bounding term frequency normalization. InProceedings of the 20th ACM international conference on Infor- mation and knowledge management. 7–16
2011
-
[157]
Jeffrey Brantingham, Nanyun Peng, and Wei Wang
Mingyu Derek Ma, Xiaoxuan Wang, Po-Nien Kung, P. Jeffrey Brantingham, Nanyun Peng, and Wei Wang. 2023. STAR: Boosting Low-Resource Event Extraction by Structure-to-Text Data Generation with Large Language Models. CoRRabs/2305.15090 (2023). arXiv:2305.15090 doi:10.48550/ARXIV.2...
-
[158]
Shengjie Ma, Chengjin Xu, Xuhui Jiang, Muzhi Li, Huaren Qu, Cehao Yang, Jiaxin Mao, and Jian Guo. 2025. Think-on-Graph 2.0: Deep and Faithful Large Language Model Reasoning with Knowledge-guided Retrieval Augmented Generation. InThe Thirteenth International Conference on Learn...
2025
-
[159]
Michael Llordes, Debasis Ganguly, Sumit Bhatia, and Chirag Agarwal. 2023. Explain like I am BM25: Interpreting a Dense Model’s Ranked-List with a Sparse Approximation. InProceedings of the 46th International ACM SIGIR Conference on Research and Development in Information Retri...
2023
-
[160]
Yubo Ma, Yixin Cao, YongChing Hong, and Aixin Sun. 2023. Large language model is not a good few-shot information extractor, but a good reranker for hard samples!arXiv preprint arXiv:2303.08559(2023)
2023 arXiv
-
[161]
Yi Luan, Jacob Eisenstein, Kristina Toutanova, and Michael Collins. 2021. Sparse, dense, and attentional representations for text retrieval.Transactions of the Association for Computational Linguistics9 (2021), 329–345
2021
-
[162]
Priyanka Mandikal and Raymond Mooney. 2024. Sparse meets dense: A hybrid approach to enhance scientific document retrieval.arXiv preprint arXiv:2401.04055(2024)
2024 arXiv
-
[163]
LINHAO LUO, Yuan-Fang Li, Reza Haf, and Shirui Pan. 2024. Reasoning on Graphs: Faithful and Interpretable Large Language Model Reasoning. InThe Twelfth International Conference on Learning Representations. https://openreview. net/forum?id=ZGNWW7xZ6Q
2024
-
[164]
Yu Meng, Jiaxin Huang, Guangyuan Wang, Zihan Wang, Chao Zhang, Yu Zhang, and Jiawei Han. 2020. Discriminative Topic Mining via Category-Name Guided Text Embedding. InWWW’20
2020
-
[165]
Ziyang Luo, Can Xu, Pu Zhao, Xiubo Geng, Chongyang Tao, Jing Ma, Qingwei Lin, and Daxin Jiang. 2023. Augmented large language models with parametric knowledge guiding.arXiv preprint arXiv:2305.04757(2023)
2023 arXiv
-
[166]
Yuanhua Lv and ChengXiang Zhai. 2011. Adaptive Term Frequency Normal- ization for BM25. InProceedings of the 20th ACM International Conference on Information and Knowledge Management (CIKM 2011). 1985–1988
2011
-
[167]
Yu Meng, Yunyi Zhang, Jiaxin Huang, Xuan Wang, Yu Zhang, Heng Ji, and Jiawei Han. 2021. Distantly-Supervised Named Entity Recognition with Noise- Robust Learning and Language Model Augmented Self-Training. InEMNLP. 10367–10378
2021
-
[168]
Yu Meng, Yunyi Zhang, Jiaxin Huang, Chenyan Xiong, Heng Ji, Chao Zhang, and Jiawei Han. 2020. Text Classification Using Label Names Only: A Language Model Self-Training Approach. InEMNLP. 9006–9017
2020
-
[169]
Yu Meng, Yunyi Zhang, Jiaxin Huang, Yu Zhang, and Jiawei Han. 2022. Topic Discovery via Latent Space Clustering of Pretrained Language Model Represen- tations. InWWW’22. 3143–3152
2022
-
[170]
Xinbei Ma, Yeyun Gong, Pengcheng He, Hai Zhao, and Nan Duan. 2023. Query rewriting in retrieval-augmented large language models. InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 5303– 5315
2023
-
[171]
Grégoire Mialon, Roberto Dessì, Maria Lomeli, Christoforos Nalmpantis, Ram Pasunuru, Roberta Raileanu, Baptiste Rozière, Timo Schick, Jane Dwivedi-Yu, Asli Celikyilmaz, et al . 2023. Augmented language models: a survey.arXiv preprint arXiv:2302.07842(2023)
2023 arXiv
-
[172]
Antonio Mallia, Joel Mackenzie, Torsten Suel, and Nicola Tonellotto. 2022. Faster learned sparse retrieval with guided traversal. InProceedings of the 45th Inter- national ACM SIGIR Conference on Research and Development in Information Retrieval. 1901–1905
2022
-
[173]
Sewon Min, Mike Lewis, Hannaneh Hajishirzi, and Luke Zettlemoyer. 2021. Noisy channel language model prompting for few-shot text classification.arXiv preprint arXiv:2108.04106(2021)
2021 arXiv
-
[174]
Dheeraj Mekala and Jingbo Shang. 2020. Contextualized Weak Supervision for Text Classification. InACL. 323–333
2020
-
[175]
Seyed Abbas Momtazi and Dietrich Klakow. 2010. Hierarchical Pitman–Yor Language Model for Information Retrieval. InProceedings of the 33rd Annual International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR 2010). 793–794
2010
-
[176]
Yu Meng, Jiaming Shen, Chao Zhang, and Jiawei Han. 2018. Weakly-Supervised Neural Text Classification. InProceedings of the 27th ACM International Confer- ence on Information and Knowledge Management. ACM, 983–992
2018
-
[177]
Yu Meng, Jiaming Shen, Chao Zhang, and Jiawei Han. 2019. Weakly-supervised hierarchical text classification. InAAAI
2019
-
[178]
Shahrzad Naseri, Jeffrey Dalton, Andrew Yates, and James Allan. 2021. Ceqe: Contextualized embeddings for query expansion. InAdvances in Information Retrieval: 43rd European Conference on IR Research, ECIR 2021, Virtual Event, March 28–April 1, 2021, Proceedings, Part I 43. Sp...
2021
-
[179]
Tri Nguyen, Mir Rosenberg, Xia Song, Jianfeng Gao, Saurabh Tiwary, Rangan Majumder, and Li Deng. 2016. Ms marco: A human-generated machine reading A Survey on Retrieval And Structuring Augmented Generation with Large Language Models KDD ’25, August 3–7, 2025, Toronto, ON, Cana...
2016
-
[180]
Jianmo Ni, Gustavo Hernandez Abrego, Noah Constant, Ji Ma, Keith B Hall, Daniel Cer, and Yinfei Yang. 2021. Sentence-t5: Scalable sentence encoders from pre-trained text-to-text models.arXiv preprint arXiv:2108.08877(2021)
2021 arXiv
-
[181]
Yu Meng, Yunyi Zhang, Jiaxin Huang, Yu Zhang, Chao Zhang, and Jiawei Han
-
[182]
InKDD’20
Hierarchical Topic Mining via Joint Spherical Tree and Text Embedding. InKDD’20
-
[183]
Rodrigo Nogueira, Jimmy Lin, and AI Epistemic. 2019. From doc2query to docTTTTTquery.Online preprint6, 2 (2019)
2019
-
[184]
Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean. 2013. Efficient esti- mation of word representations in vector space.arXiv preprint arXiv:1301.3781 (2013)
2013 arXiv
-
[185]
Zach Nussbaum, John X Morris, Brandon Duderstadt, and Andriy Mulyar. 2024. Nomic embed: Training a reproducible long context text embedder.arXiv preprint arXiv:2402.01613(2024)
2024 arXiv
-
[186]
Mandar Mitra and BB Chaudhuri. 2000. Information retrieval from documents: A survey.Information retrieval2 (2000), 141–163
2000
-
[187]
Aytuğ Onan, Serdar Korukoğlu, and Hasan Bulut. 2016. Ensemble of keyword extraction methods and classifiers in text classification.Expert Systems with Applications57 (2016), 232–247
2016
-
[188]
Eduardo Mosqueira-Rey, Elena Hernández-Pereira, David Alonso-Ríos, José Bobes-Bascarán, and Ángel Fernández-Leal. 2023. Human-in-the-loop machine learning: a state of the art.Artificial Intelligence Review56, 4 (2023), 3005–3054
2023
-
[189]
David Nadeau and Satoshi Sekine. 2007. A survey of named entity recognition and classification.Lingvisticae Investigationes30, 1 (2007), 3–26
2007
-
[190]
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022. Training language models to follow instructions with human feedback.Advances in NeurIPS35 (2022), 27730–27744
2022
-
[191]
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul F. Christiano, Jan Leik...
-
[192]
Siru Ouyang, Jiaxin Huang, Pranav Pillai, Yunyi Zhang, Yu Zhang, and Jiawei Han. 2023. Ontology Enrichment for Effective Fine-grained Entity Typing.arXiv preprint arXiv:2310.07795(2023)
2023 arXiv
-
[193]
Jianmo Ni, Chen Qu, Jing Lu, Zhuyun Dai, Gustavo Hernández Ábrego, Ji Ma, Vincent Y Zhao, Yi Luan, Keith B Hall, Ming-Wei Chang, et al. 2021. Large dual encoders are generalizable retrievers.arXiv preprint arXiv:2112.07899(2021)
2021 arXiv
-
[194]
Rodrigo Nogueira and Kyunghyun Cho. 2019. Passage Re-ranking with BERT. arXiv preprint arXiv:1901.04085(2019)
2019 arXiv
-
[195]
Shirui Pan, Linhao Luo, Yufei Wang, Chen Chen, Jiapu Wang, and Xindong Wu
-
[196]
Rodrigo Nogueira, Wei Yang, Jimmy Lin, and Kyunghyun Cho. 2019. Document expansion by query prediction.arXiv preprint arXiv:1904.08375(2019)
2019 arXiv
-
[197]
Letian Peng, Zihan Wang, and Jingbo Shang. 2023. Less than One-shot: Named Entity Recognition via Extremely Weak Supervision. InFindings of EMNLP. 13603–13616
2023
-
[198]
Abiola Obamuyide and Andreas Vlachos. 2018. Zero-shot Relation Classification as Textual Entailment. InProceedings of the First Workshop on Fact Extraction and VERification (FEVER). 72–78
2018
-
[199]
José R Pérez-Agüera, Javier Arroyo, Jane Greenberg, Joaquin Perez Iglesias, and Victor Fresno. 2010. Using BM25F for semantic search. InProceedings of the 3rd international semantic search workshop. 1–8
2010
-
[200]
Aaron van den Oord, Yazhe Li, and Oriol Vinyals. 2018. Representation learning with contrastive predictive coding.arXiv preprint arXiv:1807.03748(2018)
2018 arXiv
- [201]
-
[202]
Meng Qu, Xiang Ren, Yu Zhang, and Jiawei Han. 2018. Weakly-supervised Relation Extraction by Pattern-enhanced Embedding Learning. InWWW. 1257–1266
2018
-
[203]
Yingqi Qu, Yuchen Ding, Jing Liu, Kai Liu, Ruiyang Ren, Wayne Xin Zhao, Daxi- ang Dong, Hua Wu, and Haifeng Wang. 2020. RocketQA: An optimized training approach to dense passage retrieval for open-domain question answering.arXiv preprint arXiv:2010.08191(2020)
2020 arXiv
-
[204]
In NeurIPS
Training language models to follow instructions with human feedback. In NeurIPS
-
[205]
Manning, Stefano Ermon, and Chelsea Finn
Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning, Stefano Ermon, and Chelsea Finn. 2023. Direct Preference Optimization: Your Language Model is Secretly a Reward Model. InNeurIPS
2023
-
[206]
Oded Ovadia, Menachem Brief, Moshik Mishaeli, and Oren Elisha. 2023. Fine- tuning or retrieval? comparing knowledge injection in llms.arXiv preprint arXiv:2312.05934(2023)
2023 arXiv
-
[207]
Charles Packer, Sarah Wooders, Kevin Lin, Vivian Fang, Shishir G Patil, Ion Stoica, and Joseph E Gonzalez. 2023. MemGPT: Towards LLMs as Operating Systems.arXiv preprint arXiv:2310.08560(2023)
2023 arXiv
-
[208]
Alexander Ratner, Christopher De Sa, Sen Wu, Daniel Selsam, and Christopher Ré. 2016. Data Programming: Creating Large Training Sets, Quickly. InNIPS
2016
-
[209]
Unifying large language models and knowledge graphs: A roadmap.IEEE Transactions on Knowledge and Data Engineering36, 7 (2024), 3580–3599
2024
-
[210]
Hao Peng, Jianxin Li, Yu He, Yaopeng Liu, Mengjiao Bao, Yangqiu Song, and Qiang Yang. 2018. Large-Scale Hierarchical Text Classification with Recursively Regularized Deep Graph-CNN. InWWW
2018
-
[211]
Stephen E Robertson, Steve Walker, Susan Jones, Micheline M Hancock-Beaulieu, Mike Gatford, et al . 1995. Okapi at TREC-3.Nist Special Publication Sp109 (1995), 109
1995
-
[212]
Wenjun Peng, Guiyang Li, Yue Jiang, Zilong Wang, Dan Ou, Xiaoyi Zeng, Derong Xu, Tong Xu, and Enhong Chen. 2024. Large language model based long-tail query rewriting in taobao search. InCompanion Proceedings of the ACM Web Conference 2024. 20–28
2024
-
[213]
Dwaipayan Roy, Debjyoti Paul, Mandar Mitra, and Utpal Garain. 2016. Using word embeddings for automatic query expansion.arXiv preprint arXiv:1606.07608(2016)
2016 arXiv
-
[214]
Gabriel Poesia, Oleksandr Polozov, Vu Le, Ashish Tiwari, Gustavo Soares, Christopher Meek, and Sumit Gulwani. 2022. Synchromesh: Reliable code generation from pre-trained language models.arXiv preprint arXiv:2201.11227 (2022)
2022 arXiv
-
[215]
Jaroslav Pokorny. 2004. Web searching and information retrieval.Computing in Science & Engineering6, 4 (2004), 43–48
2004
-
[216]
Diego Sanmartin. 2024. Kg-rag: Bridging the gap between knowledge and creativity.arXiv preprint arXiv:2405.12035(2024)
2024 arXiv
-
[217]
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov
-
[218]
Alec Radford, Karthik Narasimhan, Tim Salimans, Ilya Sutskever, et al. 2018. Improving language understanding by generative pre-training. (2018)
2018
-
[219]
Jingbo Shang, Xinyang Zhang, Liyuan Liu, Sha Li, and Jiawei Han. 2020. NetTaxo: Automated Topic Taxonomy Construction from Text-Rich Network. InWWW (WWW ’20). 1908–1919
2020
-
[220]
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer.Journal of machine learning research21, 140 (2020), 1–67
2020
-
[221]
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016. Squad: 100,000+ questions for machine comprehension of text.arXiv preprint arXiv:1606.05250(2016)
2016 arXiv
-
[222]
Jiaming Shen, Wenda Qiu, Yu Meng, Jingbo Shang, Xiang Ren, and Jiawei Han
-
[223]
Wendi Ren, Yinghao Li, Hanting Su, David Kartchner, Cassie Mitchell, and Chao Zhang. 2020. Denoising Multi-Source Weak Supervision for Neural Text Classification. InEMNLP Findings
2020
-
[224]
Stephen Robertson, Hugo Zaragoza, et al . 2009. The probabilistic relevance framework: BM25 and beyond.Foundations and Trends®in Information Retrieval 3, 4 (2009), 333–389
2009
-
[225]
Jiaming Shen, Yunyi Zhang, Heng Ji, and Jiawei Han. 2021. Corpus-based Open-Domain Event Type Induction. InProceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, EMNLP 2021, Virtual Event / Punta Cana, Dominican Republic, 7-11 November, 2021, M...
2021 doi
-
[226]
Maya Rotmensch, Yoni Halpern, Abdulhakim Tlimat, Steven Horng, and David Sontag. 2017. Learning a health knowledge graph from electronic medical records.Scientific reports7, 1 (2017), 5994
2017
-
[227]
Karan Singhal, Shekoofeh Azizi, Tao Tu, S Sara Mahdavi, Jason Wei, Hyung Won Chung, Nathan Scales, Ajay Tanwani, Heather Cole-Lewis, Stephen Pfohl, et al
-
[228]
Oscar Sainz, Oier Lopez de Lacalle, Gorka Labaka, Ander Barrena, and Eneko Agirre. 2021. Label Verbalization and Entailment for Effective Zero and Few-Shot Relation Extraction. InEMNLP. 1199–1212
2021
-
[229]
Gerard Salton and Christopher Buckley. 1988. Term-weighting approaches in automatic text retrieval.Information processing & management24, 5 (1988), 513–523
1988
-
[230]
Jiashuo Sun, Chengjin Xu, Lumingyuan Tang, Saizhuo Wang, Chen Lin, Yeyun Gong, Lionel Ni, Heung-Yeung Shum, and Jian Guo. 2024. Think-on-Graph: Deep and Responsible Reasoning of Large Language Model on Knowledge Graph. InThe Twelfth International Conference on Learning Represe...
2024
-
[231]
Weiwei Sun, Lingyong Yan, Zheng Chen, Shuaiqiang Wang, Haichao Zhu, Pengjie Ren, Zhumin Chen, Dawei Yin, Maarten Rijke, and Zhaochun Ren
-
[232]
Xiaofei Sun, Xiaoya Li, Jiwei Li, Fei Wu, Shangwei Guo, Tianwei Zhang, and Guoyin Wang. 2023. Text Classification via Large Language Models. InFindings of EMNLP 2023. 8990–9005
2023
-
[233]
Jingbo Shang, Liyuan Liu, Xiang Ren, Xiaotao Gu, Teng Ren, and Jiawei Han
-
[234]
Zhiqing Sun, Xuezhi Wang, Yi Tay, Yiming Yang, and Denny Zhou. 2022. Recitation-augmented language models.arXiv preprint arXiv:2210.01296(2022)
2022 arXiv
-
[235]
Ammar Tahir. 2023. Knowledge Graph GPT. https://github.com/iAmmarTahir/ KnowledgeGraphGPT
2023
-
[236]
Yu-Ming Shang, Hongli Mao, Tian Tian, Heyan Huang, and Xian-Ling Mao. 2025. From local to global: Leveraging document graph for named entity recognition. Knowledge-Based Systems(2025), 113017
2025
-
[237]
Zhihong Shao, Yeyun Gong, Yelong Shen, Minlie Huang, Nan Duan, and Weizhu Chen. 2023. Enhancing retrieval-augmented large language models with itera- tive retrieval-generation synergy.arXiv preprint arXiv:2305.15294(2023)
2023 arXiv
-
[238]
Nandan Thakur, Nils Reimers, Andreas Rücklé, Abhishek Srivastava, and Iryna Gurevych. 2021. Beir: A heterogenous benchmark for zero-shot evaluation of information retrieval models.arXiv preprint arXiv:2104.08663(2021)
2021 arXiv
-
[239]
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. 2023. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288(2023)
2023 arXiv
-
[240]
Jiaming Shen, Zeqiu Wu, Dongming Lei, Jingbo Shang, Xiang Ren, and Jiawei Han. 2017. Setexpan: Corpus-based set expansion via context feature selection and rank ensemble. InECML-PKDD’17. 288–304
2017
-
[241]
Jiaming Shen, Zeqiu Wu, Dongming Lei, Chao Zhang, Xiang Ren, Michelle T Vanni, Brian M Sadler, and Jiawei Han. 2018. Hiexpan: Task-guided taxonomy construction by hierarchical tree expansion. InProceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery ...
2018
-
[242]
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need.Advances in NeurIPS30 (2017)
2017
-
[243]
Zhendong Shi, Jacky Keung, and Qinbao Song. 2014. An empirical study of bm25 and bm25f based feature location techniques. InProceedings of the International Workshop on Innovative Software Development Methodologies and Practices. 106– 114
2014
-
[244]
Zhen Wan, Fei Cheng, Zhuoyuan Mao, Qianying Liu, Haiyue Song, Jiwei Li, and Sadao Kurohashi. 2023. GPT-RE: In-context Learning for Relation Extraction using Large Language Models. InEMNLP. 3534–3547
2023
-
[245]
Large language models encode clinical knowledge.Nature620, 7972 (2023), 172–180
2023
-
[246]
Yangqiu Song and Dan Roth. 2014. On Dataless Hierarchical Text Classification. InAAAI
2014
-
[247]
Karen Sparck Jones. 1972. A statistical interpretation of term specificity and its application in retrieval.Journal of documentation28, 1 (1972), 11–21. KDD ’25, August 3–7, 2025, Toronto, ON, Canada Pengcheng Jiang et al
1972
-
[248]
Siyuan Wang, Zhongyu Wei, Jiarong Xu, Taishan Li, and Zhihao Fan. 2023. Uni- fying Structure Reasoning and Language Pre-Training for Complex Reasoning Tasks.IEEE/ACM Transactions on Audio, Speech, and Language Processing32 (2023), 1586–1595
2023
-
[249]
Shuohang Wang, Yichong Xu, Yuwei Fang, Yang Liu, Siqi Sun, Ruochen Xu, Chenguang Zhu, and Michael Zeng. 2022. Training data is more valuable than you think: A simple and effective method by retrieving from training data.arXiv preprint arXiv:2203.08773(2022)
2022 arXiv
-
[250]
Learning to tokenize for generative retrieval.Advances in NeurIPS36 (2023), 46345–46361
2023
-
[251]
Xintao Wang, Qianwen Yang, Yongting Qiu, Jiaqing Liang, Qianyu He, Zhouhong Gu, Yanghua Xiao, and Wei Wang. 2023. Knowledgpt: Enhanc- ing large language models with retrieval and storage access on knowledge bases. arXiv preprint arXiv:2308.11761(2023)
2023 arXiv
-
[252]
Yueqing Sun, Qi Shi, Le Qi, and Yu Zhang. 2021. JointLK: Joint reasoning with language models and knowledge graphs for commonsense question answering. arXiv preprint arXiv:2112.02732(2021)
2021 arXiv
-
[253]
Yizhong Wang, Swaroop Mishra, Pegah Alipoormolabashi, Yeganeh Kordi, Amir- reza Mirzaei, Atharva Naik, Arjun Ashok, Arut Selvan Dhanasekaran, Anjana Arunkumar, David Stap, et al. 2022. Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks. I...
2022
-
[254]
Yue Wang, Dan Qiao, Juntao Li, Jinxiong Chang, Qishen Zhang, Zhongyi Liu, Guannan Zhang, and Min Zhang. 2023. Towards Better Hierarchical Text Classification with Data Generation. InFindings of ACL. 7722–7739
2023
-
[255]
Yi Tay, Vinh Tran, Mostafa Dehghani, Jianmo Ni, Dara Bahri, Harsh Mehta, Zhen Qin, Kai Hui, Zhe Zhao, Jai Gupta, et al. 2022. Transformer memory as a differentiable search index.Advances in NeurIPS35 (2022), 21831–21843
2022
-
[256]
Rafael Teixeira de Lima, Shubham Gupta, Cesar Berrospi Ramis, Lokesh Mishra, Michele Dolfi, Peter Staar, and Panagiotis Vagenas. 2025. Know Your RAG: Dataset Taxonomy and Generation Strategies for Evaluating RAG Systems. In Proceedings of the 31st International Conference on C...
2025
-
[257]
Jonatas Wehrmann, Ricardo Cerri, and Rodrigo C. Barros. 2018. Hierarchical Multi-label Classification Networks. InICML
2018
-
[258]
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. 2022. Chain-of-thought prompting elicits rea- soning in large language models.Advances in NeurIPS35 (2022), 24824–24837
2022
-
[259]
Andrew Trotman, Antti Puurula, and Blake Burgess. 2014. Improvements to BM25 and language models examined. InProceedings of the 19th Australasian Document Computing Symposium. 58–65
2014
-
[260]
Zhucheng Tu and Sarguna Janani Padmanabhan. 2022. MIA 2022 Shared Task Submission: Leveraging Entity Representations, Dense-Sparse Hybrids, and Fusion-in-Decoder for Cross-Lingual Question Answering.arXiv preprint arXiv:2207.01940(2022)
2022 arXiv
-
[261]
Junde Wu, Jiayuan Zhu, Yunli Qi, Jingkun Chen, Min Xu, Filippo Menolascina, and Vicente Grau. 2024. Medical graph rag: Towards safe medical large language model via graph retrieval-augmented generation.arXiv preprint arXiv:2408.04187 (2024)
2024 arXiv
-
[262]
Somin Wadhwa, Silvio Amir, and Byron Wallace. 2023. Revisiting Relation Extraction in the era of Large Language Models. InACL. 15566–15589
2023
-
[263]
Jinfeng Xiao, Mohab Elkaref, Nathan Herr, Geeth De Mel, and Jiawei Han. 2023. Taxonomy-Guided Fine-Grained Entity Set Expansion. InSDM’23
2023
-
[264]
Chengyu Wang, Jianing Wang, Minghui Qiu, Jun Huang, and Ming Gao. 2021. Transprompt: Towards an automatic transferable prompting framework for few-shot text classification. InProceedings of the 2021 conference on empirical methods in natural language processing. 2792–2802
2021
-
[265]
Liang Wang, Nan Yang, Xiaolong Huang, Binxing Jiao, Linjun Yang, Daxin Jiang, Rangan Majumder, and Furu Wei. 2022. Text embeddings by weakly-supervised contrastive pre-training.arXiv preprint arXiv:2212.03533(2022)
2022 arXiv
-
[266]
Shuhe Wang, Xiaofei Sun, Xiaoya Li, Rongbin Ouyang, Fei Wu, Tianwei Zhang, Jiwei Li, and Guoyin Wang. 2023. Gpt-ner: Named entity recognition via large language models.arXiv preprint arXiv:2304.10428(2023)
2023 arXiv
-
[267]
Fangyuan Xu, Weijia Shi, and Eunsol Choi. 2023. Recomp: Improving retrieval- augmented lms with compression and selective augmentation.arXiv preprint arXiv:2310.04408(2023)
2023 arXiv
-
[268]
Mingbin Xu and Hui Jiang. 2016. A FOFE-based local detection approach for named entity recognition and mention detection.arXiv preprint arXiv:1611.00801 (2016)
2016 arXiv
-
[269]
Xuan Wang, Yingjun Guan, Yu Zhang, Qi Li, and Jiawei Han. 2020. Pattern- enhanced Named Entity Recognition with Distant Supervision. In2020 IEEE International Conference on Big Data (Big Data). 818–827
2020
-
[270]
Vikas Yadav and Steven Bethard. 2019. A survey on recent advances in named entity recognition from deep learning models.arXiv preprint arXiv:1910.11470 (2019)
2019 arXiv
-
[271]
Yujing Wang, Yingyan Hou, Haonan Wang, Ziming Miao, Shibin Wu, Qi Chen, Yuqing Xia, Chengmin Chi, Guoshuai Zhao, Zheng Liu, et al. 2022. A neural corpus indexer for document retrieval.Advances in NeurIPS35 (2022), 25600– 25614
2022
-
[272]
Zichao Yang, Diyi Yang, Chris Dyer, Xiaodong He, Alex Smola, and Eduard Hovy. 2016. Hierarchical Attention Networks for Document Classification. InProceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language ...
2016 doi
-
[273]
Zeng Yang, Linhai Zhang, and Deyu Zhou. 2022. SEE-Few: Seed, Expand and Entail for Few-shot Named Entity Recognition. InCOLING. 2540–2550
2022
-
[274]
Zihan Wang, Dheeraj Mekala, and Jingbo Shang. 2021. X-Class: Text Classifica- tion with Extremely Weak Supervision. InNAACL-HLT. 3043–3053
2021
-
[275]
Zihan Wang, Peiyi Wang, Lianzhe Huang, Xin Sun, and Houfeng Wang. 2022. Incorporating Hierarchy into Text Encoder: a Contrastive Learning Approach for Hierarchical Text Classification. InACL. 7109–7119
2022
-
[276]
Michihiro Yasunaga, Hongyu Ren, Antoine Bosselut, Percy Liang, and Jure Leskovec. 2021. QA-GNN: Reasoning with language models and knowledge graphs for question answering.NAACL(2021)
2021
-
[277]
Fanghua Ye, Meng Fang, Shenghui Li, and Emine Yilmaz. 2023. Enhancing conversational search: Large language model-aided informative query rewriting. arXiv preprint arXiv:2310.09716(2023)
2023 arXiv
-
[278]
Yilin Wen, Zifeng Wang, and Jimeng Sun. 2023. Mindmap: Knowledge graph prompting sparks graph of thoughts in large language models.arXiv preprint arXiv:2308.09729(2023)
2023 arXiv
-
[279]
Zhihao Wen and Yuan Fang. 2023. Augmenting low-resource text classification with graph-grounded pre-training and prompting. InProceedings of the 46th International ACM SIGIR Conference on Research and Development in Information Retrieval. 506–516
2023
-
[280]
Changlong Yu, Weiqi Wang, Xin Liu, Jiaxin Bai, Yangqiu Song, Zheng Li, Yifan Gao, Tianyu Cao, and Bing Yin. 2023. FolkScope: Intention Knowledge Graph Construction for E-commerce Commonsense Discovery. InFindings of ACL. 1173–1191
2023
-
[281]
Guangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han, and Mike Lewis
-
[282]
Efficient streaming language models with attention sinks.arXiv preprint arXiv:2309.17453(2023)
2023 arXiv
-
[284]
Yuxin Xiao, Zecheng Zhang, Yuning Mao, Carl Yang, and Jiawei Han. 2022. SAIS: Supervising and Augmenting Intermediate Steps for Document-Level Relation Extraction. InNAACL-HLT. 2395–2409
2022
-
[285]
Yiqing Xie, Jiaming Shen, Sha Li, Yuning Mao, and Jiawei Han. 2022. Eider: Empowering Document-level Relation Extraction with Efficient Evidence Ex- traction and Inference-stage Fusion. InFindings of ACL. 257–268
2022
-
[286]
Lee Xiong, Chenyan Xiong, Ye Li, Kwok-Fung Tang, Jialin Liu, Paul Bennett, Junaid Ahmed, and Arnold Overwijk. 2020. Approximate nearest neighbor nega- tive contrastive learning for dense text retrieval.arXiv preprint arXiv:2007.00808 (2020)
2020 arXiv
-
[289]
Zhentao Xu, Mark Jerome Cruz, Matthew Guevara, Tie Wang, Manasi Desh- pande, Xiaofeng Wang, and Zheng Li. 2024. Retrieval-augmented generation with knowledge graphs for customer service question answering. InProceedings of the 47th International ACM SIGIR Conference on Researc...
2024
-
[291]
Hang Yan, Junqi Dai, Tuo Ji, Xipeng Qiu, and Zheng Zhang. 2021. A Unified Generative Framework for Aspect-based Sentiment Analysis. InACL-IJCNLP. 2416–2429
2021
-
[294]
Michihiro Yasunaga, Armen Aghajanyan, Weijia Shi, Rich James, Jure Leskovec, Percy Liang, Mike Lewis, Luke Zettlemoyer, and Wen-tau Yih. 2022. Retrieval- augmented multimodal language modeling.arXiv preprint arXiv:2211.12561 (2022)
2022 arXiv
-
[295]
Michihiro Yasunaga, Antoine Bosselut, Hongyu Ren, Xikun Zhang, Christo- pher D Manning, Percy S Liang, and Jure Leskovec. 2022. Deep bidirectional A Survey on Retrieval And Structuring Augmented Generation with Large Language Models KDD ’25, August 3–7, 2025, Toronto, ON, Cana...
2022
-
[298]
Sizhe Zhou Heng Ji Jiawei Ha Yizhu Jiao, Sha Li. 2024. Text2DB: Integration- Aware Information Extraction with Large Language Model Agents. (2024)
2024
-
[299]
Changlong Yu, Weiqi Wang, Xin Liu, Jiaxin Bai, Yangqiu Song, Zheng Li, Yi- fan Gao, Tianyu Cao, and Bing Yin. 2022. FolkScope: Intention knowledge graph construction for e-commerce commonsense discovery.arXiv preprint arXiv:2211.08316(2022)
2022 arXiv
-
[2017]
Proximal Policy Optimization Algorithms.arXiv preprint arXiv:1707.06347 (2017)
2017 arXiv
-
[2018]
In EMNLP
Learning Named Entity Tagger using Domain-Specific Dictionary. In EMNLP
-
[2019]
InEMNLP-IJCNLP
A Little Annotation does a Lot of Good: A Study in Bootstrapping Low- resource Named Entity Recognizers. InEMNLP-IJCNLP. 5164–5174
-
[2020]
Guiding Corpus-based Set Expansion by Auxiliary Sets Generation and Co-Expansion. InWWW. 2188–2198
-
[2021]
TaxoClass: Hierarchical Multi-Label Text Classification Using Only Class Names. InNAACL
-
[2022]
ACM Transactions on Information Systems (TOIS)40, 4 (2022), 1–42
Semantic models for the first-stage retrieval: A comprehensive review. ACM Transactions on Information Systems (TOIS)40, 4 (2022), 1–42
2022
-
[2023]
Advances in NeurIPS36 (2023), 43780–43799
Lift yourself up: Retrieval-augmented text generation with self-memory. Advances in NeurIPS36 (2023), 43780–43799
2023
-
[2024]
GenRES: Rethinking Evaluation for Generative Relation Extraction in the Era of Large Language Models. InProceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), Kevin...
2024 doi
-
[7184]
doi:10.18653/v1/2024.emnlp-main.407
2024 doi
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.