REVIEW 3 major objections 6 minor 117 references
Bridging the Gap: From Ad-hoc to Proactive Search in Conversations
T0 review · 3 major / 6 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read Conversation context, converted into ad-hoc queries, lets standard retrievers handle proactive search.
desk verdict Conv2Query is a solid, useful paper with a genuine contribution, and the main circularity worry about LLaMA-family filtering does not survive contact with the results. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the Conv2Query pipeline, a conversation-to-query mapping trained on pseudo labels. For each conversation turn with a relevant document, a Doc2Query model generates candidate ad-hoc queries; QF-DC then scores each candidate by its relevance to that document and its relevance to the conversational context, using a cross-encoder-style relevance model, and selects the highest-scoring query as the training target. An LLM (Mistral-7B-Instruct in the main experiments) is fine-tuned with QLoRA to produce such a query from the conversation alone, and at inference that generated query is passed to any ad-hoc retriever. The same filtered queries also serve as positive examples for fine-tuning neural retrievers on PSC. This machinery transfers the retriever's knowledge by keeping both training and inference in the ad-hoc query format.
What would settle it
Replicate the experiments on a proactive-search dataset that also has human-written ad-hoc queries for each turn, and compare three inputs to the same retriever: the human query, Conv2Query's generated query, and the raw conversational context. If the human-query oracle does not clearly beat Conv2Query, or if a prompting-only LLM matches the full learned pipeline on held-out turns, the paper's claim that the Doc2Query-plus-QF-DC training signal is responsible for the improvement would be contradicted.
Extended reading notes
Core claim
The central claim is that the input mismatch between ad-hoc search and PSC, not the retrieval model itself, is what limits proactive retrieval quality. Conv2Query removes the mismatch by training an LLM to map conversational context to ad-hoc queries, using pseudo query targets generated by a Doc2Query model from each turn's relevant document and then filtered by QF-DC, a mechanism that keeps only queries that are both relevant to the document and aligned with the conversation. With these queries, off-the-shelf BM25, ANCE, SPLADE++, and RepLLaMA all improve significantly over using raw conversational context, and the same pseudo queries serve as better fine-tuning data than raw context. On the ProCIS and WebDisc datasets, the gains hold under both settings studied, including the harder interest-anticipation setting.
Load-bearing premise
The whole training signal rests on the assumption that Doc2Query queries generated from a turn's relevant document, after QF-DC filtering, are faithful stand-ins for the implicit search intent of that conversation; if the pseudo queries are off-target, the learned mapping and downstream retrieval gains would degrade.
Editorial extensions
If this is right
- Any ad-hoc retriever, lexical or neural, can be applied to proactive search simply by placing Conv2Query in front of it, with no retriever retraining.
- Fine-tuning retrievers on pseudo ad-hoc queries transfers pre-trained ad-hoc knowledge better than fine-tuning on raw conversational context, as shown by consistent gains for ANCE, SPLADE++, and RepLLaMA.
- Lexical retrievers such as BM25, which cannot be fine-tuned on PSC data, benefit substantially from the noise-removing query conversion.
- Larger LLMs used as the Conv2Query generator yield stronger retrieval quality, suggesting the mapping benefits from model scale.
- The prompting-only variant of Conv2Query outperforms earlier text-window and description-based methods but is itself outperformed by the fine-tuned variant, indicating that learning from filtered pseudo queries adds value beyond LLM prompting.
Reading between the lines
- The same pseudo-label recipe could be applied to other query-free retrieval settings, such as suggesting queries in retrieval-augmented generation, a direction the paper gestures at.
- Because QF-DC explicitly aligns the target query with the conversation, the framework is likely sensitive to the quality of the relevance model used for filtering; with a weaker relevance model, the advantage over query-document filtering alone may shrink.
- The paper leaves retrieval timing untouched; a natural extension would combine a learned 'when to retrieve' predictor with Conv2Query's 'what to retrieve' query generation.
- If a benchmark with human-written queries per conversation became available, Conv2Query could be evaluated against the oracle upper bound, clarifying how much headroom remains in the conversation-to-query mapping.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper addresses proactive search in conversations (PSC), where the system must retrieve relevant documents from conversational context without an explicit user query. The authors argue that off-the-shelf ad-hoc retrievers suffer from an input mismatch because they are trained on short ad-hoc queries but receive long, noisy conversational context. They propose Conv2Query, which trains an LLM (Mistral-7B) to map a conversational context to an ad-hoc query. Training targets are pseudo-queries generated from the annotated relevant documents by a Doc2Query model, filtered by a new mechanism QF-DC that scores each candidate by both query-document relevance and query-conversation relevance using RankLLaMA. At inference, the generated query is fed to any ad-hoc retriever, optionally after fine-tuning that retriever on pseudo-query/document pairs. Experiments on ProCIS and WebDisc under conversation contextualisation and interest anticipation show large, statistically significant gains over raw-context baselines, prompting-only variants, Text Window and LMGR, for BM25, ANCE, SPLADE++ and RepLLaMA. The paper also shows that QF-DC accelerates convergence and that results are robust across LLM families (3B-22B). Limitations on timing prediction and dataset diversity are acknowledged in Section 8.
Significance. If the results hold, Conv2Query is a practical and conceptually simple adaptation layer: it lets existing ad-hoc retrievers be applied to query-free conversational search without retriever fine-tuning. The paper's strengths include extensive evaluation (two datasets, two settings, four retrievers including BM25 and a state-of-the-art LLaMA retriever), a clean two-stage training-data construction pipeline with an ablation for the filter, and public release of code and data. The consistent gains across BM25, ANCE, SPLADE++, and RepLLaMA make the central empirical claim credible. However, the statistical reporting and the use of a LLaMA-family relevance model to select training targets for a LLaMA-family retriever leave open questions that should be resolved before the general claim 'any ad-hoc retriever' is fully established.
major comments (3)
- [Section 5, Tables 1 and 2] The paper reports only point estimates and marks statistical significance with a paired t-test at p < 0.05, but it does not report variances, confidence intervals, the number of paired units, the number of runs, or random seeds. Because the central claim is that Conv2Query 'significantly improves' performance across all metrics and settings, these omissions make the significance statements impossible to verify. Please report standard deviations or bootstrap confidence intervals (e.g., across at least three QLoRA training runs, or across conversations/turns), and specify whether the paired t-test is paired by conversation or by turn.
- [Section 4.2, Eqs. (3)-(4); Section 5, Implementation details] QF-DC selects the pseudo-query training targets using RankLLaMA as the relevance model frelevance, and the best-performing retriever in Tables 1 and 2 is RepLLaMA, the retriever counterpart of the same LLaMA/MS-MARCO model family. Since the training signal for Conv2Query is entirely determined by RankLLaMA's relevance scores, the selected targets may be preferentially compatible with LLaMA-family retrievers, and this potential confound is not examined in the paper. I do not regard this as fatal circularity because BM25, ANCE, and SPLADE++ also improve consistently, but the authors should test the stability of the filtered target set under an alternative relevance model (e.g., MonoT5 or TCT-ColBERT), and ideally report whether the RepLLaMA gains persist when the filter comes from a different model family.
- [Section 4.1-4.2 and Section 7.1] The pseudo-query targets are generated by Doc2Query from the relevant document and then filtered by QF-DC, but the paper does not directly validate whether the selected queries reflect the implicit conversational information need. The learning-curve comparison in Figure 4 shows that QF-DC helps relative to QF-D and Random, but it does not establish the absolute quality of the targets. A direct oracle experiment (retrieving with the filtered pseudo-queries themselves, before any learned mapping) and a short qualitative sample of generated queries would strengthen the claim that Conv2Query learns a meaningful conversation-to-query mapping rather than exploiting document-specific vocabulary.
minor comments (6)
- [Section 2.2.1 and Section 5] The reference [9] for 'Doc2Query-Llama2' appears to point to DeeperImpact (Basnet et al., 2024), which is not the Llama-2-based Doc2Query model described in the text; please verify and cite the correct source.
- [Section 5, Implementation details] The choice of truncation length 512 for conversational-context inputs and 32 for Conv2Query-generated queries is justified only as 'exceeds the average context length'; a brief sensitivity check or a statement that longer lengths produced no further gains would be helpful.
- [Section 7.1] The opening sentence of Section 7.1, 'we study the impact of query mechanism,' should read 'query filtering mechanism.'
- [Section 7.2] The text 'for llama, we use Llama-3.2-3B-Instruct' should capitalize 'Llama' for consistency with the rest of the paper.
- [Tables 1 and 2 captions] The abbreviation 'Conv2Q' appears in the table captions without being defined at first use; please define it or use 'Conv2Query' consistently throughout the captions.
- [Section 5, Baselines] The Text Window and LMGR baselines are re-implemented with RepLLaMA in place of the original retrievers; a sentence noting that this replacement was made only for comparability and that the variants were not validated against the original published systems would improve transparency.
Circularity Check
No significant circularity: Conv2Query is trained on pseudo-labels but evaluated on external relevance judgments; the filter/retriever model-family overlap is a confound, not a definitional reduction.
full rationale
The paper's derivation chain is a standard supervised pipeline: Doc2Query (an external MS MARCO-trained model) generates query candidates from curated relevant documents; QF-DC filters them using a public relevance model (RankLLaMA); Conv2Query is trained to map conversational context to the filtered query; and retrieval is evaluated on held-out sparse and dense relevance judgments. No equation in Sections 4.1-4.5 makes the evaluation target a function of the training target by construction: headline numbers are measured against dataset annotations (user-added Wikipedia/webpage links and human dense judgments), not against the pseudo-labels used for training. The overlap between RankLLaMA (filter) and RepLLaMA (evaluated retriever) is a possible confound, but it is not a definitional reduction, and improvements are also reported for BM25, ANCE, and SPLADE++, so the central claim does not collapse to a self-consistency check. Reliance on prior work (Doc2Query, RankLLaMA, ProCIS, WebDisc) is ordinary use of external artifacts, not a self-citation chain; no uniqueness theorem or unverified ansatz is imported to forbid alternatives. No circular steps were found.
Assumptions & free parameters
free parameters (6)
- Number of Doc2Query candidates n =
100
- Top-k sampling parameter k for Doc2Query =
10
- Aggregation function f_aggregate =
summation
- Retriever query truncation length =
512 (context input), 32 (generated queries)
- BM25 k1 and b =
k1=0.9, b=0.4 (ProCIS); k1=8/7, b=0.99 (WebDisc)
- QLoRA hyperparameters (learning rate, LoRA rank, etc.)
assumptions (4)
- domain assumption The Doc2Query model, pre-trained on MS MARCO, can generate queries that reveal the implicit search intents of conversations when given the relevant document.
- domain assumption The QF-DC filter, using RankLLaMA for relevance, can accurately measure both query-document and query-conversation relevance.
- domain assumption The conversational context alone is sufficient to predict a useful ad-hoc query for the relevant documents.
- domain assumption The user-added hyperlinks in the datasets serve as reliable relevance judgments.
Cite this review
Pith. "Pith review of Bridging the Gap: From Ad-hoc to Proactive Search in Conversations." pith.science (2026). https://pith.science/paper/DPVIYR33
@misc{pith2026250600983,
author = {Pith},
title = {Pith review of: Bridging the Gap: From Ad-hoc to Proactive Search in Conversations},
year = {2026},
howpublished = {\url{https://pith.science/paper/DPVIYR33}},
note = {Machine review of arXiv:2506.00983}
}
read the original abstract
Proactive search in conversations (PSC) aims to reduce user effort in formulating explicit queries by proactively retrieving useful relevant information given conversational context. Previous work in PSC either directly uses this context as input to off-the-shelf ad-hoc retrievers or further fine-tunes them on PSC data. However, ad-hoc retrievers are pre-trained on short and concise queries, while the PSC input is longer and noisier. This input mismatch between ad-hoc search and PSC limits retrieval quality. While fine-tuning on PSC data helps, its benefits remain constrained by this input gap. In this work, we propose Conv2Query, a novel conversation-to-query framework that adapts ad-hoc retrievers to PSC by bridging the input gap between ad-hoc search and PSC. Conv2Query maps conversational context into ad-hoc queries, which can either be used as input for off-the-shelf ad-hoc retrievers or for further fine-tuning on PSC data. Extensive experiments on two PSC datasets show that Conv2Query significantly improves ad-hoc retrievers' performance, both when used directly and after fine-tuning on PSC.
Figures
Reference graph
Works this paper leans on
-
[1]
Zahra Abbasiantaeb, Chuan Meng, Leif Azzopardi, and Mohammad Aliannejadi
-
[2]
T Ahmed and S Bulathwela. 2022. Towards Proactive Information Retrieval in Noisy Text with Wikipedia Concepts. In CEUR Workshop Proceedings, Vol. 3318. 1–12
2022
-
[3]
Salvatore Andolina, Valeria Orso, Hendrik Schneider, Khalil Klouche, Tuukka Ruotsalo, Luciano Gamberini, and Giulio Jacucci. 2018. Investigating Proactive Search Support in Conversations. In DIS. 1295–1307
2018
-
[6]
Negar Arabzadeh, Chuan Meng, Mohammad Aliannejadi, and Ebrahim Bagheri
-
[7]
Payal Bajaj, Daniel Campos, Nick Craswell, Li Deng, Xiaodong Liu Jianfeng Gao, Rangan Majumder, Andrew McNamara, Bhaskar Mitra, Tri Nguyen, Mir Rosen- berg, Xia Song, Alina Stoica, Saurabh Tiwary, and Tong Wang. 2016. MS MARCO: A Human Generated Machine Reading Comprehension Dataset. In NIPS
2016
-
[8]
In SIGIR-AP
Query Performance Prediction: Techniques and Applications in Modern Information Retrieval. In SIGIR-AP. 291–294
-
[9]
Soyuj Basnet, Jerry Gou, Antonio Mallia, and Torsten Suel. 2024. DeeperImpact: Optimizing Sparse Learned Index Structures. arXiv preprint arXiv:2405.17093 (2024)
arXiv 2024
-
[10]
Query Performance Prediction: Theory, Techniques and Applications. In WSDM. 991–994
Show all 117 references
-
[11]
Francesco Bonchi, Raffaele Perego, Fabrizio Silvestri, Hossein Vahabi, and Rossano Venturini. 2012. Efficient Query Recommendations in the Long Tail via Center-Piece Subgraphs. In SIGIR. 345–354
2012
-
[12]
Vevake Balaraman and Bernardo Magnini. 2020. Proactive Systems and Influ- enceable Users: Simulating Proactivity in Task-oriented Dialogues. InSEMDIAL
2020
-
[13]
Yash Butala, Siddhant Garg, Pratyay Banerjee, and Amita Misra. 2024. ProMISe: A Proactive Multi-turn Dialogue Dataset for Information-seeking Intent Resolu- tion. In EACL. 1774–1789
2024
-
[14]
Paolo Boldi, Francesco Bonchi, Carlos Castillo, Debora Donato, Aristides Gionis, and Sebastiano Vigna. 2008. The Query-flow Graph: Model and Applications. In CIKM. 609–618
2008
-
[15]
Wanyu Chen, Fei Cai, Honghui Chen, and Maarten de Rijke. 2018. Attention- based Hierarchical Neural Query Suggestion. In SIGIR. 1093–1096
2018
-
[16]
Luiz Bonifacio, Hugo Abonizio, Marzieh Fadaee, and Rodrigo Nogueira. 2022. InPars: Unsupervised Dataset Generation for Information Retrieval. In SIGIR. 2387–2392
2022
-
[17]
Zhuyun Dai, Vincent Y Zhao, Ji Ma, Yi Luan, Jianmo Ni, Jing Lu, Anton Bakalov, Kelvin Guu, Keith Hall, and Ming-Wei Chang. 2023. Promptagator: Few-shot Dense Retrieval From 8 Examples. In ICLR
2023
-
[18]
Ramraj Chandradevan, Kaustubh Dhole, and Eugene Agichtein. 2024. DUQGen: Effective Unsupervised Domain Adaptation of Neural Rankers by Diversifying Synthetic Query Generation. In NAACL, Kevin Duh, Helena Gomez, and Steven Bethard (Eds.)
2024
-
[19]
Yang Deng, Wenqiang Lei, Minlie Huang, and Tat-Seng Chua. 2023. Rethinking Conversational Agents in the Era of LLMs: Proactivity, Non-collaborativity, and Beyond. In SIGIR. 298–301
2023
-
[20]
Zhicong Cheng, Bin Gao, and Tie-Yan Liu. 2010. Actively Predicting Diverse Search Intent from User Browsing Behaviors. In WWW. 221–230
2010
-
[21]
Yang Deng, Wenqiang Lei, Wenxuan Zhang, Wai Lam, and Tat-Seng Chua. 2022. PACIFIC: Towards Proactive Conversational Question Answering over Tabular and Textual Data in Finance. In EMNLP. 6970–6984
2022
-
[22]
Mostafa Dehghani, Sascha Rothe, Enrique Alfonseca, and Pascal Fleury. 2017. Learning to Attend, Copy, and Generate for Session-Based Query Suggestion. In CIKM. 1747–1756
2017
-
[23]
Yang Deng, Lizi Liao, Zhonghua Zheng, Grace Hui Yang, and Tat-Seng Chua
-
[24]
Yang Deng, Wenqiang Lei, Wai Lam, and Tat-Seng Chua. 2023. A Survey on Proactive Dialogue Systems: Problems, Methods, and Prospects. In IJCAI. 6583– 6591
2023
-
[25]
Susan Dumais, Edward Cutrell, Raman Sarin, and Eric Horvitz. 2004. Implicit Queries (IQ) for Contextualized Search. In SIGIR. 594–594
2004
-
[26]
Yang Deng, Lizi Liao, Liang Chen, Hongru Wang, Wenqiang Lei, and Tat-Seng Chua. 2023. Prompting and Evaluating Large Language Models for Proac- tive Dialogues: Clarification, Target-guided, and Non-collaboration. In EMNLP. 10602–10621
2023
-
[27]
Henry Feild and James Allan. 2013. Task-Aware Query Recommendation. In SIGIR. 83–92
2013
-
[28]
In SIGIR
Towards Human-centered Proactive Conversational Agents. In SIGIR. 807–818
-
[29]
Tim Dettmers, Artidoro Pagnoni, Ari Holtzman, and Luke Zettlemoyer. 2023. QLoRA: Efficient Finetuning of Quantized LLMs.arXiv preprint arXiv:2305.14314 (2023)
2023 arXiv
-
[30]
Thibault Formal, Benjamin Piwowarski, and Stéphane Clinchant. 2021. SPLADE: Sparse Lexical and Expansion Model for First Stage Ranking. In SIGIR. 2288– 2292
2021
-
[31]
Desmond Elliott and Joemon M Jose. 2009. A Proactive Personalised Retrieval System. In CIKM. 1935–1938
2009
-
[32]
Mitko Gospodinov, Sean MacAvaney, and Craig Macdonald. 2023. Doc2Query–: When Less is More. In ECIR. Springer, 414–422
2023
-
[33]
Stephen Fitchett and Andy Cockburn. 2012. AccessRank: Predicting What Users Will Do Next. In SIGCHI. 2239–2242
2012
-
[34]
Thibault Formal, Carlos Lassance, Benjamin Piwowarski, and Stéphane Clin- chant. 2022. From Distillation to Hard Negative Sampling: Making Sparse Neural IR Models More Effective. In SIGIR. 2353–2359
2022
-
[35]
Yunah Jang, Kang-il Lee, Hyunkyung Bae, Seungpil Won, Hwanhee Lee, and Kyomin Jung. 2024. IterCQR: Iterative Conversational Query Reformulation with Retrieval Guidance. In EMNLP. 8121–813
2024
-
[36]
Debasis Ganguly, Dipasree Pal, Manisha Verma, and Procheta Sen. 2020. Overview of RCD-2020, the FIRE-2020 track on retrieval from conversational dialogues. In FIRE. 33–36
2020
-
[37]
Vitor Jeronymo, Luiz Bonifacio, Hugo Abonizio, Marzieh Fadaee, Roberto Lotufo, Jakub Zavrel, and Rodrigo Nogueira. 2023. InPars-v2: Large Language Mod- els as Efficient Dataset Generators for Information Retrieval. arXiv preprint arXiv:2301.01820 (2023)
2023 arXiv
-
[38]
Nam Le Hai, Thomas Gerald, Thibault Formal, Jian-Yun Nie, Benjamin Pi- wowarski, and Laure Soulier. 2023. CoSPLADE: Contextualizing SPLADE for Conversational Information Retrieval. In ECIR
2023
-
[39]
Monika Henzinger, Bay-Wei Chang, Brian Milch, and Sergey Brin. 2003. Query- Free News Search. In WWW. 1–10
2003
-
[40]
Markus Koskela, Petri Luukkonen, Tuukka Ruotsalo, Mats Sjöberg, and Patrik Floréen. 2018. Proactive Information Retrieval by Capturing Search Intent from Primary Task Context. TiiS 8, 3 (2018), 1–25
2018
-
[41]
Kalervo Järvelin and Jaana Kekäläinen. 2002. Cumulated Gain-Based Evaluation of IR Techniques. TOIS 20, 4 (2002), 422–446
2002
-
[42]
Jing Yang Lee, Seokhwan Kim, Kartik Mehta, Jiun-Yu Kao, Yu-Hsiang Lin, and Arpit Gupta. 2024. Redefining Proactivity for Information Seeking Dialogue. arXiv preprint arXiv:2410.15297 (2024)
2024 arXiv
-
[43]
Gareth JF Jones, Procheta Sen, Debasis Ganguly, and Emine Yilmaz. 2022. Work- shop on Proactive and Agent-Supported Information Retrieval (PASIR). InCIKM. 5167–5168
2022
-
[44]
Weize Kong, Rui Li, Jie Luo, Aston Zhang, Yi Chang, and James Allan. 2015. Predicting Search Intent Based on Pre-Search Context. In SIGIR. 503–512
2015
-
[45]
Lizi Liao, Grace Hui Yang, and Chirag Shah. 2023. Proactive Conversational Agents in the Post-ChatGPT World. In SIGIR. 3452–3455
2023
-
[46]
Yilong Lai, Jialong Wu, Congzhi Zhang, Haowen Sun, and Deyu Zhou. 2025. AdaCQR: Enhancing Query Reformulation for Conversational Search via Sparse and Dense Retrieval Alignment. COLING (2025)
2025
-
[47]
Sheng-Chieh Lin, Jheng-Hong Yang, and Jimmy Lin. 2021. Contextualized Query Embeddings for Conversational Search. In EMNLP. 1004–1015
2021
-
[48]
Ruirui Li, Ben Kao, Bin Bi, Reynold Cheng, and Eric Lo. 2012. DQR: A Proba- bilistic Approach to Diversified Query Recommendation. In CIKM. 16–25
2012
-
[49]
Lizi Liao, Grace Hui Yang, and Chirag Shah. 2023. Proactive Conversational Agents. In WSDM. 1244–1247
2023
-
[50]
Xueguang Ma, Liang Wang, Nan Yang, Furu Wei, and Jimmy Lin. 2024. Fine- Tuning LLaMA for Multi-Stage Text Retrieval. In SIGIR. 2421–2425
2024
-
[51]
Daniel J Liebling, Paul N Bennett, and Ryen W White. 2012. Anticipatory Search: Using Context to Initiate Search. In SIGIR. 1035–1036
2012
-
[52]
Kelong Mao, Zhicheng Dou, Bang Liu, Hongjin Qian, Fengran Mo, Xiangli Wu, Xiaohua Cheng, and Zhao Cao. 2023. Search-Oriented Conversational Query Editing. In Findings of ACL. 4160–4172
2023
-
[53]
Sheng-Chieh Lin, Jheng-Hong Yang, and Jimmy Lin. 2021. In-Batch Negatives for Knowledge Distillation with Tightly-Coupled Teachers for Dense Retrieval. In RepL4NLP-2021. 163–173
2021
-
[54]
Lili Lu, Chuan Meng, Federico Ravenda, Mohammad Aliannejadi, and Fabio Crestani. 2025. Zero-Shot and Efficient Clarification Need Prediction in Conver- sational Search. In ECIR. 389–404
2025
-
[55]
Chuan Meng, Mohammad Aliannejadi, and Maarten de Rijke. 2023. System Initiative Prediction for Multi-turn Conversational Information Seeking. In CIKM. 1807–1817. Bridging the Gap: From Ad-hoc to Proactive Search in Conversations SIGIR ’25, July 13–18, 2025, Padua, Italy
2023
-
[56]
Kelong Mao, Chenlong Deng, Haonan Chen, Fengran Mo, Zheng Liu, Tetsuya Sakai, and Zhicheng Dou. 2024. ChatRetriever: Adapting Large Language Models for Generalized and Robust Conversational Dense Retrieval. In EMNLP. 1227–1240
2024
-
[57]
Chuan Meng, Negar Arabzadeh, Arian Askari, Mohammad Aliannejadi, and Maarten de Rijke. 2025. Query Performance Prediction using Relevance Judg- ments Generated by Large Language Models. TOIS (2025)
2025
-
[58]
Kelong Mao, Zhicheng Dou, Fengran Mo, Jiewen Hou, Haonan Chen, and Hongjin Qian. 2023. Large Language Models Know Your Contextual Search Intent: A Prompting Framework for Conversational Search. In EMNLP. 1211– 1225
2023
-
[59]
Chuan Meng. 2024. Query Performance Prediction for Conversational Search and Beyond. In SIGIR. 3077–3077
2024
-
[60]
Fengran Mo, Abbas Ghaddar, Kelong Mao, Mehdi Rezagholizadeh, Boxing Chen, Qun Liu, and Jian-Yun Nie. 2024. CHIQ: Contextual History Enhancement for Improving Query Rewriting in Conversational Search. In EMNLP. 2253–2268
2024
-
[61]
Chuan Meng, Negar Arabzadeh, Mohammad Aliannejadi, and Maarten de Rijke
-
[62]
Fengran Mo, Kelong Mao, Yutao Zhu, Yihong Wu, Kaiyu Huang, and Jian-Yun Nie. 2023. ConvGQR: Generative Query Reformulation for Conversational Search. In ACL. 4998–5012
2023
-
[63]
Fengran Mo, Chuan Meng, Mohammad Aliannejadi, and Jian-Yun Nie. 2025. Conversational Search: From Fundamentals to Frontiers in the LLM Era. In SIGIR
2025
-
[64]
Chuan Meng, Guglielmo Faggioli, Mohammad Aliannejadi, Nicola Ferro, and Josiane Mothe. 2025. QPP++ 2025: Query Performance Prediction and its Appli- cations in the Era of Large Language Models. In ECIR. 319–325
2025
-
[65]
Kshitij Mishra, Azlaan Mustafa Samad, Palak Totala, and Asif Ekbal. 2022. PEPDS: A Polite and Empathetic Persuasive Dialogue System for Charity Donation. In COLING. 424–440
2022
-
[66]
Fengran Mo, Chen Qu, Kelong Mao, Tianyu Zhu, Zhan Su, Kaiyu Huang, and Jian-Yun Nie. 2024. History-Aware Conversational Dense Retrieval. Findings of ACL (2024)
2024
-
[67]
Fengran Mo, Kelong Mao, Ziliang Zhao, Hongjin Qian, Haonan Chen, Yiruo Cheng, Xiaoxi Li, Yutao Zhu, Zhicheng Dou, and Jian-Yun Nie. 2024. A Survey of Conversational Search. arXiv preprint arXiv:2410.15576 (2024)
2024 arXiv
-
[68]
Rodrigo Nogueira, Zhiying Jiang, Ronak Pradeep, and Jimmy Lin. 2020. Doc- ument Ranking with a Pretrained Sequence-to-Sequence Model. In EMNLP. 708–718
2020
-
[69]
Rodrigo Nogueira and Jimmy Lin. 2019. From doc2query to docTTTTTquery. (2019)
2019
-
[70]
Fengran Mo, Jian-Yun Nie, Kaiyu Huang, Kelong Mao, Yutao Zhu, Peng Li, and Yang Liu. 2023. Learning to Relate to Previous Turns in Conversational Search. In KDD. 1722–1732
2023
-
[71]
Fengran Mo, Chen Qu, Kelong Mao, Yihong Wu, Zhan Su, Kaiyu Huang, and Jian-Yun Nie. 2024. Aligning Query Representation with Rewritten Query and Relevance Judgments in Conversational Search. In CIKM. 1700–1710
2024
-
[72]
Dae Hoon Park, Yi Fang, Mengwen Liu, and ChengXiang Zhai. 2016. Mobile App Retrieval for Social Media Users via Inference of Implicit Intent in Social Media Text. In CIKM. 959–968
2016
-
[73]
Cristina Ioana Muntean, Franco Maria Nardini, Fabrizio Silvestri, and Marcin Sydow. 2013. Learning to Shorten Query Sessions. In WWW. 131–132
2013
-
[74]
Bradley James Rhodes and Pattie Maes. 2000. Just-In-Time Information Retrieval. IBM Systems journal 39, 3.4 (2000), 685–704
2000
-
[75]
Bradley J Rhodes and Thad Starner. 1996. Remembrance Agent: A Continuously Running Automated Information Retrieval System. In PAAM, Vol. 96. 487–495
1996
-
[76]
Rodrigo Nogueira, Wei Yang, Jimmy Lin, and Kyunghyun Cho. 2019. Document Expansion by Query Prediction. arXiv preprint arXiv:1904.08375 (2019)
2019 arXiv
-
[77]
Dipasree Pal and Debasis Ganguly. 2021. Effective Query Formulation in Con- versation Contextualization: A Query Specificity-based Approach. In ICTIR. 177–183
2021
-
[78]
Corbin Rosset, Chenyan Xiong, Xia Song, Daniel Campos, Nick Craswell, Saurabh Tiwary, and Paul Bennett. 2020. Leading Conversational Search by Suggesting Useful Questions. In The web conference. 1160–1170
2020
-
[79]
David Rau and Jaap Kamps. 2022. The Role of Complex NLP in Transformers for Text Ranking. In ICTIR. 153–160
2022
-
[80]
Procheta Sen, Debasis Ganguly, and Gareth Jones. 2018. Procrastination is the Thief of Time: Evaluating the Effectiveness of Proactive Search Systems. In SIGIR. 1157–1160
2018
-
[81]
Procheta Sen, Debasis Ganguly, and Gareth JF Jones. 2021. I Know What You Need: Investigating Document Retrieval Effectiveness with Partial Session Contexts. TOIS 40, 3 (2021), 1–30
2021
-
[82]
Stephen E Robertson, Steve Walker, Susan Jones, Micheline M Hancock-Beaulieu, Mike Gatford, et al . 1995. Okapi at TREC-3. Nist Special Publication Sp 109 (1995), 109
1995
-
[83]
Kevin Ros, Matthew Jin, Jacob Levine, and ChengXiang Zhai. 2023. Retrieving Webpages Using Online Discussions. In ICTIR. 159–168
2023
-
[84]
Anna Shtok, Oren Kurland, David Carmel, Fiana Raiber, and Gad Markovits
-
[85]
Chris Samarinas and Hamed Zamani. 2024. ProCIS: A Benchmark for Proactive Retrieval in Conversations. In SIGIR. 830–840
2024
-
[86]
Alessandro Sordoni, Yoshua Bengio, Hossein Vahabi, Christina Lioma, Jakob Grue Simonsen, and Jian-Yun Nie. 2015. A Hierarchical Recurrent Encoder- Decoder for Generative Context-Aware Query Suggestion. In CIKM. 553–562
2015
-
[87]
Yueming Sun and Yi Zhang. 2018. Conversational Recommender System. In SIGIR. 235–244
2018
-
[88]
Haizhou Shi, Zihao Xu, Hengyi Wang, Weiyi Qin, Wenyuan Wang, Yibin Wang, Zifeng Wang, Sayna Ebrahimi, and Hao Wang. 2024. Continual Learning of Large Language Models: A Comprehensive Survey. arXiv preprint arXiv:2404.16789 (2024)
2024 arXiv
-
[89]
Milad Shokouhi and Qi Guo. 2015. From Queries to Cards: Re-ranking Proactive Card Recommendations Based on Reactive Search History. In SIGIR. 695–704
2015
-
[90]
Tung Vuong, Giulio Jacucci, and Tuukka Ruotsalo. 2017. Proactive Information Retrieval via Screen Surveillance. In SIGIR. 1313–1316
2017
-
[91]
Tung Vuong, Giulio Jacucci, and Tuukka Ruotsalo. 2017. Watching inside the Screen: Digital Activity Monitoring for Task Recognition and Proactive Information Retrieval. Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies 1, 3 (2017), 1–23
2017
-
[92]
Yang Song and Qi Guo. 2016. Query-Less: Predicting Task Repetition for NextGen Proactive Search and Recommendation Engines. In WWW. 543–553
2016
-
[93]
Bin Wu, Chenyan Xiong, Maosong Sun, and Zhiyuan Liu. 2018. Query Sugges- tion with Feedback Memory Network. In Proceedings of the 2018 World Wide Web Conference. 1563–1571
2018
-
[94]
Yun Xing, Jian Kang, Aoran Xiao, Jiahao Nie, Ling Shao, and Shijian Lu. 2024. Rewrite Caption Semantics: Bridging Semantic Gaps for Language-Supervised Semantic Segmentation. NeurIPS 36 (2024)
2024
-
[95]
Anuja Tayal and Aman Tyagi. 2024. Dynamic Contexts for Generating Sugges- tion Questions in RAG Based Conversational Systems. In WWW. 1338–1341
2024
-
[96]
Tuan A Tran, Sven Schwarz, Claudia Niederée, Heiko Maus, and Nattiya Kan- habua. 2016. The Forgotten Needle in My Collections: Task-Aware Ranking of Documents in Semantic Information Space. In CHIIR. 13–22
2016
-
[97]
Liu Yang, Qi Guo, Yang Song, Sha Meng, Milad Shokouhi, Kieran McDonald, and W Bruce Croft. 2016. Modeling User Interests for Zero-Query Ranking. In ECIR. Springer, 171–184
2016
-
[98]
Liu Yang, Hamed Zamani, Yongfeng Zhang, Jiafeng Guo, and W Bruce Croft
-
[99]
Jian Wang, Yi Cheng, Dongding Lin, Chak Leong, and Wenjie Li. 2023. Target- oriented Proactive Dialogue Systems with Personalization: Problem Formulation and Dataset Curation. In EMNLP. 1132–1143
2023
-
[100]
Chanwoong Yoon, Gangwoo Kim, Byeongguk Jeon, Sungdong Kim, Yohan Jo, and Jaewoo Kang. 2024. Ask Optimal Questions: Aligning Large Language Models with Retriever’s Preference in Conversational Search. arXiv preprint arXiv:2402.11827 (2024)
2024 arXiv
-
[101]
Shi Yu, Jiahua Liu, Jingqin Yang, Chenyan Xiong, Paul Bennett, Jianfeng Gao, and Zhiyuan Liu. 2020. Few-shot Generative Conversational Query Rewriting. In SIGIR. 1933–1936
2020
-
[102]
Lee Xiong, Chenyan Xiong, Ye Li, Kwok-Fung Tang, Jialin Liu, Paul N Bennett, Junaid Ahmed, and Arnold Overwijk. 2021. Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text Retrieval. In ICLR
2021
-
[103]
Rui Yan and Dongyan Zhao. 2018. Smarter Response with Proactive Suggestion: A New Generative Neural Conversation Paradigm.. In IJCAI. 4525–4531
2018
-
[104]
Chengxiang Zhai and John Lafferty. 2001. A Study of Smoothing Methods for Language Models Applied to Ad Hoc Information Retrieval. In SIGIR. 334–342
2001
-
[105]
Wen Zhang, Lingfei Deng, Lei Zhang, and Dongrui Wu. 2022. A Survey on Negative Transfer. IEEE/CAA Journal of Automatica Sinica 10, 2 (2022), 305–329
2022
-
[106]
Xuan Zhang, Yang Deng, Zifeng Ren, See-Kiong Ng, and Tat-Seng Chua. 2024. Ask-before-Plan: Proactive Language Agents for Real-World Planning. arXiv preprint arXiv:2406.12639 (2024)
2024 arXiv
-
[107]
Fanghua Ye, Meng Fang, Shenghui Li, and Emine Yilmaz. 2023. Enhancing Con- versational Search: Large Language Model-Aided Informative Query Rewriting. In EMNLP. 5985–6006
2023
-
[108]
Yutao Zhu, Jian-Yun Nie, Kun Zhou, Pan Du, Hao Jiang, and Zhicheng Dou
-
[109]
Shengyao Zhuang, Houxing Ren, Linjun Shou, Jian Pei, Ming Gong, Guido Zuccon, and Daxin Jiang. 2022. Bridging the Gap Between Indexing and Re- trieval for Differentiable Search Index with Query Generation. arXiv preprint arXiv:2206.10128 (2022)
2022 arXiv
-
[110]
Shi Yu, Zhenghao Liu, Chenyan Xiong, Tao Feng, and Zhiyuan Liu. 2021. Few- shot Conversational Dense Retrieval. In SIGIR. 829–838
2021
-
[111]
Hansi Zeng, Chen Luo, Bowen Jin, Sheikh Muhammad Sarwar, Tianxin Wei, and Hamed Zamani. 2024. Scalable and Effective Generative Information Retrieval. In WWW. 1441–1452
2024
-
[115]
Yongfeng Zhang, Xu Chen, Qingyao Ai, Liu Yang, and W Bruce Croft. 2018. Towards Conversational Search and Recommendation. In CIKM. 177–186
2018
-
[119]
Shengyao Zhuang and Guido Zuccon. 2021. Dealing with Typos for BERT-based Passage Retrieval and Ranking. In EMNLP. 2836–2842
2021
-
[2012]
TOIS 30, 2 (2012), 1–35
Predicting Query Performance by Query-Drift Estimation. TOIS 30, 2 (2012), 1–35
2012
-
[2017]
arXiv preprint arXiv:1707.05409 (2017)
Neural Matching Models for Question Retrieval and Next Question Pre- diction in Conversation. arXiv preprint arXiv:1707.05409 (2017)
2017 arXiv
-
[2021]
In SIGIR
Proactive Retrieval-based Chatbots based on Relevant Knowledge and Goals. In SIGIR. 2000–2004
2000
-
[2023]
In SIGIR
Query Performance Prediction: From Ad-hoc to Conversational Search. In SIGIR. 2583–2593
-
[2024]
Query Performance Prediction: From Fundamentals to Advanced Tech- niques. In ECIR. 381–388
-
[2025]
Improving the Reusability of Conversational Search Test Collections. In ECIR. 196–213
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.