Pith. sign in

REVIEW 3 major objections 6 minor 117 references

Bridging the Gap: From Ad-hoc to Proactive Search in Conversations

T0 review · 3 major / 6 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read Conversation context, converted into ad-hoc queries, lets standard retrievers handle proactive search.

desk verdict Conv2Query is a solid, useful paper with a genuine contribution, and the main circularity worry about LLaMA-family filtering does not survive contact with the results. read the letter →

arxiv 2506.00983 v1 pith:DPVIYR33 submitted 2025-06-01 cs.IR cs.AIcs.CLcs.LG

classification cs.IRcs.AIcs.CLcs.LG
keywords ProactivesearchConversationalQuerypredictionDoc2QueryAd-hocretrievalPseudo-labelgenerationLLMfine-tuningInformation
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Proactive search in conversations aims to retrieve useful documents before the user formulates a query, using only the conversation as context. The paper argues that the main obstacle is an input-format gap: ad-hoc retrievers are pre-trained on short, concise queries, while conversational contexts are longer and noisier. It introduces Conv2Query, a conversation-to-query layer that rewrites conversational context into ad-hoc-style queries, and shows that this layer improves off-the-shelf lexical and neural retrievers and also improves them further after fine-tuning. The experiments on two proactive-search benchmarks cover both conversation contextualisation and interest anticipation, with statistically significant gains across the board. A sympathetic reader would care because the approach turns proactive search into a plug-in adaptation rather than requiring a new retrieval architecture.

What carries the argument

The load-bearing object is the Conv2Query pipeline, a conversation-to-query mapping trained on pseudo labels. For each conversation turn with a relevant document, a Doc2Query model generates candidate ad-hoc queries; QF-DC then scores each candidate by its relevance to that document and its relevance to the conversational context, using a cross-encoder-style relevance model, and selects the highest-scoring query as the training target. An LLM (Mistral-7B-Instruct in the main experiments) is fine-tuned with QLoRA to produce such a query from the conversation alone, and at inference that generated query is passed to any ad-hoc retriever. The same filtered queries also serve as positive examples for fine-tuning neural retrievers on PSC. This machinery transfers the retriever's knowledge by keeping both training and inference in the ad-hoc query format.

What would settle it

Replicate the experiments on a proactive-search dataset that also has human-written ad-hoc queries for each turn, and compare three inputs to the same retriever: the human query, Conv2Query's generated query, and the raw conversational context. If the human-query oracle does not clearly beat Conv2Query, or if a prompting-only LLM matches the full learned pipeline on held-out turns, the paper's claim that the Doc2Query-plus-QF-DC training signal is responsible for the improvement would be contradicted.

Watch

Extended reading notes

Core claim

The central claim is that the input mismatch between ad-hoc search and PSC, not the retrieval model itself, is what limits proactive retrieval quality. Conv2Query removes the mismatch by training an LLM to map conversational context to ad-hoc queries, using pseudo query targets generated by a Doc2Query model from each turn's relevant document and then filtered by QF-DC, a mechanism that keeps only queries that are both relevant to the document and aligned with the conversation. With these queries, off-the-shelf BM25, ANCE, SPLADE++, and RepLLaMA all improve significantly over using raw conversational context, and the same pseudo queries serve as better fine-tuning data than raw context. On the ProCIS and WebDisc datasets, the gains hold under both settings studied, including the harder interest-anticipation setting.

Load-bearing premise

The whole training signal rests on the assumption that Doc2Query queries generated from a turn's relevant document, after QF-DC filtering, are faithful stand-ins for the implicit search intent of that conversation; if the pseudo queries are off-target, the learned mapping and downstream retrieval gains would degrade.

Editorial extensions

If this is right

  • Any ad-hoc retriever, lexical or neural, can be applied to proactive search simply by placing Conv2Query in front of it, with no retriever retraining.
  • Fine-tuning retrievers on pseudo ad-hoc queries transfers pre-trained ad-hoc knowledge better than fine-tuning on raw conversational context, as shown by consistent gains for ANCE, SPLADE++, and RepLLaMA.
  • Lexical retrievers such as BM25, which cannot be fine-tuned on PSC data, benefit substantially from the noise-removing query conversion.
  • Larger LLMs used as the Conv2Query generator yield stronger retrieval quality, suggesting the mapping benefits from model scale.
  • The prompting-only variant of Conv2Query outperforms earlier text-window and description-based methods but is itself outperformed by the fine-tuned variant, indicating that learning from filtered pseudo queries adds value beyond LLM prompting.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The same pseudo-label recipe could be applied to other query-free retrieval settings, such as suggesting queries in retrieval-augmented generation, a direction the paper gestures at.
  • Because QF-DC explicitly aligns the target query with the conversation, the framework is likely sensitive to the quality of the relevance model used for filtering; with a weaker relevance model, the advantage over query-document filtering alone may shrink.
  • The paper leaves retrieval timing untouched; a natural extension would combine a learned 'when to retrieve' predictor with Conv2Query's 'what to retrieve' query generation.
  • If a benchmark with human-written queries per conversation became available, Conv2Query could be evaluated against the oracle upper bound, clarifying how much headroom remains in the conversation-to-query mapping.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The paper addresses proactive search in conversations (PSC), where the system must retrieve relevant documents from conversational context without an explicit user query. The authors argue that off-the-shelf ad-hoc retrievers suffer from an input mismatch because they are trained on short ad-hoc queries but receive long, noisy conversational context. They propose Conv2Query, which trains an LLM (Mistral-7B) to map a conversational context to an ad-hoc query. Training targets are pseudo-queries generated from the annotated relevant documents by a Doc2Query model, filtered by a new mechanism QF-DC that scores each candidate by both query-document relevance and query-conversation relevance using RankLLaMA. At inference, the generated query is fed to any ad-hoc retriever, optionally after fine-tuning that retriever on pseudo-query/document pairs. Experiments on ProCIS and WebDisc under conversation contextualisation and interest anticipation show large, statistically significant gains over raw-context baselines, prompting-only variants, Text Window and LMGR, for BM25, ANCE, SPLADE++ and RepLLaMA. The paper also shows that QF-DC accelerates convergence and that results are robust across LLM families (3B-22B). Limitations on timing prediction and dataset diversity are acknowledged in Section 8.

Significance. If the results hold, Conv2Query is a practical and conceptually simple adaptation layer: it lets existing ad-hoc retrievers be applied to query-free conversational search without retriever fine-tuning. The paper's strengths include extensive evaluation (two datasets, two settings, four retrievers including BM25 and a state-of-the-art LLaMA retriever), a clean two-stage training-data construction pipeline with an ablation for the filter, and public release of code and data. The consistent gains across BM25, ANCE, SPLADE++, and RepLLaMA make the central empirical claim credible. However, the statistical reporting and the use of a LLaMA-family relevance model to select training targets for a LLaMA-family retriever leave open questions that should be resolved before the general claim 'any ad-hoc retriever' is fully established.

major comments (3)
  1. [Section 5, Tables 1 and 2] The paper reports only point estimates and marks statistical significance with a paired t-test at p < 0.05, but it does not report variances, confidence intervals, the number of paired units, the number of runs, or random seeds. Because the central claim is that Conv2Query 'significantly improves' performance across all metrics and settings, these omissions make the significance statements impossible to verify. Please report standard deviations or bootstrap confidence intervals (e.g., across at least three QLoRA training runs, or across conversations/turns), and specify whether the paired t-test is paired by conversation or by turn.
  2. [Section 4.2, Eqs. (3)-(4); Section 5, Implementation details] QF-DC selects the pseudo-query training targets using RankLLaMA as the relevance model frelevance, and the best-performing retriever in Tables 1 and 2 is RepLLaMA, the retriever counterpart of the same LLaMA/MS-MARCO model family. Since the training signal for Conv2Query is entirely determined by RankLLaMA's relevance scores, the selected targets may be preferentially compatible with LLaMA-family retrievers, and this potential confound is not examined in the paper. I do not regard this as fatal circularity because BM25, ANCE, and SPLADE++ also improve consistently, but the authors should test the stability of the filtered target set under an alternative relevance model (e.g., MonoT5 or TCT-ColBERT), and ideally report whether the RepLLaMA gains persist when the filter comes from a different model family.
  3. [Section 4.1-4.2 and Section 7.1] The pseudo-query targets are generated by Doc2Query from the relevant document and then filtered by QF-DC, but the paper does not directly validate whether the selected queries reflect the implicit conversational information need. The learning-curve comparison in Figure 4 shows that QF-DC helps relative to QF-D and Random, but it does not establish the absolute quality of the targets. A direct oracle experiment (retrieving with the filtered pseudo-queries themselves, before any learned mapping) and a short qualitative sample of generated queries would strengthen the claim that Conv2Query learns a meaningful conversation-to-query mapping rather than exploiting document-specific vocabulary.
minor comments (6)
  1. [Section 2.2.1 and Section 5] The reference [9] for 'Doc2Query-Llama2' appears to point to DeeperImpact (Basnet et al., 2024), which is not the Llama-2-based Doc2Query model described in the text; please verify and cite the correct source.
  2. [Section 5, Implementation details] The choice of truncation length 512 for conversational-context inputs and 32 for Conv2Query-generated queries is justified only as 'exceeds the average context length'; a brief sensitivity check or a statement that longer lengths produced no further gains would be helpful.
  3. [Section 7.1] The opening sentence of Section 7.1, 'we study the impact of query mechanism,' should read 'query filtering mechanism.'
  4. [Section 7.2] The text 'for llama, we use Llama-3.2-3B-Instruct' should capitalize 'Llama' for consistency with the rest of the paper.
  5. [Tables 1 and 2 captions] The abbreviation 'Conv2Q' appears in the table captions without being defined at first use; please define it or use 'Conv2Query' consistently throughout the captions.
  6. [Section 5, Baselines] The Text Window and LMGR baselines are re-implemented with RepLLaMA in place of the original retrievers; a sentence noting that this replacement was made only for comparability and that the variants were not validated against the original published systems would improve transparency.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: Conv2Query is trained on pseudo-labels but evaluated on external relevance judgments; the filter/retriever model-family overlap is a confound, not a definitional reduction.

full rationale

The paper's derivation chain is a standard supervised pipeline: Doc2Query (an external MS MARCO-trained model) generates query candidates from curated relevant documents; QF-DC filters them using a public relevance model (RankLLaMA); Conv2Query is trained to map conversational context to the filtered query; and retrieval is evaluated on held-out sparse and dense relevance judgments. No equation in Sections 4.1-4.5 makes the evaluation target a function of the training target by construction: headline numbers are measured against dataset annotations (user-added Wikipedia/webpage links and human dense judgments), not against the pseudo-labels used for training. The overlap between RankLLaMA (filter) and RepLLaMA (evaluated retriever) is a possible confound, but it is not a definitional reduction, and improvements are also reported for BM25, ANCE, and SPLADE++, so the central claim does not collapse to a self-consistency check. Reliance on prior work (Doc2Query, RankLLaMA, ProCIS, WebDisc) is ordinary use of external artifacts, not a self-citation chain; no uniqueness theorem or unverified ansatz is imported to forbid alternatives. No circular steps were found.

Assumptions & free parameters 6 free parameters · 4 assumptions · 0 invented entities

The central claim rests on the quality of pseudo-labels generated by Doc2Query and filtered by QF-DC. These are not invented entities but the method's key dependencies. The free parameters are mostly hyperparameters that are not fully reported.

free parameters (6)
  • Number of Doc2Query candidates n = 100
    Chosen without ablation; affects the size of the pool from which QF-DC selects targets.
  • Top-k sampling parameter k for Doc2Query = 10
    Set based on prior work [69,109]; controls diversity of generated queries.
  • Aggregation function f_aggregate = summation
    Summation of query-document and query-conversation relevance scores; no ablation of other aggregation functions.
  • Retriever query truncation length = 512 (context input), 32 (generated queries)
    Chosen to exceed average context length for context input and to match ad-hoc query length for generated queries.
  • BM25 k1 and b = k1=0.9, b=0.4 (ProCIS); k1=8/7, b=0.99 (WebDisc)
    Set per dataset/setting, following prior work; represents standard parameter tuning.
  • QLoRA hyperparameters (learning rate, LoRA rank, etc.)
    Not reported in the paper; a gap for reproducibility.
assumptions (4)
  • domain assumption The Doc2Query model, pre-trained on MS MARCO, can generate queries that reveal the implicit search intents of conversations when given the relevant document.
    The entire pseudo-label generation pipeline in Section 4.1 relies on this.
  • domain assumption The QF-DC filter, using RankLLaMA for relevance, can accurately measure both query-document and query-conversation relevance.
    Section 4.2 assumes the relevance model provides reliable scores for selecting optimal training targets.
  • domain assumption The conversational context alone is sufficient to predict a useful ad-hoc query for the relevant documents.
    The task definition and the training objective (Eq. 5) assume this mapping is learnable.
  • domain assumption The user-added hyperlinks in the datasets serve as reliable relevance judgments.
    Both ProCIS and WebDisc derive sparse relevance labels from hyperlinks; the paper acknowledges this in Section 5.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Bridging the Gap: From Ad-hoc to Proactive Search in Conversations." pith.science (2026). https://pith.science/paper/DPVIYR33

@misc{pith2026250600983,
  author       = {Pith},
  title        = {Pith review of: Bridging the Gap: From Ad-hoc to Proactive Search in Conversations},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/DPVIYR33}},
  note         = {Machine review of arXiv:2506.00983}
}
read the original abstract

Proactive search in conversations (PSC) aims to reduce user effort in formulating explicit queries by proactively retrieving useful relevant information given conversational context. Previous work in PSC either directly uses this context as input to off-the-shelf ad-hoc retrievers or further fine-tunes them on PSC data. However, ad-hoc retrievers are pre-trained on short and concise queries, while the PSC input is longer and noisier. This input mismatch between ad-hoc search and PSC limits retrieval quality. While fine-tuning on PSC data helps, its benefits remain constrained by this input gap. In this work, we propose Conv2Query, a novel conversation-to-query framework that adapts ad-hoc retrievers to PSC by bridging the input gap between ad-hoc search and PSC. Conv2Query maps conversational context into ad-hoc queries, which can either be used as input for off-the-shelf ad-hoc retrievers or for further fine-tuning on PSC data. Extensive experiments on two PSC datasets show that Conv2Query significantly improves ad-hoc retrievers' performance, both when used directly and after fine-tuning on PSC.

Figures

Figures reproduced from arXiv: 2506.00983 by the authors.

Figure 1
Figure 1. Comparison between ad-hoc search and proactive [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 3
Figure 3. Prompt for interest anticipation. sample negative documents for 𝐶𝑡 following standard neural re￾trieval practices and use both positive and negative examples to fine-tune a specific ad-hoc retriever on PSC. 5 Experimental setup Research questions. Our work is steered by the following research questions: RQ1 To what extent does Conv2Query bridge the input gap be￾tween ad-hoc pre-training and PSC inference under conve… view at source ↗
Figure 5
Figure 5. Retrieval quality (MRR@10) of RepLLaMA using [PITH_FULL_IMAGE:figures/full_fig_p009_5.png] view at source ↗
Figures from the paper (1 more)
Figure 4
Figure 4. Figure 4: Conv2Query ’s learning curves on ProCIS with QF [PITH_FULL_IMAGE:figures/full_fig_p009_4.png]

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

117 extracted references · 66 canonical work pages

  1. [1]

    Zahra Abbasiantaeb, Chuan Meng, Leif Azzopardi, and Mohammad Aliannejadi

  2. [2]

    T Ahmed and S Bulathwela. 2022. Towards Proactive Information Retrieval in Noisy Text with Wikipedia Concepts. In CEUR Workshop Proceedings, Vol. 3318. 1–12

  3. [3]

    Salvatore Andolina, Valeria Orso, Hendrik Schneider, Khalil Klouche, Tuukka Ruotsalo, Luciano Gamberini, and Giulio Jacucci. 2018. Investigating Proactive Search Support in Conversations. In DIS. 1295–1307

  4. [6]

    Negar Arabzadeh, Chuan Meng, Mohammad Aliannejadi, and Ebrahim Bagheri

  5. [7]

    Payal Bajaj, Daniel Campos, Nick Craswell, Li Deng, Xiaodong Liu Jianfeng Gao, Rangan Majumder, Andrew McNamara, Bhaskar Mitra, Tri Nguyen, Mir Rosen- berg, Xia Song, Alina Stoica, Saurabh Tiwary, and Tong Wang. 2016. MS MARCO: A Human Generated Machine Reading Comprehension Dataset. In NIPS

  6. [8]

    In SIGIR-AP

    Query Performance Prediction: Techniques and Applications in Modern Information Retrieval. In SIGIR-AP. 291–294

  7. [9]

    Soyuj Basnet, Jerry Gou, Antonio Mallia, and Torsten Suel. 2024. DeeperImpact: Optimizing Sparse Learned Index Structures. arXiv preprint arXiv:2405.17093 (2024)

  8. [10]

    Query Performance Prediction: Theory, Techniques and Applications. In WSDM. 991–994

Show all 117 references
  1. [11]

    Francesco Bonchi, Raffaele Perego, Fabrizio Silvestri, Hossein Vahabi, and Rossano Venturini. 2012. Efficient Query Recommendations in the Long Tail via Center-Piece Subgraphs. In SIGIR. 345–354

  2. [12]

    Vevake Balaraman and Bernardo Magnini. 2020. Proactive Systems and Influ- enceable Users: Simulating Proactivity in Task-oriented Dialogues. InSEMDIAL

  3. [13]

    Yash Butala, Siddhant Garg, Pratyay Banerjee, and Amita Misra. 2024. ProMISe: A Proactive Multi-turn Dialogue Dataset for Information-seeking Intent Resolu- tion. In EACL. 1774–1789

  4. [14]

    Paolo Boldi, Francesco Bonchi, Carlos Castillo, Debora Donato, Aristides Gionis, and Sebastiano Vigna. 2008. The Query-flow Graph: Model and Applications. In CIKM. 609–618

  5. [15]

    Wanyu Chen, Fei Cai, Honghui Chen, and Maarten de Rijke. 2018. Attention- based Hierarchical Neural Query Suggestion. In SIGIR. 1093–1096

  6. [16]

    Luiz Bonifacio, Hugo Abonizio, Marzieh Fadaee, and Rodrigo Nogueira. 2022. InPars: Unsupervised Dataset Generation for Information Retrieval. In SIGIR. 2387–2392

  7. [17]

    Zhuyun Dai, Vincent Y Zhao, Ji Ma, Yi Luan, Jianmo Ni, Jing Lu, Anton Bakalov, Kelvin Guu, Keith Hall, and Ming-Wei Chang. 2023. Promptagator: Few-shot Dense Retrieval From 8 Examples. In ICLR

  8. [18]

    Ramraj Chandradevan, Kaustubh Dhole, and Eugene Agichtein. 2024. DUQGen: Effective Unsupervised Domain Adaptation of Neural Rankers by Diversifying Synthetic Query Generation. In NAACL, Kevin Duh, Helena Gomez, and Steven Bethard (Eds.)

  9. [19]

    Yang Deng, Wenqiang Lei, Minlie Huang, and Tat-Seng Chua. 2023. Rethinking Conversational Agents in the Era of LLMs: Proactivity, Non-collaborativity, and Beyond. In SIGIR. 298–301

  10. [20]

    Zhicong Cheng, Bin Gao, and Tie-Yan Liu. 2010. Actively Predicting Diverse Search Intent from User Browsing Behaviors. In WWW. 221–230

  11. [21]

    Yang Deng, Wenqiang Lei, Wenxuan Zhang, Wai Lam, and Tat-Seng Chua. 2022. PACIFIC: Towards Proactive Conversational Question Answering over Tabular and Textual Data in Finance. In EMNLP. 6970–6984

  12. [22]

    Mostafa Dehghani, Sascha Rothe, Enrique Alfonseca, and Pascal Fleury. 2017. Learning to Attend, Copy, and Generate for Session-Based Query Suggestion. In CIKM. 1747–1756

  13. [23]

    Yang Deng, Lizi Liao, Zhonghua Zheng, Grace Hui Yang, and Tat-Seng Chua

  14. [24]

    Yang Deng, Wenqiang Lei, Wai Lam, and Tat-Seng Chua. 2023. A Survey on Proactive Dialogue Systems: Problems, Methods, and Prospects. In IJCAI. 6583– 6591

  15. [25]

    Susan Dumais, Edward Cutrell, Raman Sarin, and Eric Horvitz. 2004. Implicit Queries (IQ) for Contextualized Search. In SIGIR. 594–594

  16. [26]

    Yang Deng, Lizi Liao, Liang Chen, Hongru Wang, Wenqiang Lei, and Tat-Seng Chua. 2023. Prompting and Evaluating Large Language Models for Proac- tive Dialogues: Clarification, Target-guided, and Non-collaboration. In EMNLP. 10602–10621

  17. [27]

    Henry Feild and James Allan. 2013. Task-Aware Query Recommendation. In SIGIR. 83–92

  18. [28]

    In SIGIR

    Towards Human-centered Proactive Conversational Agents. In SIGIR. 807–818

  19. [29]

    Tim Dettmers, Artidoro Pagnoni, Ari Holtzman, and Luke Zettlemoyer. 2023. QLoRA: Efficient Finetuning of Quantized LLMs.arXiv preprint arXiv:2305.14314 (2023)

  20. [30]

    Thibault Formal, Benjamin Piwowarski, and Stéphane Clinchant. 2021. SPLADE: Sparse Lexical and Expansion Model for First Stage Ranking. In SIGIR. 2288– 2292

  21. [31]

    Desmond Elliott and Joemon M Jose. 2009. A Proactive Personalised Retrieval System. In CIKM. 1935–1938

  22. [32]

    Mitko Gospodinov, Sean MacAvaney, and Craig Macdonald. 2023. Doc2Query–: When Less is More. In ECIR. Springer, 414–422

  23. [33]

    Stephen Fitchett and Andy Cockburn. 2012. AccessRank: Predicting What Users Will Do Next. In SIGCHI. 2239–2242

  24. [34]

    Thibault Formal, Carlos Lassance, Benjamin Piwowarski, and Stéphane Clin- chant. 2022. From Distillation to Hard Negative Sampling: Making Sparse Neural IR Models More Effective. In SIGIR. 2353–2359

  25. [35]

    Yunah Jang, Kang-il Lee, Hyunkyung Bae, Seungpil Won, Hwanhee Lee, and Kyomin Jung. 2024. IterCQR: Iterative Conversational Query Reformulation with Retrieval Guidance. In EMNLP. 8121–813

  26. [36]

    Debasis Ganguly, Dipasree Pal, Manisha Verma, and Procheta Sen. 2020. Overview of RCD-2020, the FIRE-2020 track on retrieval from conversational dialogues. In FIRE. 33–36

  27. [37]

    Vitor Jeronymo, Luiz Bonifacio, Hugo Abonizio, Marzieh Fadaee, Roberto Lotufo, Jakub Zavrel, and Rodrigo Nogueira. 2023. InPars-v2: Large Language Mod- els as Efficient Dataset Generators for Information Retrieval. arXiv preprint arXiv:2301.01820 (2023)

  28. [38]

    Nam Le Hai, Thomas Gerald, Thibault Formal, Jian-Yun Nie, Benjamin Pi- wowarski, and Laure Soulier. 2023. CoSPLADE: Contextualizing SPLADE for Conversational Information Retrieval. In ECIR

  29. [39]

    Monika Henzinger, Bay-Wei Chang, Brian Milch, and Sergey Brin. 2003. Query- Free News Search. In WWW. 1–10

  30. [40]

    Markus Koskela, Petri Luukkonen, Tuukka Ruotsalo, Mats Sjöberg, and Patrik Floréen. 2018. Proactive Information Retrieval by Capturing Search Intent from Primary Task Context. TiiS 8, 3 (2018), 1–25

  31. [41]

    Kalervo Järvelin and Jaana Kekäläinen. 2002. Cumulated Gain-Based Evaluation of IR Techniques. TOIS 20, 4 (2002), 422–446

  32. [42]

    Jing Yang Lee, Seokhwan Kim, Kartik Mehta, Jiun-Yu Kao, Yu-Hsiang Lin, and Arpit Gupta. 2024. Redefining Proactivity for Information Seeking Dialogue. arXiv preprint arXiv:2410.15297 (2024)

  33. [43]

    Gareth JF Jones, Procheta Sen, Debasis Ganguly, and Emine Yilmaz. 2022. Work- shop on Proactive and Agent-Supported Information Retrieval (PASIR). InCIKM. 5167–5168

  34. [44]

    Weize Kong, Rui Li, Jie Luo, Aston Zhang, Yi Chang, and James Allan. 2015. Predicting Search Intent Based on Pre-Search Context. In SIGIR. 503–512

  35. [45]

    Lizi Liao, Grace Hui Yang, and Chirag Shah. 2023. Proactive Conversational Agents in the Post-ChatGPT World. In SIGIR. 3452–3455

  36. [46]

    Yilong Lai, Jialong Wu, Congzhi Zhang, Haowen Sun, and Deyu Zhou. 2025. AdaCQR: Enhancing Query Reformulation for Conversational Search via Sparse and Dense Retrieval Alignment. COLING (2025)

  37. [47]

    Sheng-Chieh Lin, Jheng-Hong Yang, and Jimmy Lin. 2021. Contextualized Query Embeddings for Conversational Search. In EMNLP. 1004–1015

  38. [48]

    Ruirui Li, Ben Kao, Bin Bi, Reynold Cheng, and Eric Lo. 2012. DQR: A Proba- bilistic Approach to Diversified Query Recommendation. In CIKM. 16–25

  39. [49]

    Lizi Liao, Grace Hui Yang, and Chirag Shah. 2023. Proactive Conversational Agents. In WSDM. 1244–1247

  40. [50]

    Xueguang Ma, Liang Wang, Nan Yang, Furu Wei, and Jimmy Lin. 2024. Fine- Tuning LLaMA for Multi-Stage Text Retrieval. In SIGIR. 2421–2425

  41. [51]

    Daniel J Liebling, Paul N Bennett, and Ryen W White. 2012. Anticipatory Search: Using Context to Initiate Search. In SIGIR. 1035–1036

  42. [52]

    Kelong Mao, Zhicheng Dou, Bang Liu, Hongjin Qian, Fengran Mo, Xiangli Wu, Xiaohua Cheng, and Zhao Cao. 2023. Search-Oriented Conversational Query Editing. In Findings of ACL. 4160–4172

  43. [53]

    Sheng-Chieh Lin, Jheng-Hong Yang, and Jimmy Lin. 2021. In-Batch Negatives for Knowledge Distillation with Tightly-Coupled Teachers for Dense Retrieval. In RepL4NLP-2021. 163–173

  44. [54]

    Lili Lu, Chuan Meng, Federico Ravenda, Mohammad Aliannejadi, and Fabio Crestani. 2025. Zero-Shot and Efficient Clarification Need Prediction in Conver- sational Search. In ECIR. 389–404

  45. [55]

    Chuan Meng, Mohammad Aliannejadi, and Maarten de Rijke. 2023. System Initiative Prediction for Multi-turn Conversational Information Seeking. In CIKM. 1807–1817. Bridging the Gap: From Ad-hoc to Proactive Search in Conversations SIGIR ’25, July 13–18, 2025, Padua, Italy

  46. [56]

    Kelong Mao, Chenlong Deng, Haonan Chen, Fengran Mo, Zheng Liu, Tetsuya Sakai, and Zhicheng Dou. 2024. ChatRetriever: Adapting Large Language Models for Generalized and Robust Conversational Dense Retrieval. In EMNLP. 1227–1240

  47. [57]

    Chuan Meng, Negar Arabzadeh, Arian Askari, Mohammad Aliannejadi, and Maarten de Rijke. 2025. Query Performance Prediction using Relevance Judg- ments Generated by Large Language Models. TOIS (2025)

  48. [58]

    Kelong Mao, Zhicheng Dou, Fengran Mo, Jiewen Hou, Haonan Chen, and Hongjin Qian. 2023. Large Language Models Know Your Contextual Search Intent: A Prompting Framework for Conversational Search. In EMNLP. 1211– 1225

  49. [59]

    Chuan Meng. 2024. Query Performance Prediction for Conversational Search and Beyond. In SIGIR. 3077–3077

  50. [60]

    Fengran Mo, Abbas Ghaddar, Kelong Mao, Mehdi Rezagholizadeh, Boxing Chen, Qun Liu, and Jian-Yun Nie. 2024. CHIQ: Contextual History Enhancement for Improving Query Rewriting in Conversational Search. In EMNLP. 2253–2268

  51. [61]

    Chuan Meng, Negar Arabzadeh, Mohammad Aliannejadi, and Maarten de Rijke

  52. [62]

    Fengran Mo, Kelong Mao, Yutao Zhu, Yihong Wu, Kaiyu Huang, and Jian-Yun Nie. 2023. ConvGQR: Generative Query Reformulation for Conversational Search. In ACL. 4998–5012

  53. [63]

    Fengran Mo, Chuan Meng, Mohammad Aliannejadi, and Jian-Yun Nie. 2025. Conversational Search: From Fundamentals to Frontiers in the LLM Era. In SIGIR

  54. [64]

    Chuan Meng, Guglielmo Faggioli, Mohammad Aliannejadi, Nicola Ferro, and Josiane Mothe. 2025. QPP++ 2025: Query Performance Prediction and its Appli- cations in the Era of Large Language Models. In ECIR. 319–325

  55. [65]

    Kshitij Mishra, Azlaan Mustafa Samad, Palak Totala, and Asif Ekbal. 2022. PEPDS: A Polite and Empathetic Persuasive Dialogue System for Charity Donation. In COLING. 424–440

  56. [66]

    Fengran Mo, Chen Qu, Kelong Mao, Tianyu Zhu, Zhan Su, Kaiyu Huang, and Jian-Yun Nie. 2024. History-Aware Conversational Dense Retrieval. Findings of ACL (2024)

  57. [67]

    Fengran Mo, Kelong Mao, Ziliang Zhao, Hongjin Qian, Haonan Chen, Yiruo Cheng, Xiaoxi Li, Yutao Zhu, Zhicheng Dou, and Jian-Yun Nie. 2024. A Survey of Conversational Search. arXiv preprint arXiv:2410.15576 (2024)

  58. [68]

    Rodrigo Nogueira, Zhiying Jiang, Ronak Pradeep, and Jimmy Lin. 2020. Doc- ument Ranking with a Pretrained Sequence-to-Sequence Model. In EMNLP. 708–718

  59. [69]

    Rodrigo Nogueira and Jimmy Lin. 2019. From doc2query to docTTTTTquery. (2019)

  60. [70]

    Fengran Mo, Jian-Yun Nie, Kaiyu Huang, Kelong Mao, Yutao Zhu, Peng Li, and Yang Liu. 2023. Learning to Relate to Previous Turns in Conversational Search. In KDD. 1722–1732

  61. [71]

    Fengran Mo, Chen Qu, Kelong Mao, Yihong Wu, Zhan Su, Kaiyu Huang, and Jian-Yun Nie. 2024. Aligning Query Representation with Rewritten Query and Relevance Judgments in Conversational Search. In CIKM. 1700–1710

  62. [72]

    Dae Hoon Park, Yi Fang, Mengwen Liu, and ChengXiang Zhai. 2016. Mobile App Retrieval for Social Media Users via Inference of Implicit Intent in Social Media Text. In CIKM. 959–968

  63. [73]

    Cristina Ioana Muntean, Franco Maria Nardini, Fabrizio Silvestri, and Marcin Sydow. 2013. Learning to Shorten Query Sessions. In WWW. 131–132

  64. [74]

    Bradley James Rhodes and Pattie Maes. 2000. Just-In-Time Information Retrieval. IBM Systems journal 39, 3.4 (2000), 685–704

  65. [75]

    Bradley J Rhodes and Thad Starner. 1996. Remembrance Agent: A Continuously Running Automated Information Retrieval System. In PAAM, Vol. 96. 487–495

  66. [76]

    Rodrigo Nogueira, Wei Yang, Jimmy Lin, and Kyunghyun Cho. 2019. Document Expansion by Query Prediction. arXiv preprint arXiv:1904.08375 (2019)

  67. [77]

    Dipasree Pal and Debasis Ganguly. 2021. Effective Query Formulation in Con- versation Contextualization: A Query Specificity-based Approach. In ICTIR. 177–183

  68. [78]

    Corbin Rosset, Chenyan Xiong, Xia Song, Daniel Campos, Nick Craswell, Saurabh Tiwary, and Paul Bennett. 2020. Leading Conversational Search by Suggesting Useful Questions. In The web conference. 1160–1170

  69. [79]

    David Rau and Jaap Kamps. 2022. The Role of Complex NLP in Transformers for Text Ranking. In ICTIR. 153–160

  70. [80]

    Procheta Sen, Debasis Ganguly, and Gareth Jones. 2018. Procrastination is the Thief of Time: Evaluating the Effectiveness of Proactive Search Systems. In SIGIR. 1157–1160

  71. [81]

    Procheta Sen, Debasis Ganguly, and Gareth JF Jones. 2021. I Know What You Need: Investigating Document Retrieval Effectiveness with Partial Session Contexts. TOIS 40, 3 (2021), 1–30

  72. [82]

    Stephen E Robertson, Steve Walker, Susan Jones, Micheline M Hancock-Beaulieu, Mike Gatford, et al . 1995. Okapi at TREC-3. Nist Special Publication Sp 109 (1995), 109

  73. [83]

    Kevin Ros, Matthew Jin, Jacob Levine, and ChengXiang Zhai. 2023. Retrieving Webpages Using Online Discussions. In ICTIR. 159–168

  74. [84]

    Anna Shtok, Oren Kurland, David Carmel, Fiana Raiber, and Gad Markovits

  75. [85]

    Chris Samarinas and Hamed Zamani. 2024. ProCIS: A Benchmark for Proactive Retrieval in Conversations. In SIGIR. 830–840

  76. [86]

    Alessandro Sordoni, Yoshua Bengio, Hossein Vahabi, Christina Lioma, Jakob Grue Simonsen, and Jian-Yun Nie. 2015. A Hierarchical Recurrent Encoder- Decoder for Generative Context-Aware Query Suggestion. In CIKM. 553–562

  77. [87]

    Yueming Sun and Yi Zhang. 2018. Conversational Recommender System. In SIGIR. 235–244

  78. [88]

    Haizhou Shi, Zihao Xu, Hengyi Wang, Weiyi Qin, Wenyuan Wang, Yibin Wang, Zifeng Wang, Sayna Ebrahimi, and Hao Wang. 2024. Continual Learning of Large Language Models: A Comprehensive Survey. arXiv preprint arXiv:2404.16789 (2024)

  79. [89]

    Milad Shokouhi and Qi Guo. 2015. From Queries to Cards: Re-ranking Proactive Card Recommendations Based on Reactive Search History. In SIGIR. 695–704

  80. [90]

    Tung Vuong, Giulio Jacucci, and Tuukka Ruotsalo. 2017. Proactive Information Retrieval via Screen Surveillance. In SIGIR. 1313–1316

  81. [91]

    Tung Vuong, Giulio Jacucci, and Tuukka Ruotsalo. 2017. Watching inside the Screen: Digital Activity Monitoring for Task Recognition and Proactive Information Retrieval. Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies 1, 3 (2017), 1–23

  82. [92]

    Yang Song and Qi Guo. 2016. Query-Less: Predicting Task Repetition for NextGen Proactive Search and Recommendation Engines. In WWW. 543–553

  83. [93]

    Bin Wu, Chenyan Xiong, Maosong Sun, and Zhiyuan Liu. 2018. Query Sugges- tion with Feedback Memory Network. In Proceedings of the 2018 World Wide Web Conference. 1563–1571

  84. [94]

    Yun Xing, Jian Kang, Aoran Xiao, Jiahao Nie, Ling Shao, and Shijian Lu. 2024. Rewrite Caption Semantics: Bridging Semantic Gaps for Language-Supervised Semantic Segmentation. NeurIPS 36 (2024)

  85. [95]

    Anuja Tayal and Aman Tyagi. 2024. Dynamic Contexts for Generating Sugges- tion Questions in RAG Based Conversational Systems. In WWW. 1338–1341

  86. [96]

    Tuan A Tran, Sven Schwarz, Claudia Niederée, Heiko Maus, and Nattiya Kan- habua. 2016. The Forgotten Needle in My Collections: Task-Aware Ranking of Documents in Semantic Information Space. In CHIIR. 13–22

  87. [97]

    Liu Yang, Qi Guo, Yang Song, Sha Meng, Milad Shokouhi, Kieran McDonald, and W Bruce Croft. 2016. Modeling User Interests for Zero-Query Ranking. In ECIR. Springer, 171–184

  88. [98]

    Liu Yang, Hamed Zamani, Yongfeng Zhang, Jiafeng Guo, and W Bruce Croft

  89. [99]

    Jian Wang, Yi Cheng, Dongding Lin, Chak Leong, and Wenjie Li. 2023. Target- oriented Proactive Dialogue Systems with Personalization: Problem Formulation and Dataset Curation. In EMNLP. 1132–1143

  90. [100]

    Chanwoong Yoon, Gangwoo Kim, Byeongguk Jeon, Sungdong Kim, Yohan Jo, and Jaewoo Kang. 2024. Ask Optimal Questions: Aligning Large Language Models with Retriever’s Preference in Conversational Search. arXiv preprint arXiv:2402.11827 (2024)

  91. [101]

    Shi Yu, Jiahua Liu, Jingqin Yang, Chenyan Xiong, Paul Bennett, Jianfeng Gao, and Zhiyuan Liu. 2020. Few-shot Generative Conversational Query Rewriting. In SIGIR. 1933–1936

  92. [102]

    Lee Xiong, Chenyan Xiong, Ye Li, Kwok-Fung Tang, Jialin Liu, Paul N Bennett, Junaid Ahmed, and Arnold Overwijk. 2021. Approximate Nearest Neighbor Negative Contrastive Learning for Dense Text Retrieval. In ICLR

  93. [103]

    Rui Yan and Dongyan Zhao. 2018. Smarter Response with Proactive Suggestion: A New Generative Neural Conversation Paradigm.. In IJCAI. 4525–4531

  94. [104]

    Chengxiang Zhai and John Lafferty. 2001. A Study of Smoothing Methods for Language Models Applied to Ad Hoc Information Retrieval. In SIGIR. 334–342

  95. [105]

    Wen Zhang, Lingfei Deng, Lei Zhang, and Dongrui Wu. 2022. A Survey on Negative Transfer. IEEE/CAA Journal of Automatica Sinica 10, 2 (2022), 305–329

  96. [106]

    Xuan Zhang, Yang Deng, Zifeng Ren, See-Kiong Ng, and Tat-Seng Chua. 2024. Ask-before-Plan: Proactive Language Agents for Real-World Planning. arXiv preprint arXiv:2406.12639 (2024)

  97. [107]

    Fanghua Ye, Meng Fang, Shenghui Li, and Emine Yilmaz. 2023. Enhancing Con- versational Search: Large Language Model-Aided Informative Query Rewriting. In EMNLP. 5985–6006

  98. [108]

    Yutao Zhu, Jian-Yun Nie, Kun Zhou, Pan Du, Hao Jiang, and Zhicheng Dou

  99. [109]

    Shengyao Zhuang, Houxing Ren, Linjun Shou, Jian Pei, Ming Gong, Guido Zuccon, and Daxin Jiang. 2022. Bridging the Gap Between Indexing and Re- trieval for Differentiable Search Index with Query Generation. arXiv preprint arXiv:2206.10128 (2022)

  100. [110]

    Shi Yu, Zhenghao Liu, Chenyan Xiong, Tao Feng, and Zhiyuan Liu. 2021. Few- shot Conversational Dense Retrieval. In SIGIR. 829–838

  101. [111]

    Hansi Zeng, Chen Luo, Bowen Jin, Sheikh Muhammad Sarwar, Tianxin Wei, and Hamed Zamani. 2024. Scalable and Effective Generative Information Retrieval. In WWW. 1441–1452

  102. [115]

    Yongfeng Zhang, Xu Chen, Qingyao Ai, Liu Yang, and W Bruce Croft. 2018. Towards Conversational Search and Recommendation. In CIKM. 177–186

  103. [119]

    Shengyao Zhuang and Guido Zuccon. 2021. Dealing with Typos for BERT-based Passage Retrieval and Ranking. In EMNLP. 2836–2842

  104. [2012]

    TOIS 30, 2 (2012), 1–35

    Predicting Query Performance by Query-Drift Estimation. TOIS 30, 2 (2012), 1–35

  105. [2017]

    arXiv preprint arXiv:1707.05409 (2017)

    Neural Matching Models for Question Retrieval and Next Question Pre- diction in Conversation. arXiv preprint arXiv:1707.05409 (2017)

  106. [2021]

    In SIGIR

    Proactive Retrieval-based Chatbots based on Relevant Knowledge and Goals. In SIGIR. 2000–2004

  107. [2023]

    In SIGIR

    Query Performance Prediction: From Ad-hoc to Conversational Search. In SIGIR. 2583–2593

  108. [2024]

    Query Performance Prediction: From Fundamentals to Advanced Tech- niques. In ECIR. 381–388

  109. [2025]

    Improving the Reusability of Conversational Search Test Collections. In ECIR. 196–213

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.