REVIEW 4 major objections 5 minor 48 references
RALI@TREC iKAT 2024: Achieving Personalization via Retrieval Fusion in Conversational Search
T0 review · 4 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read Fusing personalized and neutral queries improves search ranking.
desk verdict A candid TREC system description with a sensible fusion idea, but the central claim that fusion works is not proven by the experiments as reported. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is a score-based linear-combination fusion of three BM25 ranking lists. For a document $D$, the fused score is $score(D) = \alpha_1 s_1(D) + \alpha_2 s_2(D) + \alpha_3 s_3(D)$, with fixed weights tuned on the previous year's test collection; documents missing from a list receive the score of the 1000th document in that list. This implements the 'Chorus Effect' principle, that when multiple retrieval strategies agree that a document is relevant, the agreement is treated as stronger evidence of relevance, and the top 50 documents of the fused list are then reranked by a neural reranker before evaluation.
What would settle it
Judge every unassessed document in the top 10 of the fused run for all turns, recompute NDCG@10, and compare with the same recomputation for the best non-fusion run; if the fused run no longer leads, the central claim is falsified. The authors' 9-1-3 example (5 of top 20 judged for the expanded query versus 14 for the original) shows exactly the kind of incompleteness that would matter.
Extended reading notes
Core claim
The paper's central claim is that score-based fusion of ranking lists generated from queries with different levels of personalization improves passage ranking in personalized conversational search. The mechanism is a weighted linear combination of three BM25 lists—one from a de-contextualized non-personalized rewrite, one from that rewrite expanded with a generated response, and one from a personalized rewrite that incorporates user-profile elements—followed by reranking the top 50 fused documents with a pretrained sequence-to-sequence reranker. The run with fusion and reranking achieved the strongest scores among the four submitted runs, which the authors present as evidence that retrieval fusion is effective. On the previous year's collection, however, the same fusion-then-reranking approach did not beat manual human rewrites; the paper argues this discrepancy is largely due to assessment bias, since documents retrieved by non-participant query formulations receive far fewer relevance judgments.
Load-bearing premise
The results assume the human relevance judgments are fair to the fused run's documents, yet the paper's own case study shows documents retrieved by the expanded query were far less likely to be judged than documents retrieved by the human rewrite; if that pattern holds generally, the measured ranking gains could come from incomplete judgment rather than real quality.
Editorial extensions
If this is right
- A conversational search system does not have to choose between a personalized and a non-personalized query; keeping both in the candidate mix can protect against over-personalization.
- Score-based fusion with fixed per-query-type weights is a simple, parameter-light strategy that can be tuned on an earlier test collection and applied to new turns without per-turn weight fitting.
- Documents that rank highly under both personalized and non-personalized formulations are the high-confidence candidates, so fusion can also be used as a filtering signal before expensive reranking.
- Reranking the fused list adds a further gain over fusion alone, so the two stages are complementary in the submitted pipeline.
- Because shallow pooling under-assesses non-participant query formulations, a fused run that beats participant runs on such a collection is a conservative signal of its true relative quality.
Reading between the lines
- Beyond the paper, the same fusion principle should apply to other query variants and other retrievers: what matters is that the candidate lists are diverse enough that agreement between them is informative, not the specific BM25 base.
- A direct test of the assessment-bias argument would be to have the top unassessed documents from the fused run judged and to recompute the metrics; the paper's own case study predicts the fusion advantage would grow, not shrink, under complete judgment.
- The fixed-weight design invites an adaptive extension: estimating how noisy a given turn's personalized rewrite is, and shifting weight toward the non-personalized list when profile terms look risky, could widen the gap over single-query baselines.
- For benchmark builders, the paper's bias analysis implies that shallow-pooled conversational collections should not be treated as reusable for methods that did not contribute to the pool, unless supplemental judgments are collected.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper reports RALI's participation in the TREC iKAT 2024 passage-ranking task. It identifies an 'over-personalization' problem in which adding PTKB-derived terms to an LLM-generated query rewrite can drift the query, and proposes to mitigate this by retrieving separately with three query variants: a non-personalized LLM rewrite, that rewrite expanded with a generated response, and a personalized LLM rewrite. The three BM25 ranking lists are then fused with a weighted linear combination whose weights are tuned on iKAT-23, and the fused top 50 are optionally reranked with monoT5 using the personalized rewrite as the query. Two manual runs use the provided human rewrites with BM25 plus monoT5 or RankLlama reranking. Official iKAT-24 scores are reported in Table 1; the best run (BM25 fusion + monoT5) reaches NDCG@10 47.7 versus 35.4 for manual BM25 + monoT5. The paper concludes that retrieval fusion is effective and discusses assessment bias in shallow-pooled iKAT collections, including a single-turn case study in which an expanded human rewrite has fewer assessed documents in the top 20.
Significance. Retrieval fusion is a plausible and inexpensive hedge against over-personalization, and the iKAT-24 evaluation is, with respect to the tuned weights, an out-of-collection test. The pipeline is described clearly enough to reproduce its general design, and the assessment-bias discussion raises a genuine methodological concern for future iKAT evaluations. These are real strengths. However, the central claim that fusion, rather than the automatic query variants or the reranking configuration, is responsible for the observed improvement is not established by the reported experiments. The paper's own Section 5 identifies a bias that also affects any non-submitted comparison condition, so the headline conclusion is currently under-determined. An ablation isolating the fusion operation, a report of the tuned weights, and a more systematic treatment of judgment incompleteness would materially strengthen the contribution.
major comments (4)
- [§4, Table 1] The Section 4 claim that the best run provides 'proving the effectiveness of retrieval-fusion' is not supported by Table 1. The table contains no automatic no-fusion control: both automatic runs use the same three-query fusion, while both manual runs use a single human rewrite. The automatic/manual comparison therefore varies the rewrite source, the number of query variants (including a response-expanded list), and the reranking query (the personalized rewrite for the fused run versus the human rewrite for the manual runs). The NDCG@10 gap between BM25 fusion + monoT5 (47.7) and manual BM25 + monoT5 (35.4) could be produced by any of these factors alone. Please add an ablation with an unfused automatic single-list run, or equivalently remove one list at a time, and restrict the wording to what Table 1 can show.
- [§3.2.1, Eq. (1)] The fusion weights alpha_1, alpha_2, and alpha_3 are said to be 'optimized based on performance tuning on the iKAT-23 test collection,' but their values are never reported and no sensitivity analysis is given. Without the weights, the linear-combination mechanism cannot be reproduced or interpreted. In addition, Section 4 states that on iKAT-23 the fusion-then-reranking approach did not outperform manual-rewrite runs; because iKAT-23 is the same collection used for weight tuning, that result is in-sample and cannot be used to validate the method. Please report the weights and, ideally, performance across a range of weight settings.
- [§5, Tables 2-3] The assessment-bias discussion draws on one illustrative turn (9-1-3) and on collection-level pooling statistics, which is useful context but not sufficient to support a causal reading of Table 1. The paper does not provide per-condition judged-document statistics for the actual iKAT-24 runs, and no unfused automatic variant was submitted; consequently, any post-hoc ablation on iKAT-24 would be subject to the same bias the paper documents, because documents retrieved only by the unfused variant would be less likely to have been judged. A systematic comparison of assessed-document rates in the top ranks of each condition, or re-judging of a sample of documents, is needed before the observed scores can be attributed to fusion.
- [§3.2.1, Eq. (1)] The fusion formula is a weighted sum of BM25 scores, but the paper does not state whether scores are normalized before combination. BM25 score scales differ across queries and query types, so a raw weighted sum can be dominated by whichever list happens to have larger score magnitudes, which would undermine the intended 'chorus effect.' If normalization is applied, it should be described; if not, this is a technical gap in the proposed mechanism. The defaulting of missing documents to the score of the 1000th document also needs to be reconciled with whatever normalization is used.
minor comments (5)
- [Title and abstract] The manuscript contains typographical errors, including 'Retrie val Fusion' in the title, 'and hinders search performance T his' in the abstract, and 'Rermarkably' in Section 2; it should be proofread.
- [§3.2] The text says four automatic runs were submitted but only two are described or listed; either include the other two runs in Table 1 or state clearly why they are omitted.
- [§3.2.1] The fusion equation is not numbered; numbering it would make the technical discussion and any future citations easier.
- [Table 1] Table 1 would be more informative with official run identifiers and a paired significance test, or at least a per-topic variance estimate, for the main NDCG@10 comparison.
- [§3.2.1, Related Work] The personalized rewrite prompt in Section 3.2.1 depends on the LLM4CS framework [20] and on reference [28], but [28] is not discussed in Related Work; a sentence or two describing that prior method would help readers situate the contribution.
Circularity Check
No circularity: the main iKAT-24 result is an out-of-collection evaluation after iKAT-23 tuning; self-citations are components, not load-bearing evidence.
full rationale
The paper's central claim is that score-based fusion of three LLM-rewritten queries (non-personalized, response-expanded, personalized) improves passage ranking. The fusion weights are tuned on iKAT-23 (Section 3.2.1) and the reported headline result in Table 1 is on iKAT-24, an independent test collection; the tuned weights are thus applied out-of-sample, not renamed as a prediction. The iKAT-23 re-evaluation in Section 4 is explicitly described as an auxiliary observation (fusion could not outperform manual runs there), not as evidence for the fusion claim. No equation in the paper reduces to its own input by construction: the linear-combination score is a method definition, not a fitted quantity disguised as a result. The self-citations to LLM4CS [20] and PTKB [28] supply prompt components and prior context, but the effectiveness of fusion is judged by TREC assessors on iKAT-24, so these citations are not load-bearing proof. The paper does contain a validity limitation that the authors themselves raise: Section 5 argues that shallow pooling makes iKAT collections unreliable for non-participant methods and that improved rewrites are under-assessed; this is a support/correctness concern about the "proving" language in Section 4 and about the missing no-fusion ablation, but it is not circularity. Consequently, no circular step can be quoted, and the appropriate finding is no significant circularity.
Assumptions & free parameters
free parameters (1)
- Fusion weights alpha_1, alpha_2, alpha_3 =
Not reported (tuned on iKAT-23)
assumptions (3)
- domain assumption BM25 lexical retrieval is an adequate base retriever for the three reformulated query types.
- domain assumption The iKAT test collections' pooled relevance judgments are sufficiently unbiased to compare non-participant fusion runs against participant runs.
- domain assumption GPT-4o (gpt-4o-2024-08-06) generates reliable de-contextualized and personalized rewrites using the described prompts.
Cite this review
Pith. "Pith review of RALI@TREC iKAT 2024: Achieving Personalization via Retrieval Fusion in Conversational Search." pith.science (2026). https://pith.science/paper/K77JCU4P
@misc{pith2026241207998,
author = {Pith},
title = {Pith review of: RALI@TREC iKAT 2024: Achieving Personalization via Retrieval Fusion in Conversational Search},
year = {2026},
howpublished = {\url{https://pith.science/paper/K77JCU4P}},
note = {Machine review of arXiv:2412.07998}
}
read the original abstract
The Recherche Appliquee en Linguistique Informatique (RALI) team participated in the 2024 TREC Interactive Knowledge Assistance (iKAT) Track. In personalized conversational search, effectively capturing a user's complex search intent requires incorporating both contextual information and key elements from the user profile into query reformulation. The user profile often contains many relevant pieces, and each could potentially complement the user's information needs. It is difficult to disregard any of them, whereas introducing an excessive number of these pieces risks drifting from the original query and hinders search performance. This is a challenge we denote as over-personalization. To address this, we propose different strategies by fusing ranking lists generated from the queries with different levels of personalization.
Reference graph
Works this paper leans on
-
[1]
Mohammad Aliannejadi, Zahra Abbasiantaeb, Shubham Cha tterjee, Jeffery Dal- ton, and Leif Azzopardi. 2024. Trec ikat 2023: The interacti ve knowledge assis- tance track overview. arXiv preprint arXiv:2401.01330 (2024)
arXiv 2024
-
[2]
Javed A. Aslam and Mark Montague. 2001. Models for metase arch. In Pro- ceedings of the 24th Annual International ACM SIGIR Confere nce on Research and Development in Information Retrieval (New Orleans, Louisiana, USA) (SI- GIR ’01). Association for Computing Machinery, New York, NY, USA, 27 6–284. https://doi.org/10.1145/383952.384007
arXiv 2001
-
[3]
Sebastian Bruch, Siyu Gai, and Amir Ingber. 2023. An anal ysis of fusion func- tions for hybrid retrieval. ACM Transactions on Information Systems 42, 1 (2023), 1–35
work page 2023
-
[4]
Chris Buckley, Darrin Dimmick, Ian Soboroff, and Ellen Vo orhees. 2007. Bias and the limits of pooling for large collections. Information retrieval 10 (2007), 491–508
work page 2007
-
[5]
Zhiyu Chen, Jie Zhao, Anjie Fang, Besnik Fetahu, Oleg Rok hlenko, and Shervin Malmasi. 2022. Reinforced Question Rewriting for Conversa tional Question Answering. In Proceedings of the 2022 Conference on Empirical Methods in Na t- ural Language Processing: Industry Track , Yunyao Li and Angeliki Lazaridou (Eds.). Association for Computational Linguistics,...
-
[6]
Cormack, Charles L A Clarke, and Stefan Buettch er
Gordon V. Cormack, Charles L A Clarke, and Stefan Buettch er. 2009. Recip- rocal rank fusion outperforms condorcet and individual ran k learning meth- ods. In Proceedings of the 32nd International ACM SIGIR Conference o n Re- search and Development in Information Retrieval (Boston, MA, USA) (SIGIR ’09). Association for Computing Machinery, New York, NY, U...
arXiv 2009
-
[7]
Nick Craswell, Bhaskar Mitra, Emine Yilmaz, Daniel Camp os, and Ellen M Voorhees. 2020. Overview of the TREC 2019 deep learning trac k. arXiv preprint arXiv:2003.07820 (2020)
arXiv 2020
-
[8]
Ahmed Elgohary, Denis Peskov, and Jordan Boyd-Graber. 2 019. Can You Unpack That? Learning to Rewrite Questions-in-Context. In Proceedings of the 2019 Conference on Empirical Methods in Natural Languag e Processing and the 9th International Joint Conference on Natural Langu age Processing (EMNLP-IJCNLP), Kentaro Inui, Jing Jiang, Vincent Ng, and Xiaojun W...
Show all 48 references
-
[9]
Hung-Chieh Fang, Kuo-Han Hung, Chen-Wei Huang, and Yun- Nung Chen
-
[10]
Edward Fox and Joseph Shaw. 1994. Combination of multip le searches. NIST special publication SP (1994), 243–243
1994
-
[11]
Ilias Gialampoukidis, Anastasia Moumtzidou, Dimitri s Liparas, Stefanos Vrochidis, and Ioannis Kompatsiaris. 2016. A hybrid graph-based and non-linear late fusion approach for multimedia retrieval. In 2016 14th International work- shop on content-based multimedia indexing (CBM...
2016
-
[12]
Junjie Huang, Jiarui Qin, Jianghao Lin, Ziming Feng, Yo ng Yu, and Weinan Zhang. 2024. Unleashing the Potential of Multi-Channel Fus ion in Retrieval for Personalized Recommendations. arXiv preprint arXiv:2410.16080 (2024)
2024 arXiv
-
[13]
Vladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Le wis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih. 2020. Dense Passage Retrieval for Open- Domain Question Answering. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNL...
2020 doi
-
[14]
Vaibhav Kumar and Jamie Callan. 2020. Making informati on seeking easier: An improved pipeline for conversational search. In Findings of the Association for Computational Linguistics: EMNLP 2020 . 3971–3980
2020
-
[15]
Oren Kurland and J Shane Culpepper. 2018. Fusion in info rmation retrieval: Sigir 2018 half-day tutorial. In The 41st International ACM SIGIR Conference on Research & Development in Information Retrieval . 1383–1386
2018
-
[16]
Sheng-Chieh Lin, Jheng-Hong Yang, Rodrigo Nogueira, M ing-Feng Tsai, Chuan- Ju Wang, and Jimmy Lin. 2020. Conversational question refor mulation via sequence-to-sequence architectures and pretrained langu age models. arXiv preprint arXiv:2004.01909 (2020)
2020 arXiv
-
[17]
Xueguang Ma, Liang Wang, Nan Yang, Furu Wei, and Jimmy Li n. 2024. Fine- tuning llama for multi-stage text retrieval. In Proceedings of the 47th Interna- tional ACM SIGIR Conference on Research and Development in I nformation Re- trieval. 2421–2425. RALI@TREC iKAT 2024: Achiev...
2024
-
[18]
Kelong Mao, Chenlong Deng, Haonan Chen, Fengran Mo, Zhe ng Liu, Tet- suya Sakai, and Zhicheng Dou. 2024. ChatRetriever: Adaptin g Large Lan- guage Models for Generalized and Robust Conversational Den se Retrieval. In Proceedings of the 2024 Conference on Empirical Methods in N...
2024
-
[19]
Kelong Mao, Zhicheng Dou, Bang Liu, Hongjin Qian, Fengr an Mo, Xiangli Wu, Xiaohua Cheng, and Zhao Cao. 2023. Search-oriented conversational query edit- ing. In Findings of the Association for Computational Linguistics : ACL 2023. 4160– 4172
2023
-
[21]
Kelong Mao, Hongjin Qian, Fengran Mo, Zhicheng Dou, Ban g Liu, Xiaohua Cheng, and Zhao Cao. 2023. Learning denoised and interpreta ble session repre- sentation for conversational search. In Proceedings of the ACM Web Conference
2023
-
[22]
Fengran Mo, Abbas Ghaddar, Kelong Mao, Mehdi Rezagholi zadeh, Boxing Chen, Qun Liu, and Jian-Yun Nie. 2024. CHIQ: Contextual Hist ory En- hancement for Improving Query Rewriting in Conversational Search. In Pro- ceedings of the 2024 Conference on Empirical Methods in Natu ral ...
2024
-
[23]
Fengran Mo, Kelong Mao, Ziliang Zhao, Hongjin Qian, Hao nan Chen, Yiruo Cheng, Xiaoxi Li, Yutao Zhu, Zhicheng Dou, and Jian-Yun Nie. 2024. A survey of conversational search. arXiv preprint arXiv:2410.15576 (2024)
2024 arXiv
-
[24]
Fengran Mo, Kelong Mao, Yutao Zhu, Yihong Wu, Kaiyu Huan g, and Jian-Yun Nie. 2023. ConvGQR: Generative Query Reformulation for Con versational Search. In Proceedings of the 61st Annual Meeting of the Association for Compu- tational Linguistics (Volume 1: Long Papers) , Anna R...
2023 doi
-
[25]
Fengran Mo, Jian-Yun Nie, Kaiyu Huang, Kelong Mao, Yuta o Zhu, Peng Li, and Yang Liu. 2023. Learning to relate to previous turns in conve rsational search. In Proceedings of the 29th ACM SIGKDD Conference on Knowledge Di scovery and Data Mining. 1722–1732
2023
-
[26]
Fengran Mo, Chen Qu, Kelong Mao, Yihong Wu, Zhan Su, Kaiy u Huang, and Jian-Yun Nie. 2024. Aligning query representation with rew ritten query and rel- evance judgments in conversational search. In Proceedings of the 33rd ACM In- ternational Conference on Information and Knowl...
2024
-
[27]
Fengran Mo, Bole Yi, Kelong Mao, Chen Qu, Kaiyu Huang, an d Jian-Yun Nie
-
[28]
Fengran Mo, Longxiang Zhao, Kaiyu Huang, Yue Dong, Dege n Huang, and Jian- Yun Nie. 2024. How to Leverage Personal Textual Knowledge fo r Personalized Conversational Information Retrieval. In Proceedings of the 33rd ACM Interna- tional Conference on Information and Knowledge M...
2024
-
[29]
Mark Montague and Javed A Aslam. 2002. Condorcet fusion for improved re- trieval. In Proceedings of the eleventh international conference on Inf ormation and knowledge management. 538–548
2002
-
[30]
Rodrigo Nogueira, Zhiying Jiang, Ronak Pradeep, and Ji mmy Lin. 2020. Docu- ment Ranking with a Pretrained Sequence-to-Sequence Model. In Findings of the Association for Computational Linguistics: EMNLP 2020 , Trevor Cohn, Yulan He, and Yang Liu (Eds.). Association for Computa...
2020 doi
-
[31]
Hongjin Qian and Zhicheng Dou. 2022. Explicit query rew riting for conversa- tional dense retrieval. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing . 4725–4737
2022
-
[32]
Chen Qu, Liu Yang, Cen Chen, Minghui Qiu, W Bruce Croft, a nd Mohit Iyyer
-
[33]
Ian Soboroff. 2024. Don’t Use LLMs to Make Relevance Judg ments. arXiv preprint arXiv:2409.15133 (2024)
2024 arXiv
-
[34]
Svitlana Vakulenko, Shayne Longpre, Zhucheng Tu, and R aviteja Anantha. 2021. Question rewriting for conversational question answering . In Proceedings of the 14th ACM international conference on web search and data min ing. 355–363
2021
-
[35]
Christopher C Vogt and Garrison W Cottrell. 1999. Fusio n via a linear combina- tion of scores. Information retrieval 1, 3 (1999), 151–173
1999
-
[36]
E Voorhees, Narendra K Gupta, and Ben Johnson-Laird. 19 95. The collection fusion problem. NIST SPECIAL PUBLICATION SP (1995), 95–95
1995
-
[37]
Ellen M Voorhees, Ian Soboroff, and Jimmy Lin. 2022. Can O ld TREC Col- lections Reliably Evaluate Modern Neural Retrieval Models ? arXiv preprint arXiv:2201.11086 (2022)
2022 arXiv
-
[38]
Nikos Voskarides, Dan Li, Pengjie Ren, Evangelos Kanou las, and Maarten de Rijke. 2020. Query resolution for conversational search with limited supervision. In Proceedings of the 43rd International ACM SIGIR conference o n research and development in Information Retrieval . 921–930
2020
-
[39]
Shuai Wang, Shengyao Zhuang, and Guido Zuccon. 2021. Be rt-based dense re- trievers require interpolation with bm25 for effective pass age retrieval. In Pro- ceedings of the 2021 ACM SIGIR international conference on t heory of information retrieval. 317–324
2021
-
[40]
Shengli Wu. 2012. Linear combination of component resu lts in information re- trieval. Data & Knowledge Engineering 71, 1 (2012), 114–126
2012
-
[41]
Shengli Wu, Yaxin Bi, Xiaoqin Zeng, and Lixin Han. 2009. Assigning appropriate weights for the linear combination data fusion method in inf ormation retrieval. Information Processing & Management 45, 4 (2009), 413–426
2009
-
[42]
Reitter, and Gaura v Singh Tomar
Zeqiu Wu, Yi Luan, Hannah Rashkin, D. Reitter, and Gaura v Singh Tomar. 2021. CONQRR: Conversational Query Rewriting for Retrieval with Reinforcement Learning. In Conference on Empirical Methods in Natural Language Process ing
2021
-
[43]
Fanghua Ye, Meng Fang, Shenghui Li, and Emine Yilmaz. 20 23. En- hancing Conversational Search: Large Language Model-Aide d Informa- tive Query Rewriting. In Findings of the Association for Computational Linguistics: EMNLP 2023 , Houda Bouamor, Juan Pino, and Kalika Bali (Eds....
2023 doi
-
[44]
Emine Yilmaz, Nick Craswell, Bhaskar Mitra, and Daniel Campos. 2020. On the Reliability of Test Collections for Evaluating Syste ms of Different Types. In Proceedings of the 43rd International ACM SIGIR Conference o n Re- search and Development in Information Retrieval (Virtual...
2020
-
[45]
Shi Yu, Jiahua Liu, Jingqin Yang, Chenyan Xiong, Paul Be nnett, Jianfeng Gao, and Zhiyuan Liu. 2020. Few-shot generative conversational query rewriting. In Proceedings of the 43rd International ACM SIGIR conference o n research and development in Information Retrieval . 1933–1936
2020
-
[46]
acm-jdslogo.png
Hamed Zamani, Johanne R. Trippas, Jeffrey Dalton, and Fi lip Radlinski. 2022. Conversational Information Seeking. Found. Trends Inf. Retr. 17 (2022), 244–456. https://api.semanticscholar.org/CorpusID:246210119 This figure "acm-jdslogo.png" is available in "png" format from: htt...
2022 arXiv
-
[2020]
In Proceedings of the 43rd International ACM SIGIR conference on research and development in Informa- tion Retrieval
Open-retrieval conversational question answering. In Proceedings of the 43rd International ACM SIGIR conference on research and development in Informa- tion Retrieval. 539–548
-
[2022]
In Findings of the Association for Computational Linguistics : AACL-IJCNLP 2022 , Yulan He, Heng Ji, Sujian Li, Yang Liu, and Chua-Hui Chang (Eds.)
Open-Domain Conversational Question Answering with Histori- cal Answers. In Findings of the Association for Computational Linguistics : AACL-IJCNLP 2022 , Yulan He, Heng Ji, Sujian Li, Yang Liu, and Chua-Hui Chang (Eds.). Association for Computational Linguistics, Online only,...
2022
-
[2024]
In Companion Proceedings of the ACM on Web Conference 2024
Convsdg: Session data generation for conversational search. In Companion Proceedings of the ACM on Web Conference 2024 . 1634–1642
2024
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.