REVIEW 2 major objections 6 minor 85 references
Improving GenIR Systems Based on User Feedback
T0 review · 2 major / 6 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read This survey chapter argues that improving generative information retrieval with feedback now hinges on alignment methods that see groups of inputs, because pointwise reward collection cannot make models produce outputs unique to each…
desk verdict A useful survey of feedback-driven GenIR improvement, but the RLCF necessity claim is asserted on the authors' own evidence and needs tempering. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The organizing device is a two-axis taxonomy borrowed from learning to rank: reward input (pointwise vs groupwise) and reward training (pointwise vs groupwise). On this grid, RLHF and RLAIF are pointwise-input methods, while RLCF (Reinforcement Learning from Contrastive Feedback) is the named mechanism that is groupwise in both input and training. RLCF lets the LLM generate outputs for several similar inputs at once and builds rewards by contrasting outputs across inputs, which is the mechanism claimed to teach fine-grained discrimination among near-duplicate documents.
What would settle it
Run the same base LLM through three alignments on identical reward data: pointwise RLHF, pointwise RLAIF, and groupwise-input RLCF, then test on a corpus of near-duplicate documents where the task is to generate a distinct snippet or query for each. If pointwise-aligned models achieve the same level of output distinctness and retrieval utility as the groupwise model, the chapter's central claim is wrong; if groupwise wins consistently across several base models and corpora, the claim is supported.
Extended reading notes
Core claim
On the paper's own terms, the discovery is a re-framing: user feedback for GenIR systems must be understood as coming from humans, LLM agents, and other systems, and the techniques that exploit it divide into alignment (reward collection and parameter optimization) and learning (continual learning, conversational ranking, prompt learning). The survey proposes that the pointwise/groupwise distinction from learning to rank explains why off-the-shelf alignment fails for information access: RLHF and RLAIF score outputs one prompt at a time, so they cannot teach a model to distinguish among highly similar documents, whereas groupwise input/output contrastive feedback (RLCF) can. The chapter concludes that such innovative techniques, beyond traditional feedback use, are driving GenIR evolution and lists open problems around intent understanding, limited but rich feedback, user-centric evaluation, and privacy.
Load-bearing premise
The argument assumes that what the authors see in their own studies—that rewarding one output at a time cannot teach a model to produce output unique to a given input, and that comparing several inputs side by side fixes this—holds for GenIR systems generally.
Editorial extensions
If this is right
- Alignment for information access should move from pointwise reward input to groupwise input/output paradigms, since fine-grained discrimination among similar items is a core IA need.
- Models that generate summaries, snippets, or rewrites for similar documents can be expected to produce more distinct, useful outputs after RLCF-style alignment than after standard RLHF or RLAIF.
- User feedback from LLM agents and client systems should be treated as a first-class signal in GenIR, requiring methods that handle bi-directional and multi-party interactions.
- The learning-to-rank toolbox (pairwise/listwise losses and groupwise scoring) becomes directly relevant to LLM alignment, so ranking-theoretic insights can guide future alignment methods.
- Conversational search systems can leverage LLM-generated query rewrites and session data to improve intent understanding, with continual learning as the mechanism for keeping generative systems current despite catastrophic forgetting.
Reading between the lines
- One step beyond the paper: if groupwise input is what teaches fine-grained discrimination, then RLCF should also improve multi-document summarization, recommendation explanation, or any generation task where outputs for similar inputs must differ; no such experiment is reported in the chapter.
- The chapter's widening of 'user' to agents and clients suggests a failure mode the authors list as open: self-feedback loops in which an agent's generated content becomes training data for the same system, amplifying artificial intentions; one could simulate repeated agent-system-agent rounds and measure representational drift.
- Borrowing the learning-to-rank analogy further, groupwise alignment should reduce optimization variance as well as improve discriminative quality; this is a quantitative prediction that could be checked by comparing reward-model variance across pointwise and groupwise training runs.
- A neutral benchmark for 'output uniqueness'—for example, pairwise distinctness of generated responses to near-identical inputs—would let the field test the central claim without relying on the authors' own task setups.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This chapter surveys approaches for improving generative information retrieval (GenIR) systems using user feedback. It begins by broadening the notion of 'user' to include humans, LLM agents, and clients, and distinguishes implicit and explicit feedback. Four general strategies are outlined: prompt engineering, fine-tuning, preference/intent integration, and agent-based feedback. Section 2 examines alignment, first identifying objectives shared with general LLMs (preventing harm, user intent, ethics) and objectives specific to information access (personalization, fine-grained discrimination). The authors propose a taxonomy of reward collection methods based on pointwise vs. groupwise input and pointwise vs. groupwise training, and they place RLHF, RLAIF, and RLCF within this taxonomy. Optimization methods, including PPO, DPO, and ranking-based losses, are summarized. Section 3 reviews continual learning, conversational search, and prompt learning, and Section 4 lists challenges such as user intention understanding, limited but rich feedback, user-centric evaluation, and privacy. The central thesis is that innovative feedback-driven techniques, particularly groupwise contrastive feedback alignment (RLCF), are advancing GenIR beyond traditional feedback usage.
Significance. The chapter offers a useful organizational scheme for alignment methods in IR by adapting learning-to-rank terminology, and it correctly identifies fine-grained discrimination as an underappreciated requirement in generative retrieval. The extension of 'user' to agents and clients is timely and relevant to current GenIR practice. However, the chapter's most consequential recommendation—that pointwise-input RLHF/RLAIF cannot teach fine-grained discrimination and therefore groupwise RLCF is necessary—is not established by the evidence cited and is presented too confidently. As a survey, the chapter need not contain new experiments, but it should accurately represent the strength of the evidence. The footnote in Section 2.1.2 acknowledging that the categorization is 'not inclusive' is welcome, but the main text does not always maintain that caution. If the overclaim is softened and the taxonomy is clarified, the survey would be a solid reference for researchers and practitioners.
major comments (2)
- [§2.2.1] The claim that RLHF/RLAIF's pointwise-input paradigm makes it 'difficult, if not impossible' to teach LLMs to generate outputs unique to an input is load-bearing for the chapter's recommendation of RLCF, but it is supported mainly by the authors' own SIGIR 2024 paper (ref [41]) and one related work (ref [42]), with no independent replication or neutral comparison. The 'impossible' wording is an overstatement: a pointwise-input reward model trained on labels that encode groupwise comparisons could in principle learn to reward uniqueness, because groupwise information would be carried by the labels rather than by the reward function's input. Please soften the claim to 'current pointwise-input methods typically fail to...' and discuss this alternative, or provide a formal argument why label-based pointwise approaches cannot succeed.
- [§2.2.1 and Table 1] The taxonomy's two dimensions (input and training) are not applied consistently enough for readers to place methods unambiguously. RLHF and RLAIF are described as pointwise input and potentially pairwise/groupwise in training, but Table 1's cell entries are not clearly separated by column, and RLCF is presented as the only groupwise input/output method. Section 2.2.2's ranking-based methods (RRHF, RAFT) also operate on groups of responses, so the chapter should explicitly clarify whether 'groupwise' refers to the set of input documents or the set of output responses for a single prompt. Without this distinction, the positioning of RLCF as the IR-specific groupwise alignment technique is confusing and may overstate its novelty.
minor comments (6)
- [General] The manuscript contains numerous typos that should be corrected in revision, including 'feedback feedback' (Section 1), 'forus' (Section 2), 'Micorsoft' and 'raciest' (Section 2), 'visal' (Section 1.2), and 'functiosn' (Section 2.2.1).
- [Section 2, reference [18]] The statement 'In the first technical report of ChatGPT [18]' cites the GPT-4 Technical Report, not the ChatGPT technical report; please correct the citation to the appropriate ChatGPT/InstructGPT paper or rephrase the sentence.
- [Table 1] The entries in Table 1 are not visually separated by column, making it difficult to determine which references correspond to pointwise input, pairwise/groupwise input, pointwise training, or pairwise/groupwise training; reformat the table with clear column and row labels.
- [Figure 3] Figure 3, as rendered in the text, has overlapping labels such as 'Pointwise Optimization' and 'Pairwise/Groupwise Training' with unclear arrows; redrawing the figure with separate panels for the input and training dimensions would improve readability.
- [Section 3.2] The sentence 'The massive knowledge about conversation patterns and the world of LLMs also makes it a promising end-to-end foundation to be an end-to-end foundation model for personalized conversational search systems' repeats 'end-to-end' and should be revised.
- [Section 4] The future-direction bullet on self-feedback loops is intriguing but underdeveloped; since the chapter extends the notion of user to agents and clients, one or two sentences on how to detect or mitigate feedback loops would strengthen the discussion.
Circularity Check
No circular derivation: the chapter is a survey whose RLCF claims rest on published, externally checkable studies rather than on equations that reduce to their inputs.
full rationale
This is a review chapter, not a derivation. It introduces no fitted parameters, no prediction from fitted quantities, and no imported uniqueness theorem; its taxonomy of alignment methods is explicitly borrowed from LTR literature ([46,47]) and applied as classification, not as proof. The only potential self-citation concern is in Section 2.2.1, where the motivation for RLCF ('it's difficult, if not impossible, to teach LLMs to generate outputs unique to an input without seeing and comparing with other candidate inputs', followed by 'Therefore, RLCF is proposed...') is supported by the authors' own SIGIR 2024 paper [41] and by an independent paper [42]. That support is a published, empirically testable study rather than an unexamined premise, and the chapter summarizes it rather than deriving a new result from it. The stronger impossibility statement is under-justified and is a correctness or evidence risk, not a circularity: no equation or construction forces the conclusion from the inputs. Similarly, the discussion of groupwise methods as having 'more capacity' and 'less variance' is an acknowledged transfer from learning-to-rank theory, not a re-branded derivation. Hence no circular step is present.
Assumptions & free parameters
assumptions (3)
- domain assumption The cited experimental results (e.g., RLCF effectiveness, RLHF/DPO comparisons) are accurately reported in the referenced papers.
- ad hoc to paper The taxonomy of reward collection methods (pointwise/groupwise input and training) is a useful and correct organizational scheme for alignment methods.
- domain assumption The extension of 'user' to include LLM agents and client systems is a meaningful conceptual shift for GenIR.
Cite this review
Pith. "Pith review of Improving GenIR Systems Based on User Feedback." pith.science (2026). https://pith.science/paper/I43VZVKI
@misc{pith2026250102838,
author = {Pith},
title = {Pith review of: Improving GenIR Systems Based on User Feedback},
year = {2026},
howpublished = {\url{https://pith.science/paper/I43VZVKI}},
note = {Machine review of arXiv:2501.02838}
}
read the original abstract
In this chapter, we discuss how to improve the GenIR systems based on user feedback. Before describing the approaches, it is necessary to be aware that the concept of "user" has been extended in the interactions with the GenIR systems. Different types of feedback information and strategies are also provided. Then the alignment techniques are highlighted in terms of objectives and methods. Following this, various ways of learning from user feedback in GenIR are presented, including continual learning, learning and ranking in the conversational context, and prompt learning. Through this comprehensive exploration, it becomes evident that innovative techniques are being proposed beyond traditional methods of utilizing user feedback, and contribute significantly to the evolution of GenIR in the new era. We also summarize some challenging topics and future directions that require further investigation.
Reference graph
Works this paper leans on
-
[41]
Dong, Q., Liu, Y., Ai, Q., Wu, Z., Li, H., Liu, Y., Wang, S., Yin, D., Ma, S.: Unsupervised large language model alignment for information retrieval via contrastive feedback. In: Proceedings of the 47th International ACM SIGIR Con- ference on Research and Development in Information Retrieval. SIGIR ’24, pp. 48–58. Association for Computing Machinery, New ...
arXiv 2024
-
[42]
Yoon, C., Kim, G., Jeon, B., Kim, S., Jo, Y., Kang, J.: Ask Optimal Questions: Aligning Large Language Models with Retriever’s Preference in Conversational Search (2024)
work page 2024
-
[1]
arXiv preprint arXiv:2402.08670 (2024)
Liu, Y., Wang, Y., Sun, L., Yu, P.S.: Rec-gpt4v: Multimodal recommendation 16 with large vision-language models. arXiv preprint arXiv:2402.08670 (2024)
arXiv 2024
-
[2]
In: Proceedings of the 17th ACM Conference on Recommender Systems, pp
Dai, S., Shao, N., Zhao, H., Yu, W., Si, Z., Xu, C., Sun, Z., Zhang, X., Xu, J.: Uncovering chatgpt’s capabilities in recommender systems. In: Proceedings of the 17th ACM Conference on Recommender Systems, pp. 1126–1132 (2023)
2023
-
[3]
arXiv preprint arXiv:2304.10149 (2023)
Liu, J., Liu, C., Lv, R., Zhou, K., Zhang, Y.: Is chatgpt a good recommender? a preliminary study. arXiv preprint arXiv:2304.10149 (2023)
arXiv 2023
-
[4]
arXiv preprint arXiv:2304.03153 (2023)
Wang, L., Lim, E.-P.: Zero-shot next-item recommendation using large pretrained language models. arXiv preprint arXiv:2304.03153 (2023)
arXiv 2023
-
[5]
In: Proceedings of the 16th ACM Conference on Recommender Systems, pp
Geng, S., Liu, S., Fu, Z., Ge, Y., Zhang, Y.: Recommendation as language process- ing (rlp): A unified pretrain, personalized prompt & predict paradigm (p5). In: Proceedings of the 16th ACM Conference on Recommender Systems, pp. 299–315 (2022)
2022
-
[6]
In: European Conference on Information Retrieval, pp
Hou, Y., Zhang, J., Lin, Z., Lu, H., Xie, R., McAuley, J., Zhao, W.X.: Large language models are zero-shot rankers for recommender systems. In: European Conference on Information Retrieval, pp. 364–381 (2024). Springer
2024
Show all 85 references
-
[7]
Advances in Neural Information Processing Systems 36 (2024)
Rajput, S., Mehta, N., Singh, A., Hulikal Keshavan, R., Vu, T., Heldt, L., Hong, L., Tay, Y., Tran, V., Samost, J., et al.: Recommender systems with generative retrieval. Advances in Neural Information Processing Systems 36 (2024)
2024
-
[8]
In: Proceedings of the 31st ACM International Conference on Multimedia, pp
Zhai, J., Zheng, X., Wang, C.-D., Li, H., Tian, Y.: Knowledge prompt-tuning for sequential recommendation. In: Proceedings of the 31st ACM International Conference on Multimedia, pp. 6451–6461 (2023)
2023
-
[9]
arXiv preprint arXiv:2312.02445 (2023)
Liao, J., Li, S., Yang, Z., Wu, J., Yuan, Y., Wang, X., He, X.: Llara: Aligning large language models with sequential recommenders. arXiv preprint arXiv:2312.02445 (2023)
2023 arXiv
-
[10]
arXiv preprint arXiv:2401.13870 (2024)
Luo, S., Yao, Y., He, B., Huang, Y., Zhou, A., Zhang, X., Xiao, Y., Zhan, M., Song, L.: Integrating large language models into recommendation via mutual augmentation and adaptive aggregation. arXiv preprint arXiv:2401.13870 (2024)
2024
-
[11]
arXiv preprint arXiv:2306.11114 (2023)
Petrov, A.V., Macdonald, C.: Generative sequential recommendation with gptrec. arXiv preprint arXiv:2306.11114 (2023)
2023 arXiv
-
[12]
arXiv preprint arXiv:2305.07001 (2023)
Zhang, J., Xie, R., Hou, Y., Zhao, W.X., Lin, L., Wen, J.-R.: Recommendation as instruction following: A large language model empowered recommendation approach. arXiv preprint arXiv:2305.07001 (2023)
2023 arXiv
-
[13]
arXiv preprint arXiv:2310.10108 (2023)
Zhang, A., Sheng, L., Chen, Y., Li, H., Deng, Y., Wang, X., Chua, T.-S.: On generative agents in recommendation. arXiv preprint arXiv:2310.10108 (2023)
2023 arXiv
-
[14]
arXiv preprint arXiv:2308.16505 (2023)
Huang, X., Lian, J., Lei, Y., Yao, J., Lian, D., Xie, X.: Recommender ai 17 agent: Integrating large language models for interactive recommendations. arXiv preprint arXiv:2308.16505 (2023)
2023 arXiv
-
[15]
arXiv preprint arXiv:2308.09904 (2023)
Shu, Y., Gu, H., Zhang, P., Zhang, H., Lu, T., Li, D., Gu, N.: Rah! recsys- assistant-human: A human-central recommendation framework with large lan- guage models. arXiv preprint arXiv:2308.09904 (2023)
2023 arXiv
-
[16]
arXiv preprint arXiv:2306.02552 (2023)
Wang, L., Zhang, J., Chen, X., Lin, Y., Song, R., Zhao, W.X., Wen, J.-R.: Reca- gent: A novel simulation paradigm for recommender systems. arXiv preprint arXiv:2306.02552 (2023)
2023 arXiv
-
[17]
arXiv preprint arXiv:2308.14296 (2023)
Wang, Y., Jiang, Z., Chen, Z., Yang, F., Zhou, Y., Cho, E., Fan, X., Huang, X., Lu, Y., Yang, Y.: Recmind: Large language model powered agent for recommendation. arXiv preprint arXiv:2308.14296 (2023)
2023 arXiv
-
[18]
OpenAI, :, Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Ale- man, F.L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., Avila, R., Babuschkin, I., Balaji, S., Balcom, V., Baltescu, P., Bao, H., Bavarian, M., Bel- gum, J., Bello, I., Berdine, J., Bernadett...
2023
-
[19]
In: Guyon, I., Luxburg, U.V., Bengio, S., Wallach, H., Fergus, R., Vishwanathan, S., Garnett, R
Christiano, P.F., Leike, J., Brown, T., Martic, M., Legg, S., Amodei, D.: Deep reinforcement learning from human preferences. In: Guyon, I., Luxburg, U.V., Bengio, S., Wallach, H., Fergus, R., Vishwanathan, S., Garnett, R. (eds.) Advances in Neural Information Processing Syste...
2017
-
[20]
Bai, Y., Jones, A., Ndousse, K., Askell, A., Chen, A., DasSarma, N., Drain, D., Fort, S., Ganguli, D., Henighan, T., Joseph, N., Kadavath, S., Kernion, J., Con- erly, T., El-Showk, S., Elhage, N., Hatfield-Dodds, Z., Hernandez, D., Hume, T., Johnston, S., Kravec, S., Lovitt, L...
2022
-
[21]
In: Burges, C.J., Bottou, L., Welling, M., Ghahramani, Z., Weinberger, K.Q
Mikolov, T., Sutskever, I., Chen, K., Corrado, G.S., Dean, J.: Distributed rep- resentations of words and phrases and their compositionality. In: Burges, C.J., Bottou, L., Welling, M., Ghahramani, Z., Weinberger, K.Q. (eds.) Advances in Neural Information Processing Systems, v...
2013
-
[22]
In: 2011 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp
Mikolov, T., Kombrink, S., Burget, L., ˇCernock´ y, J., Khudanpur, S.: Extensions of recurrent neural network language model. In: 2011 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 5528–5531 (2011). https://doi.org/10.1109/ICASSP.2011.5947611
2011
-
[23]
In: Guyon, I., Luxburg, U.V., Bengio, S., Wallach, H., Fergus, R., Vishwanathan, S., Garnett, R
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, L.u., Polosukhin, I.: Attention is all you need. In: Guyon, I., Luxburg, U.V., Bengio, S., Wallach, H., Fergus, R., Vishwanathan, S., Garnett, R. (eds.) Advances in Neural Information Processi...
2017
-
[24]
Peters, M.E., Ammar, W., Bhagavatula, C., Power, R.: Semi-supervised sequence tagging with bidirectional language models (2017)
2017
-
[25]
Devlin, J., Chang, M.-W., Lee, K., Toutanova, K.: BERT: Pre-training of Deep 19 Bidirectional Transformers for Language Understanding (2019)
2019
-
[26]
Vincent, J.: Twitter taught microsoft’s ai chatbot to be a racist asshole in less than a day
-
[27]
Accessed 2024-02-19
Silva, C.: It took just one weekend for meta’s new ai chatbot to become racist. Accessed 2024-02-19
2024
-
[28]
Zhang, Z., Lei, L., Wu, L., Sun, R., Huang, Y., Long, C., Liu, X., Lei, X., Tang, J., Huang, M.: SafetyBench: Evaluating the Safety of Large Language Models with Multiple Choice Questions (2023)
2023
-
[29]
In: Koyejo, S., Mohamed, S., Agarwal, A., Belgrave, D., Cho, K., Oh, A
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., Schulman, J., Hilton, J., Kelton, F., Miller, L., Simens, M., Askell, A., Welinder, P., Christiano, P.F., Leike, J., Lowe, R.: Train- ing language models to fol...
2022
-
[30]
SIGKDD Explor
Spirin, N., Han, J.: Survey on web spam detection: principles and algorithms. SIGKDD Explor. Newsl. 13(2), 50–64 (2012) https://doi.org/10.1145/2207243. 2207252
2012 doi
-
[31]
In: Proceedings of the 14th ACM International Conference on Information and Knowledge Management
Chirita, P.-A., Diederich, J., Nejdl, W.: Mailrank: using ranking for spam detec- tion. In: Proceedings of the 14th ACM International Conference on Information and Knowledge Management. CIKM ’05, pp. 373–380. Association for Comput- ing Machinery, New York, NY, USA (2005). htt...
2005 doi
-
[32]
Wolf, Y., Wies, N., Avnery, O., Levine, Y., Shashua, A.: Fundamental Limitations of Alignment in Large Language Models (2024)
2024
-
[33]
In: Proceedings of the 19th International Conference on World Wide Web
Cheng, Z., Gao, B., Liu, T.-Y.: Actively predicting diverse search intent from user browsing behaviors. In: Proceedings of the 19th International Conference on World Wide Web. WWW ’10, pp. 221–230. Association for Computing Machinery, New York, NY, USA (2010). https://doi.org/...
2010
-
[34]
In: Boughanem, M., Berrut, C., Mothe, J., Soule-Dupuy, C
Ashkan, A., Clarke, C.L.A., Agichtein, E., Guo, Q.: Classifying and characterizing query intent. In: Boughanem, M., Berrut, C., Mothe, J., Soule-Dupuy, C. (eds.) Advances in Information Retrieval, pp. 578–586. Springer, Berlin, Heidelberg (2009)
2009
-
[35]
In: Proceedings of the Eleventh ACM International Conference on Web Search and Data Mining
Su, N., He, J., Liu, Y., Zhang, M., Ma, S.: User intent, behaviour, and per- ceived satisfaction in product search. In: Proceedings of the Eleventh ACM International Conference on Web Search and Data Mining. WSDM ’18, pp. 547–555. Association for Computing Machinery, New York,...
2018
-
[36]
midterm elections
Trielli, D., Diakopoulos, N.: Partisan search behavior and google results in the 2018 u.s. midterm elections. Information, Communication & Society 25(1), 145– 161 (2022) https://doi.org/10.1080/1369118X.2020.1764605
2022
-
[37]
Proceedings of the National Academy of Sciences 112(33), 4512–4521 (2015) https://doi.org/10.1073/pnas
Epstein, R., Robertson, R.E.: The search engine manipulation effect (seme) and its possible impact on the outcomes of elections. Proceedings of the National Academy of Sciences 112(33), 4512–4521 (2015) https://doi.org/10.1073/pnas. 1419828112
2015 doi
-
[38]
ACM Trans
Teevan, J., Dumais, S.T., Horvitz, E.: Potential for personalization. ACM Trans. Comput.-Hum. Interact. 17(1) (2010) https://doi.org/10.1145/1721831.1721835
2010
-
[39]
In: Proceedings of the Sixteenth ACM Conference on Conference on Information and Knowledge Management
Sieg, A., Mobasher, B., Burke, R.: Web search personalization with ontological user profiles. In: Proceedings of the Sixteenth ACM Conference on Conference on Information and Knowledge Management. CIKM ’07, pp. 525–534. Association for Computing Machinery, New York, NY, USA (2...
2007
-
[40]
In: Proceedings of the 35th International ACM SIGIR Conference on Research and Development in Information Retrieval
Bennett, P.N., White, R.W., Chu, W., Dumais, S.T., Bailey, P., Borisyuk, F., Cui, X.: Modeling the impact of short- and long-term behavior on search per- sonalization. In: Proceedings of the 35th International ACM SIGIR Conference on Research and Development in Information Ret...
-
[43]
Ziegler, D.M., Stiennon, N., Wu, J., Brown, T.B., Radford, A., Amodei, D., Chris- tiano, P., Irving, G.: Fine-Tuning Language Models from Human Preferences (2020)
2020
-
[44]
Lee, H., Phatale, S., Mansoor, H., Mesnard, T., Ferret, J., Lu, K., Bishop, C., Hall, E., Carbune, V., Rastogi, A., Prakash, S.: RLAIF: Scaling Reinforcement Learning from Human Feedback with AI Feedback (2023)
2023
-
[45]
Yang, K., Klein, D., Celikyilmaz, A., Peng, N., Tian, Y.: RLCD: Reinforcement 21 Learning from Contrast Distillation for Language Model Alignment (2023)
2023
-
[46]
Foundations and Trends® in Information Retrieval 3(3), 225–331 (2009) https://doi.org/10.1561/ 1500000016
Liu, T.-Y.: Learning to rank for information retrieval. Foundations and Trends® in Information Retrieval 3(3), 225–331 (2009) https://doi.org/10.1561/ 1500000016
2009
-
[47]
In: Proceed- ings of the 2019 ACM SIGIR International Conference on Theory of Information Retrieval
Ai, Q., Wang, X., Bruch, S., Golbandi, N., Bendersky, M., Najork, M.: Learning groupwise multivariate scoring functions using deep neural networks. In: Proceed- ings of the 2019 ACM SIGIR International Conference on Theory of Information Retrieval. ICTIR ’19, pp. 85–92. Associ...
2019
-
[48]
Stiennon, N., Ouyang, L., Wu, J., Ziegler, D.M., Lowe, R., Voss, C., Radford, A., Amodei, D., Christiano, P.: Learning to summarize from human feedback (2022)
2022
-
[49]
In: Oh, A., Neumann, T., Globerson, A., Saenko, K., Hardt, M., Levine, S
K¨ opf, A., Kilcher, Y., R¨ utte, D., Anagnostidis, S., Tam, Z.R., Stevens, K., Barhoum, A., Nguyen, D., Stanley, O., Nagyfi, R., ES, S., Suri, S., Glushkov, D., Dantuluri, A., Maguire, A., Schuhmann, C., Nguyen, H., Mattick, A.: Openassis- tant conversations - democratizing l...
2023
-
[50]
Sun, W., Yan, L., Ma, X., Wang, S., Ren, P., Chen, Z., Yin, D., Ren, Z.: Is ChatGPT Good at Search? Investigating Large Language Models as Re-Ranking Agents (2023)
2023
-
[51]
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., Klimov, O.: Proximal Policy Optimization Algorithms (2017)
2017
-
[52]
In: Oh, A., Neumann, T., Globerson, A., Saenko, K., Hardt, M., Levine, S
Rafailov, R., Sharma, A., Mitchell, E., Manning, C.D., Ermon, S., Finn, C.: Direct preference optimization: Your language model is secretly a reward model. In: Oh, A., Neumann, T., Globerson, A., Saenko, K., Hardt, M., Levine, S. (eds.) Advances in Neural Information Processin...
2023
-
[53]
Yuan, Z., Yuan, H., Tan, C., Wang, W., Huang, S., Huang, F.: RRHF: Rank Responses to Align Language Models with Human Feedback without tears (2023)
2023
-
[54]
Dong, H., Xiong, W., Goyal, D., Zhang, Y., Chow, W., Pan, R., Diao, S., Zhang, J., Shum, K., Zhang, T.: RAFT: Reward rAnked FineTuning for Generative Foundation Model Alignment (2023)
2023
-
[55]
Schulman, J., Levine, S., Moritz, P., Jordan, M.I., Abbeel, P.: Trust Region Policy Optimization (2017)
2017
-
[56]
In: Proceedings of the 33rd ACM International Conference on Information and Knowledge Management
Chu, Z., Ai, Q., Tu, Y., Li, H., Liu, Y.: Automatic large language model evaluation 22 via peer review. In: Proceedings of the 33rd ACM International Conference on Information and Knowledge Management. CIKM ’24, pp. 384–393. Association for Computing Machinery, New York, NY, U...
2024
-
[57]
In: Proceedings of the 22nd International Conference on Machine Learning, pp
Burges, C., Shaked, T., Renshaw, E., Lazier, A., Deeds, M., Hamilton, N., Hul- lender, G.: Learning to rank using gradient descent. In: Proceedings of the 22nd International Conference on Machine Learning, pp. 89–96 (2005)
2005
-
[58]
In: The 41st International ACM SIGIR Conference on Research & Development in Information Retrieval
Ai, Q., Bi, K., Guo, J., Croft, W.B.: Learning a deep listwise context model for ranking refinement. In: The 41st International ACM SIGIR Conference on Research & Development in Information Retrieval. SIGIR ’18, pp. 135–144. Asso- ciation for Computing Machinery, New York, NY,...
2018
-
[59]
In: Proceedings of the 2019 ACM SIGIR International Conference on Theory of Information Retrieval, pp
Bruch, S., Wang, X., Bendersky, M., Najork, M.: An analysis of the softmax cross entropy loss for learning-to-rank with binary relevance. In: Proceedings of the 2019 ACM SIGIR International Conference on Theory of Information Retrieval, pp. 75–78 (2019)
2019
-
[60]
In: Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp
Liu, Y., Liu, P., Radev, D., Neubig, G.: Brio: Bringing order to abstractive sum- marization. In: Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 2890–2903 (2022)
2022
-
[61]
Springer, ??? (2022)
Chuklin, A., Markov, I., De Rijke, M.: Click Models for Web Search. Springer, ??? (2022)
2022
-
[62]
In: Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining
Zhou, G., Zhu, X., Song, C., Fan, Y., Zhu, H., Ma, X., Yan, Y., Jin, J., Li, H., Gai, K.: Deep interest network for click-through rate prediction. In: Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining. KDD ’18, pp. 1059–1068. Asso...
2018
-
[63]
In: Database Systems for Advanced Applications
Gu, L.: Ad click-through rate prediction: a survey. In: Database Systems for Advanced Applications. DASF AA 2021 International Workshops: BDQM, GDMA, MLDLDSA, MobiSocial, and MUST, Taipei, Taiwan, April 11–14, 2021, Proceedings 26, pp. 140–153 (2021). Springer
2021
-
[64]
CIKM ’08, pp
Dou, Z., Song, R., Yuan, X., Wen, J.-R.: Are click-through data adequate for learning web search rankings? In: Proceedings of the 17th ACM Conference on Information and Knowledge Management. CIKM ’08, pp. 73–82. Association for Computing Machinery, New York, NY, USA (2008). ht...
2008
-
[65]
In: Proceedings of the 15th International Conference on 23 Intelligent User Interfaces
Liu, J., Dolan, P., Pedersen, E.R.: Personalized news recommendation based on click behavior. In: Proceedings of the 15th International Conference on 23 Intelligent User Interfaces. IUI ’10, pp. 31–40. Association for Computing Machin- ery, New York, NY, USA (2010). https://do...
2010
-
[66]
Egyptian informatics journal 16(3), 261–273 (2015)
Isinkaye, F.O., Folajimi, Y.O., Ojokoh, B.A.: Recommendation systems: Prin- ciples, methods and evaluation. Egyptian informatics journal 16(3), 261–273 (2015)
2015
-
[67]
In: Proceedings of the 15th International Conference on World Wide Web, pp
Qiu, F., Cho, J.: Automatic identification of user interest for personalized search. In: Proceedings of the 15th International Conference on World Wide Web, pp. 727–736 (2006)
2006
-
[68]
In: Proceedings of the 27th ACM International Conference on Information and Knowledge Management
Ge, S., Dou, Z., Jiang, Z., Nie, J.-Y., Wen, J.-R.: Personalizing search results using hierarchical rnn with query-aware attention. In: Proceedings of the 27th ACM International Conference on Information and Knowledge Management. CIKM ’18, pp. 347–356. Association for Computin...
2018
-
[69]
https://arxiv.org/abs/2305.19860
Wu, L., Zheng, Z., Qiu, Z., Wang, H., Gu, H., Shen, T., Qin, C., Zhu, C., Zhu, H., Liu, Q., Xiong, H., Chen, E.: A Survey on Large Language Models for Recommendation (2024). https://arxiv.org/abs/2305.19860
2024 arXiv
-
[70]
arXiv preprint arXiv:2402.10548 (2024)
Zhou, Y., Zhu, Q., Jin, J., Dou, Z.: Cognitive personalized search integrat- ing large language models with an efficient memory mechanism. arXiv preprint arXiv:2402.10548 (2024)
2024 arXiv
-
[71]
arXiv preprint arXiv:2308.11432 (2023)
Wang, L., Ma, C., Feng, X., Zhang, Z., Yang, H., Zhang, J., Chen, Z., Tang, J., Chen, X., Lin, Y., et al.: A survey on large language model based autonomous agents. arXiv preprint arXiv:2308.11432 (2023)
2023 arXiv
-
[72]
IEEE Transactions on Pattern Analysis and Machine Intelligence (2024)
Wang, L., Zhang, X., Su, H., Zhu, J.: A comprehensive survey of continual learn- ing: theory, method and application. IEEE Transactions on Pattern Analysis and Machine Intelligence (2024)
2024
-
[73]
In: Proceedings of the 18th International Conference on World Wide Web, pp
Chapelle, O., Zhang, Y.: A dynamic bayesian network click model for web search ranking. In: Proceedings of the 18th International Conference on World Wide Web, pp. 1–10 (2009)
2009
-
[74]
73–82 (2008)
Dou, Z., Song, R., Yuan, X., Wen, J.-R.: Are click-through data adequate for learning web search rankings? In: Proceedings of the 17th ACM Conference on Information and Knowledge Management, pp. 73–82 (2008)
2008
-
[75]
https://arxiv.org/abs/2402.01364 24
Wu, T., Luo, L., Li, Y.-F., Pan, S., Vu, T.-T., Haffari, G.: Continual Learning for Large Language Models: A Survey (2024). https://arxiv.org/abs/2402.01364 24
2024 arXiv
-
[76]
arXiv preprint arXiv:2404.16789 (2024)
Shi, H., Xu, Z., Wang, H., Qin, W., Wang, W., Wang, Y., Wang, H.: Contin- ual learning of large language models: A comprehensive survey. arXiv preprint arXiv:2404.16789 (2024)
2024 arXiv
-
[77]
In: Bouamor, H., Pino, J., Bali, K
Mao, K., Dou, Z., Mo, F., Hou, J., Chen, H., Qian, H.: Large language models know your contextual search intent: A prompting framework for conversational search. In: Bouamor, H., Pino, J., Bali, K. (eds.) Findings of the Associa- tion for Computational Linguistics: EMNLP 2023,...
2023
-
[78]
In: Bouamor, H., Pino, J., Bali, K
Ye, F., Fang, M., Li, S., Yilmaz, E.: Enhancing conversational search: Large lan- guage model-aided informative query rewriting. In: Bouamor, H., Pino, J., Bali, K. (eds.) Findings of the Association for Computational Linguistics: EMNLP 2023, Singapore, December 6-10, 2023, pp...
2023
-
[79]
arXiv preprint arXiv:2402.07092 (2024)
Chen, H., Dou, Z., Mao, K., Liu, J., Zhao, Z.: Generalizing conversational dense retrieval via llm-cognition data augmentation. arXiv preprint arXiv:2402.07092 (2024)
2024 arXiv
-
[80]
https://www.trecikat.com/ (2023)
2023
-
[81]
arXiv preprint arXiv:2309.01157 (2023)
Li, L., Zhang, Y., Liu, D., Chen, L.: Large language models for generative recom- mendation: A survey and visionary discussions. arXiv preprint arXiv:2309.01157 (2023)
2023 arXiv
-
[82]
CoRR abs/2212.10496 (2022)
Gao, L., Ma, X., Lin, J., Callan, J.: Precise zero-shot dense retrieval without relevance labels. CoRR abs/2212.10496 (2022)
2022 arXiv
-
[83]
In: 11th International Conference on Learning Representations, ICLR 2023 (2023)
Yu, W., Iter, D., Wang, S., Xu, Y., Ju, M., Sanyal, S., Zhu, C., Zeng, M., Jiang, M.: Generate rather than retrieve: Large language models are strong context generators. In: 11th International Conference on Learning Representations, ICLR 2023 (2023)
2023
-
[84]
Advances in neural information processing systems (2020) 25
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Chi, E.H., Le, Q., Zhou, D.: Chain of thought prompting elicits reasoning in large language models. Advances in neural information processing systems (2020) 25
2020
-
[194]
https: //doi.org/10.1145/2348283.2348312
Association for Computing Machinery, New York, NY, USA (2012). https: //doi.org/10.1145/2348283.2348312
2012
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.