REVIEW 2 major objections 5 minor 57 references
Reranking optimized by LLM answer quality, not topical relevance, improves RAG without human labels.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-13 13:58 UTC pith:IJEABUGP
load-bearing objection Solid modular RL for aligning pointwise rerankers to LLM answer quality; sequential MDP + reference-greedy baseline works, gains are real but modest. the 2 major comments →
Optimizing RAG Rerankers with LLM Feedback via Reinforcement Learning
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
By casting document selection as a finite-horizon MDP and optimizing a pointwise reranker with PPO-style updates whose reward is the LLM reader's generation score (EM + F1 + Hit), RRPO produces rankings that raise final answer quality more than rankings trained on static relevance labels or listwise LLM prompts, without any human relevance annotations.
What carries the argument
ReRanking Preference Optimization (RRPO): sequential document selection under a policy derived from a pointwise scorer, trained with clipped importance sampling, a KL penalty to a fixed reference, and a reference-anchored deterministic baseline that evaluates the greedy trajectory of the reference policy instead of a learned value network.
Load-bearing premise
The scalar reward built from exact-match, F1 and answer-hit on short greedy LLM outputs is assumed to be a faithful enough proxy for true document utility that the policy gradient improves the right objective.
What would settle it
On the same candidate pools and frozen reader, replace the RRPO-trained ranking with a ranking that maximizes topical relevance (or RankZephyr listwise scores) and show that answer EM/F1 no longer improves, or that the advantage of RRPO disappears when the reward is replaced by pure topical NDCG.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces ReRanking Preference Optimization (RRPO), which reformulates document reranking for RAG as a finite-horizon MDP: a pointwise reranker policy sequentially selects an ordered subset of k documents from an initial BM25 candidate pool, receiving stepwise rewards from a frozen LLM reader’s generation quality (R_lm = EM + λ_f F1 + λ_h Hit). Training uses a PPO-style objective with GAE advantages computed against a reference-anchored deterministic baseline (greedy rollouts of a fixed pretrained reranker) rather than a learned critic, plus KL regularization and advantage normalization for stability. This eliminates human relevance labels. Experiments on HotpotQA, AmbigNQ, 2WikiMultiHopQA and MuSiQue show consistent gains over the base gte/jina/bge/qwen3 rerankers and over RankZephyr; further analyses demonstrate transfer to diverse readers (including GPT-4o, Claude, Gemini), orthogonal gains with Query2Doc, and robustness to a noisy 3B supervisor.
Significance. If the results hold, RRPO supplies a practical, modular, label-free recipe for aligning lightweight (≈0.3 B) pointwise rerankers to downstream generation utility. The sequential MDP + reference-anchored baseline construction is a clean technical contribution that sidesteps both the data bottleneck of supervised ranking and the instability of parametric critics under sparse LLM rewards. Public code, multi-architecture ablations (listwise bandit vs. standard PPO vs. RRPO), reader-transfer experiments, and stacking with query expansion strengthen the claim that the method is more than a one-off hyper-parameter tweak. Absolute gains remain modest, yet the framework is immediately usable in existing RAG pipelines and therefore of clear engineering value to the community.
major comments (2)
- §3.2.1 and Implementation Details (§4.1): the sole supervision signal is the fixed composite R_lm = EM + 1·F1 + 1·Hit evaluated on short greedy answers. No ablation is reported on alternative reward formulations (pure EM, LLM-as-judge preference, or answer-length-normalized scores). Because the policy gradient optimizes exactly this scalar, a sensitivity study is needed to confirm that the learned ranking policy captures genuine context utility rather than idiosyncrasies of the chosen metric; the current evidence (Tables 1–5, Fig. 3) is consistent but not conclusive on this point.
- Table 3 and Appendix J: the comparison with RankZephyr (7 B list-wise) and DynamicRAG (7 B joint) is informative yet asymmetric in model scale and training regime. While the authors correctly note the modularity advantage of the 0.3 B RRPO models, a controlled experiment that applies the same sequential RL objective to a list-wise or larger backbone would more cleanly isolate the contribution of the MDP formulation itself versus the choice of base architecture.
minor comments (5)
- Figure 2 caption and surrounding text: the three k_train curves are hard to distinguish in grayscale; adding markers or line styles would improve readability.
- §4.1: the precise definition of the Hit component of R_lm (1 if answer string appears, –1 otherwise) is given only in the text; placing the full formula next to the EM/F1 definitions would aid reproducibility.
- Appendix C prompt template: the multi-turn “User: passage i” format is clear, yet the system message is repeated for every dataset; a short note on whether temperature or decoding parameters differ between training and evaluation would be helpful.
- Related Work §2.3: the discussion of concurrent DPO-based methods (DynamicRAG, KnowPO, DPA-RAG) is accurate, but a one-sentence clarification that RRPO freezes the reader while those works jointly update generator parameters would sharpen the contrast.
- Typos: “pesudo codes” (Appendix A), “adapts from previous work” (Table 1 caption), and occasional inconsistent capitalization of “reranker” vs. “Reranker”.
Circularity Check
No load-bearing circularity: RRPO optimizes a pointwise policy against an external frozen-reader reward and evaluates on held-out metrics; the reference baseline is a fixed pretrained model, not defined from the learned policy.
full rationale
The paper's derivation is an empirical RL pipeline, not a first-principles claim that reduces to its own inputs. Reranking is cast as a finite-horizon MDP (Eqs. 1-5) whose actions select documents; the reward rt = R_lm(ans, response_t) is produced by a frozen LLM Reader scored against ground-truth answers via the composite EM+F1+Hit metric (Sec. 3.2.1 and Implementation Details). The reference-anchored baseline V(st) is obtained by a deterministic greedy rollout of a fixed pretrained reference policy π_ref (Eqs. 11-13), not of the policy being optimized; GAE and PPO-clip then produce an advantage that is independent of the evaluation tables. Training therefore cannot be tautological: the same datasets supply both rewards and test metrics only in the ordinary sense of supervised/RL evaluation, and the method is further stress-tested by transfer to unseen readers (GPT-4o, Claude, Gemini), orthogonal combination with Query2Doc, and distillation from a noisier 3B supervisor. No uniqueness theorem, self-citation chain, or fitted parameter is invoked to force the central claim. The modest absolute gains and the acknowledged dependence on initial recall (Limitations) are empirical limitations, not circular reductions. Score 1 reflects only the trivial shared-metric observation already noted by the reader; steps remain empty because no reduction by construction exists.
Axiom & Free-Parameter Ledger
free parameters (3)
- λ_f, λ_h (reward weights)
- k_train / k_eval
- PPO/GAE hyperparameters (γ=0.99, λ=0.95, ε=0.2, β=0.1, learning rates)
axioms (4)
- domain assumption LLM generation quality under a fixed prompt and short-answer format is a sufficient scalar reward for document utility.
- domain assumption A pointwise score model plus sequential softmax renormalization can represent useful list policies.
- ad hoc to paper Greedy rollouts of a fixed reference reranker supply a low-bias baseline for GAE under sparse LLM rewards.
- standard math PPO clip + KL to reference keeps policy updates stable for short episodes.
invented entities (1)
-
Reference-anchored deterministic baseline V(s_t)
no independent evidence
read the original abstract
Rerankers play a pivotal role in refining retrieval results for Retrieval-Augmented Generation. However, current reranking models are typically optimized on static human annotated relevance labels in isolation, decoupled from the downstream generation process. This isolation leads to a fundamental misalignment: documents identified as topically relevant by information retrieval metrics often fail to provide the actual utility required by the LLM for precise answer generation. To bridge this gap, we introduce ReRanking Preference Optimization (RRPO), a reinforcement learning framework that directly aligns reranking with the LLM's generation quality. By formulating reranking as a sequential decision-making process, RRPO optimizes for context utility using LLM feedback, thereby eliminating the need for expensive human annotations. To ensure training stability, we further introduce a reference-anchored deterministic baseline. Extensive experiments on knowledge-intensive benchmarks demonstrate that RRPO significantly outperforms strong baselines, including the powerful list-wise reranker RankZephyr. Further analysis highlights the versatility of our framework: it generalizes seamlessly to diverse readers (e.g., GPT-4o), integrates orthogonally with query expansion modules like Query2Doc, and remains robust even when trained with noisy supervisors.
Reference graph
Works this paper leans on
-
[1]
Akari Asai, Zeqiu Wu, Yizhong Wang, Avirup Sil, and Hannaneh Hajishirzi. 2024. Self-rag: Learning to retrieve, generate, and critique through self-reflection
2024
-
[2]
Sebastian Borgeaud, Arthur Mensch, Jordan Hoffmann, Trevor Cai, Eliza Rutherford, Katie Millican, George Bm Van Den Driessche, Jean-Baptiste Lespiau, Bogdan Damoc, Aidan Clark, and 1 others. 2022. Improving language models by retrieving from trillions of tokens. In International conference on machine learning, pages 2206--2240. PMLR
2022
-
[3]
Shijie Chen, Bernal Jim \'e nez Guti \'e rrez, and Yu Su. 2024. Attention in large language models yields efficient zero-shot re-rankers. arXiv preprint arXiv:2410.02642
Pith/arXiv arXiv 2024
-
[4]
Yiqun Chen, Lingyong Yan, Weiwei Sun, Xinyu Ma, Yi Zhang, Shuaiqiang Wang, Dawei Yin, Yiming Yang, and Jiaxin Mao. 2025. Improving retrieval-augmented generation through multi-agent reinforcement learning. arXiv preprint arXiv:2501.15228
arXiv 2025
-
[5]
Sukmin Cho, Soyeong Jeong, Jeong yeon Seo, and Jong C Park. 2023. Discrete prompt optimization via constrained generation for zero-shot re-ranker. In Findings of the Association for Computational Linguistics: ACL 2023, pages 960--971
2023
-
[6]
Guanting Dong, Yutao Zhu, Chenghao Zhang, Zechen Wang, Ji-Rong Wen, and Zhicheng Dou. 2025. Understand what llm needs: Dual preference alignment for retrieval-augmented generation. In Proceedings of the ACM on Web Conference 2025, pages 4206--4225
2025
-
[7]
Yunfan Gao, Yun Xiong, Xinyu Gao, Kangxiang Jia, Jinliu Pan, Yuxi Bi, Yixin Dai, Jiawei Sun, Haofen Wang, and Haofen Wang. 2023. Retrieval-augmented generation for large language models: A survey. arXiv preprint arXiv:2312.10997, 2:1
Pith/arXiv arXiv 2023
-
[8]
Xinyan Guan, Jiali Zeng, Fandong Meng, Chunlei Xin, Yaojie Lu, Hongyu Lin, Xianpei Han, Le Sun, and Jie Zhou. 2025. Deeprag: Thinking to retrieval step by step for large language models. arXiv preprint arXiv:2502.01142
Pith/arXiv arXiv 2025
-
[9]
Kelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat, and Mingwei Chang. 2020. Retrieval augmented language model pre-training. In International conference on machine learning, pages 3929--3938. PMLR
2020
-
[10]
Xanh Ho, Anh-Khoa Duong Nguyen, Saku Sugawara, and Akiko Aizawa. 2020. Constructing a multi-hop qa dataset for comprehensive evaluation of reasoning steps. arXiv preprint arXiv:2011.01060
Pith/arXiv arXiv 2020
-
[11]
Gautier Izacard and Edouard Grave. 2020. Leveraging passage retrieval with generative models for open domain question answering. arXiv preprint arXiv:2007.01282
Pith/arXiv arXiv 2020
-
[12]
Gautier Izacard, Patrick Lewis, Maria Lomeli, Lucas Hosseini, Fabio Petroni, Timo Schick, Jane Dwivedi-Yu, Armand Joulin, Sebastian Riedel, and Edouard Grave. 2023. Atlas: Few-shot learning with retrieval augmented language models. Journal of Machine Learning Research, 24(251):1--43
2023
-
[13]
Soyeong Jeong, Jinheon Baek, Sukmin Cho, Sung Ju Hwang, and Jong C Park. 2024. Adaptive-rag: Learning to adapt retrieval-augmented large language models through question complexity. arXiv preprint arXiv:2403.14403
Pith/arXiv arXiv 2024
-
[14]
Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Ye Jin Bang, Andrea Madotto, and Pascale Fung. 2023. Survey of hallucination in natural language generation. ACM computing surveys, 55(12):1--38
2023
-
[15]
Zhengbao Jiang, Frank F Xu, Luyu Gao, Zhiqing Sun, Qian Liu, Jane Dwivedi-Yu, Yiming Yang, Jamie Callan, and Graham Neubig. 2023. Active retrieval augmented generation. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 7969--7992
2023
-
[16]
Jina AI . 2024 . jina-reranker-v2-base-multilingual . https://jina.ai/models/jina-reranker-v2-base-multilingual. Accessed: 2024-06-25
2024
-
[17]
Vladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih. 2020. https://doi.org/10.18653/v1/2020.emnlp-main.550 Dense passage retrieval for open-domain question answering . In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 6769--6781, Online. Ass...
-
[18]
Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, and 1 others. 2019. Natural questions: a benchmark for question answering research. Transactions of the Association for Computational Linguistics, 7:453--466
2019
-
[19]
Gonzalez, Hao Zhang, and Ion Stoica
Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph E. Gonzalez, Hao Zhang, and Ion Stoica. 2023. Efficient memory management for large language model serving with pagedattention. In Proceedings of the ACM SIGOPS 29th Symposium on Operating Systems Principles
2023
-
[20]
Dahyun Lee, Yongrae Jo, Haeju Park, and Moontae Lee. 2025. https://doi.org/10.18653/v1/2025.acl-long.861 Shifting from ranking to set selection for retrieval augmented generation . In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 17606--17619, Vienna, Austria. Association for Computa...
-
[21]
u ttler, Mike Lewis, Wen-tau Yih, Tim Rockt\
Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich K\" u ttler, Mike Lewis, Wen-tau Yih, Tim Rockt\" a schel, Sebastian Riedel, and Douwe Kiela. 2020. Retrieval-augmented generation for knowledge-intensive nlp tasks. In Proceedings of the 34th International Conference on Neural Information Processing Sy...
2020
-
[22]
Xi Victoria Lin, Xilun Chen, Mingda Chen, Weijia Shi, Maria Lomeli, Richard James, Pedro Rodriguez, Jacob Kahn, Gergely Szilvasy, Mike Lewis, and 1 others. 2023. Ra-dit: Retrieval-augmented dual instruction tuning. In The Twelfth International Conference on Learning Representations
2023
-
[23]
Zihan Liu, Wei Ping, Rajarshi Roy, Peng Xu, Chankyu Lee, Mohammad Shoeybi, and Bryan Catanzaro. 2024. Chatqa: Surpassing gpt-4 on conversational qa and rag. Advances in Neural Information Processing Systems, 37:15416--15459
2024
-
[24]
Xinbei Ma, Yeyun Gong, Pengcheng He, Hai Zhao, and Nan Duan. 2023. https://doi.org/10.18653/v1/2023.emnlp-main.322 Query rewriting in retrieval-augmented large language models . In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 5303--5315, Singapore. Association for Computational Linguistics
-
[25]
Xueguang Ma, Liang Wang, Nan Yang, Furu Wei, and Jimmy Lin. 2024. Fine-tuning llama for multi-stage text retrieval. In Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval, pages 2421--2425
2024
-
[26]
Sewon Min, Julian Michael, Hannaneh Hajishirzi, and Luke Zettlemoyer. 2020. Ambigqa: Answering ambiguous open-domain questions. arXiv preprint arXiv:2004.10645
Pith/arXiv arXiv 2020
-
[27]
Niklas Muennighoff, SU Hongjin, Liang Wang, Nan Yang, Furu Wei, Tao Yu, Amanpreet Singh, and Douwe Kiela. 2024. Generative representational instruction tuning. In ICLR 2024 Workshop: How Far Are We From AGI
2024
-
[28]
Reiichiro Nakano, Jacob Hilton, Suchir Balaji, Jeff Wu, Long Ouyang, Christina Kim, Christopher Hesse, Shantanu Jain, Vineet Kosaraju, William Saunders, and 1 others. 2021. Webgpt: Browser-assisted question-answering with human feedback. arXiv preprint arXiv:2112.09332
Pith/arXiv arXiv 2021
-
[29]
Rodrigo Nogueira and Kyunghyun Cho. 2019. Passage re-ranking with bert. arXiv preprint arXiv:1901.04085
Pith/arXiv arXiv 2019
-
[30]
Rodrigo Nogueira, Wei Yang, Kyunghyun Cho, and Jimmy Lin. 2019. Multi-stage document ranking with bert. arXiv preprint arXiv:1910.14424
Pith/arXiv arXiv 2019
-
[31]
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, and 1 others. 2022. Training language models to follow instructions with human feedback. Advances in neural information processing systems, 35:27730--27744
2022
-
[32]
Ronak Pradeep, Sahel Sharifymoghaddam, and Jimmy Lin. 2023. Rankzephyr: Effective and robust zero-shot listwise reranking is a breeze! arXiv preprint arXiv:2312.02724
Pith/arXiv arXiv 2023
-
[33]
Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn. 2023. Direct preference optimization: Your language model is secretly a reward model. Advances in neural information processing systems, 36:53728--53741
2023
-
[34]
Revanth Gangi Reddy, JaeHyeok Doo, Yifei Xu, Md Arafat Sultan, Deevya Swain, Avirup Sil, and Heng Ji. 2024. First: Faster improved listwise reranking with single token decoding. arXiv preprint arXiv:2406.15657
Pith/arXiv arXiv 2024
-
[35]
Stephen Robertson and Hugo Zaragoza. 2009. https://doi.org/10.1561/1500000019 The probabilistic relevance framework: Bm25 and beyond . Foundations and Trends® in Information Retrieval, 3(4):333--389
-
[36]
John Schulman, Philipp Moritz, Sergey Levine, Michael Jordan, and Pieter Abbeel. 2015. High-dimensional continuous control using generalized advantage estimation. arXiv preprint arXiv:1506.02438
Pith/arXiv arXiv 2015
-
[37]
Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Junxiao Song, Xiao Bi, Haowei Zhang, Mingchuan Zhang, YK Li, Yang Wu, and 1 others. 2024. Deepseekmath: Pushing the limits of mathematical reasoning in open language models. arXiv preprint arXiv:2402.03300
Pith/arXiv arXiv 2024
-
[38]
Kurt Shuster, Spencer Poff, Moya Chen, Douwe Kiela, and Jason Weston. 2021. Retrieval augmentation reduces hallucination in conversation. arXiv preprint arXiv:2104.07567
Pith/arXiv arXiv 2021
-
[39]
Weihang Su, Yichen Tang, Qingyao Ai, Zhijing Wu, and Yiqun Liu. 2024. https://doi.org/10.18653/v1/2024.acl-long.702 DRAGIN : Dynamic retrieval augmented generation based on the real-time information needs of large language models . In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 129...
-
[40]
Jiashuo Sun, Xianrui Zhong, Sizhe Zhou, and Jiawei Han. 2025. https://arxiv.org/abs/2505.07233 Dynamicrag: Leveraging outputs of large language model as feedback for dynamic reranking in retrieval-augmented generation . Preprint, arXiv:2505.07233
Pith/arXiv arXiv 2025
-
[41]
Weiwei Sun, Lingyong Yan, Xinyu Ma, Shuaiqiang Wang, Pengjie Ren, Zhumin Chen, Dawei Yin, and Zhaochun Ren. 2023. Is chatgpt good at search? investigating large language models as re-ranking agents. arXiv preprint arXiv:2304.09542
Pith/arXiv arXiv 2023
-
[42]
Qwen Team. 2024. Qwen2 technical report. arXiv preprint arXiv:2407.10671, 2
Pith/arXiv arXiv 2024
-
[43]
Harsh Trivedi, Niranjan Balasubramanian, Tushar Khot, and Ashish Sabharwal. 2022 a . Interleaving retrieval with chain-of-thought reasoning for knowledge-intensive multi-step questions. arXiv preprint arXiv:2212.10509
Pith/arXiv arXiv 2022
-
[44]
Harsh Trivedi, Niranjan Balasubramanian, Tushar Khot, and Ashish Sabharwal. 2022 b . https://doi.org/10.1162/TACL\_A\_00475 Musique: Multihop questions via single-hop question composition . Trans. Assoc. Comput. Linguistics, 10:539--554
doi:10.1162/tacl 2022
-
[45]
Liang Wang, Nan Yang, and Furu Wei. 2023 a . https://doi.org/10.18653/v1/2023.emnlp-main.585 Query2doc: Query expansion with large language models . In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 9414--9423, Singapore. Association for Computational Linguistics
-
[46]
Xinyuan Wang, Chenxi Li, Zhen Wang, Fan Bai, Haotian Luo, Jiayou Zhang, Nebojsa Jojic, Eric P Xing, and Zhiting Hu. 2023 b . Promptagent: Strategic planning with language models enables expert-level prompt optimization. arXiv preprint arXiv:2310.16427
Pith/arXiv arXiv 2023
-
[47]
Shitao Xiao, Zheng Liu, Peitian Zhang, and Niklas Muennighoff. 2023. https://arxiv.org/abs/2309.07597 C-pack: Packaged resources to advance general chinese embedding . Preprint, arXiv:2309.07597
Pith/arXiv arXiv 2023
-
[48]
Fangyuan Xu, Weijia Shi, and Eunsol Choi. 2024. Recomp: Improving retrieval-augmented lms with context compression and selective augmentation. In The Twelfth International Conference on Learning Representations
2024
-
[49]
Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William W Cohen, Ruslan Salakhutdinov, and Christopher D Manning. 2018. Hotpotqa: A dataset for diverse, explainable multi-hop question answering. arXiv preprint arXiv:1809.09600
Pith/arXiv arXiv 2018
-
[50]
Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao. 2023. React: Synergizing reasoning and acting in language models. In International Conference on Learning Representations (ICLR)
2023
-
[51]
Xiaowei Yuan, Zhao Yang, Yequan Wang, Jun Zhao, and Kang Liu. 2024. Improving zero-shot llm re-ranker with risk minimization. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 17967--17983
2024
-
[52]
Ye Yuan, Mohammad Amin Shabani, and Siqi Liu. 2025. Embedding-based context-aware reranker. arXiv preprint arXiv:2510.13329
arXiv 2025
-
[53]
Le Zhang, Bo Wang, Xipeng Qiu, Siva Reddy, and Aishwarya Agrawal. 2025 a . Rearank: Reasoning re-ranking agent via reinforcement learning. arXiv preprint arXiv:2505.20046
Pith/arXiv arXiv 2025
-
[54]
Ruizhe Zhang, Yongxin Xu, Yuzhen Xiao, Runchuan Zhu, Xinke Jiang, Xu Chu, Junfeng Zhao, and Yasha Wang. 2025 b . https://doi.org/10.1609/aaai.v39i24.34783 Knowpo: Knowledge-aware preference optimization for controllable knowledge selection in retrieval-augmented language models . Proceedings of the AAAI Conference on Artificial Intelligence, 39(24):25895--25903
-
[55]
Xin Zhang, Yanzhao Zhang, Dingkun Long, Wen Xie, Ziqi Dai, Jialong Tang, Huan Lin, Baosong Yang, Pengjun Xie, Fei Huang, and 1 others. 2024. mgte: Generalized long-context text representation and reranking models for multilingual text retrieval. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing: Industry Track, page...
2024
-
[56]
Yanzhao Zhang, Mingxin Li, Dingkun Long, Xin Zhang, Huan Lin, Baosong Yang, Pengjun Xie, An Yang, Dayiheng Liu, Junyang Lin, Fei Huang, and Jingren Zhou. 2025 c . Qwen3 embedding: Advancing text embedding and reranking through foundation models. arXiv preprint arXiv:2506.05176
Pith/arXiv arXiv 2025
-
[57]
Yutao Zhu, Peitian Zhang, Chenghao Zhang, Yifei Chen, Binyu Xie, Zheng Liu, Ji-Rong Wen, and Zhicheng Dou. 2024. Inters: unlocking the power of large language models in search with instruction tuning. arXiv preprint arXiv:2401.06532
Pith/arXiv arXiv 2024
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.