Pith. sign in

REVIEW 2 major objections 5 minor 57 references

Reranking optimized by LLM answer quality, not topical relevance, improves RAG without human labels.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-13 13:58 UTC pith:IJEABUGP

load-bearing objection Solid modular RL for aligning pointwise rerankers to LLM answer quality; sequential MDP + reference-greedy baseline works, gains are real but modest. the 2 major comments →

arxiv 2604.02091 v2 pith:IJEABUGP submitted 2026-04-02 cs.CL cs.AIcs.IR

Optimizing RAG Rerankers with LLM Feedback via Reinforcement Learning

classification cs.CL cs.AIcs.IR
keywords Retrieval-Augmented Generationrerankingreinforcement learningLLM feedbackpreference optimizationcontext utilityRAG alignment
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Standard RAG rerankers are trained on human relevance labels that ignore the downstream language model. Documents that look topically on-topic often do not help the model produce a correct answer. This paper claims the fix is to treat reranking itself as a short sequential decision process and train the reranker with reinforcement learning whose reward is the quality of the answer the frozen LLM actually generates. A reference-anchored deterministic baseline replaces an unstable learned critic, so training stays stable. The resulting policy, RRPO, beats strong supervised and listwise baselines on multi-hop and open-domain QA, transfers to unseen readers including closed models, and still works when the feedback comes from a smaller, noisier supervisor.

Core claim

By casting document selection as a finite-horizon MDP and optimizing a pointwise reranker with PPO-style updates whose reward is the LLM reader's generation score (EM + F1 + Hit), RRPO produces rankings that raise final answer quality more than rankings trained on static relevance labels or listwise LLM prompts, without any human relevance annotations.

What carries the argument

ReRanking Preference Optimization (RRPO): sequential document selection under a policy derived from a pointwise scorer, trained with clipped importance sampling, a KL penalty to a fixed reference, and a reference-anchored deterministic baseline that evaluates the greedy trajectory of the reference policy instead of a learned value network.

Load-bearing premise

The scalar reward built from exact-match, F1 and answer-hit on short greedy LLM outputs is assumed to be a faithful enough proxy for true document utility that the policy gradient improves the right objective.

What would settle it

On the same candidate pools and frozen reader, replace the RRPO-trained ranking with a ranking that maximizes topical relevance (or RankZephyr listwise scores) and show that answer EM/F1 no longer improves, or that the advantage of RRPO disappears when the reward is replaced by pure topical NDCG.

Watch this falsifier — get emailed when new claim-graph text bears on it.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 5 minor

Summary. The paper introduces ReRanking Preference Optimization (RRPO), which reformulates document reranking for RAG as a finite-horizon MDP: a pointwise reranker policy sequentially selects an ordered subset of k documents from an initial BM25 candidate pool, receiving stepwise rewards from a frozen LLM reader’s generation quality (R_lm = EM + λ_f F1 + λ_h Hit). Training uses a PPO-style objective with GAE advantages computed against a reference-anchored deterministic baseline (greedy rollouts of a fixed pretrained reranker) rather than a learned critic, plus KL regularization and advantage normalization for stability. This eliminates human relevance labels. Experiments on HotpotQA, AmbigNQ, 2WikiMultiHopQA and MuSiQue show consistent gains over the base gte/jina/bge/qwen3 rerankers and over RankZephyr; further analyses demonstrate transfer to diverse readers (including GPT-4o, Claude, Gemini), orthogonal gains with Query2Doc, and robustness to a noisy 3B supervisor.

Significance. If the results hold, RRPO supplies a practical, modular, label-free recipe for aligning lightweight (≈0.3 B) pointwise rerankers to downstream generation utility. The sequential MDP + reference-anchored baseline construction is a clean technical contribution that sidesteps both the data bottleneck of supervised ranking and the instability of parametric critics under sparse LLM rewards. Public code, multi-architecture ablations (listwise bandit vs. standard PPO vs. RRPO), reader-transfer experiments, and stacking with query expansion strengthen the claim that the method is more than a one-off hyper-parameter tweak. Absolute gains remain modest, yet the framework is immediately usable in existing RAG pipelines and therefore of clear engineering value to the community.

major comments (2)
  1. §3.2.1 and Implementation Details (§4.1): the sole supervision signal is the fixed composite R_lm = EM + 1·F1 + 1·Hit evaluated on short greedy answers. No ablation is reported on alternative reward formulations (pure EM, LLM-as-judge preference, or answer-length-normalized scores). Because the policy gradient optimizes exactly this scalar, a sensitivity study is needed to confirm that the learned ranking policy captures genuine context utility rather than idiosyncrasies of the chosen metric; the current evidence (Tables 1–5, Fig. 3) is consistent but not conclusive on this point.
  2. Table 3 and Appendix J: the comparison with RankZephyr (7 B list-wise) and DynamicRAG (7 B joint) is informative yet asymmetric in model scale and training regime. While the authors correctly note the modularity advantage of the 0.3 B RRPO models, a controlled experiment that applies the same sequential RL objective to a list-wise or larger backbone would more cleanly isolate the contribution of the MDP formulation itself versus the choice of base architecture.
minor comments (5)
  1. Figure 2 caption and surrounding text: the three k_train curves are hard to distinguish in grayscale; adding markers or line styles would improve readability.
  2. §4.1: the precise definition of the Hit component of R_lm (1 if answer string appears, –1 otherwise) is given only in the text; placing the full formula next to the EM/F1 definitions would aid reproducibility.
  3. Appendix C prompt template: the multi-turn “User: passage i” format is clear, yet the system message is repeated for every dataset; a short note on whether temperature or decoding parameters differ between training and evaluation would be helpful.
  4. Related Work §2.3: the discussion of concurrent DPO-based methods (DynamicRAG, KnowPO, DPA-RAG) is accurate, but a one-sentence clarification that RRPO freezes the reader while those works jointly update generator parameters would sharpen the contrast.
  5. Typos: “pesudo codes” (Appendix A), “adapts from previous work” (Table 1 caption), and occasional inconsistent capitalization of “reranker” vs. “Reranker”.

Circularity Check

0 steps flagged

No load-bearing circularity: RRPO optimizes a pointwise policy against an external frozen-reader reward and evaluates on held-out metrics; the reference baseline is a fixed pretrained model, not defined from the learned policy.

full rationale

The paper's derivation is an empirical RL pipeline, not a first-principles claim that reduces to its own inputs. Reranking is cast as a finite-horizon MDP (Eqs. 1-5) whose actions select documents; the reward rt = R_lm(ans, response_t) is produced by a frozen LLM Reader scored against ground-truth answers via the composite EM+F1+Hit metric (Sec. 3.2.1 and Implementation Details). The reference-anchored baseline V(st) is obtained by a deterministic greedy rollout of a fixed pretrained reference policy π_ref (Eqs. 11-13), not of the policy being optimized; GAE and PPO-clip then produce an advantage that is independent of the evaluation tables. Training therefore cannot be tautological: the same datasets supply both rewards and test metrics only in the ordinary sense of supervised/RL evaluation, and the method is further stress-tested by transfer to unseen readers (GPT-4o, Claude, Gemini), orthogonal combination with Query2Doc, and distillation from a noisier 3B supervisor. No uniqueness theorem, self-citation chain, or fitted parameter is invoked to force the central claim. The modest absolute gains and the acknowledged dependence on initial recall (Limitations) are empirical limitations, not circular reductions. Score 1 reflects only the trivial shared-metric observation already noted by the reader; steps remain empty because no reduction by construction exists.

Axiom & Free-Parameter Ledger

3 free parameters · 4 axioms · 1 invented entities

The central claim rests on standard RL machinery plus a small set of design choices (composite reward weights, short-horizon sequential selection, deterministic reference baseline) that are not derived from first principles but chosen for stability and convenience. No new physical or mathematical entities are postulated; free parameters are ordinary RL and reward hyperparameters.

free parameters (3)
  • λ_f, λ_h (reward weights)
    Set by hand to 1.0 in R_lm = EM + λ_f F1 + λ_h Hit; directly scale the training signal.
  • k_train / k_eval
    Number of documents selected per episode (3 for HotpotQA, 5 for AmbigNQ); chosen by ablation, not derived.
  • PPO/GAE hyperparameters (γ=0.99, λ=0.95, ε=0.2, β=0.1, learning rates)
    Standard RL knobs adjusted for the short-horizon discrete-reward setting; affect stability and final policy.
axioms (4)
  • domain assumption LLM generation quality under a fixed prompt and short-answer format is a sufficient scalar reward for document utility.
    Invoked throughout Section 3 when defining r_t = R_lm(ans, response_t); never proved, only validated empirically.
  • domain assumption A pointwise score model plus sequential softmax renormalization can represent useful list policies.
    Policy definition in Eqs. (3)–(5); assumes combinatorial set utility can be optimized without an explicit listwise scorer.
  • ad hoc to paper Greedy rollouts of a fixed reference reranker supply a low-bias baseline for GAE under sparse LLM rewards.
    Core of the reference-anchored baseline (Section 3.2.3); replaces a learned critic by construction.
  • standard math PPO clip + KL to reference keeps policy updates stable for short episodes.
    Standard RLHF practice cited from Ouyang et al. and related work; used in Eq. (8).
invented entities (1)
  • Reference-anchored deterministic baseline V(s_t) no independent evidence
    purpose: Provide a critic-free value estimate by evaluating greedy reference trajectories with the same LLM reward.
    Defined in Eqs. (11)–(13); new algorithmic construct local to this paper, not an external physical entity.

pith-pipeline@v1.1.0-grok45 · 23494 in / 2778 out tokens · 20852 ms · 2026-07-13T13:58:11.961699+00:00 · methodology

0 comments
read the original abstract

Rerankers play a pivotal role in refining retrieval results for Retrieval-Augmented Generation. However, current reranking models are typically optimized on static human annotated relevance labels in isolation, decoupled from the downstream generation process. This isolation leads to a fundamental misalignment: documents identified as topically relevant by information retrieval metrics often fail to provide the actual utility required by the LLM for precise answer generation. To bridge this gap, we introduce ReRanking Preference Optimization (RRPO), a reinforcement learning framework that directly aligns reranking with the LLM's generation quality. By formulating reranking as a sequential decision-making process, RRPO optimizes for context utility using LLM feedback, thereby eliminating the need for expensive human annotations. To ensure training stability, we further introduce a reference-anchored deterministic baseline. Extensive experiments on knowledge-intensive benchmarks demonstrate that RRPO significantly outperforms strong baselines, including the powerful list-wise reranker RankZephyr. Further analysis highlights the versatility of our framework: it generalizes seamlessly to diverse readers (e.g., GPT-4o), integrates orthogonally with query expansion modules like Query2Doc, and remains robust even when trained with noisy supervisors.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

57 extracted references · 2 canonical work pages

  1. [1]

    Akari Asai, Zeqiu Wu, Yizhong Wang, Avirup Sil, and Hannaneh Hajishirzi. 2024. Self-rag: Learning to retrieve, generate, and critique through self-reflection

  2. [2]

    Sebastian Borgeaud, Arthur Mensch, Jordan Hoffmann, Trevor Cai, Eliza Rutherford, Katie Millican, George Bm Van Den Driessche, Jean-Baptiste Lespiau, Bogdan Damoc, Aidan Clark, and 1 others. 2022. Improving language models by retrieving from trillions of tokens. In International conference on machine learning, pages 2206--2240. PMLR

  3. [3]

    Shijie Chen, Bernal Jim \'e nez Guti \'e rrez, and Yu Su. 2024. Attention in large language models yields efficient zero-shot re-rankers. arXiv preprint arXiv:2410.02642

  4. [4]

    Yiqun Chen, Lingyong Yan, Weiwei Sun, Xinyu Ma, Yi Zhang, Shuaiqiang Wang, Dawei Yin, Yiming Yang, and Jiaxin Mao. 2025. Improving retrieval-augmented generation through multi-agent reinforcement learning. arXiv preprint arXiv:2501.15228

  5. [5]

    Sukmin Cho, Soyeong Jeong, Jeong yeon Seo, and Jong C Park. 2023. Discrete prompt optimization via constrained generation for zero-shot re-ranker. In Findings of the Association for Computational Linguistics: ACL 2023, pages 960--971

  6. [6]

    Guanting Dong, Yutao Zhu, Chenghao Zhang, Zechen Wang, Ji-Rong Wen, and Zhicheng Dou. 2025. Understand what llm needs: Dual preference alignment for retrieval-augmented generation. In Proceedings of the ACM on Web Conference 2025, pages 4206--4225

  7. [7]

    Yunfan Gao, Yun Xiong, Xinyu Gao, Kangxiang Jia, Jinliu Pan, Yuxi Bi, Yixin Dai, Jiawei Sun, Haofen Wang, and Haofen Wang. 2023. Retrieval-augmented generation for large language models: A survey. arXiv preprint arXiv:2312.10997, 2:1

  8. [8]

    Xinyan Guan, Jiali Zeng, Fandong Meng, Chunlei Xin, Yaojie Lu, Hongyu Lin, Xianpei Han, Le Sun, and Jie Zhou. 2025. Deeprag: Thinking to retrieval step by step for large language models. arXiv preprint arXiv:2502.01142

  9. [9]

    Kelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat, and Mingwei Chang. 2020. Retrieval augmented language model pre-training. In International conference on machine learning, pages 3929--3938. PMLR

  10. [10]

    Xanh Ho, Anh-Khoa Duong Nguyen, Saku Sugawara, and Akiko Aizawa. 2020. Constructing a multi-hop qa dataset for comprehensive evaluation of reasoning steps. arXiv preprint arXiv:2011.01060

  11. [11]

    Gautier Izacard and Edouard Grave. 2020. Leveraging passage retrieval with generative models for open domain question answering. arXiv preprint arXiv:2007.01282

  12. [12]

    Gautier Izacard, Patrick Lewis, Maria Lomeli, Lucas Hosseini, Fabio Petroni, Timo Schick, Jane Dwivedi-Yu, Armand Joulin, Sebastian Riedel, and Edouard Grave. 2023. Atlas: Few-shot learning with retrieval augmented language models. Journal of Machine Learning Research, 24(251):1--43

  13. [13]

    Soyeong Jeong, Jinheon Baek, Sukmin Cho, Sung Ju Hwang, and Jong C Park. 2024. Adaptive-rag: Learning to adapt retrieval-augmented large language models through question complexity. arXiv preprint arXiv:2403.14403

  14. [14]

    Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Ye Jin Bang, Andrea Madotto, and Pascale Fung. 2023. Survey of hallucination in natural language generation. ACM computing surveys, 55(12):1--38

  15. [15]

    Zhengbao Jiang, Frank F Xu, Luyu Gao, Zhiqing Sun, Qian Liu, Jane Dwivedi-Yu, Yiming Yang, Jamie Callan, and Graham Neubig. 2023. Active retrieval augmented generation. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 7969--7992

  16. [16]

    Jina AI . 2024 . jina-reranker-v2-base-multilingual . https://jina.ai/models/jina-reranker-v2-base-multilingual. Accessed: 2024-06-25

  17. [17]

    Vladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih. 2020. https://doi.org/10.18653/v1/2020.emnlp-main.550 Dense passage retrieval for open-domain question answering . In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 6769--6781, Online. Ass...

  18. [18]

    Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, and 1 others. 2019. Natural questions: a benchmark for question answering research. Transactions of the Association for Computational Linguistics, 7:453--466

  19. [19]

    Gonzalez, Hao Zhang, and Ion Stoica

    Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph E. Gonzalez, Hao Zhang, and Ion Stoica. 2023. Efficient memory management for large language model serving with pagedattention. In Proceedings of the ACM SIGOPS 29th Symposium on Operating Systems Principles

  20. [20]

    Dahyun Lee, Yongrae Jo, Haeju Park, and Moontae Lee. 2025. https://doi.org/10.18653/v1/2025.acl-long.861 Shifting from ranking to set selection for retrieval augmented generation . In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 17606--17619, Vienna, Austria. Association for Computa...

  21. [21]

    u ttler, Mike Lewis, Wen-tau Yih, Tim Rockt\

    Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich K\" u ttler, Mike Lewis, Wen-tau Yih, Tim Rockt\" a schel, Sebastian Riedel, and Douwe Kiela. 2020. Retrieval-augmented generation for knowledge-intensive nlp tasks. In Proceedings of the 34th International Conference on Neural Information Processing Sy...

  22. [22]

    Xi Victoria Lin, Xilun Chen, Mingda Chen, Weijia Shi, Maria Lomeli, Richard James, Pedro Rodriguez, Jacob Kahn, Gergely Szilvasy, Mike Lewis, and 1 others. 2023. Ra-dit: Retrieval-augmented dual instruction tuning. In The Twelfth International Conference on Learning Representations

  23. [23]

    Zihan Liu, Wei Ping, Rajarshi Roy, Peng Xu, Chankyu Lee, Mohammad Shoeybi, and Bryan Catanzaro. 2024. Chatqa: Surpassing gpt-4 on conversational qa and rag. Advances in Neural Information Processing Systems, 37:15416--15459

  24. [24]

    Xinbei Ma, Yeyun Gong, Pengcheng He, Hai Zhao, and Nan Duan. 2023. https://doi.org/10.18653/v1/2023.emnlp-main.322 Query rewriting in retrieval-augmented large language models . In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 5303--5315, Singapore. Association for Computational Linguistics

  25. [25]

    Xueguang Ma, Liang Wang, Nan Yang, Furu Wei, and Jimmy Lin. 2024. Fine-tuning llama for multi-stage text retrieval. In Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval, pages 2421--2425

  26. [26]

    Sewon Min, Julian Michael, Hannaneh Hajishirzi, and Luke Zettlemoyer. 2020. Ambigqa: Answering ambiguous open-domain questions. arXiv preprint arXiv:2004.10645

  27. [27]

    Niklas Muennighoff, SU Hongjin, Liang Wang, Nan Yang, Furu Wei, Tao Yu, Amanpreet Singh, and Douwe Kiela. 2024. Generative representational instruction tuning. In ICLR 2024 Workshop: How Far Are We From AGI

  28. [28]

    Reiichiro Nakano, Jacob Hilton, Suchir Balaji, Jeff Wu, Long Ouyang, Christina Kim, Christopher Hesse, Shantanu Jain, Vineet Kosaraju, William Saunders, and 1 others. 2021. Webgpt: Browser-assisted question-answering with human feedback. arXiv preprint arXiv:2112.09332

  29. [29]

    Rodrigo Nogueira and Kyunghyun Cho. 2019. Passage re-ranking with bert. arXiv preprint arXiv:1901.04085

  30. [30]

    Rodrigo Nogueira, Wei Yang, Kyunghyun Cho, and Jimmy Lin. 2019. Multi-stage document ranking with bert. arXiv preprint arXiv:1910.14424

  31. [31]

    Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, and 1 others. 2022. Training language models to follow instructions with human feedback. Advances in neural information processing systems, 35:27730--27744

  32. [32]

    Ronak Pradeep, Sahel Sharifymoghaddam, and Jimmy Lin. 2023. Rankzephyr: Effective and robust zero-shot listwise reranking is a breeze! arXiv preprint arXiv:2312.02724

  33. [33]

    Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn. 2023. Direct preference optimization: Your language model is secretly a reward model. Advances in neural information processing systems, 36:53728--53741

  34. [34]

    Revanth Gangi Reddy, JaeHyeok Doo, Yifei Xu, Md Arafat Sultan, Deevya Swain, Avirup Sil, and Heng Ji. 2024. First: Faster improved listwise reranking with single token decoding. arXiv preprint arXiv:2406.15657

  35. [35]

    Stephen Robertson and Hugo Zaragoza. 2009. https://doi.org/10.1561/1500000019 The probabilistic relevance framework: Bm25 and beyond . Foundations and Trends® in Information Retrieval, 3(4):333--389

  36. [36]

    John Schulman, Philipp Moritz, Sergey Levine, Michael Jordan, and Pieter Abbeel. 2015. High-dimensional continuous control using generalized advantage estimation. arXiv preprint arXiv:1506.02438

  37. [37]

    Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Junxiao Song, Xiao Bi, Haowei Zhang, Mingchuan Zhang, YK Li, Yang Wu, and 1 others. 2024. Deepseekmath: Pushing the limits of mathematical reasoning in open language models. arXiv preprint arXiv:2402.03300

  38. [38]

    Kurt Shuster, Spencer Poff, Moya Chen, Douwe Kiela, and Jason Weston. 2021. Retrieval augmentation reduces hallucination in conversation. arXiv preprint arXiv:2104.07567

  39. [39]

    Weihang Su, Yichen Tang, Qingyao Ai, Zhijing Wu, and Yiqun Liu. 2024. https://doi.org/10.18653/v1/2024.acl-long.702 DRAGIN : Dynamic retrieval augmented generation based on the real-time information needs of large language models . In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 129...

  40. [40]

    Jiashuo Sun, Xianrui Zhong, Sizhe Zhou, and Jiawei Han. 2025. https://arxiv.org/abs/2505.07233 Dynamicrag: Leveraging outputs of large language model as feedback for dynamic reranking in retrieval-augmented generation . Preprint, arXiv:2505.07233

  41. [41]

    Weiwei Sun, Lingyong Yan, Xinyu Ma, Shuaiqiang Wang, Pengjie Ren, Zhumin Chen, Dawei Yin, and Zhaochun Ren. 2023. Is chatgpt good at search? investigating large language models as re-ranking agents. arXiv preprint arXiv:2304.09542

  42. [42]

    Qwen Team. 2024. Qwen2 technical report. arXiv preprint arXiv:2407.10671, 2

  43. [43]

    Harsh Trivedi, Niranjan Balasubramanian, Tushar Khot, and Ashish Sabharwal. 2022 a . Interleaving retrieval with chain-of-thought reasoning for knowledge-intensive multi-step questions. arXiv preprint arXiv:2212.10509

  44. [44]

    Harsh Trivedi, Niranjan Balasubramanian, Tushar Khot, and Ashish Sabharwal. 2022 b . https://doi.org/10.1162/TACL\_A\_00475 Musique: Multihop questions via single-hop question composition . Trans. Assoc. Comput. Linguistics, 10:539--554

  45. [45]

    Liang Wang, Nan Yang, and Furu Wei. 2023 a . https://doi.org/10.18653/v1/2023.emnlp-main.585 Query2doc: Query expansion with large language models . In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 9414--9423, Singapore. Association for Computational Linguistics

  46. [46]

    Xinyuan Wang, Chenxi Li, Zhen Wang, Fan Bai, Haotian Luo, Jiayou Zhang, Nebojsa Jojic, Eric P Xing, and Zhiting Hu. 2023 b . Promptagent: Strategic planning with language models enables expert-level prompt optimization. arXiv preprint arXiv:2310.16427

  47. [47]

    Shitao Xiao, Zheng Liu, Peitian Zhang, and Niklas Muennighoff. 2023. https://arxiv.org/abs/2309.07597 C-pack: Packaged resources to advance general chinese embedding . Preprint, arXiv:2309.07597

  48. [48]

    Fangyuan Xu, Weijia Shi, and Eunsol Choi. 2024. Recomp: Improving retrieval-augmented lms with context compression and selective augmentation. In The Twelfth International Conference on Learning Representations

  49. [49]

    Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William W Cohen, Ruslan Salakhutdinov, and Christopher D Manning. 2018. Hotpotqa: A dataset for diverse, explainable multi-hop question answering. arXiv preprint arXiv:1809.09600

  50. [50]

    Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao. 2023. React: Synergizing reasoning and acting in language models. In International Conference on Learning Representations (ICLR)

  51. [51]

    Xiaowei Yuan, Zhao Yang, Yequan Wang, Jun Zhao, and Kang Liu. 2024. Improving zero-shot llm re-ranker with risk minimization. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 17967--17983

  52. [52]

    Ye Yuan, Mohammad Amin Shabani, and Siqi Liu. 2025. Embedding-based context-aware reranker. arXiv preprint arXiv:2510.13329

  53. [53]

    Le Zhang, Bo Wang, Xipeng Qiu, Siva Reddy, and Aishwarya Agrawal. 2025 a . Rearank: Reasoning re-ranking agent via reinforcement learning. arXiv preprint arXiv:2505.20046

  54. [54]

    Ruizhe Zhang, Yongxin Xu, Yuzhen Xiao, Runchuan Zhu, Xinke Jiang, Xu Chu, Junfeng Zhao, and Yasha Wang. 2025 b . https://doi.org/10.1609/aaai.v39i24.34783 Knowpo: Knowledge-aware preference optimization for controllable knowledge selection in retrieval-augmented language models . Proceedings of the AAAI Conference on Artificial Intelligence, 39(24):25895--25903

  55. [55]

    Xin Zhang, Yanzhao Zhang, Dingkun Long, Wen Xie, Ziqi Dai, Jialong Tang, Huan Lin, Baosong Yang, Pengjun Xie, Fei Huang, and 1 others. 2024. mgte: Generalized long-context text representation and reranking models for multilingual text retrieval. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing: Industry Track, page...

  56. [56]

    Yanzhao Zhang, Mingxin Li, Dingkun Long, Xin Zhang, Huan Lin, Baosong Yang, Pengjun Xie, An Yang, Dayiheng Liu, Junyang Lin, Fei Huang, and Jingren Zhou. 2025 c . Qwen3 embedding: Advancing text embedding and reranking through foundation models. arXiv preprint arXiv:2506.05176

  57. [57]

    Yutao Zhu, Peitian Zhang, Chenghao Zhang, Yifei Chen, Binyu Xie, Zheng Liu, Ji-Rong Wen, and Zhicheng Dou. 2024. Inters: unlocking the power of large language models in search with instruction tuning. arXiv preprint arXiv:2401.06532