Pith. sign in

Paper Citation Record · LEDGER

Optimizing RAG Rerankers with LLM Feedback via Reinforcement Learning

As of 19 August 2026, this Paper Citation Record lists 57 of 57 outbound references and 0 inbound Pith citation observations for arXiv:2604.02091.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.02091 v2

Coverage vector

measured 57 of 57 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-13T13:59:01.287449Z

measured 57 of 57 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

57 of 57 outbound references displayed

  • verified exact2
  • verified fuzzy0
  • unresolved54
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e319dbd0-c789-41db-bc1c-27ef7513cfc1 · outbound

This paper cites an unresolved cited work.

Optimizing RAG Rerankers with LLM Feedback via Reinforcement Learning Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-13T13:59:01.287449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T13:59:01.287449Z digest=sha256:c99ec8082d10e89cc839ab704a6f559f7533db4485125f1b89b94f7e056cd644

Observation d64b3c60-9a7d-49f3-ae14-76661596701f · outbound

This paper cites an unresolved cited work.

Optimizing RAG Rerankers with LLM Feedback via Reinforcement Learning Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-13T13:59:01.287449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T13:59:01.287449Z digest=sha256:9ea10d285e8228af91f7cb605d2b397f8685d70ff4efae47968f82dc10bebe91

Observation 91744e54-eac7-42f6-9732-e298f256bb8f · outbound

This paper cites Attention in Large Language Models Yields Efficient Zero-Shot Re-Rankers.

Optimizing RAG Rerankers with LLM Feedback via Reinforcement Learning Attention in Large Language Models Yields Efficient Zero-Shot Re-Rankers

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-13T13:59:01.287449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T13:59:01.287449Z digest=sha256:b0d31c758502cbc1df9234ae36a3eb7fee4ae68edb31e847fae6715f23417bdb

Observation b8db992d-2cfa-45c7-a029-a14ef9c750f6 · outbound

This paper cites an unresolved cited work.

Optimizing RAG Rerankers with LLM Feedback via Reinforcement Learning Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-13T13:59:01.287449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T13:59:01.287449Z digest=sha256:1c4dd59ff79ba6c291780d18c598fc0bbbd50eda2aa3d66c447c8e3f32712515

Observation db6996dc-f1c9-4f85-91ba-6b71ff16808e · outbound

This paper cites an unresolved cited work.

Optimizing RAG Rerankers with LLM Feedback via Reinforcement Learning Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-13T13:59:01.287449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T13:59:01.287449Z digest=sha256:30f6a265b283fa0afd004d6d1bde62b6d81516161eba92098b4f5af9d8da9f2c

Observation a2181c2b-5f0b-4fdb-b041-619478788764 · outbound

This paper cites an unresolved cited work.

Optimizing RAG Rerankers with LLM Feedback via Reinforcement Learning Unresolved cited work

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-13T13:59:01.287449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T13:59:01.287449Z digest=sha256:9e25db1cedebe3ceb4f90384eaeb09ca830759bd265692492c4fa6fcd27b2658

Observation 5d32d157-23e2-4cac-9cd5-e7149cc2f6b1 · outbound

This paper cites Retrieval-Augmented Generation for Large Language Models: A Survey.

Optimizing RAG Rerankers with LLM Feedback via Reinforcement Learning Retrieval-Augmented Generation for Large Language Models: A Survey

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-13T13:59:01.287449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T13:59:01.287449Z digest=sha256:9d0fd4406aba1ee8b01896f74370e6f2b563347212724a25e787979cc5e81b17

Observation 9aaa5580-a00e-4da4-9507-541ad90ce0bc · outbound

This paper cites DeepRAG: Thinking to Retrieve Step by Step for Large Language Models.

Optimizing RAG Rerankers with LLM Feedback via Reinforcement Learning DeepRAG: Thinking to Retrieve Step by Step for Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-13T13:59:01.287449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T13:59:01.287449Z digest=sha256:44e7518dddbe40d3ca35ef74bcbb40c4ae5725e3641b8aed7c1b10459606c9b0

Observation 2631c07a-d806-4ac4-a165-0921c27eab26 · outbound

This paper cites an unresolved cited work.

Optimizing RAG Rerankers with LLM Feedback via Reinforcement Learning Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-13T13:59:01.287449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T13:59:01.287449Z digest=sha256:391eb1efa39fd8b6f802f773668dca9bc1183821a92f0927a65a0e77671c49ac

Observation 8fc202cb-85db-4836-96c9-7cbe4707b7b6 · outbound

This paper cites Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps.

Optimizing RAG Rerankers with LLM Feedback via Reinforcement Learning Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-13T13:59:01.287449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T13:59:01.287449Z digest=sha256:f3a85ac4fb79ec2b2a17b7804f203fa821941cfea316f77394a45fdb0214e55c

Observation 2dd75243-53c7-4e5b-b69b-6975849bc3b5 · outbound

This paper cites Leveraging Passage Retrieval with Generative Models for Open Domain Question Answering.

Optimizing RAG Rerankers with LLM Feedback via Reinforcement Learning Leveraging Passage Retrieval with Generative Models for Open Domain Question Answering

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-13T13:59:01.287449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T13:59:01.287449Z digest=sha256:ed25e587243df7681e93353fddf10b1fa75c86b42ba4d94705bbbb653bb6fe28

Observation ffb2735e-d440-47bf-b948-c8d9734e8cd0 · outbound

This paper cites an unresolved cited work.

Optimizing RAG Rerankers with LLM Feedback via Reinforcement Learning Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-13T13:59:01.287449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T13:59:01.287449Z digest=sha256:a5ef9ae3c9ae4de3cceece76614e9ca474404d70f63d2d497b1af0ddc76f0d76

Observation 66bf8b90-d21b-4c01-96a1-555573eadf98 · outbound

This paper cites Adaptive-RAG: Learning to Adapt Retrieval-Augmented Large Language Models through Question Complexity.

Optimizing RAG Rerankers with LLM Feedback via Reinforcement Learning Adaptive-RAG: Learning to Adapt Retrieval-Augmented Large Language Models through Question Complexity

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-13T13:59:01.287449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T13:59:01.287449Z digest=sha256:5c344cd6f690e9f7f83c2e13e35dd75fd64ede457f2383fc99cec2a7f45baf85

Observation 57fb572e-0ade-40f1-816e-2b929b1b04d1 · outbound

This paper cites an unresolved cited work.

Optimizing RAG Rerankers with LLM Feedback via Reinforcement Learning Unresolved cited work

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-13T13:59:01.287449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T13:59:01.287449Z digest=sha256:65349241553d92d8b7e94467c0638170b1c1061674712c1f07539a869f004b2d

Observation 521d7dc9-1e27-43e7-9b4e-769c3123052e · outbound

This paper cites an unresolved cited work.

Optimizing RAG Rerankers with LLM Feedback via Reinforcement Learning Unresolved cited work

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-13T13:59:01.287449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T13:59:01.287449Z digest=sha256:f629075a0b50e332cda6682f9577a5860d0c04d896d8d87fa8caa6a5e66b3b75

Observation f2bf93c0-5b4f-4fb6-a5ab-ed82c5e3b010 · outbound

This paper cites an unresolved cited work.

Optimizing RAG Rerankers with LLM Feedback via Reinforcement Learning Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-13T13:59:01.287449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T13:59:01.287449Z digest=sha256:a36f71d220733a4fc020d69695be6132fa8e14eb440e8af4a1835dd7faf1020b

Observation 8e44e9ec-47e5-49ce-a5be-7cbe6e32a579 · outbound

This paper cites an unresolved cited work.

Optimizing RAG Rerankers with LLM Feedback via Reinforcement Learning Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-13T13:59:01.287449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T13:59:01.287449Z digest=sha256:531f7df634aceb24d5efced7752cc059ef7ff9c7245531c440ad80564f854e51

Observation f16c8cb4-a615-4083-a195-b28e8a44f26d · outbound

This paper cites an unresolved cited work.

Optimizing RAG Rerankers with LLM Feedback via Reinforcement Learning Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-13T13:59:01.287449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T13:59:01.287449Z digest=sha256:d5f667d4b3763a75427a3bc101727c38d8ead2917b4fb4347e84960f0d95c0bd

Observation b1e5531d-bd7e-4d62-94ec-5dc2274fa37f · outbound

This paper cites Gonzalez, Hao Zhang, and Ion Stoica.

Optimizing RAG Rerankers with LLM Feedback via Reinforcement Learning Gonzalez, Hao Zhang, and Ion Stoica

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-13T13:59:01.287449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T13:59:01.287449Z digest=sha256:7a44c35b6fc018319a7d21f0789e91894db3777da637cf64620655398275d218

Observation 3259635c-63d2-4752-a955-e0eda334a52f · outbound

This paper cites an unresolved cited work.

Optimizing RAG Rerankers with LLM Feedback via Reinforcement Learning Unresolved cited work

Reference 20

Resolution
verified exact
doi, observed 2026-07-13T13:59:38.986618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-13T13:59:01.287449Z digest=sha256:bab7a2e1d05c350bfe4fe0a9f69fcb8506cc9024f872f812bd0c481075a65b12

Observation 98d2654a-bef3-4f9b-8e0b-6b9a6bb4e4c9 · outbound

This paper cites u ttler, Mike Lewis, Wen-tau Yih, Tim Rockt\.

Optimizing RAG Rerankers with LLM Feedback via Reinforcement Learning u ttler, Mike Lewis, Wen-tau Yih, Tim Rockt\

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-13T13:59:01.287449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T13:59:01.287449Z digest=sha256:6db974ce617fc970dd980730af8ba39723c053c76116bc73296a606eed3cce54

Observation ebb46ea2-80a8-420e-84cd-430ba0764a6c · outbound

This paper cites an unresolved cited work.

Optimizing RAG Rerankers with LLM Feedback via Reinforcement Learning Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-13T13:59:01.287449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T13:59:01.287449Z digest=sha256:15e1aea34a9395d25c4ba685ec57fab547da2c6b942d2a1feb6257493699440e

Observation e371d8d8-0baa-409c-b2a1-d63aa8fb4861 · outbound

This paper cites an unresolved cited work.

Optimizing RAG Rerankers with LLM Feedback via Reinforcement Learning Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-13T13:59:01.287449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T13:59:01.287449Z digest=sha256:6df8bae775442c94c09e073ecded0044465489581d9cdee53262c575ac59db85

Observation 38e2699b-0ae9-4ccd-aa75-7094988120bc · outbound

This paper cites an unresolved cited work.

Optimizing RAG Rerankers with LLM Feedback via Reinforcement Learning Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-13T13:59:01.287449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T13:59:01.287449Z digest=sha256:60e53de96c4753f31903bae03816dfa1c18a7870bcb49337cc1190f1446e3fe8

Observation 531003a4-cd1b-4e7e-b626-afd8c859b197 · outbound

This paper cites an unresolved cited work.

Optimizing RAG Rerankers with LLM Feedback via Reinforcement Learning Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-13T13:59:01.287449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T13:59:01.287449Z digest=sha256:13301e2459919cf39da420b0be04109cfa5309d2189600600313e601329f72dd

Observation b0284e23-730a-41ba-b72c-ad9ed24d3480 · outbound

This paper cites AmbigQA: Answering Ambiguous Open-domain Questions.

Optimizing RAG Rerankers with LLM Feedback via Reinforcement Learning AmbigQA: Answering Ambiguous Open-domain Questions

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-13T13:59:01.287449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T13:59:01.287449Z digest=sha256:0e20ee829c561b41ed9215a79a51308379b5ef618582dc0b94c0e46139a13af3

Observation 5405dbae-be00-47ae-8d3f-70f46d10a6bb · outbound

This paper cites an unresolved cited work.

Optimizing RAG Rerankers with LLM Feedback via Reinforcement Learning Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-13T13:59:01.287449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T13:59:01.287449Z digest=sha256:dc409087c04216111f3819485b5247859a07aafe3a73b0f7a9a7f0db39fb7c02

Observation cdc17a1c-5daa-4bfc-9034-c5d2ced7933a · outbound

This paper cites WebGPT: Browser-assisted question-answering with human feedback.

Optimizing RAG Rerankers with LLM Feedback via Reinforcement Learning WebGPT: Browser-assisted question-answering with human feedback

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-13T13:59:01.287449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T13:59:01.287449Z digest=sha256:6702d683107208cee4002afb26797df40521892654aff2e2900a438b65c7b8c2

Observation ac505bf9-f590-4b74-bf48-8f4cc96274ec · outbound

This paper cites Passage Re-ranking with BERT.

Optimizing RAG Rerankers with LLM Feedback via Reinforcement Learning Passage Re-ranking with BERT

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-13T13:59:01.287449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T13:59:01.287449Z digest=sha256:f3c854b06e6a4de6d3ed9b024b3da8346debd91d08ad1fa8c3b47296cc06c428

Observation d073a16f-e56f-4ac4-9403-90d940f5c569 · outbound

This paper cites Multi-Stage Document Ranking with BERT.

Optimizing RAG Rerankers with LLM Feedback via Reinforcement Learning Multi-Stage Document Ranking with BERT

Reference 30

Resolution
unresolved
no resolver link, observed 2026-07-13T13:59:01.287449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T13:59:01.287449Z digest=sha256:12b19da6d64da43eb8c48367bbbe4a7bbcd7d67b994ea8f8286c008217dbf31d

Observation d0a56791-c706-46fb-9e00-47c56d7b762a · outbound

This paper cites an unresolved cited work.

Optimizing RAG Rerankers with LLM Feedback via Reinforcement Learning Unresolved cited work

Reference 31

Resolution
unresolved
no resolver link, observed 2026-07-13T13:59:01.287449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T13:59:01.287449Z digest=sha256:50491803e066aebc7ba27b6d31b20537f9b12e2d7a25bbd20f8a0b0ea7ed15f2

Observation bfc6cbe7-d292-40c1-b696-2169ba8a52e1 · outbound

This paper cites RankZephyr: Effective and Robust Zero-Shot Listwise Reranking is a Breeze!.

Optimizing RAG Rerankers with LLM Feedback via Reinforcement Learning RankZephyr: Effective and Robust Zero-Shot Listwise Reranking is a Breeze!

Reference 32

Resolution
unresolved
no resolver link, observed 2026-07-13T13:59:01.287449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T13:59:01.287449Z digest=sha256:581cd08666af0824308757d60b3a41426d6a0a74e94f6e6fe13dfb472f2f5889

Observation 02c0f16a-ee4a-4503-8588-f657852e92d7 · outbound

This paper cites an unresolved cited work.

Optimizing RAG Rerankers with LLM Feedback via Reinforcement Learning Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-13T13:59:01.287449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T13:59:01.287449Z digest=sha256:c8d1127825751d87930aaeafe1966af0ef98c8ca8cf468f7ef290db5ca25bd8f

Observation 5f7e9cc9-11f8-4523-99e4-1e16235a9300 · outbound

This paper cites FIRST: Faster Improved Listwise Reranking with Single Token Decoding.

Optimizing RAG Rerankers with LLM Feedback via Reinforcement Learning FIRST: Faster Improved Listwise Reranking with Single Token Decoding

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-13T13:59:01.287449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T13:59:01.287449Z digest=sha256:bca53d1dda255d200a5c852b2f60c170c1aedd59f79fc4a4d3158d4c30022a4d

Observation a4d113d1-f455-44fa-9f62-b9ee1c9951e2 · outbound

This paper cites an unresolved cited work.

Optimizing RAG Rerankers with LLM Feedback via Reinforcement Learning Unresolved cited work

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-13T13:59:01.287449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T13:59:01.287449Z digest=sha256:036cdad71d4362cda1e9eed5e31b4ad594e6aa814c07d043772deca193293009

Observation 7c898c14-49db-4b9f-a007-3f6ebbba6adf · outbound

This paper cites High-Dimensional Continuous Control Using Generalized Advantage Estimation.

Optimizing RAG Rerankers with LLM Feedback via Reinforcement Learning High-Dimensional Continuous Control Using Generalized Advantage Estimation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-13T13:59:01.287449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T13:59:01.287449Z digest=sha256:dfb0cab7495689ad7f2fa0495c5af8edd58d69db8bb42166a445bd2634b22aa8

Observation 830f22c7-3991-4bf5-b515-0992798518c1 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Optimizing RAG Rerankers with LLM Feedback via Reinforcement Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-13T13:59:01.287449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T13:59:01.287449Z digest=sha256:7f3274a763d41a46e3efe1886ed5b723946260c3bf5bc078ad55e0913940b56d

Observation ea2c8636-a071-4899-90b7-3e6e9d9c83b1 · outbound

This paper cites Retrieval Augmentation Reduces Hallucination in Conversation.

Optimizing RAG Rerankers with LLM Feedback via Reinforcement Learning Retrieval Augmentation Reduces Hallucination in Conversation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-07-13T13:59:01.287449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T13:59:01.287449Z digest=sha256:069d5dff78a10a4a75141b6721e83f27a5399420bb8de89fedae91e25c7cc4a5

Observation 0e646110-e67d-49df-b641-55a4bea42bd9 · outbound

This paper cites an unresolved cited work.

Optimizing RAG Rerankers with LLM Feedback via Reinforcement Learning Unresolved cited work

Reference 39

Resolution
unresolved
no resolver link, observed 2026-07-13T13:59:01.287449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T13:59:01.287449Z digest=sha256:c7705d2eb249dfd12bb47ff1cbc2858e0ae344019882add03307c6a81af3f32c

Observation c5c6cba6-58a2-4529-a767-ee34a853a36c · outbound

This paper cites DynamicRAG: Leveraging Outputs of Large Language Model as Feedback for Dynamic Reranking in Retrieval-Augmented Generation.

Optimizing RAG Rerankers with LLM Feedback via Reinforcement Learning DynamicRAG: Leveraging Outputs of Large Language Model as Feedback for Dynamic Reranking in Retrieval-Augmented Generation

Reference 40

Resolution
unresolved
no resolver link, observed 2026-07-13T13:59:01.287449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T13:59:01.287449Z digest=sha256:ec1072555433637fa02934ae270c52c839d57830c8e8cba219f098d33cf84409

Observation d80fef6f-11d2-40af-894b-de80f63b820d · outbound

This paper cites Is ChatGPT Good at Search? Investigating Large Language Models as Re-Ranking Agents.

Optimizing RAG Rerankers with LLM Feedback via Reinforcement Learning Is ChatGPT Good at Search? Investigating Large Language Models as Re-Ranking Agents

Reference 41

Resolution
unresolved
no resolver link, observed 2026-07-13T13:59:01.287449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T13:59:01.287449Z digest=sha256:101450022cb15c3f9b68a835d0aa8eb28ec675577be0b67185e4ea88be592d7f

Observation 6f6c0135-5b84-4663-8c3f-8bd3a42d0358 · outbound

This paper cites Qwen2 Technical Report.

Optimizing RAG Rerankers with LLM Feedback via Reinforcement Learning Qwen2 Technical Report

Reference 42

Resolution
unresolved
no resolver link, observed 2026-07-13T13:59:01.287449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T13:59:01.287449Z digest=sha256:56aa494c80fd4ac02c2c269c16115282b168db8ad3bd0f06b7f1f4a4b2824e2f

Observation 4875af82-6871-46ed-b30f-54e2222d9230 · outbound

This paper cites Interleaving Retrieval with Chain-of-Thought Reasoning for Knowledge-Intensive Multi-Step Questions.

Optimizing RAG Rerankers with LLM Feedback via Reinforcement Learning Interleaving Retrieval with Chain-of-Thought Reasoning for Knowledge-Intensive Multi-Step Questions

Reference 43

Resolution
unresolved
no resolver link, observed 2026-07-13T13:59:01.287449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T13:59:01.287449Z digest=sha256:b85ca96379198d87a42581a59b48885c7a8d32b1394cb504081b0c2fefa07e39

Observation 778fc0f8-4e93-4bec-ade4-4b16e201a1d8 · outbound

This paper cites an unresolved cited work.

Optimizing RAG Rerankers with LLM Feedback via Reinforcement Learning Unresolved cited work

Reference 44

Resolution
malformed identifier
no resolver link, observed 2026-07-13T13:59:01.287449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T13:59:01.287449Z digest=sha256:d5fc565b614e8f6de1fdc03658cf55c69675e9294bb753a586061f8644f77c97

Observation 5323cb47-3a6a-45ac-89b6-538237a64e13 · outbound

This paper cites an unresolved cited work.

Optimizing RAG Rerankers with LLM Feedback via Reinforcement Learning Unresolved cited work

Reference 45

Resolution
unresolved
no resolver link, observed 2026-07-13T13:59:01.287449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T13:59:01.287449Z digest=sha256:c52e1ee3a7739fd52db4e20932a7060289c07fe5f33554c4ce10d773549fb05b

Observation 75fd3e11-57cf-4541-b5d2-b577384457e7 · outbound

This paper cites PromptAgent: Strategic Planning with Language Models Enables Expert-level Prompt Optimization.

Optimizing RAG Rerankers with LLM Feedback via Reinforcement Learning PromptAgent: Strategic Planning with Language Models Enables Expert-level Prompt Optimization

Reference 46

Resolution
unresolved
no resolver link, observed 2026-07-13T13:59:01.287449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T13:59:01.287449Z digest=sha256:920b2bd7588d84b4f2f320785e00827694925d7e4a2b7c85d4eee449383cc0b0

Observation c4954a0d-e10d-4e80-8280-6129688e4989 · outbound

This paper cites C-Pack: Packed Resources For General Chinese Embeddings.

Optimizing RAG Rerankers with LLM Feedback via Reinforcement Learning C-Pack: Packed Resources For General Chinese Embeddings

Reference 47

Resolution
unresolved
no resolver link, observed 2026-07-13T13:59:01.287449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T13:59:01.287449Z digest=sha256:8b1583912469027209c30eada565cf27f8876c6722b9768904b32a402edc6069

Observation 5475d41e-0df9-47d3-b6e8-2ba75bcf861b · outbound

This paper cites an unresolved cited work.

Optimizing RAG Rerankers with LLM Feedback via Reinforcement Learning Unresolved cited work

Reference 48

Resolution
unresolved
no resolver link, observed 2026-07-13T13:59:01.287449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T13:59:01.287449Z digest=sha256:4ecae0743f4138abd4b90150e3c7cc20663d25eff11c66f0afbbd974ebfa75ca

Observation f7776387-9543-4468-a19e-5fa2955252a2 · outbound

This paper cites HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering.

Optimizing RAG Rerankers with LLM Feedback via Reinforcement Learning HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering

Reference 49

Resolution
unresolved
no resolver link, observed 2026-07-13T13:59:01.287449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T13:59:01.287449Z digest=sha256:756139a4d6b9d1c70c9ef170fab3bd9c73a794b3220639a43ad1beb7d3da20dd

Observation 4db4453b-27ea-4b04-b281-ffbc8d386f01 · outbound

This paper cites an unresolved cited work.

Optimizing RAG Rerankers with LLM Feedback via Reinforcement Learning Unresolved cited work

Reference 50

Resolution
unresolved
no resolver link, observed 2026-07-13T13:59:01.287449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T13:59:01.287449Z digest=sha256:81d4b117fd223f8f3a7761909007392c4e7ef38d31011dec162a7effbca4c882

Observation f89c5535-6543-42aa-be52-24ed84fa887e · outbound

This paper cites an unresolved cited work.

Optimizing RAG Rerankers with LLM Feedback via Reinforcement Learning Unresolved cited work

Reference 51

Resolution
unresolved
no resolver link, observed 2026-07-13T13:59:01.287449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T13:59:01.287449Z digest=sha256:7b5a058f1e51bb69fb8f5290d9c54e35d03adb479d610ac259ffc8628626a0de

Observation dd0f40a5-8536-44cf-a258-89ff43212f08 · outbound

This paper cites an unresolved cited work.

Optimizing RAG Rerankers with LLM Feedback via Reinforcement Learning Unresolved cited work

Reference 52

Resolution
unresolved
no resolver link, observed 2026-07-13T13:59:01.287449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T13:59:01.287449Z digest=sha256:c85e7fd7d540521d85bd1523aa20b1bded2f3113fd6da06a484631b4aab125b9

Observation 9e5b552c-b100-400b-a8f4-79f2f8db55a4 · outbound

This paper cites REARANK: Reasoning Re-ranking Agent via Reinforcement Learning.

Optimizing RAG Rerankers with LLM Feedback via Reinforcement Learning REARANK: Reasoning Re-ranking Agent via Reinforcement Learning

Reference 53

Resolution
unresolved
no resolver link, observed 2026-07-13T13:59:01.287449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T13:59:01.287449Z digest=sha256:163790d3412e4e12cc849430dc0a5b633b573725e0c951761b5463c541e2162b

Observation 86ab9ee2-68d9-4402-a0e9-ca2c83ba7f37 · outbound

This paper cites an unresolved cited work.

Optimizing RAG Rerankers with LLM Feedback via Reinforcement Learning Unresolved cited work

Reference 54

Resolution
verified exact
doi, observed 2026-07-13T13:59:38.992267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-13T13:59:01.287449Z digest=sha256:41a4c8bbdcaa218811652394541de0c9de62ec2b39aa586a45c615df280b83f5

Observation 0e94330a-cf02-47e1-baa5-9192a790102e · outbound

This paper cites an unresolved cited work.

Optimizing RAG Rerankers with LLM Feedback via Reinforcement Learning Unresolved cited work

Reference 55

Resolution
unresolved
no resolver link, observed 2026-07-13T13:59:01.287449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T13:59:01.287449Z digest=sha256:4b21a0f258d35a5078443a86edd008a7ea33f981723542e48fca8c2b8e282c14

Observation 8c5d59e8-6cc1-41d4-8909-9c5cacc56aa0 · outbound

This paper cites Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models.

Optimizing RAG Rerankers with LLM Feedback via Reinforcement Learning Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-07-13T13:59:01.287449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T13:59:01.287449Z digest=sha256:f207fb4a360e22a70920f40960a73f1c0a8dcb3719bf9a8a7e293a9c4e1ac2ab

Observation 954e8775-b5a0-49a2-ae9d-60be74f45334 · outbound

This paper cites INTERS: Unlocking the Power of Large Language Models in Search with Instruction Tuning.

Optimizing RAG Rerankers with LLM Feedback via Reinforcement Learning INTERS: Unlocking the Power of Large Language Models in Search with Instruction Tuning

Reference 57

Resolution
unresolved
no resolver link, observed 2026-07-13T13:59:01.287449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T13:59:01.287449Z digest=sha256:cad3dc7a233f29af724134e617a6ddb9ecb20bb92b56fbc884746fabfbc2abe0

Pith citing papers

No inbound Pith citation observations are available.