Pith. sign in

Paper Citation Record · LEDGER

Sailing by the Stars: A Survey on Reward Models and Learning Strategies for Learning from Rewards

As of 20 August 2026, this Paper Citation Record lists 24 of 24 outbound references and 9 inbound Pith citation observations for arXiv:2505.02686.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.02686 v2

Coverage vector

measured 24 of 24 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:47:59.942231Z

measured 33 of 33 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T12:58:02.233748Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T07:59:40.123100Z

Reference resolution

24 of 24 outbound references displayed

  • verified exact1
  • verified fuzzy1
  • unresolved21
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 02fa8ef8-e6f9-482a-aaa8-a6a12c1465fb · outbound

This paper cites Concrete Problems in AI Safety.

Sailing by the Stars: A Survey on Reward Models and Learning Strategies for Learning from Rewards Concrete Problems in AI Safety

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T00:47:59.844459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:47:59.844459Z digest=sha256:04904f1a88d0c19cad4d494aeafa7ee3e74fec639382c42762f9e11919281858

Observation 5568689e-7b10-47bb-8d07-1fb35de0a9ee · outbound

This paper cites Safe RLHF: Safe Reinforcement Learning from Human Feedback.

Sailing by the Stars: A Survey on Reward Models and Learning Strategies for Learning from Rewards Safe RLHF: Safe Reinforcement Learning from Human Feedback

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T00:47:59.853822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:47:59.853822Z digest=sha256:9857dbdc0adcad123825c48605efeca237fef37b09a3bf87653ea296c83ae851

Observation e1290362-3a1f-4b42-8629-2ef1c05a7f2d · outbound

This paper cites CRITIC: Large Language Models Can Self-Correct with Tool-Interactive Critiquing.

Sailing by the Stars: A Survey on Reward Models and Learning Strategies for Learning from Rewards CRITIC: Large Language Models Can Self-Correct with Tool-Interactive Critiquing

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T00:47:59.858369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:47:59.858369Z digest=sha256:9f7070d7233ae5dd01b8438617b545afe4c5d24efd1a58ca0475d2be710494af

Observation 334a5b08-4e21-4690-9374-120a8a4265a3 · outbound

This paper cites In The Twelfth International Conference on Learning Representations.

Sailing by the Stars: A Survey on Reward Models and Learning Strategies for Learning from Rewards In The Twelfth International Conference on Learning Representations

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:48:00.553206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T00:47:59.866533Z digest=sha256:95cbb0dcc7707effbd8fe96638e4615a2a0c36baaa0b388c4957b659811ab352

Observation af1a094b-f262-4281-a1bf-9e227f1dd2af · outbound

This paper cites Training Language Models to Self-Correct via Reinforcement Learning.

Sailing by the Stars: A Survey on Reward Models and Learning Strategies for Learning from Rewards Training Language Models to Self-Correct via Reinforcement Learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-16T00:47:59.869987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:47:59.869987Z digest=sha256:0dfcc1c20a38c0c69009e31b39dc2e743e9c7ee48d3954aa14a48d46923aa090

Observation 3108e031-008a-4345-9fa0-ee2c5a40d36b · outbound

This paper cites In The Twelfth Inter- national Conference on Learning Representations.

Sailing by the Stars: A Survey on Reward Models and Learning Strategies for Learning from Rewards In The Twelfth Inter- national Conference on Learning Representations

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T00:47:59.873946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:47:59.873946Z digest=sha256:217fc10b0395dae4011f23dadf63172aec366066b0fa5c9a1a5ec82b78825585

Observation e5e5818d-8133-40f4-894d-5a13cdce061a · outbound

This paper cites Generative Reward Models.

Sailing by the Stars: A Survey on Reward Models and Learning Strategies for Learning from Rewards Generative Reward Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T00:47:59.877697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:47:59.877697Z digest=sha256:a2f03f9049f33a0d0f4158bbff415dabebb3c5515f9df6bb1faa7eb52ed1a70f

Observation 704c9448-cedf-4e2e-a8a0-3e46cadbbbfc · outbound

This paper cites GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models.

Sailing by the Stars: A Survey on Reward Models and Learning Strategies for Learning from Rewards GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T00:47:59.882096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:47:59.882096Z digest=sha256:87b7f2fe5a63a069c565f1660895c7b395b55f7706ab498b63b20824ee6df500

Observation 76c2ce46-1a46-40e1-8947-a9eb9308992a · outbound

This paper cites GPT-4 Technical Report.

Sailing by the Stars: A Survey on Reward Models and Learning Strategies for Learning from Rewards GPT-4 Technical Report

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T00:47:59.886298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:47:59.886298Z digest=sha256:a020bf59fb219dfa4a57a1c9e84e3591b49ff636155e32b505fd7a045b4994a1

Observation 1ca95a0f-f308-4e95-a3d5-653fc7fefe31 · outbound

This paper cites Ensembling Large Language Models with Process Reward-Guided Tree Search for Better Complex Reasoning.

Sailing by the Stars: A Survey on Reward Models and Learning Strategies for Learning from Rewards Ensembling Large Language Models with Process Reward-Guided Tree Search for Better Complex Reasoning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T00:47:59.894438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:47:59.894438Z digest=sha256:4859d1dbbfaeccae920d83571325271fb0ed658ec17d31f78467785fc3224d65

Observation 60ce98d7-9f47-43ad-a4f6-72dbacb48a92 · outbound

This paper cites Phenomenal Yet Puzzling: Testing Inductive Reasoning Capabilities of Language Models with Hypothesis Refinement.

Sailing by the Stars: A Survey on Reward Models and Learning Strategies for Learning from Rewards Phenomenal Yet Puzzling: Testing Inductive Reasoning Capabilities of Language Models with Hypothesis Refinement

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T00:47:59.898459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:47:59.898459Z digest=sha256:b89fc7bb62b6ff37365b7feabcb41f8fe7532d481c0fb1685de3d6dda86518a6

Observation 2609e571-1858-40f8-bdff-b0766b5e2a48 · outbound

This paper cites Towards Cost-Effective Reward Guided Text Generation.

Sailing by the Stars: A Survey on Reward Models and Learning Strategies for Learning from Rewards Towards Cost-Effective Reward Guided Text Generation

Reference 16

Resolution
metadata mismatch
local_arxiv, observed 2026-08-16T00:48:00.222120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T00:47:59.902347Z digest=sha256:6217d7af565e8180d340cfdca34f78a6adf567e59e2592b933f59bfa1d4167c4

Observation 668e6ea7-8665-4257-bf2b-4cca0dc65295 · outbound

This paper cites arXiv preprint arXiv:2407.04615.

Sailing by the Stars: A Survey on Reward Models and Learning Strategies for Learning from Rewards arXiv preprint arXiv:2407.04615

Reference 18

Resolution
verified exact
raw_fallback, observed 2026-08-16T00:48:00.185769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T00:47:59.916083Z digest=sha256:7c879986c5df1d0f7d87189c253fd3d93bb0a82515b2365e55ecb845309bf726

Observation 823e038e-2e65-4052-babd-e4499e8ef3eb · outbound

This paper cites mDPO: Conditional Preference Optimization for Multimodal Large Language Models.

Sailing by the Stars: A Survey on Reward Models and Learning Strategies for Learning from Rewards mDPO: Conditional Preference Optimization for Multimodal Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T00:47:59.920486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:47:59.920486Z digest=sha256:4d31db41a172e46724cde336d52fd3490a478508585349b087cf37d47922792c

Observation bd1e7523-399a-4c2d-beb4-2658e816daa7 · outbound

This paper cites SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution.

Sailing by the Stars: A Survey on Reward Models and Learning Strategies for Learning from Rewards SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-16T00:47:59.925135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:47:59.925135Z digest=sha256:c11afe29c32ae0ff8ea3ef1c95b6cbbb2d101396c7082e548b570ac8ca97ad49

Observation 94e9bbb2-fb5c-44c9-b6d2-e1aa7f4966f6 · outbound

This paper cites On-Policy Self-Alignment with Fine-grained Knowledge Feedback for Hallucination Mitigation.

Sailing by the Stars: A Survey on Reward Models and Learning Strategies for Learning from Rewards On-Policy Self-Alignment with Fine-grained Knowledge Feedback for Hallucination Mitigation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T00:47:59.929137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:47:59.929137Z digest=sha256:55b0f3258844f733e2322e01327455c6328522b39f1286df0f24643302a69588

Observation 617e7a27-5af9-45f3-91ed-3234000289e1 · outbound

This paper cites Advances in Neural Information Processing Systems, 36:41618–41650.

Sailing by the Stars: A Survey on Reward Models and Learning Strategies for Learning from Rewards Advances in Neural Information Processing Systems, 36:41618–41650

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-16T00:47:59.934306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:47:59.934306Z digest=sha256:5f717815b372e7d53ff98f4bc467c9f5d93c214db31d82988b26417a174ca6f8

Observation f885e547-7aa9-4779-bbd9-71218d29e38d · outbound

This paper cites Not All Rollouts are Useful: Down-Sampling Rollouts in LLM Reinforcement Learning.

Sailing by the Stars: A Survey on Reward Models and Learning Strategies for Learning from Rewards Not All Rollouts are Useful: Down-Sampling Rollouts in LLM Reinforcement Learning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T00:47:59.938066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:47:59.938066Z digest=sha256:1ae19c83edf9d1855f8e71ed68d2bf13b7ec47fef115d7d08e4d6505bf308a46

Observation 2394fba1-3d86-4feb-b22a-07787f631ac8 · outbound

This paper cites Multimodal RewardBench: Holistic Evaluation of Reward Models for Vision Language Models.

Sailing by the Stars: A Survey on Reward Models and Learning Strategies for Learning from Rewards Multimodal RewardBench: Holistic Evaluation of Reward Models for Vision Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T00:47:59.942231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:47:59.942231Z digest=sha256:e2f67c6b7cb634160f8890cdfa31567e7d9b9e10aaa5d5fe1b5e59448d4593d1

Observation ba5e3fe7-c7ea-4b47-9b40-552cf6b86b5d · outbound

This paper cites Self-critiquing models for assisting human evaluators.

Sailing by the Stars: A Survey on Reward Models and Learning Strategies for Learning from Rewards Self-critiquing models for assisting human evaluators

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-16T00:47:59.907338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:47:59.907338Z digest=sha256:24f53f231a50b927c0f8b7c05e7a5b8ff03dca0b116adb767ab26ae1d63d9a09

Observation d380dc8c-6a17-472f-89ba-b03aa0e6d3e8 · outbound

This paper cites The Effects of Reward Misspecification: Mapping and Mitigating Misaligned Models.

Sailing by the Stars: A Survey on Reward Models and Learning Strategies for Learning from Rewards The Effects of Reward Misspecification: Mapping and Mitigating Misaligned Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-16T00:47:59.890088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:47:59.890088Z digest=sha256:ff9debb0b8546525a38e9d9d48e6c52df91fbf4706d6b0dc173238539b3c58b6

Observation c64dff9b-1e48-44a6-97d1-c7c9e877a38b · outbound

This paper cites Tuning Large Multimodal Models for Videos using Reinforcement Learning from AI Feedback.

Sailing by the Stars: A Survey on Reward Models and Learning Strategies for Learning from Rewards Tuning Large Multimodal Models for Videos using Reinforcement Learning from AI Feedback

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-16T00:47:59.839673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:47:59.839673Z digest=sha256:f8d5c1df417570c7bce21df9852c26785fc361c45d860cf37ae1fac1fa7ed186

Observation d3d93e4c-8845-44e2-8a86-98ef30b624a8 · outbound

This paper cites Critique-out-Loud Reward Models.

Sailing by the Stars: A Survey on Reward Models and Learning Strategies for Learning from Rewards Critique-out-Loud Reward Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-16T00:47:59.849111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:47:59.849111Z digest=sha256:b07bee3c8490be116d66f2da89ed73a663374fa69d4ca66c9bf93130c15ffba3

Observation f65b86e6-2d33-46a5-b87e-37a75716c1b9 · outbound

This paper cites rStar-Math: Small LLMs Can Master Math Reasoning with Self-Evolved Deep Thinking.

Sailing by the Stars: A Survey on Reward Models and Learning Strategies for Learning from Rewards rStar-Math: Small LLMs Can Master Math Reasoning with Self-Evolved Deep Thinking

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-16T00:47:59.862170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:47:59.862170Z digest=sha256:59853a9e136544bc8c4874cbfefea01994302f845aa0a107db22697561bacce8

Pith citing papers

Observation cc987656-7d73-4688-9723-f97ce0a6b9d0 · inbound

AntiLeakBench: Preventing Data Contamination by Automatically Constructing Benchmarks with Updated Real-World Knowledge cites this paper.

AntiLeakBench: Preventing Data Contamination by Automatically Constructing Benchmarks with Updated Real-World Knowledge Sailing by the Stars: A Survey on Reward Models and Learning Strategies for Learning from Rewards

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T12:58:02.233748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:58:02.233748Z digest=sha256:21006b368ba497c255bc73eab4929f366bbf7926a87b75f4018c71c937e9aa54

Observation 60d052a9-dcb5-45c1-9584-344c8260366a · inbound

SCOPE: Compress Mathematical Reasoning Steps for Efficient Automated Process Annotation cites this paper.

SCOPE: Compress Mathematical Reasoning Steps for Efficient Automated Process Annotation Sailing by the Stars: A Survey on Reward Models and Learning Strategies for Learning from Rewards

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T15:39:19.201927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:39:19.201927Z digest=sha256:e35b143cd08412ef48e10753bc8fddbba60bcc674044317685947163e5cc0d40

Observation b08205ac-ba61-43a9-a7d5-dbcb2efed3b2 · inbound

Training LLMs for EHR-Based Reasoning Tasks via Reinforcement Learning cites this paper.

Training LLMs for EHR-Based Reasoning Tasks via Reinforcement Learning Sailing by the Stars: A Survey on Reward Models and Learning Strategies for Learning from Rewards

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T12:40:13.070016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:40:13.070016Z digest=sha256:78699be5e08ceab40ac81d242af1655f8377779ad6b996e3dcb2aa889e2fd8f8

Observation af4c72b2-bc22-40cc-85b2-9cb08932bbd9 · inbound

ReAgent-V: A Reward-Driven Multi-Agent Framework for Video Understanding cites this paper.

ReAgent-V: A Reward-Driven Multi-Agent Framework for Video Understanding Sailing by the Stars: A Survey on Reward Models and Learning Strategies for Learning from Rewards

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T11:52:04.806540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:52:04.806540Z digest=sha256:43059a4f1fc2a734eb413aadaee318c80fe24e78327c6a902e68786b49f0af82

Observation ba0db544-a3dc-49a4-99ab-304f93204ed0 · inbound

Towards Storage-Efficient Visual Document Retrieval: An Empirical Study on Reducing Patch-Level Embeddings cites this paper.

Towards Storage-Efficient Visual Document Retrieval: An Empirical Study on Reducing Patch-Level Embeddings Sailing by the Stars: A Survey on Reward Models and Learning Strategies for Learning from Rewards

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T10:34:49.797401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:34:49.797401Z digest=sha256:8b8647335f380d759e9f56a93b93667c6d63f2735ce455141161c22a591dcae2

Observation a9257381-076a-420b-acae-22465db6a9cc · inbound

PEER: Unified Process-Outcome Reinforcement Learning for Structured Empathetic Reasoning cites this paper.

PEER: Unified Process-Outcome Reinforcement Learning for Structured Empathetic Reasoning Sailing by the Stars: A Survey on Reward Models and Learning Strategies for Learning from Rewards

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-18T23:16:53.795613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-18T23:16:13.165715Z digest=sha256:7a37d1974632e8d3b3546ed5d607ef7d00b2a77a97b3cd67bd0ee5ae4e5deb73

Observation fc378cd4-d001-44b5-98fd-4880a08afbb8 · inbound

Unsupervised Hallucination Detection by Inspecting Reasoning Processes cites this paper.

Unsupervised Hallucination Detection by Inspecting Reasoning Processes Sailing by the Stars: A Survey on Reward Models and Learning Strategies for Learning from Rewards

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-04T18:25:37.408457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T18:25:37.408457Z digest=sha256:1f39ef509ad5e25031aacf72596ca698062f5e486aa1208932ce4931ae9fb019

Observation 44eb936b-cc01-450d-a763-d02363a1ccd7 · inbound

Trust Region On-Policy Distillation cites this paper.

Trust Region On-Policy Distillation Sailing by the Stars: A Survey on Reward Models and Learning Strategies for Learning from Rewards

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-07-01T20:56:13.574362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-28T17:38:50.313305Z digest=sha256:5a7ae359372b250777afb8c834367df9afa901f3417139965ebb34753890c569

Observation 5ceacf62-af66-4d4e-9ce2-3b095fa951ad · inbound

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning cites this paper.

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning Sailing by the Stars: A Survey on Reward Models and Learning Strategies for Learning from Rewards

Reference 228

Resolution
verified exact
arxiv_id, observed 2026-07-04T07:59:40.124708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-26T12:15:08.304150Z digest=sha256:d6629a04c3a668fc8959ebf342b0fa872fb43328d580c56ad69df2813790d064