Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T00:47:59.942231Z
Paper Citation Record · LEDGER
As of 20 August 2026, this Paper Citation Record lists 24 of 24 outbound references and 9 inbound Pith citation observations for arXiv:2505.02686.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T00:47:59.942231Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-11T12:58:02.233748Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T07:59:40.123100Z
24 of 24 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 02fa8ef8-e6f9-482a-aaa8-a6a12c1465fb · outbound
Sailing by the Stars: A Survey on Reward Models and Learning Strategies for Learning from Rewards Concrete Problems in AI Safety
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5568689e-7b10-47bb-8d07-1fb35de0a9ee · outbound
Sailing by the Stars: A Survey on Reward Models and Learning Strategies for Learning from Rewards Safe RLHF: Safe Reinforcement Learning from Human Feedback
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e1290362-3a1f-4b42-8629-2ef1c05a7f2d · outbound
Sailing by the Stars: A Survey on Reward Models and Learning Strategies for Learning from Rewards CRITIC: Large Language Models Can Self-Correct with Tool-Interactive Critiquing
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 334a5b08-4e21-4690-9374-120a8a4265a3 · outbound
Sailing by the Stars: A Survey on Reward Models and Learning Strategies for Learning from Rewards In The Twelfth International Conference on Learning Representations
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation af1a094b-f262-4281-a1bf-9e227f1dd2af · outbound
Sailing by the Stars: A Survey on Reward Models and Learning Strategies for Learning from Rewards Training Language Models to Self-Correct via Reinforcement Learning
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3108e031-008a-4345-9fa0-ee2c5a40d36b · outbound
Sailing by the Stars: A Survey on Reward Models and Learning Strategies for Learning from Rewards In The Twelfth Inter- national Conference on Learning Representations
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e5e5818d-8133-40f4-894d-5a13cdce061a · outbound
Sailing by the Stars: A Survey on Reward Models and Learning Strategies for Learning from Rewards Generative Reward Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 704c9448-cedf-4e2e-a8a0-3e46cadbbbfc · outbound
Sailing by the Stars: A Survey on Reward Models and Learning Strategies for Learning from Rewards GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 76c2ce46-1a46-40e1-8947-a9eb9308992a · outbound
Sailing by the Stars: A Survey on Reward Models and Learning Strategies for Learning from Rewards GPT-4 Technical Report
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ca95a0f-f308-4e95-a3d5-653fc7fefe31 · outbound
Sailing by the Stars: A Survey on Reward Models and Learning Strategies for Learning from Rewards Ensembling Large Language Models with Process Reward-Guided Tree Search for Better Complex Reasoning
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 60ce98d7-9f47-43ad-a4f6-72dbacb48a92 · outbound
Sailing by the Stars: A Survey on Reward Models and Learning Strategies for Learning from Rewards Phenomenal Yet Puzzling: Testing Inductive Reasoning Capabilities of Language Models with Hypothesis Refinement
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2609e571-1858-40f8-bdff-b0766b5e2a48 · outbound
Sailing by the Stars: A Survey on Reward Models and Learning Strategies for Learning from Rewards Towards Cost-Effective Reward Guided Text Generation
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 668e6ea7-8665-4257-bf2b-4cca0dc65295 · outbound
Sailing by the Stars: A Survey on Reward Models and Learning Strategies for Learning from Rewards arXiv preprint arXiv:2407.04615
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 823e038e-2e65-4052-babd-e4499e8ef3eb · outbound
Sailing by the Stars: A Survey on Reward Models and Learning Strategies for Learning from Rewards mDPO: Conditional Preference Optimization for Multimodal Large Language Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd1e7523-399a-4c2d-beb4-2658e816daa7 · outbound
Sailing by the Stars: A Survey on Reward Models and Learning Strategies for Learning from Rewards SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94e9bbb2-fb5c-44c9-b6d2-e1aa7f4966f6 · outbound
Sailing by the Stars: A Survey on Reward Models and Learning Strategies for Learning from Rewards On-Policy Self-Alignment with Fine-grained Knowledge Feedback for Hallucination Mitigation
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 617e7a27-5af9-45f3-91ed-3234000289e1 · outbound
Sailing by the Stars: A Survey on Reward Models and Learning Strategies for Learning from Rewards Advances in Neural Information Processing Systems, 36:41618–41650
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f885e547-7aa9-4779-bbd9-71218d29e38d · outbound
Sailing by the Stars: A Survey on Reward Models and Learning Strategies for Learning from Rewards Not All Rollouts are Useful: Down-Sampling Rollouts in LLM Reinforcement Learning
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2394fba1-3d86-4feb-b22a-07787f631ac8 · outbound
Sailing by the Stars: A Survey on Reward Models and Learning Strategies for Learning from Rewards Multimodal RewardBench: Holistic Evaluation of Reward Models for Vision Language Models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba5e3fe7-c7ea-4b47-9b40-552cf6b86b5d · outbound
Sailing by the Stars: A Survey on Reward Models and Learning Strategies for Learning from Rewards Self-critiquing models for assisting human evaluators
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d380dc8c-6a17-472f-89ba-b03aa0e6d3e8 · outbound
Sailing by the Stars: A Survey on Reward Models and Learning Strategies for Learning from Rewards The Effects of Reward Misspecification: Mapping and Mitigating Misaligned Models
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c64dff9b-1e48-44a6-97d1-c7c9e877a38b · outbound
Sailing by the Stars: A Survey on Reward Models and Learning Strategies for Learning from Rewards Tuning Large Multimodal Models for Videos using Reinforcement Learning from AI Feedback
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3d93e4c-8845-44e2-8a86-98ef30b624a8 · outbound
Sailing by the Stars: A Survey on Reward Models and Learning Strategies for Learning from Rewards Critique-out-Loud Reward Models
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f65b86e6-2d33-46a5-b87e-37a75716c1b9 · outbound
Sailing by the Stars: A Survey on Reward Models and Learning Strategies for Learning from Rewards rStar-Math: Small LLMs Can Master Math Reasoning with Self-Evolved Deep Thinking
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc987656-7d73-4688-9723-f97ce0a6b9d0 · inbound
AntiLeakBench: Preventing Data Contamination by Automatically Constructing Benchmarks with Updated Real-World Knowledge Sailing by the Stars: A Survey on Reward Models and Learning Strategies for Learning from Rewards
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 60d052a9-dcb5-45c1-9584-344c8260366a · inbound
SCOPE: Compress Mathematical Reasoning Steps for Efficient Automated Process Annotation Sailing by the Stars: A Survey on Reward Models and Learning Strategies for Learning from Rewards
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b08205ac-ba61-43a9-a7d5-dbcb2efed3b2 · inbound
Training LLMs for EHR-Based Reasoning Tasks via Reinforcement Learning Sailing by the Stars: A Survey on Reward Models and Learning Strategies for Learning from Rewards
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af4c72b2-bc22-40cc-85b2-9cb08932bbd9 · inbound
ReAgent-V: A Reward-Driven Multi-Agent Framework for Video Understanding Sailing by the Stars: A Survey on Reward Models and Learning Strategies for Learning from Rewards
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba0db544-a3dc-49a4-99ab-304f93204ed0 · inbound
Towards Storage-Efficient Visual Document Retrieval: An Empirical Study on Reducing Patch-Level Embeddings Sailing by the Stars: A Survey on Reward Models and Learning Strategies for Learning from Rewards
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9257381-076a-420b-acae-22465db6a9cc · inbound
PEER: Unified Process-Outcome Reinforcement Learning for Structured Empathetic Reasoning Sailing by the Stars: A Survey on Reward Models and Learning Strategies for Learning from Rewards
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation fc378cd4-d001-44b5-98fd-4880a08afbb8 · inbound
Unsupervised Hallucination Detection by Inspecting Reasoning Processes Sailing by the Stars: A Survey on Reward Models and Learning Strategies for Learning from Rewards
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44eb936b-cc01-450d-a763-d02363a1ccd7 · inbound
Trust Region On-Policy Distillation Sailing by the Stars: A Survey on Reward Models and Learning Strategies for Learning from Rewards
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 5ceacf62-af66-4d4e-9ce2-3b095fa951ad · inbound
Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning Sailing by the Stars: A Survey on Reward Models and Learning Strategies for Learning from Rewards
Reference 228
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.