Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T12:59:49.168492Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 0 inbound Pith citation observations for arXiv:2505.23247.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T12:59:49.168492Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
43 of 43 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 77bbfb63-640e-41ac-880e-8354745ec0f1 · outbound
Accelerating RLHF Training with Reward Variance Increase GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5cd1c3b0-8545-40f8-9e76-ab1e7b05cc95 · outbound
Accelerating RLHF Training with Reward Variance Increase Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5df03760-eb3a-4895-bc6e-e7617d7c849e · outbound
Accelerating RLHF Training with Reward Variance Increase Biderman, H
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 009e307f-77ea-4bc4-b906-2638e8a9c2b4 · outbound
Accelerating RLHF Training with Reward Variance Increase On the Opportunities and Risks of Foundation Models
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 873fd355-ecdc-4e96-bc55-42f8314b04fd · outbound
Accelerating RLHF Training with Reward Variance Increase Brown, B
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 478a1110-21f8-48ec-80f5-02d5da0601ff · outbound
Accelerating RLHF Training with Reward Variance Increase Busa-Fekete, B
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c28f5f59-6cf7-43b9-96d3-637382b39030 · outbound
Accelerating RLHF Training with Reward Variance Increase The Accuracy Paradox in RLHF: When Better Reward Models Don't Yield Better Language Models
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 54bf522b-4f9c-4680-a28f-706e37191467 · outbound
Accelerating RLHF Training with Reward Variance Increase Unresolved cited work
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c363396a-0a42-4644-b8a4-9fe8775af82d · outbound
Accelerating RLHF Training with Reward Variance Increase Chujie, S
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b97eb7be-0ff2-497d-988d-50a8efd04ff0 · outbound
Accelerating RLHF Training with Reward Variance Increase Unresolved cited work
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 87614c65-c5fd-4e18-bdb3-896c4ca1e193 · outbound
Accelerating RLHF Training with Reward Variance Increase Unresolved cited work
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 997c44ba-346c-42fc-858a-3666c5699c72 · outbound
Accelerating RLHF Training with Reward Variance Increase Unresolved cited work
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b1a7a1bb-a0c5-47c3-ae39-46bc7d4c07d0 · outbound
Accelerating RLHF Training with Reward Variance Increase Gr¨unbaum, V
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e13f9c7f-146b-428b-ae70-b1175beb77b0 · outbound
Accelerating RLHF Training with Reward Variance Increase DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 61f26049-22d0-4e45-9147-e926ac8b41d2 · outbound
Accelerating RLHF Training with Reward Variance Increase An Overview of Large Language Models for Statisticians
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38fc1af3-733a-47dc-80a4-1c3105557b5d · outbound
Accelerating RLHF Training with Reward Variance Increase RewardBench: Evaluating Reward Models for Language Modeling
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7a1ad49-e044-4c75-b602-64ebd181ce73 · outbound
Accelerating RLHF Training with Reward Variance Increase Unresolved cited work
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 703955ed-8188-4880-bbc3-5b459386055e · outbound
Accelerating RLHF Training with Reward Variance Increase DeepSeek-V3 Technical Report
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51262a20-3615-4009-b319-8df2320193ab · outbound
Accelerating RLHF Training with Reward Variance Increase How do Large Language Models Navigate Conflicts between Honesty and Helpfulness?
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a18dfff-a744-402e-a18a-80706f1dd73e · outbound
Accelerating RLHF Training with Reward Variance Increase Understanding R1-Zero-Like Training: A Critical Perspective
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b06221c-1394-4095-9e18-8c0c2f95f4fd · outbound
Accelerating RLHF Training with Reward Variance Increase Ouyang, J
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 64e7dd7c-78dd-4bae-8327-e00e335e3915 · outbound
Accelerating RLHF Training with Reward Variance Increase Radford, K
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 75223c18-a2d3-4934-943e-de025d093d7c · outbound
Accelerating RLHF Training with Reward Variance Increase Rafailov, A
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cfbcce77-e685-4005-8254-0f9945216123 · outbound
Accelerating RLHF Training with Reward Variance Increase Razin, Z
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation affcfb43-06b0-4c29-b9f0-a9397c72fec3 · outbound
Accelerating RLHF Training with Reward Variance Increase High-Dimensional Continuous Control Using Generalized Advantage Estimation
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f4287c15-8172-459b-863a-289302a6ce41 · outbound
Accelerating RLHF Training with Reward Variance Increase Proximal Policy Optimization Algorithms
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 119a85bf-f936-4372-865c-ad5fec2e324e · outbound
Accelerating RLHF Training with Reward Variance Increase DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef203d19-aad7-492f-8405-b79b262047f3 · outbound
Accelerating RLHF Training with Reward Variance Increase Understanding the performance gap between online and offline alignment algorithms
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7bb1405-a6c5-4f81-b012-9c8576dee5c9 · outbound
Accelerating RLHF Training with Reward Variance Increase Gemini: A Family of Highly Capable Multimodal Models
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 126dde1d-c5fd-4ec2-810a-70f7a6e82833 · outbound
Accelerating RLHF Training with Reward Variance Increase LLaMA: Open and Efficient Foundation Language Models
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c098633a-f25b-4ad0-bd18-8f66ba24e76e · outbound
Accelerating RLHF Training with Reward Variance Increase Towards Safety and Helpfulness Balanced Responses via Controllable Large Language Models
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4be40304-6593-490b-9633-e738ab1320a2 · outbound
Accelerating RLHF Training with Reward Variance Increase Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ffdd6de7-4d7a-4546-80f4-97f8a1324b4b · outbound
Accelerating RLHF Training with Reward Variance Increase Balancing Truthfulness and Informativeness with Uncertainty-Aware Instruction Fine-Tuning
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e5783d80-1ead-4e57-9135-51ec6639f0f9 · outbound
Accelerating RLHF Training with Reward Variance Increase Unresolved cited work
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5a983e2c-f021-444c-964b-9088cb1550be · outbound
Accelerating RLHF Training with Reward Variance Increase Not All Rollouts are Useful: Down-Sampling Rollouts in LLM Reinforcement Learning
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb94bad4-608f-4d77-a0c3-982c36af98d4 · outbound
Accelerating RLHF Training with Reward Variance Increase Qwen3 Technical Report
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89d871bd-b7f4-4f56-9a54-030a5549edef · outbound
Accelerating RLHF Training with Reward Variance Increase Unresolved cited work
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation aa45b29e-9039-478e-b74b-7f210b2826c2 · outbound
Accelerating RLHF Training with Reward Variance Increase DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4bc96b5a-387e-48fa-9098-d040458b63c4 · outbound
Accelerating RLHF Training with Reward Variance Increase Zhang and C
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6882edae-8191-403f-89d4-5bfd3b936fea · outbound
Accelerating RLHF Training with Reward Variance Increase A Survey of Large Language Models
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b043d61d-bef4-4eca-8649-e5a056901647 · outbound
Accelerating RLHF Training with Reward Variance Increase Secrets of RLHF in Large Language Models Part I: PPO
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 61a6078d-05bc-40bf-9c16-cd974259432c · outbound
Accelerating RLHF Training with Reward Variance Increase Fine-Tuning Language Models from Human Preferences
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d408927-a50c-40c7-8209-de3ac31fc322 · outbound
Accelerating RLHF Training with Reward Variance Increase Unresolved cited work
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
No inbound Pith citation observations are available.