Pith. sign in

Paper Citation Record · LEDGER

Accelerating RLHF Training with Reward Variance Increase

As of 8 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 0 inbound Pith citation observations for arXiv:2505.23247.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.23247 v2

Coverage vector

measured 43 of 43 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:59:49.168492Z

measured 43 of 43 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

43 of 43 outbound references displayed

  • verified exact2
  • verified fuzzy8
  • unresolved33
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 77bbfb63-640e-41ac-880e-8354745ec0f1 · outbound

This paper cites GPT-4 Technical Report.

Accelerating RLHF Training with Reward Variance Increase GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:44.286507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:44.286507Z digest=sha256:9d73c834658fb2c1745146ece0929028e9c89a547244f3c1207242fa23cb2fbe

Observation 5cd1c3b0-8545-40f8-9e76-ab1e7b05cc95 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Accelerating RLHF Training with Reward Variance Increase Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:44.338823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:44.338823Z digest=sha256:666687d274ed4d9f53f3a46d201bd0d6eb020b0b9f93bb1685545b4fbd4680e0

Observation 5df03760-eb3a-4895-bc6e-e7617d7c849e · outbound

This paper cites Biderman, H.

Accelerating RLHF Training with Reward Variance Increase Biderman, H

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:52.930185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:59:44.398065Z digest=sha256:494835a54ec734c1d1dda300ca53c9bc490097a68aaeaad951dc2fd906fd837a

Observation 009e307f-77ea-4bc4-b906-2638e8a9c2b4 · outbound

This paper cites On the Opportunities and Risks of Foundation Models.

Accelerating RLHF Training with Reward Variance Increase On the Opportunities and Risks of Foundation Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:44.493191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:44.493191Z digest=sha256:8a3744e0991cfb1797522386f1c8f89c1003a2b9e2848b4c02f34e1decafab78

Observation 873fd355-ecdc-4e96-bc55-42f8314b04fd · outbound

This paper cites Brown, B.

Accelerating RLHF Training with Reward Variance Increase Brown, B

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:52.702380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:59:44.558034Z digest=sha256:d613a14337fca9155313adbbd114ec783b88d8ed2e9384578fba53eb3b2c0fbb

Observation 478a1110-21f8-48ec-80f5-02d5da0601ff · outbound

This paper cites Busa-Fekete, B.

Accelerating RLHF Training with Reward Variance Increase Busa-Fekete, B

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:52.492367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:59:44.669918Z digest=sha256:0591ad2062cc076698aedde7d6bcf95581638c7eed03a3776b930e58de4a95b4

Observation c28f5f59-6cf7-43b9-96d3-637382b39030 · outbound

This paper cites The Accuracy Paradox in RLHF: When Better Reward Models Don't Yield Better Language Models.

Accelerating RLHF Training with Reward Variance Increase The Accuracy Paradox in RLHF: When Better Reward Models Don't Yield Better Language Models

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:59:50.084647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:59:44.776384Z digest=sha256:a4a3b8ef3170deb74ec024f9f52d973d8c7f1ec8d05b97672e06a44a2f2aaa90

Observation 54bf522b-4f9c-4680-a28f-706e37191467 · outbound

This paper cites an unresolved cited work.

Accelerating RLHF Training with Reward Variance Increase Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:59:52.339589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:59:44.880939Z digest=sha256:329df2a41c70f0d9eec027b503a39db54070e5b4d68b013e71f0e0368d60b6c0

Observation c363396a-0a42-4644-b8a4-9fe8775af82d · outbound

This paper cites Chujie, S.

Accelerating RLHF Training with Reward Variance Increase Chujie, S

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:52.201484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:59:44.987046Z digest=sha256:fa367b8e15c91435885f5f905055176cc3015b9aabe55176457956c8cfacfc17

Observation b97eb7be-0ff2-497d-988d-50a8efd04ff0 · outbound

This paper cites an unresolved cited work.

Accelerating RLHF Training with Reward Variance Increase Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:59:51.988803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:59:45.089926Z digest=sha256:be6daf1b48bcde4ae7d4f3a4ceef42ae6346094f22360cb3b89bd5c0a965d486

Observation 87614c65-c5fd-4e18-bdb3-896c4ca1e193 · outbound

This paper cites an unresolved cited work.

Accelerating RLHF Training with Reward Variance Increase Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:59:51.807241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:59:45.227730Z digest=sha256:f290592405149c47a7ff384c3eebdf3b4ba9464615918b8d645f677de8129533

Observation 997c44ba-346c-42fc-858a-3666c5699c72 · outbound

This paper cites an unresolved cited work.

Accelerating RLHF Training with Reward Variance Increase Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:59:51.534506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:59:45.334159Z digest=sha256:1aae3524130ab1b368c2c1f3f2437b2b81a3ee0e9bbec3554ab6cd52cf8f6660

Observation b1a7a1bb-a0c5-47c3-ae39-46bc7d4c07d0 · outbound

This paper cites Gr¨unbaum, V.

Accelerating RLHF Training with Reward Variance Increase Gr¨unbaum, V

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:51.371181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:59:45.466150Z digest=sha256:47cf66a41ac693dc37a63d7797f384bc63081f22f9ae55f5f03927deafc1bbec

Observation e13f9c7f-146b-428b-ae70-b1175beb77b0 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Accelerating RLHF Training with Reward Variance Increase DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:45.553164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:45.553164Z digest=sha256:68db8d4007f01e7d9cabc54d029bf67034d1601c4b64aac79a2c8d6fa71be8e3

Observation 61f26049-22d0-4e45-9147-e926ac8b41d2 · outbound

This paper cites An Overview of Large Language Models for Statisticians.

Accelerating RLHF Training with Reward Variance Increase An Overview of Large Language Models for Statisticians

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:45.680041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:45.680041Z digest=sha256:96a2efd6be43242be1c2facaa77fc5b07690bb1e89b08ae018345dddc617134e

Observation 38fc1af3-733a-47dc-80a4-1c3105557b5d · outbound

This paper cites RewardBench: Evaluating Reward Models for Language Modeling.

Accelerating RLHF Training with Reward Variance Increase RewardBench: Evaluating Reward Models for Language Modeling

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:45.856856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:45.856856Z digest=sha256:c478a3b3fbb3f42e53e9cf13ca8c04e066c8e830c70338ee8a75b14441616088

Observation d7a1ad49-e044-4c75-b602-64ebd181ce73 · outbound

This paper cites an unresolved cited work.

Accelerating RLHF Training with Reward Variance Increase Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:45.973358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:45.973358Z digest=sha256:b1f651eef1a28e801fa6417540840a7a0c122fc12c7119a13f926f7f167d6378

Observation 703955ed-8188-4880-bbc3-5b459386055e · outbound

This paper cites DeepSeek-V3 Technical Report.

Accelerating RLHF Training with Reward Variance Increase DeepSeek-V3 Technical Report

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:46.078140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:46.078140Z digest=sha256:840b2b7118a981c7f3ae4592da188d9bd5641f740b9a0be83aedbe2d8245858a

Observation 51262a20-3615-4009-b319-8df2320193ab · outbound

This paper cites How do Large Language Models Navigate Conflicts between Honesty and Helpfulness?.

Accelerating RLHF Training with Reward Variance Increase How do Large Language Models Navigate Conflicts between Honesty and Helpfulness?

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:46.192236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:46.192236Z digest=sha256:2a06d67aacde2baf014f30b11062efdc45116337916e86e8130e628cb6ee59b4

Observation 2a18dfff-a744-402e-a18a-80706f1dd73e · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

Accelerating RLHF Training with Reward Variance Increase Understanding R1-Zero-Like Training: A Critical Perspective

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:46.332920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:46.332920Z digest=sha256:d774fe9bdfd4cd2831d486bee4f5b16e7fbeff4ab0e0bfdcd2073eb778691f26

Observation 3b06221c-1394-4095-9e18-8c0c2f95f4fd · outbound

This paper cites Ouyang, J.

Accelerating RLHF Training with Reward Variance Increase Ouyang, J

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:51.205344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:59:46.436502Z digest=sha256:325cf8d3c39b155d5c98485a18fa925b93b9569021a9111ab6c9f40534e4f121

Observation 64e7dd7c-78dd-4bae-8327-e00e335e3915 · outbound

This paper cites Radford, K.

Accelerating RLHF Training with Reward Variance Increase Radford, K

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:50.956294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:59:46.506583Z digest=sha256:296f512e6d5c05419b9a9622516ec78cf580b7f682e71163b1da2e71be64aa70

Observation 75223c18-a2d3-4934-943e-de025d093d7c · outbound

This paper cites Rafailov, A.

Accelerating RLHF Training with Reward Variance Increase Rafailov, A

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:50.800422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:59:46.639735Z digest=sha256:8664b5d4ab8b45bfd019097263679d55c0b80c042b956c1a8b3778729a28c9c4

Observation cfbcce77-e685-4005-8254-0f9945216123 · outbound

This paper cites Razin, Z.

Accelerating RLHF Training with Reward Variance Increase Razin, Z

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:46.741897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:46.741897Z digest=sha256:b8d4cc16ecf832f9f185c55554fb562ec41238473ef7f8d12819cea595b7ef7b

Observation affcfb43-06b0-4c29-b9f0-a9397c72fec3 · outbound

This paper cites High-Dimensional Continuous Control Using Generalized Advantage Estimation.

Accelerating RLHF Training with Reward Variance Increase High-Dimensional Continuous Control Using Generalized Advantage Estimation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:46.842272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:46.842272Z digest=sha256:633471f330bfdfae5760b983380f055c53bb251e5aa2f4ad02c46af413edf98d

Observation f4287c15-8172-459b-863a-289302a6ce41 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Accelerating RLHF Training with Reward Variance Increase Proximal Policy Optimization Algorithms

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:46.991866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:46.991866Z digest=sha256:3217ea0bc333055f7c064aa8139156b1bc5d8ad105318cf489d07ce54473ad0c

Observation 119a85bf-f936-4372-865c-ad5fec2e324e · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Accelerating RLHF Training with Reward Variance Increase DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:47.080154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:47.080154Z digest=sha256:a73e120a5d16bd0378ed7524f1e9fd47bbcd5f0cd0ea6936458f7d96a8ba8817

Observation ef203d19-aad7-492f-8405-b79b262047f3 · outbound

This paper cites Understanding the performance gap between online and offline alignment algorithms.

Accelerating RLHF Training with Reward Variance Increase Understanding the performance gap between online and offline alignment algorithms

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:47.229540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:47.229540Z digest=sha256:08ca25934bdaa4b61bbcc03746aa93ee0f2ff203904c3a765644e8b2ea395874

Observation e7bb1405-a6c5-4f81-b012-9c8576dee5c9 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Accelerating RLHF Training with Reward Variance Increase Gemini: A Family of Highly Capable Multimodal Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:47.359724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:47.359724Z digest=sha256:8f3187edf8e24786d3726574bff9da0318cc97d694fea8286726c8b3653601d7

Observation 126dde1d-c5fd-4ec2-810a-70f7a6e82833 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Accelerating RLHF Training with Reward Variance Increase LLaMA: Open and Efficient Foundation Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:47.506865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:47.506865Z digest=sha256:60e277b44ce2851a61dbed08efe253cda3ab1130433fc8defb8ed51c317d3360

Observation c098633a-f25b-4ad0-bd18-8f66ba24e76e · outbound

This paper cites Towards Safety and Helpfulness Balanced Responses via Controllable Large Language Models.

Accelerating RLHF Training with Reward Variance Increase Towards Safety and Helpfulness Balanced Responses via Controllable Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:47.666105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:47.666105Z digest=sha256:77bc1d5d0850925a165b106ea32ec925b54e6354aa5952576f6411aefb331114

Observation 4be40304-6593-490b-9633-e738ab1320a2 · outbound

This paper cites Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts.

Accelerating RLHF Training with Reward Variance Increase Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:47.775977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:47.775977Z digest=sha256:b7ead964d3298b5e5ad257458de6f024e0cc19d42e9a7c04cf92d6cfe86b460b

Observation ffdd6de7-4d7a-4546-80f4-97f8a1324b4b · outbound

This paper cites Balancing Truthfulness and Informativeness with Uncertainty-Aware Instruction Fine-Tuning.

Accelerating RLHF Training with Reward Variance Increase Balancing Truthfulness and Informativeness with Uncertainty-Aware Instruction Fine-Tuning

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:59:49.583436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:59:47.876252Z digest=sha256:225be229048d90b18f2fcd49e502648edd5f80d002f61fe9597aea0e0a79cab6

Observation e5783d80-1ead-4e57-9135-51ec6639f0f9 · outbound

This paper cites an unresolved cited work.

Accelerating RLHF Training with Reward Variance Increase Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:59:50.617804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:59:48.013858Z digest=sha256:1aa9280864af52e3c35b195e8ff897d369b23f35021568fb7bafca519eb63a4c

Observation 5a983e2c-f021-444c-964b-9088cb1550be · outbound

This paper cites Not All Rollouts are Useful: Down-Sampling Rollouts in LLM Reinforcement Learning.

Accelerating RLHF Training with Reward Variance Increase Not All Rollouts are Useful: Down-Sampling Rollouts in LLM Reinforcement Learning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:48.111769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:48.111769Z digest=sha256:183bfbe0d802013a17c7b4ecfa07dc173ca9f6296d862d5b5f623a5420ceb9d8

Observation fb94bad4-608f-4d77-a0c3-982c36af98d4 · outbound

This paper cites Qwen3 Technical Report.

Accelerating RLHF Training with Reward Variance Increase Qwen3 Technical Report

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:48.283646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:48.283646Z digest=sha256:c60bb77e2af206fc546ec1b5050199baa8468a6a1f425ff30201260a5a6849a1

Observation 89d871bd-b7f4-4f56-9a54-030a5549edef · outbound

This paper cites an unresolved cited work.

Accelerating RLHF Training with Reward Variance Increase Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:59:50.416612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:59:48.421561Z digest=sha256:0dd2266f895568fb43f8b5ae34b78c7aadd3428570ef3667e686d91180b81bd9

Observation aa45b29e-9039-478e-b74b-7f210b2826c2 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Accelerating RLHF Training with Reward Variance Increase DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:48.554997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:48.554997Z digest=sha256:57923003254ad683c88bda3b113925bd1090a6363e8219834de2581687ad87bf

Observation 4bc96b5a-387e-48fa-9098-d040458b63c4 · outbound

This paper cites Zhang and C.

Accelerating RLHF Training with Reward Variance Increase Zhang and C

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:48.646323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:48.646323Z digest=sha256:8702463fa62032125b4d0d89e71b90ebb13b51e43e05b2ed18c488920d2ccc25

Observation 6882edae-8191-403f-89d4-5bfd3b936fea · outbound

This paper cites A Survey of Large Language Models.

Accelerating RLHF Training with Reward Variance Increase A Survey of Large Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:48.798191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:48.798191Z digest=sha256:1d35a950d19924a7baf2b1ad6b080411a435e4f28de9d6180ce28239e149753c

Observation b043d61d-bef4-4eca-8649-e5a056901647 · outbound

This paper cites Secrets of RLHF in Large Language Models Part I: PPO.

Accelerating RLHF Training with Reward Variance Increase Secrets of RLHF in Large Language Models Part I: PPO

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:48.911300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:48.911300Z digest=sha256:d26e2e16decfdf16c0955be3dff1084efe1cbd107e25946f1999db960858197d

Observation 61a6078d-05bc-40bf-9c16-cd974259432c · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

Accelerating RLHF Training with Reward Variance Increase Fine-Tuning Language Models from Human Preferences

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:49.025901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:49.025901Z digest=sha256:fdfab2dc9e905034c2e7dc62a4fcd36dbde154a32aa973ca61aa0528e24e6ad7

Observation 4d408927-a50c-40c7-8209-de3ac31fc322 · outbound

This paper cites an unresolved cited work.

Accelerating RLHF Training with Reward Variance Increase Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:59:50.328967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:59:49.168492Z digest=sha256:5150be26ec61497555654a18cd6d1ab5721e3c38ca2d408ce1e5465a4425ce35

Pith citing papers

No inbound Pith citation observations are available.