Pith. sign in

Paper Citation Record · LEDGER

Accelerating RLHF Training with Reward Variance Increase

As of 8 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 0 inbound Pith citation observations for arXiv:2505.23247.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.23247 v2

Coverage vector

measured 43 of 43 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:59:49.168492Z

measured 43 of 43 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

43 of 43 outbound references displayed

  • verified exact2
  • verified fuzzy8
  • unresolved33
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 77bbfb63-640e-41ac-880e-8354745ec0f1 · outbound

This paper cites GPT-4 Technical Report.

Accelerating RLHF Training with Reward Variance Increase GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:44.286507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:44.286507Z digest=sha256:9d73c834658fb2c1745146ece0929028e9c89a547244f3c1207242fa23cb2fbe

Observation 5cd1c3b0-8545-40f8-9e76-ab1e7b05cc95 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Accelerating RLHF Training with Reward Variance Increase Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:44.338823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:44.338823Z digest=sha256:666687d274ed4d9f53f3a46d201bd0d6eb020b0b9f93bb1685545b4fbd4680e0

Observation 5df03760-eb3a-4895-bc6e-e7617d7c849e · outbound

This paper cites Biderman, H.

Accelerating RLHF Training with Reward Variance Increase Biderman, H

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:52.930185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:59:44.398065Z digest=sha256:b6603131814c06244f704abb6223d995de0737c6457fb39ecb8b0c1b7d64f0fd

Observation 009e307f-77ea-4bc4-b906-2638e8a9c2b4 · outbound

This paper cites On the Opportunities and Risks of Foundation Models.

Accelerating RLHF Training with Reward Variance Increase On the Opportunities and Risks of Foundation Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:44.493191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:44.493191Z digest=sha256:8a3744e0991cfb1797522386f1c8f89c1003a2b9e2848b4c02f34e1decafab78

Observation 873fd355-ecdc-4e96-bc55-42f8314b04fd · outbound

This paper cites Brown, B.

Accelerating RLHF Training with Reward Variance Increase Brown, B

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:52.702380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:59:44.558034Z digest=sha256:749f291864a507f702ad65265d30807b83320774a9ad7cb2b4d5c73b3d5abdde

Observation 478a1110-21f8-48ec-80f5-02d5da0601ff · outbound

This paper cites Busa-Fekete, B.

Accelerating RLHF Training with Reward Variance Increase Busa-Fekete, B

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:52.492367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:59:44.669918Z digest=sha256:bf73a175c52ba54dac2e6c04a4e25ac26813699baf15430d7c0d9cd1cd2612c6

Observation c28f5f59-6cf7-43b9-96d3-637382b39030 · outbound

This paper cites The Accuracy Paradox in RLHF: When Better Reward Models Don't Yield Better Language Models.

Accelerating RLHF Training with Reward Variance Increase The Accuracy Paradox in RLHF: When Better Reward Models Don't Yield Better Language Models

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:59:50.084647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:59:44.776384Z digest=sha256:05acec4b4ccdbfc76f290d487fef251392ea07881dc71a5749d3a227894a537f

Observation 54bf522b-4f9c-4680-a28f-706e37191467 · outbound

This paper cites an unresolved cited work.

Accelerating RLHF Training with Reward Variance Increase Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:59:52.339589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:59:44.880939Z digest=sha256:36219844f33a595a2e1598d1a27d0376076e4c1e04af0f56047503d276c9e77f

Observation c363396a-0a42-4644-b8a4-9fe8775af82d · outbound

This paper cites Chujie, S.

Accelerating RLHF Training with Reward Variance Increase Chujie, S

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:52.201484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:59:44.987046Z digest=sha256:682901e821c9d2c8ab416c30ed6f4eb942c53adbfb76e4e8fc71e8ad956cf852

Observation b97eb7be-0ff2-497d-988d-50a8efd04ff0 · outbound

This paper cites an unresolved cited work.

Accelerating RLHF Training with Reward Variance Increase Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:59:51.988803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:59:45.089926Z digest=sha256:c845210cca658e9dce3542ea02703635869847cec1c35a86fc65937b9ad7a6e4

Observation 87614c65-c5fd-4e18-bdb3-896c4ca1e193 · outbound

This paper cites an unresolved cited work.

Accelerating RLHF Training with Reward Variance Increase Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:59:51.807241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:59:45.227730Z digest=sha256:1aa10103beec72337ea93c6d9ef9684ba4e1b38da251afa5e632df124979295f

Observation 997c44ba-346c-42fc-858a-3666c5699c72 · outbound

This paper cites an unresolved cited work.

Accelerating RLHF Training with Reward Variance Increase Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:59:51.534506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:59:45.334159Z digest=sha256:4b2ed4e5959062fd152be11e8f221255f6be38a38476371b77144f34b896b801

Observation b1a7a1bb-a0c5-47c3-ae39-46bc7d4c07d0 · outbound

This paper cites Gr¨unbaum, V.

Accelerating RLHF Training with Reward Variance Increase Gr¨unbaum, V

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:51.371181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:59:45.466150Z digest=sha256:3c9631fda2ce8c1c462f9d03302b0389aac2652ffc0ae950a06711be5e42e43f

Observation e13f9c7f-146b-428b-ae70-b1175beb77b0 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Accelerating RLHF Training with Reward Variance Increase DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:45.553164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:45.553164Z digest=sha256:68db8d4007f01e7d9cabc54d029bf67034d1601c4b64aac79a2c8d6fa71be8e3

Observation 61f26049-22d0-4e45-9147-e926ac8b41d2 · outbound

This paper cites An Overview of Large Language Models for Statisticians.

Accelerating RLHF Training with Reward Variance Increase An Overview of Large Language Models for Statisticians

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:45.680041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:45.680041Z digest=sha256:96a2efd6be43242be1c2facaa77fc5b07690bb1e89b08ae018345dddc617134e

Observation 38fc1af3-733a-47dc-80a4-1c3105557b5d · outbound

This paper cites RewardBench: Evaluating Reward Models for Language Modeling.

Accelerating RLHF Training with Reward Variance Increase RewardBench: Evaluating Reward Models for Language Modeling

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:45.856856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:45.856856Z digest=sha256:c478a3b3fbb3f42e53e9cf13ca8c04e066c8e830c70338ee8a75b14441616088

Observation d7a1ad49-e044-4c75-b602-64ebd181ce73 · outbound

This paper cites an unresolved cited work.

Accelerating RLHF Training with Reward Variance Increase Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:45.973358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:45.973358Z digest=sha256:b1f651eef1a28e801fa6417540840a7a0c122fc12c7119a13f926f7f167d6378

Observation 703955ed-8188-4880-bbc3-5b459386055e · outbound

This paper cites DeepSeek-V3 Technical Report.

Accelerating RLHF Training with Reward Variance Increase DeepSeek-V3 Technical Report

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:46.078140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:46.078140Z digest=sha256:840b2b7118a981c7f3ae4592da188d9bd5641f740b9a0be83aedbe2d8245858a

Observation 51262a20-3615-4009-b319-8df2320193ab · outbound

This paper cites How do Large Language Models Navigate Conflicts between Honesty and Helpfulness?.

Accelerating RLHF Training with Reward Variance Increase How do Large Language Models Navigate Conflicts between Honesty and Helpfulness?

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:46.192236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:46.192236Z digest=sha256:2a06d67aacde2baf014f30b11062efdc45116337916e86e8130e628cb6ee59b4

Observation 2a18dfff-a744-402e-a18a-80706f1dd73e · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

Accelerating RLHF Training with Reward Variance Increase Understanding R1-Zero-Like Training: A Critical Perspective

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:46.332920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:46.332920Z digest=sha256:d774fe9bdfd4cd2831d486bee4f5b16e7fbeff4ab0e0bfdcd2073eb778691f26

Observation 3b06221c-1394-4095-9e18-8c0c2f95f4fd · outbound

This paper cites Ouyang, J.

Accelerating RLHF Training with Reward Variance Increase Ouyang, J

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:51.205344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:59:46.436502Z digest=sha256:787515080cebe8d5299818c18ae93f85e5923f69fd4f138a652b250c66750f14

Observation 64e7dd7c-78dd-4bae-8327-e00e335e3915 · outbound

This paper cites Radford, K.

Accelerating RLHF Training with Reward Variance Increase Radford, K

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:50.956294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:59:46.506583Z digest=sha256:991e4554afcf50326fa9a2ee9cf289673a612c8a5f6c624eb7b8286fa96fe9dc

Observation 75223c18-a2d3-4934-943e-de025d093d7c · outbound

This paper cites Rafailov, A.

Accelerating RLHF Training with Reward Variance Increase Rafailov, A

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:59:50.800422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:59:46.639735Z digest=sha256:789baab21d8772b240884e8ad99bfc2279d272cffec0b2539284c27c9f92df86

Observation cfbcce77-e685-4005-8254-0f9945216123 · outbound

This paper cites Razin, Z.

Accelerating RLHF Training with Reward Variance Increase Razin, Z

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:46.741897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:46.741897Z digest=sha256:b8d4cc16ecf832f9f185c55554fb562ec41238473ef7f8d12819cea595b7ef7b

Observation affcfb43-06b0-4c29-b9f0-a9397c72fec3 · outbound

This paper cites High-Dimensional Continuous Control Using Generalized Advantage Estimation.

Accelerating RLHF Training with Reward Variance Increase High-Dimensional Continuous Control Using Generalized Advantage Estimation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:46.842272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:46.842272Z digest=sha256:633471f330bfdfae5760b983380f055c53bb251e5aa2f4ad02c46af413edf98d

Observation f4287c15-8172-459b-863a-289302a6ce41 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Accelerating RLHF Training with Reward Variance Increase Proximal Policy Optimization Algorithms

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:46.991866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:46.991866Z digest=sha256:3217ea0bc333055f7c064aa8139156b1bc5d8ad105318cf489d07ce54473ad0c

Observation 119a85bf-f936-4372-865c-ad5fec2e324e · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Accelerating RLHF Training with Reward Variance Increase DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:47.080154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:47.080154Z digest=sha256:a73e120a5d16bd0378ed7524f1e9fd47bbcd5f0cd0ea6936458f7d96a8ba8817

Observation ef203d19-aad7-492f-8405-b79b262047f3 · outbound

This paper cites Understanding the performance gap between online and offline alignment algorithms.

Accelerating RLHF Training with Reward Variance Increase Understanding the performance gap between online and offline alignment algorithms

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:47.229540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:47.229540Z digest=sha256:08ca25934bdaa4b61bbcc03746aa93ee0f2ff203904c3a765644e8b2ea395874

Observation e7bb1405-a6c5-4f81-b012-9c8576dee5c9 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Accelerating RLHF Training with Reward Variance Increase Gemini: A Family of Highly Capable Multimodal Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:47.359724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:47.359724Z digest=sha256:8f3187edf8e24786d3726574bff9da0318cc97d694fea8286726c8b3653601d7

Observation 126dde1d-c5fd-4ec2-810a-70f7a6e82833 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Accelerating RLHF Training with Reward Variance Increase LLaMA: Open and Efficient Foundation Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:47.506865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:47.506865Z digest=sha256:60e277b44ce2851a61dbed08efe253cda3ab1130433fc8defb8ed51c317d3360

Observation c098633a-f25b-4ad0-bd18-8f66ba24e76e · outbound

This paper cites Towards Safety and Helpfulness Balanced Responses via Controllable Large Language Models.

Accelerating RLHF Training with Reward Variance Increase Towards Safety and Helpfulness Balanced Responses via Controllable Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:47.666105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:47.666105Z digest=sha256:77bc1d5d0850925a165b106ea32ec925b54e6354aa5952576f6411aefb331114

Observation 4be40304-6593-490b-9633-e738ab1320a2 · outbound

This paper cites Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts.

Accelerating RLHF Training with Reward Variance Increase Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:47.775977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:47.775977Z digest=sha256:b7ead964d3298b5e5ad257458de6f024e0cc19d42e9a7c04cf92d6cfe86b460b

Observation ffdd6de7-4d7a-4546-80f4-97f8a1324b4b · outbound

This paper cites Balancing Truthfulness and Informativeness with Uncertainty-Aware Instruction Fine-Tuning.

Accelerating RLHF Training with Reward Variance Increase Balancing Truthfulness and Informativeness with Uncertainty-Aware Instruction Fine-Tuning

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:59:49.583436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:59:47.876252Z digest=sha256:350129bd5db864678d8c7fbdfaecd860be292e9415a6aa6b787f7c52211b7196

Observation e5783d80-1ead-4e57-9135-51ec6639f0f9 · outbound

This paper cites an unresolved cited work.

Accelerating RLHF Training with Reward Variance Increase Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:59:50.617804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:59:48.013858Z digest=sha256:b23e56725fcdd7fc68f4ff61213f426bc6d452c12a2a37aeca04ae0f391aa99b

Observation 5a983e2c-f021-444c-964b-9088cb1550be · outbound

This paper cites Not All Rollouts are Useful: Down-Sampling Rollouts in LLM Reinforcement Learning.

Accelerating RLHF Training with Reward Variance Increase Not All Rollouts are Useful: Down-Sampling Rollouts in LLM Reinforcement Learning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:48.111769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:48.111769Z digest=sha256:183bfbe0d802013a17c7b4ecfa07dc173ca9f6296d862d5b5f623a5420ceb9d8

Observation fb94bad4-608f-4d77-a0c3-982c36af98d4 · outbound

This paper cites Qwen3 Technical Report.

Accelerating RLHF Training with Reward Variance Increase Qwen3 Technical Report

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:48.283646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:48.283646Z digest=sha256:c60bb77e2af206fc546ec1b5050199baa8468a6a1f425ff30201260a5a6849a1

Observation 89d871bd-b7f4-4f56-9a54-030a5549edef · outbound

This paper cites an unresolved cited work.

Accelerating RLHF Training with Reward Variance Increase Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:59:50.416612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:59:48.421561Z digest=sha256:f2a310f5a2777bbcea7dafaeaa8c23fc5196fe45781d11aa05a7099b404224d3

Observation aa45b29e-9039-478e-b74b-7f210b2826c2 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Accelerating RLHF Training with Reward Variance Increase DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:48.554997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:48.554997Z digest=sha256:57923003254ad683c88bda3b113925bd1090a6363e8219834de2581687ad87bf

Observation 4bc96b5a-387e-48fa-9098-d040458b63c4 · outbound

This paper cites Zhang and C.

Accelerating RLHF Training with Reward Variance Increase Zhang and C

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:48.646323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:48.646323Z digest=sha256:8702463fa62032125b4d0d89e71b90ebb13b51e43e05b2ed18c488920d2ccc25

Observation 6882edae-8191-403f-89d4-5bfd3b936fea · outbound

This paper cites A Survey of Large Language Models.

Accelerating RLHF Training with Reward Variance Increase A Survey of Large Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:48.798191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:48.798191Z digest=sha256:1d35a950d19924a7baf2b1ad6b080411a435e4f28de9d6180ce28239e149753c

Observation b043d61d-bef4-4eca-8649-e5a056901647 · outbound

This paper cites Secrets of RLHF in Large Language Models Part I: PPO.

Accelerating RLHF Training with Reward Variance Increase Secrets of RLHF in Large Language Models Part I: PPO

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:48.911300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:48.911300Z digest=sha256:d26e2e16decfdf16c0955be3dff1084efe1cbd107e25946f1999db960858197d

Observation 61a6078d-05bc-40bf-9c16-cd974259432c · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

Accelerating RLHF Training with Reward Variance Increase Fine-Tuning Language Models from Human Preferences

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:49.025901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:49.025901Z digest=sha256:fdfab2dc9e905034c2e7dc62a4fcd36dbde154a32aa973ca61aa0528e24e6ad7

Observation 4d408927-a50c-40c7-8209-de3ac31fc322 · outbound

This paper cites an unresolved cited work.

Accelerating RLHF Training with Reward Variance Increase Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:59:50.328967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:59:49.168492Z digest=sha256:0696efab1a5d7cb8ce709b2c49f863d9e7641e221abc78311ea56f489922d421

Pith citing papers

No inbound Pith citation observations are available.