Pith. sign in

Paper Citation Record · LEDGER

WARM: On the Benefits of Weight Averaged Reward Models

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 13 inbound Pith citation observations for arXiv:2401.12187.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2401.12187 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 13 of 13 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:01:05.649860Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

3
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation a2b21913-dd8a-401e-bf9e-3fe44b655998 · inbound

Learning a Pessimistic Reward Model in RLHF cites this paper.

Learning a Pessimistic Reward Model in RLHF WARM: On the Benefits of Weight Averaged Reward Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:05.649860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:05.649860Z digest=sha256:a0c4b3b1c9476daff04d0d399ce35db451da4e1c86cfe19dd7fe01f0b1354677

Observation 024add06-7af1-4b65-8d0d-cbf5366dfd3d · inbound

Teach a Reward Model to Correct Itself: Reward Guided Adversarial Failure Discovery for Robust Reward Modeling cites this paper.

Teach a Reward Model to Correct Itself: Reward Guided Adversarial Failure Discovery for Robust Reward Modeling WARM: On the Benefits of Weight Averaged Reward Models

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-19T05:17:05.814125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-19T05:16:22.274580Z digest=sha256:880b6f0f831542ac8baddf58935d503f728386c46b770dbbec18eddabc7a274f

Observation 608ab5de-8333-4751-8d8c-66a9fe2ffdbd · inbound

Bradley-Terry and Multi-Objective Reward Modeling Are Complementary cites this paper.

Bradley-Terry and Multi-Objective Reward Modeling Are Complementary WARM: On the Benefits of Weight Averaged Reward Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:28.316657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:28.316657Z digest=sha256:907f6b04846fb170bfa87df8e7aecc52234da112b734ae80104ae9fa39b50b7a

Observation c96758dd-57fc-40a8-9c3f-1eaab651fd41 · inbound

Tiny Reward Models cites this paper.

Tiny Reward Models WARM: On the Benefits of Weight Averaged Reward Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.178654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.178654Z digest=sha256:20853db715eb10d943d12e814559fcd3568405c818753bbd42b9547c0841cec3

Observation e2c4a7cc-89a7-40d8-941d-96d86ef9ce1e · inbound

Off-Policy Corrected Reward Modeling for Reinforcement Learning from Human Feedback cites this paper.

Off-Policy Corrected Reward Modeling for Reinforcement Learning from Human Feedback WARM: On the Benefits of Weight Averaged Reward Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T15:36:05.446280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:36:05.446280Z digest=sha256:7725b55623c5db899149a08fb3cbb2117a93dfb20f8c4c4823f69e2b9f7db047

Observation 5650bcbd-ea5c-4a08-9b87-388f4fe98063 · inbound

Towards Reliable, Uncertainty-Aware Alignment cites this paper.

Towards Reliable, Uncertainty-Aware Alignment WARM: On the Benefits of Weight Averaged Reward Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:16.470462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:40:16.470462Z digest=sha256:2d6f04acddb38489360aee49fdcdf830c680e922cd1ad2a3568d1940496c355d

Observation b34333c6-eb98-413a-9e3b-ea7b1f8d8977 · inbound

Mitigating Multimodal Hallucination via Phase-wise Self-reward cites this paper.

Mitigating Multimodal Hallucination via Phase-wise Self-reward WARM: On the Benefits of Weight Averaged Reward Models

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:38:43.298125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T05:10:45.144421Z digest=sha256:b05aec6dbc8300c130a1e841be6a511c82945d9593232a3c108f179ae930d830

Observation 4dc419e5-df9d-4ad3-ac14-335e43c02c16 · inbound

Palette: A Modular, Controllable, and Efficient Framework for On-demand Authorized Safety Alignment Relaxation in LLMs cites this paper.

Palette: A Modular, Controllable, and Efficient Framework for On-demand Authorized Safety Alignment Relaxation in LLMs WARM: On the Benefits of Weight Averaged Reward Models

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-06-30T16:04:52.576970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T16:03:12.728352Z digest=sha256:d21a2943607e11a9278dbca707187277d4f28bc05d0146eea33cc93ce94169fd

Observation eb16ecb7-3933-47a1-9e97-7eaa7e4bb1dc · inbound

DynaCF: Mitigating Shortcut Learning in Reward Models via Dynamic Counterfactual Sensitivity cites this paper.

DynaCF: Mitigating Shortcut Learning in Reward Models via Dynamic Counterfactual Sensitivity WARM: On the Benefits of Weight Averaged Reward Models

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-06-27T17:31:06.966629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-27T17:26:17.072017Z digest=sha256:645b2cd63c4c79a2ee34d4cfaba43762862ccb702e06bd2fad505b70f91306f6

Observation 6fc38837-3c15-43a3-bcf9-70fdafaa011f · inbound

Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization cites this paper.

Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization WARM: On the Benefits of Weight Averaged Reward Models

Reference 69

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T01:37:30.528193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-27T16:26:34.918099Z digest=sha256:b0ef0e919c5d0b2494efd9a429d2be181b362e43b534dc26de0736b03898e481

Observation 7de95927-6c1b-4e41-9256-d45d2ccd6eb2 · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay WARM: On the Benefits of Weight Averaged Reward Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-07-11T13:53:36.775836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T13:53:36.775836Z digest=sha256:83f07c83f63fc68de7eeb5404420ad6d4e9d0bbd6584a18725d5e95ad2412e6a

Observation 42f096b6-45c0-42cb-8ad9-9578d5e75310 · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay WARM: On the Benefits of Weight Averaged Reward Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:37.580679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:37.580679Z digest=sha256:a2b9bc76df411a09f5d92ee667dc47211867bd583302ecca992cb5ca390abcb3

Observation d164ed71-f4cc-46b1-90dc-f45e60b3ff22 · inbound

Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL cites this paper.

Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL WARM: On the Benefits of Weight Averaged Reward Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-03T10:57:58.293232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:57:58.293232Z digest=sha256:638bfda537559ec68282aaae90a947dd7dc836e31963d1c22d502c3c7049ab76