Pith. sign in

Paper Citation Record · LEDGER

Reward Shaping to Mitigate Reward Hacking in RLHF

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 22 inbound Pith citation observations for arXiv:2502.18770.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.18770 v6

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 22 of 22 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 22 of 22 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:21:41.769204Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation d48a086d-209a-4c6b-8926-5cf6b54b797e · inbound

Supervising the search process produces reliable and generalizable information-seeking agents cites this paper.

Supervising the search process produces reliable and generalizable information-seeking agents Reward Shaping to Mitigate Reward Hacking in RLHF

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-08-07T01:35:47.288139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-23T02:18:27.204122Z digest=sha256:bc7b2f1a41729a0035354e11eea06175dd8e60eca398e1d430e134016ccde5ea

Observation 58c6321b-7665-4423-8b27-86437a460e6d · inbound

From System 1 to System 2: A Survey of Reasoning Large Language Models cites this paper.

From System 1 to System 2: A Survey of Reasoning Large Language Models Reward Shaping to Mitigate Reward Hacking in RLHF

Reference 170

Resolution
verified exact
arxiv_id, observed 2026-08-07T01:35:47.288139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T01:36:23.845366Z digest=sha256:c1e0601213189c688c631a9eac98e3683b9a44a6eed906accb9061503b14be08

Observation e947c3dd-4fd3-4e12-876b-9d57f2af9dbf · inbound

Enhancing Tool Learning in Large Language Models with Hierarchical Error Checklists cites this paper.

Enhancing Tool Learning in Large Language Models with Hierarchical Error Checklists Reward Shaping to Mitigate Reward Hacking in RLHF

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T13:21:41.769204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:21:41.769204Z digest=sha256:2769b1fe4cc854adf4aea7628ec11641cfcbd69f236b5a5ee1f2bea96ce59d59

Observation 5c3df9b4-277a-4135-b4c4-75ffb9785030 · inbound

RewardAnything: Generalizable Principle-Following Reward Models cites this paper.

RewardAnything: Generalizable Principle-Following Reward Models Reward Shaping to Mitigate Reward Hacking in RLHF

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:06.897214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:04:06.897214Z digest=sha256:914a96da2ce851789bf0d3b7805d6a163ebe111d88c22cd6532383e5d8aa90da

Observation f1f2f615-d4f8-4f43-a686-c92943506c4b · inbound

Self-Rewarding Vision-Language Model via Reasoning Decomposition cites this paper.

Self-Rewarding Vision-Language Model via Reasoning Decomposition Reward Shaping to Mitigate Reward Hacking in RLHF

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-08-07T01:35:47.288139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T21:03:31.606674Z digest=sha256:d25ee244863fd29e60afec293a8436bedb3eef98cc1baff83493df59430bddef

Observation 1e5c7b63-c8d2-41ef-ba7f-41347e6abf8c · inbound

Factored Causal Representation Learning for Robust Reward Modeling in RLHF cites this paper.

Factored Causal Representation Learning for Robust Reward Modeling in RLHF Reward Shaping to Mitigate Reward Hacking in RLHF

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-08-07T01:35:47.288139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T14:18:33.768962Z digest=sha256:8db528851a0383a17b6caaba8bb483116a9a011c9c8a2e52a2234037ae1fd02a

Observation f423b48c-8cd7-4e5b-85fb-08c49e2b8e17 · inbound

Rethinking the Comparison Unit in Sequence-Level Reinforcement Learning: An Equal-Length Paired Training Framework from Loss Correction to Sample Construction cites this paper.

Rethinking the Comparison Unit in Sequence-Level Reinforcement Learning: An Equal-Length Paired Training Framework from Loss Correction to Sample Construction Reward Shaping to Mitigate Reward Hacking in RLHF

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-08-07T01:35:47.288139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-10T06:23:14.905706Z digest=sha256:2b21627ba29a285b3ceeda1ce9dd0d436880b1ff5a17eea270d9f5dbe256f9fc

Observation e0aa7ce6-4486-46d5-b0d7-cff2b43fe244 · inbound

Measure Twice, Click Once: Co-evolving Proposer and Visual Critic via Reinforcement Learning for GUI Grounding cites this paper.

Measure Twice, Click Once: Co-evolving Proposer and Visual Critic via Reinforcement Learning for GUI Grounding Reward Shaping to Mitigate Reward Hacking in RLHF

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-08-07T01:35:47.288139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-09T23:05:05.251150Z digest=sha256:bbb25b469d8a179dd289dfdb1faff4d6167611aacc89e71b58668345edb7ac23

Observation a496afa7-0edb-44c0-95fa-ff04bd50f10b · inbound

Towards Robust LLM Post-Training: Automatic Failure Management for Reinforcement Fine-Tuning cites this paper.

Towards Robust LLM Post-Training: Automatic Failure Management for Reinforcement Fine-Tuning Reward Shaping to Mitigate Reward Hacking in RLHF

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-08-07T01:35:47.288139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T18:19:41.302164Z digest=sha256:0d186a62f24532ae97d297986c264450be684c1b3f9c6314808e030e5594b99c

Observation 31879369-3523-4a1d-82b4-7a711302bdb3 · inbound

Optimal Transport for LLM Reward Modeling from Noisy Preference cites this paper.

Optimal Transport for LLM Reward Modeling from Noisy Preference Reward Shaping to Mitigate Reward Hacking in RLHF

Reference 260

Resolution
metadata mismatch
arxiv_id, observed 2026-08-07T01:35:47.288139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-08T14:10:49.634358Z digest=sha256:c405cb848b1597f82236469c66b5563560f73d1083d7e71523208cbdef6d50e3

Observation 28cb54ed-920c-419c-8d23-2482c8cc6dad · inbound

Latent Personality Alignment: Improving Harmlessness Without Mentioning Harms cites this paper.

Latent Personality Alignment: Improving Harmlessness Without Mentioning Harms Reward Shaping to Mitigate Reward Hacking in RLHF

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-08-07T01:35:47.288139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T01:46:49.586630Z digest=sha256:71aa09537b9c2e2bffcc11a3c89700776b52376a34d86c61aa3aa0d35dddbf70

Observation 4409e985-2190-4e31-b2d7-ef560d706b26 · inbound

Variance-aware Reward Modeling with Anchor Guidance cites this paper.

Variance-aware Reward Modeling with Anchor Guidance Reward Shaping to Mitigate Reward Hacking in RLHF

Reference 49

Resolution
metadata mismatch
arxiv_id, observed 2026-08-07T01:35:47.288139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-13T05:11:05.546092Z digest=sha256:f1cd6a6643a46d7fbaa0f2f601585cac114d5a1ffd139d326a5b25830a5ce7e4

Observation b3fc049c-ece3-453c-8748-63065fc9883b · inbound

Reward Hacking in Rubric-Based Reinforcement Learning cites this paper.

Reward Hacking in Rubric-Based Reinforcement Learning Reward Shaping to Mitigate Reward Hacking in RLHF

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-08-07T01:35:47.288139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T04:08:51.578772Z digest=sha256:03ab166d4143154d7745462528b077269b97d4a8621d1a95b98790ddcdb8aff8

Observation c98675b8-a2c6-45e8-bd29-ddb4c5b20399 · inbound

Diagnosing Training Inference Mismatch in LLM Reinforcement Learning cites this paper.

Diagnosing Training Inference Mismatch in LLM Reinforcement Learning Reward Shaping to Mitigate Reward Hacking in RLHF

Reference 51

Resolution
metadata mismatch
arxiv_id, observed 2026-08-07T01:35:47.288139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-15T02:05:44.813341Z digest=sha256:8a7ddf02cbf2fb1d950d4f5bf372d445239930f5fad68ef1baea201d4de29ecc

Observation f03a1e6c-f8f9-45b8-b88b-552164449677 · inbound

General Preference Reinforcement Learning cites this paper.

General Preference Reinforcement Learning Reward Shaping to Mitigate Reward Hacking in RLHF

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-08-07T01:35:47.288139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T12:43:56.522345Z digest=sha256:41038365af47fade531c307eb6c80cfc2f25fdb4015f10c2bddd213b89ed425c

Observation c4ec532b-ea35-4495-8772-1c98695e82ae · inbound

General Preference Reinforcement Learning cites this paper.

General Preference Reinforcement Learning Reward Shaping to Mitigate Reward Hacking in RLHF

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-08-07T01:35:47.288139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T07:50:00.963837Z digest=sha256:2474fc458589e815641b3ae5bb369f1be360704f7b3522ba7a6c366da0fdb671

Observation 67adf7ce-7aa5-4b17-8c48-c0279bdd6785 · inbound

General Preference Reinforcement Learning cites this paper.

General Preference Reinforcement Learning Reward Shaping to Mitigate Reward Hacking in RLHF

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-08-07T01:35:47.288139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-22T09:24:39.228616Z digest=sha256:9129c46f36b14cdc69fe916553a8c5f22c6218ccc1c846441c352d7a02a54290

Observation 23180915-ff3b-4ee9-a737-6d9700bc46a9 · inbound

Representation-Aware Advantage Estimation: Your Reward Model Provides More Than A Scalar Output cites this paper.

Representation-Aware Advantage Estimation: Your Reward Model Provides More Than A Scalar Output Reward Shaping to Mitigate Reward Hacking in RLHF

Reference 32

Resolution
metadata mismatch
arxiv_id, observed 2026-08-07T01:35:47.288139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-27T13:39:17.701196Z digest=sha256:58bc29da067989292fd650c577f52b90bdbe471125396254898fdec589c976bf

Observation 542a30f0-8e1a-428d-8eaf-16dd502c6c43 · inbound

Attention Amnesia in Hybrid LLMs: When CoT Fine-Tuning Breaks Long-Range Recall, and How to Fix It cites this paper.

Attention Amnesia in Hybrid LLMs: When CoT Fine-Tuning Breaks Long-Range Recall, and How to Fix It Reward Shaping to Mitigate Reward Hacking in RLHF

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-08-07T01:35:47.288139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-27T13:08:57.218711Z digest=sha256:27de15fcfaa0b5e21580e1908dd783dc3fdcbd4d2b8a7742b1fcbfb0b020cf3e

Observation f25422c1-50d3-4e78-8a25-94556b030c76 · inbound

From Trainee to Trainer: LLM-Designed Training Environment for RL with Multi-Agent Reasoning cites this paper.

From Trainee to Trainer: LLM-Designed Training Environment for RL with Multi-Agent Reasoning Reward Shaping to Mitigate Reward Hacking in RLHF

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-08-07T01:35:47.288139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-27T00:59:50.038405Z digest=sha256:d2d18fd9f1042ec30dbd19ce7904df8c8f045a68333699fbca7a44ac6e39d942

Observation ae675e5a-4070-45e6-a112-598783ebdaa7 · inbound

Uncertainty-Aware Reward Modeling for Stable RLHF cites this paper.

Uncertainty-Aware Reward Modeling for Stable RLHF Reward Shaping to Mitigate Reward Hacking in RLHF

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-08-07T01:35:47.288139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T18:14:33.673995Z digest=sha256:a75c9a4a3056c72570b0dcec96238fc5df8ad114af6b4c4c9cb87eebfd6c7b2f

Observation 0c0ba185-f333-4895-adb9-702110f2cef1 · inbound

OPERA: Aligning Open-Ended Reasoning via Objective Perplexity-based Reinforcement Learning cites this paper.

OPERA: Aligning Open-Ended Reasoning via Objective Perplexity-based Reinforcement Learning Reward Shaping to Mitigate Reward Hacking in RLHF

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-08-07T01:35:47.288139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-25T21:00:37.797967Z digest=sha256:c0381b4c8c7715a8f5447e2fcd7c8122ebd97f6689ce291f4d81d27086899faa