Pith. sign in

Paper Citation Record · LEDGER

Reward Shaping to Mitigate Reward Hacking in RLHF

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 22 inbound Pith citation observations for arXiv:2502.18770.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.18770 v6

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 22 of 22 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 22 of 22 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:21:41.769204Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation d48a086d-209a-4c6b-8926-5cf6b54b797e · inbound

Supervising the search process produces reliable and generalizable information-seeking agents cites this paper.

Supervising the search process produces reliable and generalizable information-seeking agents Reward Shaping to Mitigate Reward Hacking in RLHF

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-08-07T01:35:47.288139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-23T02:18:27.204122Z digest=sha256:a38dfba42032ec6ba32f235039b2971f4217a7d59a34bea3653b4b30a57a3326

Observation 58c6321b-7665-4423-8b27-86437a460e6d · inbound

From System 1 to System 2: A Survey of Reasoning Large Language Models cites this paper.

From System 1 to System 2: A Survey of Reasoning Large Language Models Reward Shaping to Mitigate Reward Hacking in RLHF

Reference 170

Resolution
verified exact
arxiv_id, observed 2026-08-07T01:35:47.288139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T01:36:23.845366Z digest=sha256:023dab47aaf3fa9c64a82b4b1ad2fd199708c01f36a49d741c9ecc229e194ee9

Observation e947c3dd-4fd3-4e12-876b-9d57f2af9dbf · inbound

Enhancing Tool Learning in Large Language Models with Hierarchical Error Checklists cites this paper.

Enhancing Tool Learning in Large Language Models with Hierarchical Error Checklists Reward Shaping to Mitigate Reward Hacking in RLHF

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T13:21:41.769204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:21:41.769204Z digest=sha256:86dbafe34d08897a9adf27e970ebe1c90b97af5043091acdb794ac8bb979695d

Observation 5c3df9b4-277a-4135-b4c4-75ffb9785030 · inbound

RewardAnything: Generalizable Principle-Following Reward Models cites this paper.

RewardAnything: Generalizable Principle-Following Reward Models Reward Shaping to Mitigate Reward Hacking in RLHF

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:06.897214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:04:06.897214Z digest=sha256:4cc0613480ce9ed621ca35e52ee2e3c3aa226ca37cdd369dde4ada9bbe2e26fd

Observation f1f2f615-d4f8-4f43-a686-c92943506c4b · inbound

Self-Rewarding Vision-Language Model via Reasoning Decomposition cites this paper.

Self-Rewarding Vision-Language Model via Reasoning Decomposition Reward Shaping to Mitigate Reward Hacking in RLHF

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-08-07T01:35:47.288139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T21:03:31.606674Z digest=sha256:d0d9e4b9b8f5fcfe5cef472e91599e5449e75fa848e96441bc4e0b5359cf8b6a

Observation 1e5c7b63-c8d2-41ef-ba7f-41347e6abf8c · inbound

Factored Causal Representation Learning for Robust Reward Modeling in RLHF cites this paper.

Factored Causal Representation Learning for Robust Reward Modeling in RLHF Reward Shaping to Mitigate Reward Hacking in RLHF

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-08-07T01:35:47.288139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T14:18:33.768962Z digest=sha256:35f23bc8140b1b925f001233a54a94fcf187978e1dd96474c825ae8cf5ded60a

Observation f423b48c-8cd7-4e5b-85fb-08c49e2b8e17 · inbound

Rethinking the Comparison Unit in Sequence-Level Reinforcement Learning: An Equal-Length Paired Training Framework from Loss Correction to Sample Construction cites this paper.

Rethinking the Comparison Unit in Sequence-Level Reinforcement Learning: An Equal-Length Paired Training Framework from Loss Correction to Sample Construction Reward Shaping to Mitigate Reward Hacking in RLHF

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-08-07T01:35:47.288139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-10T06:23:14.905706Z digest=sha256:83997d14cc1a034ef9bdf5742c230341e11868336d394b74be9a488e7d0cbfdc

Observation e0aa7ce6-4486-46d5-b0d7-cff2b43fe244 · inbound

Measure Twice, Click Once: Co-evolving Proposer and Visual Critic via Reinforcement Learning for GUI Grounding cites this paper.

Measure Twice, Click Once: Co-evolving Proposer and Visual Critic via Reinforcement Learning for GUI Grounding Reward Shaping to Mitigate Reward Hacking in RLHF

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-08-07T01:35:47.288139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-09T23:05:05.251150Z digest=sha256:9b7cb9cc139d8c24532471e6b183b9c9eec713204eb011ac7347042e513116bb

Observation a496afa7-0edb-44c0-95fa-ff04bd50f10b · inbound

Towards Robust LLM Post-Training: Automatic Failure Management for Reinforcement Fine-Tuning cites this paper.

Towards Robust LLM Post-Training: Automatic Failure Management for Reinforcement Fine-Tuning Reward Shaping to Mitigate Reward Hacking in RLHF

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-08-07T01:35:47.288139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T18:19:41.302164Z digest=sha256:f9e617bc6473f571d9ff0e7ee260ef23aaa1291e0168fdc2d90733b9c4472fd6

Observation 31879369-3523-4a1d-82b4-7a711302bdb3 · inbound

Optimal Transport for LLM Reward Modeling from Noisy Preference cites this paper.

Optimal Transport for LLM Reward Modeling from Noisy Preference Reward Shaping to Mitigate Reward Hacking in RLHF

Reference 260

Resolution
metadata mismatch
arxiv_id, observed 2026-08-07T01:35:47.288139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-08T14:10:49.634358Z digest=sha256:1468c93ba88eecb5682d3f6f8a94b4e78435bd635a39229d66d5912fcb58df6c

Observation 28cb54ed-920c-419c-8d23-2482c8cc6dad · inbound

Latent Personality Alignment: Improving Harmlessness Without Mentioning Harms cites this paper.

Latent Personality Alignment: Improving Harmlessness Without Mentioning Harms Reward Shaping to Mitigate Reward Hacking in RLHF

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-08-07T01:35:47.288139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T01:46:49.586630Z digest=sha256:393a731f60e035316c44f934f43d5ef3883fe09815f9f75b29b063dd814ccc6a

Observation 4409e985-2190-4e31-b2d7-ef560d706b26 · inbound

Variance-aware Reward Modeling with Anchor Guidance cites this paper.

Variance-aware Reward Modeling with Anchor Guidance Reward Shaping to Mitigate Reward Hacking in RLHF

Reference 49

Resolution
metadata mismatch
arxiv_id, observed 2026-08-07T01:35:47.288139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-13T05:11:05.546092Z digest=sha256:829c172c52735ff74231963f05d44ca772d1283965dfdc3919af81137e506ad8

Observation b3fc049c-ece3-453c-8748-63065fc9883b · inbound

Reward Hacking in Rubric-Based Reinforcement Learning cites this paper.

Reward Hacking in Rubric-Based Reinforcement Learning Reward Shaping to Mitigate Reward Hacking in RLHF

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-08-07T01:35:47.288139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T04:08:51.578772Z digest=sha256:7ed05fd39ef2e95eb992f2c022c0a24d714f51a39ea75f39c3f620353dd3e7ab

Observation c98675b8-a2c6-45e8-bd29-ddb4c5b20399 · inbound

Diagnosing Training Inference Mismatch in LLM Reinforcement Learning cites this paper.

Diagnosing Training Inference Mismatch in LLM Reinforcement Learning Reward Shaping to Mitigate Reward Hacking in RLHF

Reference 51

Resolution
metadata mismatch
arxiv_id, observed 2026-08-07T01:35:47.288139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-15T02:05:44.813341Z digest=sha256:72c4e2df25e4a02412f27908344c88029ff8f65501e4958118a13f830a02df19

Observation f03a1e6c-f8f9-45b8-b88b-552164449677 · inbound

General Preference Reinforcement Learning cites this paper.

General Preference Reinforcement Learning Reward Shaping to Mitigate Reward Hacking in RLHF

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-08-07T01:35:47.288139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T12:43:56.522345Z digest=sha256:f41f20dee92233173f9075d1af27b2087d43942dea6b231677343f0c52658630

Observation c4ec532b-ea35-4495-8772-1c98695e82ae · inbound

General Preference Reinforcement Learning cites this paper.

General Preference Reinforcement Learning Reward Shaping to Mitigate Reward Hacking in RLHF

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-08-07T01:35:47.288139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T07:50:00.963837Z digest=sha256:5b472cea5a217dab86a4ca047c3c0f40dc2f34bd27629249c1f2d20365ef9f50

Observation 67adf7ce-7aa5-4b17-8c48-c0279bdd6785 · inbound

General Preference Reinforcement Learning cites this paper.

General Preference Reinforcement Learning Reward Shaping to Mitigate Reward Hacking in RLHF

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-08-07T01:35:47.288139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-22T09:24:39.228616Z digest=sha256:ea4de2c16c9bf52c40b07a530382929b3aa50f50869e37fe5204a148e1f11782

Observation 23180915-ff3b-4ee9-a737-6d9700bc46a9 · inbound

Representation-Aware Advantage Estimation: Your Reward Model Provides More Than A Scalar Output cites this paper.

Representation-Aware Advantage Estimation: Your Reward Model Provides More Than A Scalar Output Reward Shaping to Mitigate Reward Hacking in RLHF

Reference 32

Resolution
metadata mismatch
arxiv_id, observed 2026-08-07T01:35:47.288139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-27T13:39:17.701196Z digest=sha256:83d79b27ba978535bbf66ef8fbce64aade675f6f91c658b51e7903cbe023934e

Observation 542a30f0-8e1a-428d-8eaf-16dd502c6c43 · inbound

Attention Amnesia in Hybrid LLMs: When CoT Fine-Tuning Breaks Long-Range Recall, and How to Fix It cites this paper.

Attention Amnesia in Hybrid LLMs: When CoT Fine-Tuning Breaks Long-Range Recall, and How to Fix It Reward Shaping to Mitigate Reward Hacking in RLHF

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-08-07T01:35:47.288139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-27T13:08:57.218711Z digest=sha256:96e49a1cf26de1911a11c1e6c2f4cd27dc6421f4866a9157feeaa481f1628623

Observation f25422c1-50d3-4e78-8a25-94556b030c76 · inbound

From Trainee to Trainer: LLM-Designed Training Environment for RL with Multi-Agent Reasoning cites this paper.

From Trainee to Trainer: LLM-Designed Training Environment for RL with Multi-Agent Reasoning Reward Shaping to Mitigate Reward Hacking in RLHF

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-08-07T01:35:47.288139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-27T00:59:50.038405Z digest=sha256:f24a2ab2c25d2d7dfe3bb80ba726d31782a9243143ea91c47ccd7d9c3a19a96e

Observation ae675e5a-4070-45e6-a112-598783ebdaa7 · inbound

Uncertainty-Aware Reward Modeling for Stable RLHF cites this paper.

Uncertainty-Aware Reward Modeling for Stable RLHF Reward Shaping to Mitigate Reward Hacking in RLHF

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-08-07T01:35:47.288139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T18:14:33.673995Z digest=sha256:4c35d89c351d26236661e6d3a2da419cc9bbe7995311d4e615c1d3ccbaa46752

Observation 0c0ba185-f333-4895-adb9-702110f2cef1 · inbound

OPERA: Aligning Open-Ended Reasoning via Objective Perplexity-based Reinforcement Learning cites this paper.

OPERA: Aligning Open-Ended Reasoning via Objective Perplexity-based Reinforcement Learning Reward Shaping to Mitigate Reward Hacking in RLHF

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-08-07T01:35:47.288139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-25T21:00:37.797967Z digest=sha256:745004c3d389eb84dcc0d318bb412822111212db20ee78a7e7df7c32d476a9a1