Pith. sign in

Paper Citation Record · LEDGER

Directional Alignment Mitigates Reward Hacking in Reinforcement Learning for Language Models

As of 4 August 2026, this Paper Citation Record lists 8 of 8 outbound references and 0 inbound Pith citation observations for arXiv:2605.25189.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.25189 v1

Coverage vector

measured 8 of 8 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-30T12:22:49.708718Z

measured 8 of 8 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

8 of 8 outbound references displayed

  • verified exact8
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 07a8bcc9-1ba9-474a-bab3-e838e28b545f · outbound

This paper cites Gradient Regularization Mitigates Reward Hacking in Reinforcement Learning from Human Feedback and Verifiable Rewards.

Directional Alignment Mitigates Reward Hacking in Reinforcement Learning for Language Models Gradient Regularization Mitigates Reward Hacking in Reinforcement Learning from Human Feedback and Verifiable Rewards

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-07-07T01:16:06.081071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T12:22:49.708718Z digest=sha256:a1790474284b258725f6aae8ec906352dd3440952d11451ca623eb7abaa3299a

Observation ac1388ec-7d60-4724-96f0-db8c0098503e · outbound

This paper cites On grpo collapse in search-r1: The lazy likelihood- displacement death spiral.

Directional Alignment Mitigates Reward Hacking in Reinforcement Learning for Language Models On grpo collapse in search-r1: The lazy likelihood- displacement death spiral

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-06-30T12:24:39.621034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T12:22:49.708718Z digest=sha256:dbc34a2785f8eb722f9477e71285a8bd1b53471b763e60432c2547a2cfdf2c49

Observation 99887f51-e9c0-45e4-899e-329e22056f11 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Directional Alignment Mitigates Reward Hacking in Reinforcement Learning for Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-06-30T12:24:39.617075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T12:22:49.708718Z digest=sha256:9d1ead1e5c7974f774127a3203e712124fbaf045d09edcece26d065e66fcb70d

Observation 67f17294-4440-485d-adf9-44ac7c497b7c · outbound

This paper cites GARDO: Reinforcing diffusion models without reward hacking.

Directional Alignment Mitigates Reward Hacking in Reinforcement Learning for Language Models GARDO: Reinforcing diffusion models without reward hacking

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-06-30T12:24:39.636165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T12:22:49.708718Z digest=sha256:84fa464c1aaf7cc302c6951cddda9a5354a49b8bcc828c3236231318b239c6d1

Observation 9340113c-aa56-4a4b-8ee1-294f5aa71c62 · outbound

This paper cites Correlated Proxies: A New Definition and Improved Mitigation for Reward Hacking.

Directional Alignment Mitigates Reward Hacking in Reinforcement Learning for Language Models Correlated Proxies: A New Definition and Improved Mitigation for Reward Hacking

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-06-30T12:24:39.629009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T12:22:49.708718Z digest=sha256:fe4c58e1059259d02ceded74c590cd03de524d6bf8dcde4a6dc78244eba6d969

Observation b9b50137-25e4-46c4-8d25-5ef029d8dfd0 · outbound

This paper cites Generalist Reward Models: Found Inside Large Language Models.

Directional Alignment Mitigates Reward Hacking in Reinforcement Learning for Language Models Generalist Reward Models: Found Inside Large Language Models

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-06-30T12:24:39.632631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T12:22:49.708718Z digest=sha256:a20017a62a07734d8c09a274e3eb6e883a714b8c6290c51ea373ef5f0caa1e21

Observation 2ab2e77b-637f-4c56-bf79-50d71119482c · outbound

This paper cites Inference-time scaling for generalist reward modeling.

Directional Alignment Mitigates Reward Hacking in Reinforcement Learning for Language Models Inference-time scaling for generalist reward modeling

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-06-30T12:24:39.643213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T12:22:49.708718Z digest=sha256:0a55e54872cd4367553b94ac4417968e1bc082859aaf5a60f24fc0e7e9b7d94a

Observation 57d7f00b-157e-4cc2-a731-ce0ac2d98dd7 · outbound

This paper cites Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning.

Directional Alignment Mitigates Reward Hacking in Reinforcement Learning for Language Models Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-06-30T12:24:39.625126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-30T12:22:49.708718Z digest=sha256:f89cdd9550f4b90fff2b9caf4cbbd91a7fb6e71659e2e8cec56a871402167225

Pith citing papers

No inbound Pith citation observations are available.