Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-30T12:22:49.708718Z
Paper Citation Record · LEDGER
As of 4 August 2026, this Paper Citation Record lists 8 of 8 outbound references and 0 inbound Pith citation observations for arXiv:2605.25189.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-30T12:22:49.708718Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
8 of 8 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 07a8bcc9-1ba9-474a-bab3-e838e28b545f · outbound
Directional Alignment Mitigates Reward Hacking in Reinforcement Learning for Language Models Gradient Regularization Mitigates Reward Hacking in Reinforcement Learning from Human Feedback and Verifiable Rewards
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ac1388ec-7d60-4724-96f0-db8c0098503e · outbound
Directional Alignment Mitigates Reward Hacking in Reinforcement Learning for Language Models On grpo collapse in search-r1: The lazy likelihood- displacement death spiral
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 99887f51-e9c0-45e4-899e-329e22056f11 · outbound
Directional Alignment Mitigates Reward Hacking in Reinforcement Learning for Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 67f17294-4440-485d-adf9-44ac7c497b7c · outbound
Directional Alignment Mitigates Reward Hacking in Reinforcement Learning for Language Models GARDO: Reinforcing diffusion models without reward hacking
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 9340113c-aa56-4a4b-8ee1-294f5aa71c62 · outbound
Directional Alignment Mitigates Reward Hacking in Reinforcement Learning for Language Models Correlated Proxies: A New Definition and Improved Mitigation for Reward Hacking
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b9b50137-25e4-46c4-8d25-5ef029d8dfd0 · outbound
Directional Alignment Mitigates Reward Hacking in Reinforcement Learning for Language Models Generalist Reward Models: Found Inside Large Language Models
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 2ab2e77b-637f-4c56-bf79-50d71119482c · outbound
Directional Alignment Mitigates Reward Hacking in Reinforcement Learning for Language Models Inference-time scaling for generalist reward modeling
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 57d7f00b-157e-4cc2-a731-ce0ac2d98dd7 · outbound
Directional Alignment Mitigates Reward Hacking in Reinforcement Learning for Language Models Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
No inbound Pith citation observations are available.