Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 11 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 4 inbound Pith citation observations for arXiv:2504.06141.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-06-28T17:38:50.313305Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-03T00:37:30.094389Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 635917d5-f7e0-4dcc-9d61-51d134ef460d · inbound
Beyond Semantic Manipulation: Token-Space Attacks on Reward Models Adversarial Training of Reward Models
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 363569e9-b234-4a8b-b550-6b2e2d3d7a8f · inbound
Wasserstein Distributionally Robust Regret Optimization for Reinforcement Learning from Human Feedback Adversarial Training of Reward Models
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c8872970-de65-442b-b457-dbf3cb7d0445 · inbound
Trust Region On-Policy Distillation Adversarial Training of Reward Models
Reference 151
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation df2503a7-ef0b-46f4-a1cd-9aa9b5a54620 · inbound
The Hidden Bias of Process Reward Models:PRISM for Rewarding the Right Reasoning Adversarial Training of Reward Models
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.