Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-10T13:50:47.902911Z
Paper Citation Record · LEDGER
As of 1 August 2026, this Paper Citation Record lists 5 of 5 outbound references and 0 inbound Pith citation observations for arXiv:2604.13833.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-10T13:50:47.902911Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-01T06:32:01.292127+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
5 of 5 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 93163713-4833-40c2-ad31-24fd5d384d74 · outbound
Robust Reward Modeling for Large Language Models via Causal Decomposition WARP: On the Benefits of Weight Averaged Rewarded Policies
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation f5703def-dd90-49d8-8885-be064697dd6a · outbound
Robust Reward Modeling for Large Language Models via Causal Decomposition A Long Way to Go: Investigating Length Correlations in RLHF
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation cfc6e527-9550-452b-97d5-ed00910a49c1 · outbound
Robust Reward Modeling for Large Language Models via Causal Decomposition Beyond Reward Hacking: Causal Rewards for Large Language Model Alignment
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation f45a0ccb-bc98-4615-be12-c9d49c4529aa · outbound
Robust Reward Modeling for Large Language Models via Causal Decomposition Unresolved cited work
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
Observation e9e7b4a4-8b3d-4c0a-b2c0-a5d4468ba555 · outbound
Robust Reward Modeling for Large Language Models via Causal Decomposition Assumption 3(Top-K Margin Condition).For si =P f(w i) and ideal Top-K indices Jwi, there existsδ >0such that: min j∈Jwi min t /∈Jwi (|si,j| − |si,t|)≥δ
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.
No inbound Pith citation observations are available.