Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-10T16:07:37.727362Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 14 of 14 outbound references and 0 inbound Pith citation observations for arXiv:2604.11119.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-10T16:07:37.727362Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
14 of 14 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 9fe813c5-2311-4ee6-b64b-4304dbba9dbb · outbound
DDO-RM: Distribution-Level Policy Improvement after Reward Learning Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a71e353f-fce8-49a7-82f9-98e65f4df5fb · outbound
DDO-RM: Distribution-Level Policy Improvement after Reward Learning ultrafeedback\_binarized dataset card
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6b8df648-ba88-4c28-ade9-287fbe59950f · outbound
DDO-RM: Distribution-Level Policy Improvement after Reward Learning Manning, and Chelsea Finn
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e774069b-5869-47aa-b09f-0851d665515b · outbound
DDO-RM: Distribution-Level Policy Improvement after Reward Learning Proximal Policy Optimization Algorithms
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6eaf6fdd-984d-48b1-9cca-df5fa6d6a2e7 · outbound
DDO-RM: Distribution-Level Policy Improvement after Reward Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 24d229a0-a0d3-4d10-9b52-4edad85d2f5c · outbound
DDO-RM: Distribution-Level Policy Improvement after Reward Learning Mirror-descent and nonlinear projected subgradient methods for convex optimization
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 13a09cf1-48c2-45de-a702-b4035b7b9b97 · outbound
DDO-RM: Distribution-Level Policy Improvement after Reward Learning KTO: Model Alignment as Prospect Theoretic Optimization
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b0f7a24a-88aa-4950-8c18-3c6c9de38163 · outbound
DDO-RM: Distribution-Level Policy Improvement after Reward Learning ORPO: Monolithic Preference Optimization without Reference Model
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2b5b1c53-1c7b-4fbf-8ab0-42dd4309dfcf · outbound
DDO-RM: Distribution-Level Policy Improvement after Reward Learning Unresolved cited work
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cb1bff44-c5c2-4ff6-9103-d2dacdc9e0a6 · outbound
DDO-RM: Distribution-Level Policy Improvement after Reward Learning SimPO: Simple Preference Optimization with a Reference-Free Reward
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e5e266bb-eecc-4ba8-ac98-e3223b31d4ea · outbound
DDO-RM: Distribution-Level Policy Improvement after Reward Learning Problem Complexity and Method Efficiency in Optimization
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e92a85a6-eee3-4791-a216-4083fd9ff0c8 · outbound
DDO-RM: Distribution-Level Policy Improvement after Reward Learning Training language models to follow instructions with human feedback
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 101ab7f2-b80d-4379-a155-f6ba02affb5d · outbound
DDO-RM: Distribution-Level Policy Improvement after Reward Learning Understanding the Performance Gap in Preference Learning: A Dichotomy of RLHF and DPO
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4bd2963f-eea1-4449-a814-bc3c150e2952 · outbound
DDO-RM: Distribution-Level Policy Improvement after Reward Learning DDO-RM LLM preference benchmark
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
No inbound Pith citation observations are available.