Pith. sign in

Paper Citation Record · LEDGER

DDO-RM: Distribution-Level Policy Improvement after Reward Learning

As of 7 August 2026, this Paper Citation Record lists 14 of 14 outbound references and 0 inbound Pith citation observations for arXiv:2604.11119.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.11119 v2

Coverage vector

measured 14 of 14 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-10T16:07:37.727362Z

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

14 of 14 outbound references displayed

  • verified exact7
  • verified fuzzy6
  • unresolved1
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9fe813c5-2311-4ee6-b64b-4304dbba9dbb · outbound

This paper cites Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling.

DDO-RM: Distribution-Level Policy Improvement after Reward Learning Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-15T17:45:18.440424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T16:07:37.727362Z digest=sha256:53a44245289ee62f1424937120497faa36be326f0615f4ad1a05545ce158395d

Observation a71e353f-fce8-49a7-82f9-98e65f4df5fb · outbound

This paper cites ultrafeedback\_binarized dataset card.

DDO-RM: Distribution-Level Policy Improvement after Reward Learning ultrafeedback\_binarized dataset card

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T16:54:59.129727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T16:07:37.727362Z digest=sha256:51f3a07638d8be7d1327d183226822cbe861c1976f7cd64d01b0bd9f3c0dcba0

Observation 6b8df648-ba88-4c28-ade9-287fbe59950f · outbound

This paper cites Manning, and Chelsea Finn.

DDO-RM: Distribution-Level Policy Improvement after Reward Learning Manning, and Chelsea Finn

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T16:54:59.123616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T16:07:37.727362Z digest=sha256:aa29ab8a3cca606f8b65ba34e3f4cdaabb4a758fe4a43031e48203cc2f259cdc

Observation e774069b-5869-47aa-b09f-0851d665515b · outbound

This paper cites Proximal Policy Optimization Algorithms.

DDO-RM: Distribution-Level Policy Improvement after Reward Learning Proximal Policy Optimization Algorithms

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-11T09:16:03.592025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T16:07:37.727362Z digest=sha256:a61e64c1bf943ea472ea5f2a30acb9723703a56cf3d52c1ce03cbb4a9b2b4366

Observation 6eaf6fdd-984d-48b1-9cca-df5fa6d6a2e7 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

DDO-RM: Distribution-Level Policy Improvement after Reward Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-11T09:16:03.575220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T16:07:37.727362Z digest=sha256:5e4ed726974bd5b39c5a9fbb0353d38f27b0c1bd4681a5336834b67a26e08860

Observation 24d229a0-a0d3-4d10-9b52-4edad85d2f5c · outbound

This paper cites Mirror-descent and nonlinear projected subgradient methods for convex optimization.

DDO-RM: Distribution-Level Policy Improvement after Reward Learning Mirror-descent and nonlinear projected subgradient methods for convex optimization

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T16:54:59.113349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T16:07:37.727362Z digest=sha256:6a01d499c3844c1145774f649153eeb149e1cff305aec9939e7b16b90b6355c5

Observation 13a09cf1-48c2-45de-a702-b4035b7b9b97 · outbound

This paper cites KTO: Model Alignment as Prospect Theoretic Optimization.

DDO-RM: Distribution-Level Policy Improvement after Reward Learning KTO: Model Alignment as Prospect Theoretic Optimization

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-12T12:17:53.598200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T16:07:37.727362Z digest=sha256:6c81606b1a1c434216afaf61723cb43bde7ad11e852dc406a5d79de217c65fa4

Observation b0f7a24a-88aa-4950-8c18-3c6c9de38163 · outbound

This paper cites ORPO: Monolithic Preference Optimization without Reference Model.

DDO-RM: Distribution-Level Policy Improvement after Reward Learning ORPO: Monolithic Preference Optimization without Reference Model

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-16T09:34:04.958044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T16:07:37.727362Z digest=sha256:d3983abfe8686ddc75d4bcbe70a7c85e3a27405df41b615bbcb43ed1bb156e7e

Observation 2b5b1c53-1c7b-4fbf-8ab0-42dd4309dfcf · outbound

This paper cites an unresolved cited work.

DDO-RM: Distribution-Level Policy Improvement after Reward Learning Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-05-17T16:54:59.116709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T16:07:37.727362Z digest=sha256:2a925ac2f9a736bc93837b4adf78ee5e07f84eca189fbd092d65da169b8fe1cf

Observation cb1bff44-c5c2-4ff6-9103-d2dacdc9e0a6 · outbound

This paper cites SimPO: Simple Preference Optimization with a Reference-Free Reward.

DDO-RM: Distribution-Level Policy Improvement after Reward Learning SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:16:03.613104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T16:07:37.727362Z digest=sha256:4e1d3aee6424e4d4226a0d325dc59db04655566c0e7610402a55e8b239afa366

Observation e5e266bb-eecc-4ba8-ac98-e3223b31d4ea · outbound

This paper cites Problem Complexity and Method Efficiency in Optimization.

DDO-RM: Distribution-Level Policy Improvement after Reward Learning Problem Complexity and Method Efficiency in Optimization

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T16:54:59.132827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T16:07:37.727362Z digest=sha256:99580446df12086ae91c72b44e8b6a571037283b0b026a78a91f975e7a9ca6f5

Observation e92a85a6-eee3-4791-a216-4083fd9ff0c8 · outbound

This paper cites Training language models to follow instructions with human feedback.

DDO-RM: Distribution-Level Policy Improvement after Reward Learning Training language models to follow instructions with human feedback

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T16:54:59.120202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T16:07:37.727362Z digest=sha256:cabca94a7e9d632ad87b4379fe36266d84fcc47fc0a968becd9f1a7a53aab905

Observation 101ab7f2-b80d-4379-a155-f6ba02affb5d · outbound

This paper cites Understanding the Performance Gap in Preference Learning: A Dichotomy of RLHF and DPO.

DDO-RM: Distribution-Level Policy Improvement after Reward Learning Understanding the Performance Gap in Preference Learning: A Dichotomy of RLHF and DPO

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-11T09:16:03.619193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T16:07:37.727362Z digest=sha256:e9910bf73d8a6e6f2445159aa830e1cf6b5760a628ecaaae9383695804660531

Observation 4bd2963f-eea1-4449-a814-bc3c150e2952 · outbound

This paper cites DDO-RM LLM preference benchmark.

DDO-RM: Distribution-Level Policy Improvement after Reward Learning DDO-RM LLM preference benchmark

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T16:54:59.126662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T16:07:37.727362Z digest=sha256:b833dcd1a64b98f2fd7f67629c73b03f8e79c98f12e5f170108695fa03f9bc11

Pith citing papers

No inbound Pith citation observations are available.