Pith. sign in

Paper Citation Record · LEDGER

Robust Preference Optimization through Reward Model Distillation

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 14 inbound Pith citation observations for arXiv:2405.19316.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2405.19316 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 14 of 14 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:20:00.991880Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-21T13:10:10.421442Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 797bfae3-4e3f-4202-ba83-994136bb68fa · inbound

Frictional Agent Alignment Framework: Slow Down and Don't Break Things cites this paper.

Frictional Agent Alignment Framework: Slow Down and Don't Break Things Robust Preference Optimization through Reward Model Distillation

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:00.991880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:20:00.991880Z digest=sha256:022da73254f698a3f67043f18430e1226316fd03d82de8d3cadb411960dd77e6

Observation 1acd4098-80ed-4623-bbf1-898a97b2c34f · inbound

Risk-aware Direct Preference Optimization under Nested Risk Measure cites this paper.

Risk-aware Direct Preference Optimization under Nested Risk Measure Robust Preference Optimization through Reward Model Distillation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:19.345978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:19.345978Z digest=sha256:ce4328eb7a2be3e0811de12aab09e6f9278d0414ee0bba49fd6cad3a251f6930

Observation f499fe4f-e977-411a-88ef-e517db30b68f · inbound

Learning a Pessimistic Reward Model in RLHF cites this paper.

Learning a Pessimistic Reward Model in RLHF Robust Preference Optimization through Reward Model Distillation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:04.861147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:04.861147Z digest=sha256:4af91706054f8643762d4b96943c42f9954e6869fea214015666e9ca95628429

Observation 05043c1f-ced1-415c-9229-4faccb5a5452 · inbound

On Symmetric Losses for Robust Policy Optimization with Noisy Preferences cites this paper.

On Symmetric Losses for Robust Policy Optimization with Noisy Preferences Robust Preference Optimization through Reward Model Distillation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T12:27:54.536018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:27:54.536018Z digest=sha256:de33e878a75fb749c6ff3bc4e33ea9e470299904fbe50fd25d28e476e104b901

Observation 8d036d27-18f1-40b4-9b72-d2a38a1bdf8a · inbound

Evaluating the Effectiveness of Direct Preference Optimization for Personalizing German Automatic Text Simplifications for Persons with Intellectual Disabilities cites this paper.

Evaluating the Effectiveness of Direct Preference Optimization for Personalizing German Automatic Text Simplifications for Persons with Intellectual Disabilities Robust Preference Optimization through Reward Model Distillation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T20:55:25.113671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:55:25.113671Z digest=sha256:758c873c206763a831ca084bec1ac326661fa4e566bd7ea469c922a60a022b0c

Observation 803e3cde-bdbf-4e6e-849f-6e2179ecd459 · inbound

Robust Single-Stage Fully Sparse 3D Object Detection via Detachable Latent Diffusion cites this paper.

Robust Single-Stage Fully Sparse 3D Object Detection via Detachable Latent Diffusion Robust Preference Optimization through Reward Model Distillation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T04:36:13.397543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:36:13.397543Z digest=sha256:abd8684b5355c41fb929f1de5a008da67a854bdff954a62ad3a919db7a616b04

Observation 8134630d-ed0f-48c1-b116-42f1f1a20d78 · inbound

MPO: Multidimensional Preference Optimization for Language Model-based Text-to-Speech cites this paper.

MPO: Multidimensional Preference Optimization for Language Model-based Text-to-Speech Robust Preference Optimization through Reward Model Distillation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:34.255346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:24:34.255346Z digest=sha256:0a40c2d09b1f6eacf0d5f1a401fac296d35447476f9e9a3df482309c5db2e663

Observation 76f6a24b-e98e-4788-ac89-2b54f12811da · inbound

Multiplayer Nash Preference Optimization cites this paper.

Multiplayer Nash Preference Optimization Robust Preference Optimization through Reward Model Distillation

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:11:23.953881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T13:09:54.433720Z digest=sha256:ccf3276a9c38455a8369fbc8e69a60febe735c6eaedde433cde7bc8c31e9cba1

Observation ee0af6d8-2c8e-43f7-9456-befc13bb6f26 · inbound

LLM Harms: A Taxonomy and Discussion cites this paper.

LLM Harms: A Taxonomy and Discussion Robust Preference Optimization through Reward Model Distillation

Reference 217

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:31:24.845144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-17T00:29:07.951709Z digest=sha256:e1669663e814b61e87c5683c85dc907ba97eb445327f4a5110925b816fbe558c

Observation d8754ce3-7ed1-4c12-b701-f451f8677142 · inbound

LLM Harms: A Taxonomy and Discussion cites this paper.

LLM Harms: A Taxonomy and Discussion Robust Preference Optimization through Reward Model Distillation

Reference 217

Resolution
unresolved
no resolver link, observed 2026-08-03T18:19:30.780009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:19:30.780009Z digest=sha256:ba5e745a5c3b269408b82233bb94f5f6feb7ee2260f54b736970f254a213a09c

Observation 7e62ac6f-5cce-443d-a546-a63ae0fd0e8c · inbound

Provably avoiding over-optimization in Direct Preference Optimization without knowing the data distribution cites this paper.

Provably avoiding over-optimization in Direct Preference Optimization without knowing the data distribution Robust Preference Optimization through Reward Model Distillation

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:37:28.585241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T06:35:30.479542Z digest=sha256:5c1d38c79417f6d108f41caf8c00d0cf3222ede4fb26eb5eba3637709496466f

Observation 76156812-159e-4015-bb05-c44fe90c4f7e · inbound

Provably avoiding over-optimization in Direct Preference Optimization without knowing the data distribution cites this paper.

Provably avoiding over-optimization in Direct Preference Optimization without knowing the data distribution Robust Preference Optimization through Reward Model Distillation

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-21T13:10:10.423914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-21T13:06:54.002248Z digest=sha256:fe5bc21ae0a9a492a069dd02437d931336e7910b8bdc264815a3d278a7260a7b

Observation 8da4c509-4e9c-464b-b81b-2995af2b488f · inbound

Generating Place-Based Compromises Between Two Points of View cites this paper.

Generating Place-Based Compromises Between Two Points of View Robust Preference Optimization through Reward Model Distillation

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-11T22:01:12.122390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T03:36:31.695964Z digest=sha256:103c34d7af29fd505674f1e631c5a50b95a8d1071b89c86d5858afd7f808ce64

Observation 52382d59-bd63-41ce-87da-ffbec13e15e0 · inbound

Normalized Rewards for Preference Optimization cites this paper.

Normalized Rewards for Preference Optimization Robust Preference Optimization through Reward Model Distillation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T10:01:57.377453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:01:57.377453Z digest=sha256:fa7ee8277f825292d699ec0cc238301ea9d3fdaa99a1dc761c38be2bcef1d2fa