Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T05:03:59.969588Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 13 of 13 outbound references and 0 inbound Pith citation observations for arXiv:2506.08965.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T05:03:59.969588Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
13 of 13 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 03fe4b88-0c49-479c-a137-bec123b343a1 · outbound
GFRIEND: Generative Few-shot Reward Inference through EfficieNt DPO Unresolved cited work
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 76d43501-a5be-4718-a772-01626ecea597 · outbound
GFRIEND: Generative Few-shot Reward Inference through EfficieNt DPO strong signal
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 45c4b94b-5f19-4f9b-8836-fea49deded49 · outbound
GFRIEND: Generative Few-shot Reward Inference through EfficieNt DPO Weak Accept or Strong Reject vs
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e7bb3a19-2414-47b3-bf2d-6a8c61a5c1b6 · outbound
GFRIEND: Generative Few-shot Reward Inference through EfficieNt DPO Generative Reward Models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0346ca06-8dc5-42b6-8bac-5b01371268fd · outbound
GFRIEND: Generative Few-shot Reward Inference through EfficieNt DPO Beyond Scalar Reward Model: Learning Generative Judge from Preference Data
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6921bfe3-fb3a-4a1c-9b50-f04875cc2178 · outbound
GFRIEND: Generative Few-shot Reward Inference through EfficieNt DPO The model thus converges to an optimal solu- tion consistent with the true preference order- ing (up to an additive scaling)
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0d90e109-0ec1-4a4d-a75e-8c517cb91d5b · outbound
GFRIEND: Generative Few-shot Reward Inference through EfficieNt DPO This retains the positive gradient properties 14 of logistic log-likelihood, allowing standard optimizers to converge efficiently
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7e971c68-9613-45c3-98c3-6249acb37590 · outbound
GFRIEND: Generative Few-shot Reward Inference through EfficieNt DPO Proximal Policy Optimization Algorithms
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 93599e44-d838-4697-8c2d-62358cf345b5 · outbound
GFRIEND: Generative Few-shot Reward Inference through EfficieNt DPO Training Deep Nets with Sublinear Memory Cost
Reference 1992
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a49f408-87ef-4be4-8efc-4928a420e7ba · outbound
GFRIEND: Generative Few-shot Reward Inference through EfficieNt DPO A General Language Assistant as a Laboratory for Alignment
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c48772e5-ded9-4861-8e38-87baf9bd560c · outbound
GFRIEND: Generative Few-shot Reward Inference through EfficieNt DPO Training language models to follow instructions with human feedback
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98fe3752-c817-4a50-880f-dea6985cabb7 · outbound
GFRIEND: Generative Few-shot Reward Inference through EfficieNt DPO Samyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, and Yuxiong He
Reference 2023
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 717b82e3-abeb-4939-8f27-cbe14a2d3d99 · outbound
GFRIEND: Generative Few-shot Reward Inference through EfficieNt DPO Critique-out-Loud Reward Models
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.