Pith. sign in

Paper Citation Record · LEDGER

GFRIEND: Generative Few-shot Reward Inference through EfficieNt DPO

As of 8 August 2026, this Paper Citation Record lists 13 of 13 outbound references and 0 inbound Pith citation observations for arXiv:2506.08965.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.08965 v1

Coverage vector

measured 13 of 13 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:03:59.969588Z

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

13 of 13 outbound references displayed

  • verified exact0
  • verified fuzzy5
  • unresolved8
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 03fe4b88-0c49-479c-a137-bec123b343a1 · outbound

This paper cites an unresolved cited work.

GFRIEND: Generative Few-shot Reward Inference through EfficieNt DPO Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:04:00.872967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:03:59.629212Z digest=sha256:95805455a861be5cda83b682d2bbeccdf3d9bac7f97b0d7f0d01f29a62a84c93

Observation 76d43501-a5be-4718-a772-01626ecea597 · outbound

This paper cites strong signal.

GFRIEND: Generative Few-shot Reward Inference through EfficieNt DPO strong signal

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:04:00.712230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:03:59.701979Z digest=sha256:186f3c1b812c4717bea64886dfe16fdafcdf5ae5d86e7c3f3a35439323b964ae

Observation 45c4b94b-5f19-4f9b-8836-fea49deded49 · outbound

This paper cites Weak Accept or Strong Reject vs.

GFRIEND: Generative Few-shot Reward Inference through EfficieNt DPO Weak Accept or Strong Reject vs

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:04:00.271775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:03:59.969588Z digest=sha256:5410aa5abe268d699877a8a6d6e3d154e94418088787451cb34244d407fb70e6

Observation e7bb3a19-2414-47b3-bf2d-6a8c61a5c1b6 · outbound

This paper cites Generative Reward Models.

GFRIEND: Generative Few-shot Reward Inference through EfficieNt DPO Generative Reward Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:59.316747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:59.316747Z digest=sha256:009e36e5b463c2638b547b596c9725f19b26ada744feceefd1563fb0d24984e7

Observation 0346ca06-8dc5-42b6-8bac-5b01371268fd · outbound

This paper cites Beyond Scalar Reward Model: Learning Generative Judge from Preference Data.

GFRIEND: Generative Few-shot Reward Inference through EfficieNt DPO Beyond Scalar Reward Model: Learning Generative Judge from Preference Data

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:59.531765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:59.531765Z digest=sha256:1d836e6790f19eab1962358f9d810ad7794e1c338e9dd47e33344e96e0443e23

Observation 6921bfe3-fb3a-4a1c-9b50-f04875cc2178 · outbound

This paper cites The model thus converges to an optimal solu- tion consistent with the true preference order- ing (up to an additive scaling).

GFRIEND: Generative Few-shot Reward Inference through EfficieNt DPO The model thus converges to an optimal solu- tion consistent with the true preference order- ing (up to an additive scaling)

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:04:00.559616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:03:59.800755Z digest=sha256:79e7bd91a99942520e2dcd544fa08026c35fa041cb59149a50075fe6d085d59e

Observation 0d90e109-0ec1-4a4d-a75e-8c517cb91d5b · outbound

This paper cites This retains the positive gradient properties 14 of logistic log-likelihood, allowing standard optimizers to converge efficiently.

GFRIEND: Generative Few-shot Reward Inference through EfficieNt DPO This retains the positive gradient properties 14 of logistic log-likelihood, allowing standard optimizers to converge efficiently

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:04:00.426144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:03:59.861614Z digest=sha256:477255f9cda0a3d15a7a5cb19b081169f57597e9afb21bf7109d2aaf3aac4759

Observation 7e971c68-9613-45c3-98c3-6249acb37590 · outbound

This paper cites Proximal Policy Optimization Algorithms.

GFRIEND: Generative Few-shot Reward Inference through EfficieNt DPO Proximal Policy Optimization Algorithms

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:59.448741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:59.448741Z digest=sha256:639a67ab15bc9843a3fa85720362cf355fd0ad8a0f061ef1b41a72e8590925ee

Observation 93599e44-d838-4697-8c2d-62358cf345b5 · outbound

This paper cites Training Deep Nets with Sublinear Memory Cost.

GFRIEND: Generative Few-shot Reward Inference through EfficieNt DPO Training Deep Nets with Sublinear Memory Cost

Reference 1992

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:59.265968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:59.265968Z digest=sha256:b7e75dc82df971039b2096ba276d6b86ab82b2e0cecb72519e8ed01e28aa2f66

Observation 0a49f408-87ef-4be4-8efc-4928a420e7ba · outbound

This paper cites A General Language Assistant as a Laboratory for Alignment.

GFRIEND: Generative Few-shot Reward Inference through EfficieNt DPO A General Language Assistant as a Laboratory for Alignment

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:59.109692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:59.109692Z digest=sha256:f9ff8256032e292641ef9cd26066d8e474b7eeeb9448001592e64850e0f5b4c6

Observation c48772e5-ded9-4861-8e38-87baf9bd560c · outbound

This paper cites Training language models to follow instructions with human feedback.

GFRIEND: Generative Few-shot Reward Inference through EfficieNt DPO Training language models to follow instructions with human feedback

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:59.185692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:59.185692Z digest=sha256:8c20e7920d28e6d2bea1d3804eab0fb44828bad986ecad5fddb708e7c812d388

Observation 98fe3752-c817-4a50-880f-dea6985cabb7 · outbound

This paper cites Samyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, and Yuxiong He.

GFRIEND: Generative Few-shot Reward Inference through EfficieNt DPO Samyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, and Yuxiong He

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:04:01.019646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:03:59.380040Z digest=sha256:f7378a4ba755d7dce93dac7f375225057b6afbfcdb2ce847f7c78ceb10823ff4

Observation 717b82e3-abeb-4939-8f27-cbe14a2d3d99 · outbound

This paper cites Critique-out-Loud Reward Models.

GFRIEND: Generative Few-shot Reward Inference through EfficieNt DPO Critique-out-Loud Reward Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:59.046421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:59.046421Z digest=sha256:f91620e82eb16c6720447bf9a1dd697b03f0c187bbb62aa1f170711a9d6098e8

Pith citing papers

No inbound Pith citation observations are available.