Pith. sign in

Paper Citation Record · LEDGER

Design Considerations in Offline Preference-based RL

As of 9 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 2 inbound Pith citation observations for arXiv:2502.06861.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.06861 v1

Coverage vector

measured 21 of 21 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T19:40:42.080296Z

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-21T13:06:54.002248Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-21T13:10:10.404876Z

Reference resolution

21 of 21 outbound references displayed

  • verified exact0
  • verified fuzzy1
  • unresolved20
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 18182556-cba3-4140-bd73-4b789af87ceb · outbound

This paper cites Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs.

Design Considerations in Offline Preference-based RL Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-08T19:40:41.977517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:40:41.977517Z digest=sha256:be1eafa9c80f4b6212e03bf0f9ec12b2eb771a26867da399e0f05449a0593119

Observation 5c30b055-fb90-4bf3-8124-fbb6fabce50c · outbound

This paper cites an unresolved cited work.

Design Considerations in Offline Preference-based RL Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-08T19:40:42.415603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T19:40:42.074455Z digest=sha256:23e72fdd630c64f41e75b6dc2947b04ca60fbba71a3c82c9f940932f5af5d84f

Observation 89e4aabb-baae-45b5-af0f-9e2d9e9d6917 · outbound

This paper cites Robust Preference Optimization through Reward Model Distillation.

Design Considerations in Offline Preference-based RL Robust Preference Optimization through Reward Model Distillation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T19:40:41.999377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:40:41.999377Z digest=sha256:4c09415948321fc6d00340eaa8f3fa43076b2b1e7b5c2269aabdb92ae3188748

Observation 28b3bc5c-75f7-4b8e-bf00-8d71f0b107ff · outbound

This paper cites Correcting the Mythos of KL-Regularization: Direct Alignment without Overoptimization via Chi-Squared Preference Optimization.

Design Considerations in Offline Preference-based RL Correcting the Mythos of KL-Regularization: Direct Alignment without Overoptimization via Chi-Squared Preference Optimization

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T19:40:42.004209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:40:42.004209Z digest=sha256:a862cad279605ee40b55f348db7b7ab7919ea93ea13da7be4eb1c0cabac1d432

Observation 8e476658-38ac-4c5e-8c3c-b49441490b4c · outbound

This paper cites Provably Mitigating Overoptimization in RLHF: Your SFT Loss is Implicitly an Adversarial Regularizer.

Design Considerations in Offline Preference-based RL Provably Mitigating Overoptimization in RLHF: Your SFT Loss is Implicitly an Adversarial Regularizer

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T19:40:42.009598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:40:42.009598Z digest=sha256:76cf2d1422aac0fce1c79a165ad1a66165973f5dba2faaf03c255cc6f1219086

Observation 9ba1fac7-9632-4932-a5e4-4a15ee9fb486 · outbound

This paper cites Nash Learning from Human Feedback.

Design Considerations in Offline Preference-based RL Nash Learning from Human Feedback

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T19:40:42.019294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:40:42.019294Z digest=sha256:dd3b07f6f5ccb367d2996e723c6ede9576169bbb35eb61e3aa3022754ac00206

Observation c94c4030-9263-4763-91fc-8508caf05beb · outbound

This paper cites Disentangling Length from Quality in Direct Preference Optimization.

Design Considerations in Offline Preference-based RL Disentangling Length from Quality in Direct Preference Optimization

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T19:40:42.029675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:40:42.029675Z digest=sha256:80e7746cc583fc404e759f22bce6f0deb20e51ee958930a5bf34fbd3641c14d7

Observation 5f84b809-9c0b-464b-9ae6-7e16dff56c16 · outbound

This paper cites Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms.

Design Considerations in Offline Preference-based RL Scaling Laws for Reward Model Overoptimization in Direct Alignment Algorithms

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T19:40:42.034684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:40:42.034684Z digest=sha256:1f2b2108aca17e29759356b4b7c028c3fc4879f1899a7f46efeaacc88a304d8d

Observation 4639f06b-92d8-410e-aac7-ed4c6e12b946 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Design Considerations in Offline Preference-based RL Gemini: A Family of Highly Capable Multimodal Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T19:40:42.049982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:40:42.049982Z digest=sha256:1f0250eadfbe3f708fa3378e043d9218d2a530d299084efc3c11402aeed6ce05

Observation a64d54ee-d14d-4af9-a98e-368242dbd36e · outbound

This paper cites Is RLHF More Difficult than Standard RL?.

Design Considerations in Offline Preference-based RL Is RLHF More Difficult than Standard RL?

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T19:40:42.054396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:40:42.054396Z digest=sha256:de5bf3f6fe515f352511a135604143cdda155f48120b1571f7e9b44d6b2a183e

Observation 3c00348f-ff9c-439e-a0ce-a04c3d03c991 · outbound

This paper cites SLiC-HF: Sequence Likelihood Calibration with Human Feedback.

Design Considerations in Offline Preference-based RL SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-08T19:40:42.069597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:40:42.069597Z digest=sha256:3bdef3d0a8d6f13aad6c7e28a9160ba924a80e58d1ced1bb06641a6eb8c586f1

Observation 5641bb44-ced0-44c8-8857-b9c0b6d82cc5 · outbound

This paper cites The optimizer used is Adafactor with learning rate that is constant with a linear warm-up for 2000 steps and a base rate of 1e −.

Design Considerations in Offline Preference-based RL The optimizer used is Adafactor with learning rate that is constant with a linear warm-up for 2000 steps and a base rate of 1e −

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:40:42.399596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T19:40:42.080296Z digest=sha256:ae67bc18e7bb91f919d6fc0346e17ffc5e5f782b12bcd2e0c299495f71b66840

Observation 02e5e5a0-dffe-4b30-822d-13b599384c68 · outbound

This paper cites Value-Incentivized Preference Optimization: A Unified Approach to Online and Offline RLHF.

Design Considerations in Offline Preference-based RL Value-Incentivized Preference Optimization: A Unified Approach to Online and Offline RLHF

Reference 1952

Resolution
unresolved
no resolver link, observed 2026-08-08T19:40:41.989242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:40:41.989242Z digest=sha256:bd521b15b127fef14b76f8cf2b6cb02c95945c574210a71dacfc187033a587fc

Observation e8739783-b312-4542-905a-f34c8373c5f4 · outbound

This paper cites Calibrating Sequence likelihood Improves Conditional Language Generation.

Design Considerations in Offline Preference-based RL Calibrating Sequence likelihood Improves Conditional Language Generation

Reference 2007

Resolution
unresolved
no resolver link, observed 2026-08-08T19:40:42.064438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:40:42.064438Z digest=sha256:9a8e905db5c85577dab6dec1578c5a96431f68ae4af6978ceb1199020a8f27a9

Observation a37fcd4e-f981-4ab9-8bd2-ce2ed2866e01 · outbound

This paper cites On Regularization via Early Stopping for Least Squares Regression.

Design Considerations in Offline Preference-based RL On Regularization via Early Stopping for Least Squares Regression

Reference 2010

Resolution
unresolved
no resolver link, observed 2026-08-08T19:40:42.039660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:40:42.039660Z digest=sha256:4efb846af3fa33ec8ef02b8f7b96db99ea7dad33f3d99f8b92deb8d5c0d98dbf

Observation caabea31-3e40-41bf-b61e-2ab71ce38f96 · outbound

This paper cites SimPO: Simple Preference Optimization with a Reference-Free Reward.

Design Considerations in Offline Preference-based RL SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 2012

Resolution
unresolved
no resolver link, observed 2026-08-08T19:40:42.014455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:40:42.014455Z digest=sha256:8af82d53d4d5594db32f754ddcc242e2c06a9f328d3814a50f186e6cba50ef41

Observation d202202d-2a25-428d-938b-5161964e17b7 · outbound

This paper cites KTO: Model Alignment as Prospect Theoretic Optimization.

Design Considerations in Offline Preference-based RL KTO: Model Alignment as Prospect Theoretic Optimization

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-08T19:40:41.994465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:40:41.994465Z digest=sha256:9604c5d5af0b9a575e6c503e72c269aa3bcbc1a0457d7736885591fcb54292ba

Observation 1daba942-17dd-4a26-bfe4-8731dcd449fd · outbound

This paper cites A Minimaximalist Approach to Reinforcement Learning from Human Feedback.

Design Considerations in Offline Preference-based RL A Minimaximalist Approach to Reinforcement Learning from Human Feedback

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-08T19:40:42.044924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:40:42.044924Z digest=sha256:215275431250d2cb83239a677f7b4b5d0d184dab7fbd0cdf310ddcb7955c06ee

Observation 7d1f0652-e606-4e17-8087-744fdbbf5409 · outbound

This paper cites Contrastive Preference Optimization: Pushing the Boundaries of LLM Performance in Machine Translation.

Design Considerations in Offline Preference-based RL Contrastive Preference Optimization: Pushing the Boundaries of LLM Performance in Machine Translation

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-08T19:40:42.059588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:40:42.059588Z digest=sha256:05dc79ebb59cc4b8469bef2f98544de87616421033d42f0f1bbaf572c132fd85

Observation 89d56662-a529-4de4-8ddc-6401e9a9a7d6 · outbound

This paper cites Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive.

Design Considerations in Offline Preference-based RL Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-08T19:40:42.024176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:40:42.024176Z digest=sha256:f0178bb409884884bac945c09e6f5e4238addab1cf3800efc1c72d0d3bed4da5

Observation 09c3038c-a61c-4307-9b25-f209c322ef1c · outbound

This paper cites Direct Preference Optimization with an Offset.

Design Considerations in Offline Preference-based RL Direct Preference Optimization with an Offset

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-08T19:40:41.983451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:40:41.983451Z digest=sha256:d96c36be3b77680e32d8eaf6f6d453bca02f2542666bc16f2b12dab02b1cdfa7

Pith citing papers

Observation 3fe1c548-1e61-48a0-9c5f-c0a9d67d8b9e · inbound

Provably avoiding over-optimization in Direct Preference Optimization without knowing the data distribution cites this paper.

Provably avoiding over-optimization in Direct Preference Optimization without knowing the data distribution Design Considerations in Offline Preference-based RL

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:37:28.516891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T06:35:30.479542Z digest=sha256:bafddf6769b082122114812e4a59e1a66552324a0638388c1294447ba355017d

Observation 6a5f2044-b3f1-4b49-9a79-dc6d9ad3e799 · inbound

Provably avoiding over-optimization in Direct Preference Optimization without knowing the data distribution cites this paper.

Provably avoiding over-optimization in Direct Preference Optimization without knowing the data distribution Design Considerations in Offline Preference-based RL

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-21T13:10:10.406878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T13:06:54.002248Z digest=sha256:2b7353863e8b9960706722837ab77f1aae8a2e8cd7bb6021335e72977fe5758f