Pith. sign in

Paper Citation Record · LEDGER

Generalized Preference Optimization: A Unified Approach to Offline Alignment

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 21 inbound Pith citation observations for arXiv:2402.05749.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.05749 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 21 of 21 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 21 of 21 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T22:20:13.819384Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-23T03:52:29.455696Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 57ef5985-b31d-4716-bd89-10a02249b3a5 · inbound

BRiTE: Bootstrapping Reinforced Thinking Process to Enhance Language Model Reasoning cites this paper.

BRiTE: Bootstrapping Reinforced Thinking Process to Enhance Language Model Reasoning Generalized Preference Optimization: A Unified Approach to Offline Alignment

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-09T22:20:13.819384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T22:20:13.819384Z digest=sha256:fdcf529770543f8fd39b31c229c4f7af188dc92d09041b9b0f207a62f23c499c

Observation 2cd7047a-8241-4f99-aed5-d87a2d392949 · inbound

On Almost Surely Safe Alignment of Large Language Models at Inference-Time cites this paper.

On Almost Surely Safe Alignment of Large Language Models at Inference-Time Generalized Preference Optimization: A Unified Approach to Offline Alignment

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-09T16:18:40.588740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:18:40.588740Z digest=sha256:318431f0dd93401b4c59a734cabaee06bc753fb6abd1016a50965805c5e1de7f

Observation 47600c21-1a3d-4c32-8cd1-207366753daa · inbound

The Differences Between Direct Alignment Algorithms are a Blur cites this paper.

The Differences Between Direct Alignment Algorithms are a Blur Generalized Preference Optimization: A Unified Approach to Offline Alignment

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-23T03:52:29.458777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-23T03:50:03.720389Z digest=sha256:6bab0020c9ae020f8ab273343267d261bca92e44002060f2eaa77102cf68abc0

Observation 38788a5f-748a-40cb-8c14-7866e2068cf4 · inbound

PIPA: Preference Alignment as Prior-Informed Statistical Estimation cites this paper.

PIPA: Preference Alignment as Prior-Informed Statistical Estimation Generalized Preference Optimization: A Unified Approach to Offline Alignment

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-08T18:10:53.433659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:10:53.433659Z digest=sha256:82d2d41b771c37d051dc517e236774434203355c6058eb41d236d4061f4ed608

Observation 1708aeca-820c-4d29-93d4-736cf1eb84b1 · inbound

DPO-Shift: Shifting the Distribution of Direct Preference Optimization cites this paper.

DPO-Shift: Shifting the Distribution of Direct Preference Optimization Generalized Preference Optimization: A Unified Approach to Offline Alignment

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-08T12:17:45.991858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:17:45.991858Z digest=sha256:fa26e225f3fedf91d55bdb5c6edc1597cd8a5643f267aa9293724804c64c0f21

Observation 35c96d81-5e34-4029-b468-edaa857d42f6 · inbound

Fino1: On the Transferability of Reasoning-Enhanced LLMs and Reinforcement Learning to Finance cites this paper.

Fino1: On the Transferability of Reasoning-Enhanced LLMs and Reinforcement Learning to Finance Generalized Preference Optimization: A Unified Approach to Offline Alignment

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-08T10:25:13.056866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T10:25:13.056866Z digest=sha256:3b961184e7546fc033b539e3cedb4132af3883fa0a62647fb56c8a24dc1f9432

Observation 400d8bdf-8572-4ceb-ba4e-f027564c672b · inbound

Learning a Pessimistic Reward Model in RLHF cites this paper.

Learning a Pessimistic Reward Model in RLHF Generalized Preference Optimization: A Unified Approach to Offline Alignment

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:06.110662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:06.110662Z digest=sha256:9d82696b7fc1a036a801045e521475ec5d7e537a9bd1e1043299d8982fd306c9

Observation 31f00f59-19a3-4c2b-a2f4-8491679d2c76 · inbound

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment cites this paper.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Generalized Preference Optimization: A Unified Approach to Offline Alignment

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T13:45:01.361604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:45:01.361604Z digest=sha256:0b77d2f385bd2d8f571e242437f740ca683075e97dc598c33b8e636316abc3ef

Observation 2e6ff718-c6fc-46f7-9aea-f0932bbcd6dd · inbound

OpenReview Should be Protected and Leveraged as a Community Asset for Research in the Era of Large Language Models cites this paper.

OpenReview Should be Protected and Leveraged as a Community Asset for Research in the Era of Large Language Models Generalized Preference Optimization: A Unified Approach to Offline Alignment

Reference 132

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:46.360499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:31:46.360499Z digest=sha256:b72b4d50ad801c8677ad44ef0ba4b52f2e151969f755c33669fca94fbc64d6b9

Observation ba1dbab5-2129-42cb-9ed6-8fe0fa6d62b7 · inbound

Learning Parametric Distributions from Samples and Preferences cites this paper.

Learning Parametric Distributions from Samples and Preferences Generalized Preference Optimization: A Unified Approach to Offline Alignment

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:13.529395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:50:13.529395Z digest=sha256:f6e2c7f51c99a0ab4e66d67a4294584fd5eed9ef05b23cea42bb975e79c87a4d

Observation 21d8ba4c-9c58-46cc-b33b-f6537323faba · inbound

On Symmetric Losses for Robust Policy Optimization with Noisy Preferences cites this paper.

On Symmetric Losses for Robust Policy Optimization with Noisy Preferences Generalized Preference Optimization: A Unified Approach to Offline Alignment

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T12:27:56.555683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:27:56.555683Z digest=sha256:35efae49f8a944876186189e33433d704848adf38f3de040adf4065ffb1eb58b

Observation 6bf9cade-a35f-49df-802a-f4f1d1c79197 · inbound

Explicit Preference Optimization: No Need for an Implicit Reward Model cites this paper.

Explicit Preference Optimization: No Need for an Implicit Reward Model Generalized Preference Optimization: A Unified Approach to Offline Alignment

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T05:40:18.016343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:40:18.016343Z digest=sha256:1753fb0ae8061c5888bdad17afa944db546d6c69b3c037afc104914912805ccc

Observation 9fc43729-2309-44a1-b104-6dd6dd3a7ff4 · inbound

On a few pitfalls in KL divergence gradient estimation for RL cites this paper.

On a few pitfalls in KL divergence gradient estimation for RL Generalized Preference Optimization: A Unified Approach to Offline Alignment

Reference 1999

Resolution
unresolved
no resolver link, observed 2026-08-07T04:53:38.578236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:53:38.578236Z digest=sha256:3427fc475e7459632d5b67502f26df542335a55764d11bf0810fcccf632b630a

Observation 5f28af2f-0126-4227-86bd-402bdf9fda61 · inbound

Curr-RLCER:Curriculum Reinforcement Learning For Coherence Explainable Recommendation cites this paper.

Curr-RLCER:Curriculum Reinforcement Learning For Coherence Explainable Recommendation Generalized Preference Optimization: A Unified Approach to Offline Alignment

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T22:35:49.941183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T19:43:11.384787Z digest=sha256:5b387fbbe94939b004c4c0d0251a30d341c66fce131e1b99176edd44652a8f24

Observation bfc0acf3-1e5c-4465-ba88-4d4026a4de31 · inbound

Response Time Enhances Alignment with Heterogeneous Preferences cites this paper.

Response Time Enhances Alignment with Heterogeneous Preferences Generalized Preference Optimization: A Unified Approach to Offline Alignment

Reference 163

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T04:46:00.197978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-11T01:04:26.288913Z digest=sha256:686c9212c1666cc283601ff40599d7f0c3c19268adfec394310e2999e7a63f24

Observation 1590a79e-fe7d-455b-b731-62dfdfc85523 · inbound

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching cites this paper.

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching Generalized Preference Optimization: A Unified Approach to Offline Alignment

Reference 139

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T04:57:17.175932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-13T04:55:55.013900Z digest=sha256:d765097c0e45a654001fdfba078e660441e2bcfa5f7381019fbd527c5887821f

Observation 3b2bee7a-53ff-49ac-9411-5ca08a07456f · inbound

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching cites this paper.

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching Generalized Preference Optimization: A Unified Approach to Offline Alignment

Reference 139

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T05:45:06.522644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-15T05:41:10.714594Z digest=sha256:a63d1d48672d930f35dcea268d0ecaad995be1e3516d44266bfa3bd2eb0d8e6d

Observation f2eb89b7-8cf7-4aaa-be79-a61664a91d56 · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay Generalized Preference Optimization: A Unified Approach to Offline Alignment

Reference 190

Resolution
unresolved
no resolver link, observed 2026-07-11T13:53:36.775836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T13:53:36.775836Z digest=sha256:d40e4156e76d7b9b0c606d394caa4daaf60e62a9afa373afbdd62dc83b0807c3

Observation 0e16a44a-39cc-4b67-92e7-b47ce1f28de8 · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay Generalized Preference Optimization: A Unified Approach to Offline Alignment

Reference 191

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:54.675387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:54.675387Z digest=sha256:7e0a596d6dedc9071c4fb638403cce19af8aa009f6523157033f5812abe845ca

Observation 9cfa9a43-235a-4937-86fe-7ae51fa1af33 · inbound

Normalized Rewards for Preference Optimization cites this paper.

Normalized Rewards for Preference Optimization Generalized Preference Optimization: A Unified Approach to Offline Alignment

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T10:01:58.919069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:01:58.919069Z digest=sha256:4a084683acda1cd6c09087e9c12ff47ff412e0a23ab02d07212ea03a0a675f46

Observation d8a343d5-9c94-4a8b-b3da-b9f381382b9a · inbound

Test-Time Scaling via Error Localization cites this paper.

Test-Time Scaling via Error Localization Generalized Preference Optimization: A Unified Approach to Offline Alignment

Reference 105

Resolution
unresolved
no resolver link, observed 2026-08-01T07:28:28.247565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:28:28.247565Z digest=sha256:517b1f3a7a52dfd5c748a74a2e2a68d4be04d630c7df6d77691e24092f645b07