Pith. sign in

Paper Citation Record · LEDGER

LiPO: Listwise Preference Optimization through Learning-to-Rank

As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 20 inbound Pith citation observations for arXiv:2402.01878.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.01878 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 20 of 20 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:24:12.662588Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-23T21:23:27.413513Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 2ebd1d91-ecbf-45fb-bfd3-e5688bb9c815 · inbound

UNA: A Unified Supervised Framework for Efficient LLM Alignment Across Feedback Types cites this paper.

UNA: A Unified Supervised Framework for Efficient LLM Alignment Across Feedback Types LiPO: Listwise Preference Optimization through Learning-to-Rank

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-23T21:23:27.417555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-23T21:22:36.970101Z digest=sha256:afd2bd1221c1498e1bfdb043acc25411323334ee56fc30021f38373ecd03801c

Observation 0bb642c8-3734-47cb-8b39-7bd02e952bdb · inbound

Reward Modeling with Ordinal Feedback: Wisdom of the Crowd cites this paper.

Reward Modeling with Ordinal Feedback: Wisdom of the Crowd LiPO: Listwise Preference Optimization through Learning-to-Rank

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T17:15:47.880969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:15:47.880969Z digest=sha256:c1fb9f841d7e88cd87a7512c1ae5e193878edac02bdc7f52f8c84b24c4ef03d1

Observation ecebbdea-9c34-4716-8e01-db7c56e17d37 · inbound

ROSE: A Reward-Oriented Data Selection Framework for LLM Task-Specific Instruction Tuning cites this paper.

ROSE: A Reward-Oriented Data Selection Framework for LLM Task-Specific Instruction Tuning LiPO: Listwise Preference Optimization through Learning-to-Rank

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T05:13:41.491585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T05:13:41.491585Z digest=sha256:a7f0cd56889e06ca38918970553c1613795f97c2b514f6419183dd69308e836c

Observation 52f1e475-40f3-437a-a27d-c611f82978de · inbound

Controllable Protein Sequence Generation with LLM Preference Optimization cites this paper.

Controllable Protein Sequence Generation with LLM Preference Optimization LiPO: Listwise Preference Optimization through Learning-to-Rank

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T14:46:51.815631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:46:51.815631Z digest=sha256:bde8e690073777686c9574cd6975489d6ac185a036e0c71ea3848f8968f836fb

Observation 17aeefa1-009b-4909-aa4d-091165f021e5 · inbound

The Differences Between Direct Alignment Algorithms are a Blur cites this paper.

The Differences Between Direct Alignment Algorithms are a Blur LiPO: Listwise Preference Optimization through Learning-to-Rank

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-23T03:52:29.529197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-23T03:50:03.720389Z digest=sha256:942681f4aec14f24b26e0e9e9038cfa29a764d4a5dba457c2014d561b1ea8e60

Observation 8b1e524f-e854-4e74-b820-d5a26e47e936 · inbound

LLM Alignment as Retriever Optimization: An Information Retrieval Perspective cites this paper.

LLM Alignment as Retriever Optimization: An Information Retrieval Perspective LiPO: Listwise Preference Optimization through Learning-to-Rank

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-09T04:08:51.724809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T04:08:51.724809Z digest=sha256:094cdd547cd43ca8170fa7207b3792eeee24555301a56d5edc129da5d0a8bfa7

Observation f00dd238-8ddf-480e-a080-9759f83a6233 · inbound

PerPO: Perceptual Preference Optimization via Discriminative Rewarding cites this paper.

PerPO: Perceptual Preference Optimization via Discriminative Rewarding LiPO: Listwise Preference Optimization through Learning-to-Rank

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-09T06:01:13.306315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T06:01:13.306315Z digest=sha256:d773f51e1e4a849ec3df3684190df77ef95786b321f3a3433f567e015db97470

Observation f7881d25-80d0-457c-a75b-53286b7bad60 · inbound

A Comprehensive Survey in LLM(-Agent) Full Stack Safety: Data, Training and Deployment cites this paper.

A Comprehensive Survey in LLM(-Agent) Full Stack Safety: Data, Training and Deployment LiPO: Listwise Preference Optimization through Learning-to-Rank

Reference 221

Resolution
unresolved
no resolver link, observed 2026-08-16T11:24:12.662588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:24:12.662588Z digest=sha256:ba71a3bb27f4d1b1dada921993600c7047ade1afa4057ec1262a1b074e6821c6

Observation ac0a7142-37ff-468f-868a-eb0e608aa41a · inbound

Advancing LLM Safe Alignment with Safety Representation Ranking cites this paper.

Advancing LLM Safe Alignment with Safety Representation Ranking LiPO: Listwise Preference Optimization through Learning-to-Rank

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:18.276005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:18.276005Z digest=sha256:3044f8c74dcafb200be648f089926d58393d8f640c23cde3e318fc02e3a6d251

Observation 6852c46f-e6fb-4d63-972d-aac640c962ee · inbound

MPO: Multilingual Safety Alignment via Reward Gap Optimization cites this paper.

MPO: Multilingual Safety Alignment via Reward Gap Optimization LiPO: Listwise Preference Optimization through Learning-to-Rank

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T14:57:28.962243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:57:28.962243Z digest=sha256:3e1265f66f79be91b5eba78393cff1b9736818e252947a19733ac047a9b09823

Observation 8d7eba4a-3ff9-4c93-ad87-bddebd038360 · inbound

LPOI: Listwise Preference Optimization for Vision Language Models cites this paper.

LPOI: Listwise Preference Optimization for Vision Language Models LiPO: Listwise Preference Optimization through Learning-to-Rank

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T13:45:55.816724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:45:55.816724Z digest=sha256:b4208b9cb15d789f46d8249a71d7874f8073d8c656a5471330b1970a2f17c143

Observation c104ef4b-df02-4545-b873-a02691f147ea · inbound

HSCR: Hierarchical Self-Contrastive Rewarding for Aligning Medical Vision Language Models cites this paper.

HSCR: Hierarchical Self-Contrastive Rewarding for Aligning Medical Vision Language Models LiPO: Listwise Preference Optimization through Learning-to-Rank

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T12:00:48.923951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:00:48.923951Z digest=sha256:536488a7b3383bcd291f8497cd9f966954bc62e9bb300f4557eb2871ea731c41

Observation 98d5ad29-7c6a-41bb-900b-d36ab5f2d551 · inbound

LeanPO: Lean Preference Optimization for Likelihood Alignment in Video-LLMs cites this paper.

LeanPO: Lean Preference Optimization for Likelihood Alignment in Video-LLMs LiPO: Listwise Preference Optimization through Learning-to-Rank

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:48.212275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:48.212275Z digest=sha256:6252743607af6a582f3b657b0bc2b352a18e708238e714c830e48b4497927db8

Observation 2fe16626-45d2-4e09-a520-a493d26b41fe · inbound

Reward Models in Deep Reinforcement Learning: A Survey cites this paper.

Reward Models in Deep Reinforcement Learning: A Survey LiPO: Listwise Preference Optimization through Learning-to-Rank

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-15T19:38:50.674235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:38:50.674235Z digest=sha256:9e5f1a2a78bb3fa9f25e9e8dc72cceb8fdeb6a10a2cc52f05b1e7365332df054

Observation 9bfe8b37-794b-4f8e-8281-d7588622099f · inbound

HAEPO: History-Aggregated Exploratory Policy Optimization cites this paper.

HAEPO: History-Aggregated Exploratory Policy Optimization LiPO: Listwise Preference Optimization through Learning-to-Rank

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T16:13:08.310472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:13:08.310472Z digest=sha256:ea595d2aa4314b5f96134de8ea762ec7f8b885c9ed33322d4e78cbb82ce3c679

Observation 582509b3-ae83-4d8b-ad9a-5603ce94dccc · inbound

Threshold-Guided Optimization for Visual Generative Models cites this paper.

Threshold-Guided Optimization for Visual Generative Models LiPO: Listwise Preference Optimization through Learning-to-Rank

Reference 30

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T17:46:07.941014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-08T17:14:36.632493Z digest=sha256:7e78c43206d4a6b814b0520ed76cc060af2184b15d80181e60eb49925db6ee6f

Observation 1838de3e-0a41-455a-b996-f0c3cb00713a · inbound

Response Time Enhances Alignment with Heterogeneous Preferences cites this paper.

Response Time Enhances Alignment with Heterogeneous Preferences LiPO: Listwise Preference Optimization through Learning-to-Rank

Reference 160

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T04:46:00.020414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-11T01:04:26.288913Z digest=sha256:eb01454e9360a9a97fa56e9fe91dbe243f3ae6000918b079fe1574636c2ead7b

Observation c99c18e0-2630-4461-99d9-37a51bcdf627 · inbound

MentalThink: Shaping Thoughts in Mental SVG World cites this paper.

MentalThink: Shaping Thoughts in Mental SVG World LiPO: Listwise Preference Optimization through Learning-to-Rank

Reference 231

Resolution
unresolved
no resolver link, observed 2026-07-12T01:50:59.184754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T01:50:59.184754Z digest=sha256:6fafb93e28f4c8d0d4c7cf46b7de153ac7142c7f72e38b0a78e2d23ef374a14a

Observation b3380034-14bd-4832-8393-397a7c40bb0f · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay LiPO: Listwise Preference Optimization through Learning-to-Rank

Reference 88

Resolution
unresolved
no resolver link, observed 2026-07-11T13:53:36.775836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T13:53:36.775836Z digest=sha256:eba903de81466d23012ce7f39c440d0402f3f23e013f03710466748d7537fce4

Observation 910aa229-e412-462f-b930-696349ee7d06 · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay LiPO: Listwise Preference Optimization through Learning-to-Rank

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:41.165786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:41.165786Z digest=sha256:fee32eb8763efa1dc47708ce3a48c764718e6d157cb4ed9975db9eaf1b4964b7