Pith. sign in

Paper Citation Record · LEDGER

Contrastive Preference Learning: Learning from Human Feedback without RL

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 10 inbound Pith citation observations for arXiv:2310.13639.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2310.13639 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 10 of 10 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:08:05.178661Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

2
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation db549edc-196a-4039-b3d5-0c454212e1a4 · inbound

Direct Preference Optimization for Primitive-Enabled Hierarchical RL: A Bilevel Approach cites this paper.

Direct Preference Optimization for Primitive-Enabled Hierarchical RL: A Bilevel Approach Contrastive Preference Learning: Learning from Human Feedback without RL

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-23T18:33:19.704972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-23T18:30:38.818711Z digest=sha256:396ab9d3e9ae4b525b9ca24e74af2d434644fdd995c68853deb8ef1e77e7e6ba

Observation 331e7915-e8e4-4789-b9a3-9d870694a02d · inbound

On Monotonicity in AI Alignment cites this paper.

On Monotonicity in AI Alignment Contrastive Preference Learning: Learning from Human Feedback without RL

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T05:08:05.178661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:08:05.178661Z digest=sha256:87cd9745742098196d7cf8f2fff52b3467413bf137d6898802287ede19abb261

Observation ba3d006b-be9a-4b58-923b-160fdf816342 · inbound

Robot-Gated Interactive Imitation Learning with Adaptive Intervention Mechanism cites this paper.

Robot-Gated Interactive Imitation Learning with Adaptive Intervention Mechanism Contrastive Preference Learning: Learning from Human Feedback without RL

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T05:02:05.531884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:02:05.531884Z digest=sha256:f90362305b78df78dc4f918992ad69a1c44d1b32ab503d9567ead07bc1faccb5

Observation 2be71f98-e9ce-498d-9c6c-6f9980561ab2 · inbound

Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities cites this paper.

Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities Contrastive Preference Learning: Learning from Human Feedback without RL

Reference 1994

Resolution
unresolved
no resolver link, observed 2026-08-06T16:34:25.031856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:34:25.031856Z digest=sha256:b546a49b4fbaeb1d5319e0e5ff357400aeeebb8cb99f6b80d15224b0e8647f3c

Observation cea6677b-c46f-4e5a-bab4-70b51806097d · inbound

$M^2PO$: Multi-Perspective Multi-Pair Preference Optimization for Machine Translation cites this paper.

$M^2PO$: Multi-Perspective Multi-Pair Preference Optimization for Machine Translation Contrastive Preference Learning: Learning from Human Feedback without RL

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T09:50:04.047978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T09:50:04.047978Z digest=sha256:c331b223832626b96ded9e8d7344a8296993a29f8e107eadc712a2160970f784

Observation dd8b1b72-f44e-4bb1-b20a-643e8a72ff31 · inbound

Your Self-Play Algorithm is Secretly an Adversarial Imitator: Understanding LLM Self-Play through the Lens of Imitation Learning cites this paper.

Your Self-Play Algorithm is Secretly an Adversarial Imitator: Understanding LLM Self-Play through the Lens of Imitation Learning Contrastive Preference Learning: Learning from Human Feedback without RL

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-03T05:49:24.586560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:49:24.586560Z digest=sha256:053c1254f22ff94d5726ede57f4cd1a6a74843ed6fe7720c63c31820584e0898

Observation 960af314-a171-4893-b5e7-fac0da6c4744 · inbound

A Regret Minimization Framework on Preference Learning in Large Language Models cites this paper.

A Regret Minimization Framework on Preference Learning in Large Language Models Contrastive Preference Learning: Learning from Human Feedback without RL

Reference 1

Resolution
malformed identifier
arxiv_id, observed 2026-07-03T01:27:30.773321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T16:34:19.323865Z digest=sha256:eec5b55e30531e3af3986ad53ea5967b7abdca271c51ba9f84b4dd534a2b7ec4

Observation 42ec9632-a473-4111-bb90-6f7d88dd068c · inbound

UniIntervene: Agentic Intervention for Efficient Real-World Reinforcement Learning cites this paper.

UniIntervene: Agentic Intervention for Efficient Real-World Reinforcement Learning Contrastive Preference Learning: Learning from Human Feedback without RL

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-03T10:58:02.723259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T09:46:59.746745Z digest=sha256:e77013c4f98b6241a821eaf172b5a8559acb60864eed726bd0e1514021db876c

Observation 9f84ff1c-b804-45e5-a04b-554849bd0ed0 · inbound

Contextualizing Biological Language Models across Modalities via Logit-Space Contrastive Alignment cites this paper.

Contextualizing Biological Language Models across Modalities via Logit-Space Contrastive Alignment Contrastive Preference Learning: Learning from Human Feedback without RL

Reference 94

Resolution
verified exact
arxiv_id, observed 2026-06-26T21:40:08.481236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T21:31:07.417680Z digest=sha256:16a19f10b2b07cfc5997dce4ad7927136303994c35462e5d313ab87591ab2b55

Observation 382bfbfe-7516-4062-ae85-35d2c6448ef3 · inbound

S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF cites this paper.

S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF Contrastive Preference Learning: Learning from Human Feedback without RL

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T13:58:42.282065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:58:42.282065Z digest=sha256:208a14c686e8d3bf8b6302e20397dbea44c785fe0b7d96277fd945bdf1c79d7a