Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 10 inbound Pith citation observations for arXiv:2310.13639.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T05:08:05.178661Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
0 of 0 outbound references displayed
External citation measurements
2
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
No outbound reference observations are available for this paper version.
Observation db549edc-196a-4039-b3d5-0c454212e1a4 · inbound
Direct Preference Optimization for Primitive-Enabled Hierarchical RL: A Bilevel Approach Contrastive Preference Learning: Learning from Human Feedback without RL
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 331e7915-e8e4-4789-b9a3-9d870694a02d · inbound
On Monotonicity in AI Alignment Contrastive Preference Learning: Learning from Human Feedback without RL
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba3d006b-be9a-4b58-923b-160fdf816342 · inbound
Robot-Gated Interactive Imitation Learning with Adaptive Intervention Mechanism Contrastive Preference Learning: Learning from Human Feedback without RL
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2be71f98-e9ce-498d-9c6c-6f9980561ab2 · inbound
Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities Contrastive Preference Learning: Learning from Human Feedback without RL
Reference 1994
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cea6677b-c46f-4e5a-bab4-70b51806097d · inbound
$M^2PO$: Multi-Perspective Multi-Pair Preference Optimization for Machine Translation Contrastive Preference Learning: Learning from Human Feedback without RL
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd8b1b72-f44e-4bb1-b20a-643e8a72ff31 · inbound
Your Self-Play Algorithm is Secretly an Adversarial Imitator: Understanding LLM Self-Play through the Lens of Imitation Learning Contrastive Preference Learning: Learning from Human Feedback without RL
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 960af314-a171-4893-b5e7-fac0da6c4744 · inbound
A Regret Minimization Framework on Preference Learning in Large Language Models Contrastive Preference Learning: Learning from Human Feedback without RL
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 42ec9632-a473-4111-bb90-6f7d88dd068c · inbound
UniIntervene: Agentic Intervention for Efficient Real-World Reinforcement Learning Contrastive Preference Learning: Learning from Human Feedback without RL
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9f84ff1c-b804-45e5-a04b-554849bd0ed0 · inbound
Contextualizing Biological Language Models across Modalities via Logit-Space Contrastive Alignment Contrastive Preference Learning: Learning from Human Feedback without RL
Reference 94
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 382bfbfe-7516-4062-ae85-35d2c6448ef3 · inbound
S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF Contrastive Preference Learning: Learning from Human Feedback without RL
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.