Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T14:39:22.659889Z
Paper Citation Record · LEDGER
As of 11 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 0 inbound Pith citation observations for arXiv:2501.15109.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T14:39:22.659889Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
21 of 21 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 4cb1b3ab-7199-469d-8200-1ade5491316f · outbound
Clear Preferences Leave Traces: Reference Model-Guided Sampling for Preference Learning Enhancing Chat Language Models by Scaling High-quality Instructional Conversations
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e204ea03-bab1-452c-9ab4-91cbacbf35f0 · outbound
Clear Preferences Leave Traces: Reference Model-Guided Sampling for Preference Learning The Llama 3 Herd of Models
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 12b3a973-da04-4eee-9de3-ba597e5fcc34 · outbound
Clear Preferences Leave Traces: Reference Model-Guided Sampling for Preference Learning KTO: Model Alignment as Prospect Theoretic Optimization
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 778039a5-bf25-4cb2-a026-57d5f990d648 · outbound
Clear Preferences Leave Traces: Reference Model-Guided Sampling for Preference Learning Impact of Preference Noise on the Alignment Performance of Generative Language Models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96893a3d-cb9f-4d0f-8f3b-37b860da39a1 · outbound
Clear Preferences Leave Traces: Reference Model-Guided Sampling for Preference Learning Towards Comprehensive Preference Data Collection for Reward Modeling
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation eb394745-d226-4486-8f3d-34dce92d2ff2 · outbound
Clear Preferences Leave Traces: Reference Model-Guided Sampling for Preference Learning In Proceedings of the 2023 Conference on Em- pirical Methods in Natural Language Processing , 9187–
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 6374e27e-381f-473f-b62b-d321476f14b0 · outbound
Clear Preferences Leave Traces: Reference Model-Guided Sampling for Preference Learning One-Shot Safety Alignment for Large Language Models via Optimal Dualization
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd44d406-4f46-49e3-be8d-c303673f55ad · outbound
Clear Preferences Leave Traces: Reference Model-Guided Sampling for Preference Learning Unpacking DPO and PPO: Disentangling Best Practices for Learning from Preference Feedback
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9f28792-0da2-4476-8710-094a050b3e65 · outbound
Clear Preferences Leave Traces: Reference Model-Guided Sampling for Preference Learning Mistral 7B
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5348c521-0ea4-4238-b262-f561f98c6ff4 · outbound
Clear Preferences Leave Traces: Reference Model-Guided Sampling for Preference Learning A Survey on Human Preference Learning for Large Language Models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a792d65-1c92-4d01-9816-18e06630e5f7 · outbound
Clear Preferences Leave Traces: Reference Model-Guided Sampling for Preference Learning Spread Preference Annotation: Direct Preference Judgment for Efficient LLM Alignment
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 624eec9c-f1d9-41f5-b2a7-b2eedc21ec2d · outbound
Clear Preferences Leave Traces: Reference Model-Guided Sampling for Preference Learning From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d3bb5d8-2270-436a-ab72-33a6e358a940 · outbound
Clear Preferences Leave Traces: Reference Model-Guided Sampling for Preference Learning Statistical Rejection Sampling Improves Preference Optimization
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df0fb837-b648-4b15-9f59-646e4f21287a · outbound
Clear Preferences Leave Traces: Reference Model-Guided Sampling for Preference Learning Enhancing LLM Safety via Constrained Direct Preference Optimization
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b442214-fe35-4712-80ef-f6a06e762b45 · outbound
Clear Preferences Leave Traces: Reference Model-Guided Sampling for Preference Learning Adapting Large Language Models for Content Moderation: Pitfalls in Data Engineering and Supervised Fine-tuning
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7dacf60-509c-4f74-a326-58e84b29d054 · outbound
Clear Preferences Leave Traces: Reference Model-Guided Sampling for Preference Learning SimPO: Simple Preference Optimization with a Reference-Free Reward
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c152b506-7f2b-42cf-8134-ac543e973363 · outbound
Clear Preferences Leave Traces: Reference Model-Guided Sampling for Preference Learning Filtered Direct Preference Optimization
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d5707e30-3e9c-4230-8576-05f25e428de8 · outbound
Clear Preferences Leave Traces: Reference Model-Guided Sampling for Preference Learning Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a60a2d6f-201c-476f-8f32-95acc3e726bd · outbound
Clear Preferences Leave Traces: Reference Model-Guided Sampling for Preference Learning AlphaDPO: Adaptive Reward Margin for Direct Preference Optimization
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0c11f7f-e6a5-45f9-af30-f6f10cd69af9 · outbound
Clear Preferences Leave Traces: Reference Model-Guided Sampling for Preference Learning UltraFeedback: Boosting Language Models with Scaled AI Feedback
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86a11f2f-4317-4f7d-8b69-1c7f9e50bc3c · outbound
Clear Preferences Leave Traces: Reference Model-Guided Sampling for Preference Learning Provably Robust DPO: Aligning Language Models with Noisy Feedback
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.