Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T05:24:21.201297Z
Paper Citation Record · LEDGER
As of 11 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 1 inbound Pith citation observation for arXiv:2412.17696.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T05:24:21.201297Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-04T17:40:08.618174Z
A source-named dated measurement, never combined with another source.
Source: cited_works
43 of 43 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation d0a1f17e-f51c-40d3-8ab6-001ac9e1eba5 · outbound
Understanding the Logic of Direct Preference Alignment through Logic DPO and reference approaches For DPO we see a simi- lar derivation
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 0111fb54-4587-4207-b997-8cb7afef2f24 · outbound
Understanding the Logic of Direct Preference Alignment through Logic Camels in a Changing Climate: Enhancing LM Adaptation with Tulu 2
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf35de68-9326-4805-99dc-8c8d70e18ea5 · outbound
Understanding the Logic of Direct Preference Alignment through Logic Prompting is programming: A query language for large language models
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 6ad0fa82-4bd7-4133-908a-cde191a52c5e · outbound
Understanding the Logic of Direct Preference Alignment through Logic However, the semantics of the resulting formulas are less transparent and often hidden in the weights
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 2d8a9011-4bef-4877-8b94-0c90a79650be · outbound
Understanding the Logic of Direct Preference Alignment through Logic Unresolved cited work
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 5e46e22c-3252-447b-9901-8dcdbbbb3437 · outbound
Understanding the Logic of Direct Preference Alignment through Logic (2024)), all of which were originally implemented using the logistic log-loss, i.e., each ℓx = − log σ(βρθ)
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation c66a9439-ef3d-441a-b65a-e1e8c1fb8e9e · outbound
Understanding the Logic of Direct Preference Alignment through Logic Declarative Design of Neural Predicates in Neuro-Symbolic Systems
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 0c66f357-b5a3-40e5-9fa6-e302bcf1e5d7 · outbound
Understanding the Logic of Direct Preference Alignment through Logic New Desiderata for Direct Preference Optimization
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation abb00165-1fc0-45e3-902b-a15b12efbb6e · outbound
Understanding the Logic of Direct Preference Alignment through Logic Unresolved cited work
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation bce74ebd-ad1f-496b-b3b1-031884d5bdb2 · outbound
Understanding the Logic of Direct Preference Alignment through Logic DSPy: Compiling Declarative Language Model Calls into Self-Improving Pipelines
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eafd6341-2023-40b8-b3e6-2f8de101f702 · outbound
Understanding the Logic of Direct Preference Alignment through Logic What Makes Good Data for Alignment? A Comprehensive Study of Automatic Data Selection in Instruction Tuning
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation abec1c88-e4f0-44a8-b698-8f6fa80a6cc9 · outbound
Understanding the Logic of Direct Preference Alignment through Logic Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06962f4c-82bd-4b5b-bec6-339adcad19d5 · outbound
Understanding the Logic of Direct Preference Alignment through Logic Online DPO: Online Direct Preference Optimization with Fast-Slow Chasing
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a751ee2-beb9-44da-abe4-69ee1e1a6251 · outbound
Understanding the Logic of Direct Preference Alignment through Logic Unintentional Unalignment: Likelihood Displacement in Direct Preference Optimization
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f5e94b0-964b-422b-95cc-a1d42b3718ab · outbound
Understanding the Logic of Direct Preference Alignment through Logic Logic of Differentiable Logics: Towards a Uniform Semantics of DL
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 10e3b6cb-bd84-427e-8659-3fdf57c1791d · outbound
Understanding the Logic of Direct Preference Alignment through Logic On the Independence Assumption in Neurosymbolic Learning
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 284ccd89-8ef2-440e-bfc6-ee470834b2dd · outbound
Understanding the Logic of Direct Preference Alignment through Logic Aligning Large Language Models with Human: A Survey
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4901e4e7-b1e8-459a-bf61-8b71d7051ea0 · outbound
Understanding the Logic of Direct Preference Alignment through Logic Contrastive Preference Optimization: Pushing the Boundaries of LLM Performance in Machine Translation
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40eda98e-d5ba-4dc1-a581-bec958921811 · outbound
Understanding the Logic of Direct Preference Alignment through Logic Direct Preference Knowledge Distillation for Large Language Models
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 850d7631-4d98-485c-b575-5144a537d229 · outbound
Understanding the Logic of Direct Preference Alignment through Logic RRHF: Rank Responses to Align Language Models with Human Feedback without tears
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 288e9c2b-27e3-4ba6-a0e5-93f5ba2565a2 · outbound
Understanding the Logic of Direct Preference Alignment through Logic SLiC-HF: Sequence Likelihood Calibration with Human Feedback
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 07cb74ed-9771-4d1b-9b57-abe0b32ff339 · outbound
Understanding the Logic of Direct Preference Alignment through Logic Fine-Tuning Language Models from Human Preferences
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 08ac64bf-a997-4ce5-8d86-31eba4898ee9 · outbound
Understanding the Logic of Direct Preference Alignment through Logic Original losses Further details of the original losses in Table 2, along with other variants such as R-DPO (Park et al., 2024), ODPO (Amini et al.,
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 79cc82fc-f00b-456e-ae12-ce0e46fa8dbc · outbound
Understanding the Logic of Direct Preference Alignment through Logic Unresolved cited work
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 760683a3-978d-4327-ae60-b26e18cb20ff · outbound
Understanding the Logic of Direct Preference Alignment through Logic Unresolved cited work
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation dddd4f5a-9429-49cc-bc0b-28ae4e408065 · outbound
Understanding the Logic of Direct Preference Alignment through Logic Figure 8 shows the Boolean semantics of DPO/SimPO and some novel variants based on the ref- erence form of ORPO (ℓORPO-ref), qfUNL (ℓqfUNL-ref) and l5 (ℓl5-ref)
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation df7b99b7-a451-476c-a0ac-ecfe48709ca6 · outbound
Understanding the Logic of Direct Preference Alignment through Logic Specifically, we focus on losses around the known lossℓCPO, which we treat as a natural baseline to compare against
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 2df5a295-1b62-4a41-8878-ca9d34ff8621 · outbound
Understanding the Logic of Direct Preference Alignment through Logic While these experiments are small scale and limited in scope, they are merely meant to suggest possible uses our frame- work and open questions
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 277d8756-1265-4cb3-90e4-4bf1a4c3bef1 · outbound
Understanding the Logic of Direct Preference Alignment through Logic To avoid repeating the process of instruction tuning, we started from the trained Qwen model released in the TRL library6
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 5e9f422a-dc72-4b1f-a4c9-e8100f6224a7 · outbound
Understanding the Logic of Direct Preference Alignment through Logic This suggests that different types of preference data rely on a different semantics of preference, which requires a tuning approach that’s tailored to those differences
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 2ea43b70-4807-4d59-b405-1be9ba8fffb7 · outbound
Understanding the Logic of Direct Preference Alignment through Logic Unresolved cited work
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 7173bdc4-e1e1-48a7-85ae-4dbfe67fac49 · outbound
Understanding the Logic of Direct Preference Alignment through Logic Self-Exploring Language Models: Active Preference Elicitation for Online Alignment
Reference 1975
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 042ff44e-30a7-4c63-b41f-c437002ba784 · outbound
Understanding the Logic of Direct Preference Alignment through Logic Generalized Preference Optimization: A Unified Approach to Offline Alignment
Reference 1977
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7aefa3f-ddfd-4ddf-9c9e-581dfea35e95 · outbound
Understanding the Logic of Direct Preference Alignment through Logic Language Model Cascades
Reference 2007
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b9b2a69-ef90-4f56-89ff-cc6b22f408c4 · outbound
Understanding the Logic of Direct Preference Alignment through Logic Insights into Alignment: Evaluating DPO and its Variants Across Multiple Tasks
Reference 2015
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5008a5d1-b1f4-4458-9452-ba3077d2eac9 · outbound
Understanding the Logic of Direct Preference Alignment through Logic Adversarially Regularising Neural NLI Models to Integrate Logical Background Knowledge
Reference 2017
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation a47bfd00-65e2-47a2-840c-31ccc875dd04 · outbound
Understanding the Logic of Direct Preference Alignment through Logic Hybrid Preferences: Learning to Route Instances for Human vs. AI Feedback
Reference 2018
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ac67d50-851b-4d1c-bfda-2c50bd113595 · outbound
Understanding the Logic of Direct Preference Alignment through Logic Preference Tuning with Human Feedback on Language, Speech, and Vision Tasks: A Survey
Reference 2019
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6df2947-5ffe-4330-ab02-943c86d44356 · outbound
Understanding the Logic of Direct Preference Alignment through Logic A General Theoretical Paradigm to Understand Learning from Human Preferences
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 196dde86-df55-4ed1-a5cf-279d0f442537 · outbound
Understanding the Logic of Direct Preference Alignment through Logic Direct Language Model Alignment from Online AI Feedback
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f4ac2847-9f4d-432d-84e5-379e57c986ee · outbound
Understanding the Logic of Direct Preference Alignment through Logic Logic Tensor Networks for Semantic Image Interpretation
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f56f1ffc-df57-47c5-93e2-aabd9cd11047 · outbound
Understanding the Logic of Direct Preference Alignment through Logic Qwen Technical Report
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5efe21c0-8cd9-42a8-9303-558f084c5ea0 · outbound
Understanding the Logic of Direct Preference Alignment through Logic Logically Consistent Language Models via Neuro-Symbolic Integration
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1256f83f-8b8c-48b4-952c-e40239160c53 · inbound
LLM Enhancement with Domain Expert Mental Model to Reduce LLM Hallucination with Causal Prompt Engineering Understanding the Logic of Direct Preference Alignment through Logic
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.