Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T12:16:32.070018Z
Paper Citation Record · LEDGER
As of 12 August 2026, this Paper Citation Record lists 87 of 87 outbound references and 3 inbound Pith citation observations for arXiv:2412.14516.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T12:16:32.070018Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-08T12:17:46.012360Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-19T04:42:04.899091Z
87 of 87 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation a3043f17-7c1a-4907-9627-c1f5b1960ec7 · outbound
Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a53bf06-8323-4cfd-805a-41592a79ce75 · outbound
Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Training language models to follow instructions with human feedback
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1c5331e-62a6-4b14-8f7e-9d06a1242baf · outbound
Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Learning to summarize with human feedback
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c21db7c7-e901-4ace-aa9d-f03d596f7616 · outbound
Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Deep reinforcement learning from human preferences
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ab6c605-1107-4c47-ba23-174ba61bb750 · outbound
Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Proximal Policy Optimization Algorithms
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c2e03700-8331-4b9a-9237-0832dace122f · outbound
Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Implementation matters in deep policy gradients: A case study on ppo and trpo
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c8364072-296c-4bfa-9f37-7b5b687dd2c0 · outbound
Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Direct preference optimization: Your language model is secretly a reward model
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de430066-de2d-42a1-aab7-fe0e94031552 · outbound
Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment A general theoretical paradigm to understand learning from human preferences
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 252bde81-682b-48cd-b034-cb66fe3413f1 · outbound
Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment SLiC-HF: Sequence Likelihood Calibration with Human Feedback
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1260564c-4074-4593-9df5-ec7ae10168ed · outbound
Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Preference Fine-Tuning of LLMs Should Leverage Suboptimal, On-Policy Data
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d69ef4d6-85ad-494f-a7e1-c9a4554f450f · outbound
Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99b3e0ae-f37d-4861-986f-cdd089439c71 · outbound
Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Advancing LLM Reasoning Generalists with Preference Trees
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6796dcc0-2ce7-4394-805f-483abf14982d · outbound
Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Learning word vectors for sentiment analysis
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4cd60717-3a3a-459c-8c01-687001bbea5c · outbound
Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Tl; dr: Mining reddit to learn automatic summarization
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation edadd08e-1c6a-43ed-998e-18509afea875 · outbound
Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment A framework for few-shot language model evaluation, 12 2023
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 65c40c16-36fe-473b-ad35-f090bd37ad5b · outbound
Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Nash Learning from Human Feedback
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ec0e0994-acd2-4e42-a38b-9f8c176824b2 · outbound
Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Statistical rejection sampling improves preference optimization
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 70cce962-625c-40dd-ad7e-be59950a25ca · outbound
Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Self-Rewarding Language Models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0762ebe0-d00d-443c-ba4f-58f67230247e · outbound
Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6819dafb-6643-4018-b926-430f6136ac17 · outbound
Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 087d7327-30ac-4fe0-9c68-53e750be321b · outbound
Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Direct Language Model Alignment from Online AI Feedback
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7690f04-8146-4675-90d9-319bb5331e75 · outbound
Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Simper: Simple preference fine-tuning without hyperparameters by perplexity optimization
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 07476bde-f31a-476d-8daa-12c33e203891 · outbound
Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Policy Optimization in RLHF: The Impact of Out-of-preference Data
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f3207bb2-2819-4a05-88bf-8338026d52c2 · outbound
Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Is DPO Superior to PPO for LLM Alignment? A Comprehensive Study
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 00fe9418-3368-48a4-9310-8b5c3d35de57 · outbound
Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Provably Robust DPO: Aligning Language Models with Noisy Feedback
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c89ca9c3-a17a-4814-871d-11c438522a3a · outbound
Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Noise Contrastive Alignment of Language Models with Explicit Rewards
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4960576a-471c-4581-801a-9b1107e24467 · outbound
Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment COPR: Continual Human Preference Learning via Optimal Policy Regularization
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c4044b74-173e-4ab9-8a56-2590f91da535 · outbound
Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Towards Efficient Exact Optimization of Language Model Alignment
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 793ac3b3-1a83-4aea-9f02-c5230f473f25 · outbound
Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Noise contrastive estimation and negative sampling for condi- tional models: Consistency and statistical efficiency
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation ea2b61fe-61ac-47a8-9026-a9ad5900318c · outbound
Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Towards Analyzing and Understanding the Limitations of DPO: A Theoretical Perspective
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53babd5a-45e5-4f9d-8b08-949cb10d1c75 · outbound
Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment From $r$ to $Q^*$: Your Language Model is Secretly a Q-Function
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f4659949-2b58-4b5b-8532-ade787f3a012 · outbound
Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment DPO Meets PPO: Reinforced Token Optimization for RLHF
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df554fa1-4c4e-4e47-8839-3d33f48fea7a · outbound
Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Token-level Direct Preference Optimization
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1343b00a-f14a-4d03-ae9d-80f0692c6f1a · outbound
Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment A general offline reinforcement learning framework for interac- tive recommendation
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 6f6fb95b-bc85-4058-9a5f-7997b2df7c71 · outbound
Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment On calibration of modern neural networks
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 64781a79-dc7c-4d9e-82a9-962a674930b6 · outbound
Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Scale calibration of deep ranking models
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation dfdcb30e-710b-4eb5-9e07-495b1f6bcc7d · outbound
Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Calibrated model-based deep reinforcement learning
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 5714eafd-90fb-415e-b599-e941e60eb9b0 · outbound
Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Cal-ql: Calibrated offline rl pre-training for efficient online fine-tuning
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8437b2c8-1af5-4ddd-9ab1-d93fa3439987 · outbound
Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment On the Calibration of Large Language Models and Alignment
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9312365d-b111-474d-b068-e16fa8b1bee6 · outbound
Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Language Models (Mostly) Know What They Know
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 56d76232-f930-4dd0-af5b-52442956b335 · outbound
Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Rank analysis of incomplete block designs: I
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 83463c92-b4e6-4ccd-8540-2d4bb110ed37 · outbound
Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Generalized Preference Optimization: A Unified Approach to Offline Alignment
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4256de8-453d-45da-95c8-e554d4f05985 · outbound
Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment How to Leverage Demonstration Data in Alignment for Large Language Model? A Self-Imitation Learning Perspective
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0943c1fb-e0db-48ed-9cf5-9f167b4771ef · outbound
Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment SimPO: Simple Preference Optimization with a Reference-Free Reward
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac1a2811-daad-4d69-bb65-a0193be5fac8 · outbound
Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Reinforced Self-Training (ReST) for Language Modeling
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25e9db84-767a-48ea-b70d-e03cd68fe5a9 · outbound
Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Openchat: Advancing open-source language models with mixed-quality data
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 84b5bfc4-f444-45dd-ac05-f508bab73996 · outbound
Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Fine-Tuning Language Models with Advantage-Induced Policy Alignment
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa582487-b8bf-4e08-9c55-7a73db68d3b8 · outbound
Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Machine learning: a probabilistic perspective
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42d83879-5ccc-4995-84eb-62af187d97d4 · outbound
Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Improving Policy Gradient by Exploring Under-appreciated Rewards
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a03ab81-e9a4-4ed8-a15e-156dba954253 · outbound
Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Learning how to propagate messages in graph neural networks
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 00ecfb45-e50e-47c1-b8fa-c9e7af06eb87 · outbound
Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Learning to generalize from sparse and underspecified rewards
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation f71633dc-4925-4aad-8ec1-3b59b4422e85 · outbound
Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Decoupled self-supervised learning for graphs
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation f57ea6f7-d3af-4c71-b8aa-df3d3e2ac0ef · outbound
Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment UltraFeedback: Boosting Language Models with Scaled AI Feedback
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a888593-3615-4760-a93b-621af2650ef9 · outbound
Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Zephyr: Direct Distillation of LM Alignment
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8842ab19-1cfe-4140-910d-d92946430a09 · outbound
Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a528a66-0fca-47d9-91f4-689ff083e0ae · outbound
Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53e1e8af-cde6-4422-ac5d-4e030c0db960 · outbound
Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Instruction-Following Evaluation for Large Language Models
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38d0ece0-5b9d-4acd-89ef-12140dc30d40 · outbound
Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f242457-b2f7-4380-9e7b-3bcb877d3602 · outbound
Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment GPQA: A Graduate-Level Google-Proof Q&A Benchmark
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 085c13ee-d299-423b-b2c1-408a20437e10 · outbound
Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Training Verifiers to Solve Math Word Problems
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f68d5328-0d1d-41ed-86d2-480ad6135bdf · outbound
Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Measuring Mathematical Problem Solving With the MATH Dataset
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 959e2286-93db-4f1d-bf5b-cc7288c8181c · outbound
Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Alpacaeval: An automatic evaluator of instruction-following models, 2023
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5c75b4a-1225-4dd4-bbf8-58a1b9a2280f · outbound
Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Pythia: A suite for analyzing large language models across training and scaling
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3eecb998-3173-402e-b8d0-ea74583ca919 · outbound
Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Language models are unsupervised multitask learners
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation d2f9f718-aa82-40fd-96c9-61af13d31330 · outbound
Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Contrastive Preference Optimization: Pushing the Boundaries of LLM Performance in Machine Translation
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc65e0f0-fbbf-4f39-a91c-6616e4399414 · outbound
Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Beyond reverse kl: Generalizing direct preference optimization with diverse divergence constraints
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 650830fe-ea77-4a4c-b4dd-a0dc74095e6d · outbound
Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Iterative Reasoning Preference Optimization
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33bd544e-d5b4-4168-b0a0-581e3d3d6058 · outbound
Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Information, divergence and risk for binary experiments
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation da37e0fa-1e0f-437f-9935-12465a4d6df0 · outbound
Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Reward augmented maximum likelihood for neural structured prediction
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 05b24dda-9261-44f0-9a92-abeb00905921 · outbound
Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Clustering with bregman divergences
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 41168029-a0b7-4045-82c1-31903ff9efa5 · outbound
Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 484e4191-a869-4b78-bbf9-af172e8cac7c · outbound
Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Understanding the Effects of RLHF on LLM Generalisation and Diversity
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8210680e-89c7-418e-bff8-ce1a1fcc093d · outbound
Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Guidelines: • The answer NA means that the abstract and introduction do not include the claims made in the paper
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 26608078-d248-42dd-abf9-23a4eb8d7b11 · outbound
Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Limitations
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 4884cd23-33a9-4595-8b34-3bae026efaf8 · outbound
Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Guidelines: • The answer NA means that the paper does not include theoretical results
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 4826bd82-96f4-4751-875e-7cee43bda8b5 · outbound
Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Guidelines: • The answer NA means that the paper does not include experiments
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c0db1cac-3a60-4de3-8d6a-1daad2ad7748 · outbound
Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment • Please see the NeurIPS code and data submission guidelines ( https://nips.cc/ public/guides/CodeSubmissionPolicy) for more details
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation aec86a50-2589-4464-aa65-e396144e2d65 · outbound
Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Guidelines: • The answer NA means that the paper does not include experiments
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 06642fd1-12b3-4dcd-bb92-f67df389f24a · outbound
Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Guidelines: • The answer NA means that the paper does not include experiments
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 28738ff5-4694-48a8-8d58-87fd63ab9128 · outbound
Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Guidelines: • The answer NA means that the paper does not include experiments
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 30282f20-6d4c-4f28-aed4-c4af1e1fea3c · outbound
Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Guidelines: • The answer NA means that the authors have not reviewed the NeurIPS Code of Ethics
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation f9cea911-8a11-41e2-917e-369700efd3ee · outbound
Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Guidelines: • The answer NA means that there is no societal impact of the work performed
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 9bb2e117-02d4-49b7-b2ec-819966320e54 · outbound
Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Guidelines: • The answer NA means that the paper poses no such risks
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea19d25d-37db-4e40-9eea-312ff696fdf3 · outbound
Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Guidelines: • The answer NA means that the paper does not use existing assets
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 76ebee73-8f0e-4919-9cc8-415d8e5c5bb6 · outbound
Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Guidelines: • The answer NA means that the paper does not release new assets
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5df0804e-650e-45aa-91b6-34563bce2854 · outbound
Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects
Reference 86
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea1687f3-4a6e-4d5d-aeef-083d0bd918a7 · outbound
Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 5983e2f6-f1c3-49aa-999c-9598cfd3d96f · inbound
DPO-Shift: Shifting the Distribution of Direct Preference Optimization Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b664141-2b19-4be9-bc82-8fb1de2a8e68 · inbound
Bridging Brains and Machines: A Unified Frontier in Neuroscience, Artificial Intelligence, and Neuromorphic Systems Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment
Reference 150
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 10e735fa-5805-42b3-a9da-57eff2c8f526 · inbound
Gradient-Gated DPO: Stabilizing Preference Optimization in Language Models Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.