Pith. sign in

Paper Citation Record · LEDGER

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment

As of 20 August 2026, this Paper Citation Record lists 87 of 87 outbound references and 4 inbound Pith citation observations for arXiv:2412.14516.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.14516 v1

Coverage vector

measured 87 of 87 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T12:16:32.070018Z

measured 91 of 91 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T23:58:38.854695Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-19T04:42:04.899091Z

Reference resolution

87 of 87 outbound references displayed

  • verified exact1
  • verified fuzzy30
  • unresolved56
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a3043f17-7c1a-4907-9627-c1f5b1960ec7 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.787908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.787908Z digest=sha256:edcf09c1bf18a851dfc2daee323e52d151dc03d47fba07ad52611d61f1b4d699

Observation 7a53bf06-8323-4cfd-805a-41592a79ce75 · outbound

This paper cites Training language models to follow instructions with human feedback.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Training language models to follow instructions with human feedback

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.792312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.792312Z digest=sha256:76ab2a4752c9c107fe34e81f4fc892ad37c99713437b97a942f4a5231e368c76

Observation f1c5331e-62a6-4b14-8f7e-9d06a1242baf · outbound

This paper cites Learning to summarize with human feedback.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Learning to summarize with human feedback

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.795485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.795485Z digest=sha256:e1251bcb07a02ea024402b69f257ed8c5bc6bca3f4e3a9b224130593b450aac2

Observation c21db7c7-e901-4ace-aa9d-f03d596f7616 · outbound

This paper cites Deep reinforcement learning from human preferences.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Deep reinforcement learning from human preferences

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.798253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.798253Z digest=sha256:b068710684a09894ff90023e2ea2e0feb63ae574cecd7c8fbc63b8eb53e0d5c8

Observation 4ab6c605-1107-4c47-ba23-174ba61bb750 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Proximal Policy Optimization Algorithms

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.800935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.800935Z digest=sha256:8dc5fd727e94d1d208ba60e95c4892f7dd4685d886e04b677f45027f338bc5f3

Observation c2e03700-8331-4b9a-9237-0832dace122f · outbound

This paper cites Implementation matters in deep policy gradients: A case study on ppo and trpo.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Implementation matters in deep policy gradients: A case study on ppo and trpo

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.805323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.805323Z digest=sha256:b94d5f2abf5d5b44dddb7d9001883a71794eb866b42d46801bbeea2e61621a8b

Observation c8364072-296c-4bfa-9f37-7b5b687dd2c0 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Direct preference optimization: Your language model is secretly a reward model

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.809910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.809910Z digest=sha256:9993a5877c2a9faf1b0fe023fdf583a5f354d00fcafabb1d873fbbdec1e4f220

Observation de430066-de2d-42a1-aab7-fe0e94031552 · outbound

This paper cites A general theoretical paradigm to understand learning from human preferences.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment A general theoretical paradigm to understand learning from human preferences

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.813699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.813699Z digest=sha256:7dcbc3c19a1c3915aeed0dca11fb16dd5d93331941f1405d794580d25b388866

Observation 252bde81-682b-48cd-b034-cb66fe3413f1 · outbound

This paper cites SLiC-HF: Sequence Likelihood Calibration with Human Feedback.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.816513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.816513Z digest=sha256:ae5b0065aaff0dd829253641585ae75b884c6d67e4c29e47cbe42408b995e881

Observation 1260564c-4074-4593-9df5-ec7ae10168ed · outbound

This paper cites Preference Fine-Tuning of LLMs Should Leverage Suboptimal, On-Policy Data.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Preference Fine-Tuning of LLMs Should Leverage Suboptimal, On-Policy Data

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.820225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.820225Z digest=sha256:ba51c48ce15674452b1df703b16dd73ad30f6df7bd62290c30f444c7d92c7be4

Observation d69ef4d6-85ad-494f-a7e1-c9a4554f450f · outbound

This paper cites Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.823443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.823443Z digest=sha256:58a7b14508ea99e803477517f97f1cb6b81f35e002ec537189c66b1f1397baf0

Observation 99b3e0ae-f37d-4861-986f-cdd089439c71 · outbound

This paper cites Advancing LLM Reasoning Generalists with Preference Trees.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Advancing LLM Reasoning Generalists with Preference Trees

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.827786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.827786Z digest=sha256:4374454ca0124eb294ad665a7666d85d58d00f0cdfa142e03041a1d309fe87ce

Observation 6796dcc0-2ce7-4394-805f-483abf14982d · outbound

This paper cites Learning word vectors for sentiment analysis.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Learning word vectors for sentiment analysis

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.831179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.831179Z digest=sha256:772d337b6b42120fc4713be97272fde2cc91d61d447cf79ce1fe29980067403b

Observation 4cd60717-3a3a-459c-8c01-687001bbea5c · outbound

This paper cites Tl; dr: Mining reddit to learn automatic summarization.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Tl; dr: Mining reddit to learn automatic summarization

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:16:32.711917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T12:16:31.834698Z digest=sha256:472a3ccb19dce653ca63dd33b0a0ff4681793c3b70468c773e9a9e3d814b9d13

Observation edadd08e-1c6a-43ed-998e-18509afea875 · outbound

This paper cites A framework for few-shot language model evaluation, 12 2023.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment A framework for few-shot language model evaluation, 12 2023

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.838747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.838747Z digest=sha256:3e156ff114b5fc48beb241f57e84dbbe3b09a9ade39e29953757da194165758b

Observation 65c40c16-36fe-473b-ad35-f090bd37ad5b · outbound

This paper cites Nash Learning from Human Feedback.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Nash Learning from Human Feedback

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.841413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.841413Z digest=sha256:51d74ab4bd8894c7c632c69dc39949c129b8e7153f49a459ab89c27cc115efbd

Observation ec0e0994-acd2-4e42-a38b-9f8c176824b2 · outbound

This paper cites Statistical rejection sampling improves preference optimization.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Statistical rejection sampling improves preference optimization

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:16:32.700617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T12:16:31.844889Z digest=sha256:6da61cb686a1befef8d0f9a87f7f66418f5ab3b51a130decf01c7c7a27cc03b9

Observation 70cce962-625c-40dd-ad7e-be59950a25ca · outbound

This paper cites Self-Rewarding Language Models.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Self-Rewarding Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.847944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.847944Z digest=sha256:9ea753c6acafc37a56ae3052f4b7d7307960061f2d48e59b1a7a56ba45835485

Observation 0762ebe0-d00d-443c-ba4f-58f67230247e · outbound

This paper cites Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.851281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.851281Z digest=sha256:b5ab086fe1fd37b5fc350a6a4f3f5e2859c86eafa9c58827b1217eaf6413b632

Observation 6819dafb-6643-4018-b926-430f6136ac17 · outbound

This paper cites Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.855210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.855210Z digest=sha256:ef1cbd9b30ced5ecba3556f66d5de75a41e8dbc09b949908f994dd844cac797f

Observation 087d7327-30ac-4fe0-9c68-53e750be321b · outbound

This paper cites Direct Language Model Alignment from Online AI Feedback.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Direct Language Model Alignment from Online AI Feedback

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.859337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.859337Z digest=sha256:689bae26898fbb6b8e16e63c642b52c2303158ef76828c8f0d7331e8a05c5963

Observation a7690f04-8146-4675-90d9-319bb5331e75 · outbound

This paper cites Simper: Simple preference fine-tuning without hyperparameters by perplexity optimization.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Simper: Simple preference fine-tuning without hyperparameters by perplexity optimization

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:16:32.693022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T12:16:31.863229Z digest=sha256:640b4d5a69779a977718793c89a84f0c0e6c6978a56aebfa177be6ae89700644

Observation 07476bde-f31a-476d-8daa-12c33e203891 · outbound

This paper cites Policy Optimization in RLHF: The Impact of Out-of-preference Data.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Policy Optimization in RLHF: The Impact of Out-of-preference Data

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.865955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.865955Z digest=sha256:37d404657d3f96b571d3bf31ae91eff9cdb283c8f7acfbe4f3da250263e961fa

Observation f3207bb2-2819-4a05-88bf-8338026d52c2 · outbound

This paper cites Is DPO Superior to PPO for LLM Alignment? A Comprehensive Study.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Is DPO Superior to PPO for LLM Alignment? A Comprehensive Study

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.868865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.868865Z digest=sha256:f0b5b687a8b8d093999194effedd48232813b63950daea3622f923567032ff8f

Observation 00fe9418-3368-48a4-9310-8b5c3d35de57 · outbound

This paper cites Provably Robust DPO: Aligning Language Models with Noisy Feedback.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Provably Robust DPO: Aligning Language Models with Noisy Feedback

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.872076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.872076Z digest=sha256:8390609048c025019690982a9cfd11db29e27558b2179abbac5f19d0e5e6d82a

Observation c89ca9c3-a17a-4814-871d-11c438522a3a · outbound

This paper cites Noise Contrastive Alignment of Language Models with Explicit Rewards.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Noise Contrastive Alignment of Language Models with Explicit Rewards

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.875936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.875936Z digest=sha256:834e7a6efd192a7ccc2868de768b22bdf7306490f126842bbb21e912fb13d6e4

Observation 4960576a-471c-4581-801a-9b1107e24467 · outbound

This paper cites COPR: Continual Human Preference Learning via Optimal Policy Regularization.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment COPR: Continual Human Preference Learning via Optimal Policy Regularization

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-08-11T12:16:32.256071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T12:16:31.879139Z digest=sha256:10c63d6ff9ca568d306a7e1bb08362bba8b6fe94d4ea21db89f05e2ad0d0d7bc

Observation c4044b74-173e-4ab9-8a56-2590f91da535 · outbound

This paper cites Towards Efficient Exact Optimization of Language Model Alignment.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Towards Efficient Exact Optimization of Language Model Alignment

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.882201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.882201Z digest=sha256:678101026fd884d988bf613dbc78799bc0e0efca8329c0fe5b3cc0a38536e4e9

Observation 793ac3b3-1a83-4aea-9f02-c5230f473f25 · outbound

This paper cites Noise contrastive estimation and negative sampling for condi- tional models: Consistency and statistical efficiency.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Noise contrastive estimation and negative sampling for condi- tional models: Consistency and statistical efficiency

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:16:32.684069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T12:16:31.885458Z digest=sha256:b69f50003a880b352449920e9b33b409ba6e480b5c288ecb3f2ad505ff79bd15

Observation ea2b61fe-61ac-47a8-9026-a9ad5900318c · outbound

This paper cites Towards Analyzing and Understanding the Limitations of DPO: A Theoretical Perspective.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Towards Analyzing and Understanding the Limitations of DPO: A Theoretical Perspective

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.888937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.888937Z digest=sha256:ce8a18d2eab64bba2e95049714104fb70bb5e89a6025ad0d43f5657b91a537d0

Observation 53babd5a-45e5-4f9d-8b08-949cb10d1c75 · outbound

This paper cites From $r$ to $Q^*$: Your Language Model is Secretly a Q-Function.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment From $r$ to $Q^*$: Your Language Model is Secretly a Q-Function

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.891745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.891745Z digest=sha256:aec7e9b3e6fcb51dbe66b18f52c283229c513642382598d73f569afde0a2edb3

Observation f4659949-2b58-4b5b-8532-ade787f3a012 · outbound

This paper cites DPO Meets PPO: Reinforced Token Optimization for RLHF.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment DPO Meets PPO: Reinforced Token Optimization for RLHF

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.895081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.895081Z digest=sha256:9c5f52dbcef7828f8b72ae275ed9f316ad151c965a3386817d0312e7cf627784

Observation df554fa1-4c4e-4e47-8839-3d33f48fea7a · outbound

This paper cites Token-level Direct Preference Optimization.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Token-level Direct Preference Optimization

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.898358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.898358Z digest=sha256:dd2212df453661787841dcce2058c57f0ba6da2ac5725d657205cbc5ffc4fc83

Observation 1343b00a-f14a-4d03-ae9d-80f0692c6f1a · outbound

This paper cites A general offline reinforcement learning framework for interac- tive recommendation.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment A general offline reinforcement learning framework for interac- tive recommendation

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:16:32.674683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T12:16:31.903506Z digest=sha256:89a3b97951f098a1aa832fa25be73f359f57f851df9540629fa7634165bccc88

Observation 6f6fb95b-bc85-4058-9a5f-7997b2df7c71 · outbound

This paper cites On calibration of modern neural networks.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment On calibration of modern neural networks

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:16:32.664634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T12:16:31.906673Z digest=sha256:e9adf6d41ab6619377077f6ea7485bc4c10d28dff35572f09cf885ab036cd77e

Observation 64781a79-dc7c-4d9e-82a9-962a674930b6 · outbound

This paper cites Scale calibration of deep ranking models.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Scale calibration of deep ranking models

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:16:32.655401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T12:16:31.910137Z digest=sha256:eea7a3d812554659d7d4acb6a3dec915d6d66b5c960fbde33478401eb051f60e

Observation dfdcb30e-710b-4eb5-9e07-495b1f6bcc7d · outbound

This paper cites Calibrated model-based deep reinforcement learning.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Calibrated model-based deep reinforcement learning

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:16:32.646066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T12:16:31.912682Z digest=sha256:b0c55d18c25f2f6f8fa1a761f6d6e390d40cca835fdd5710b1a145fa92f753cc

Observation 5714eafd-90fb-415e-b599-e941e60eb9b0 · outbound

This paper cites Cal-ql: Calibrated offline rl pre-training for efficient online fine-tuning.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Cal-ql: Calibrated offline rl pre-training for efficient online fine-tuning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.916373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.916373Z digest=sha256:914d6354a7ff87739b19a1b733d5b00486a6c565339f09a100876695921c3316

Observation 8437b2c8-1af5-4ddd-9ab1-d93fa3439987 · outbound

This paper cites On the Calibration of Large Language Models and Alignment.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment On the Calibration of Large Language Models and Alignment

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.919947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.919947Z digest=sha256:e99ca786abb3a36e3687e60b8b506b6263027f1108effe659886c76c05126e63

Observation 9312365d-b111-474d-b068-e16fa8b1bee6 · outbound

This paper cites Language Models (Mostly) Know What They Know.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Language Models (Mostly) Know What They Know

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.922990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.922990Z digest=sha256:c50c2951e54ddc7fbe50ce60f7a1fc08218706ccd8be796b052d5a4d5568786f

Observation 56d76232-f930-4dd0-af5b-52442956b335 · outbound

This paper cites Rank analysis of incomplete block designs: I.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Rank analysis of incomplete block designs: I

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:16:32.631213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T12:16:31.927134Z digest=sha256:92af389fe298b536b8213264e47cbe737be879e29cd4fac421c352a95b42483a

Observation 83463c92-b4e6-4ccd-8540-2d4bb110ed37 · outbound

This paper cites Generalized Preference Optimization: A Unified Approach to Offline Alignment.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Generalized Preference Optimization: A Unified Approach to Offline Alignment

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.930130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.930130Z digest=sha256:d31a9b56106292577bb50bf121f5eed9e2e5fc6ca5f9a6003c9a94a5c56b6a82

Observation a4256de8-453d-45da-95c8-e554d4f05985 · outbound

This paper cites How to Leverage Demonstration Data in Alignment for Large Language Model? A Self-Imitation Learning Perspective.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment How to Leverage Demonstration Data in Alignment for Large Language Model? A Self-Imitation Learning Perspective

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.933188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.933188Z digest=sha256:f7b1a402c34f3cbc2537b3f7ff0e908921e407557aeba5f65aebac8e00379762

Observation 0943c1fb-e0db-48ed-9cf5-9f167b4771ef · outbound

This paper cites SimPO: Simple Preference Optimization with a Reference-Free Reward.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.937213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.937213Z digest=sha256:8b0b705c5a1b5ec36d4ba393b51c1b2202295299c10bc29574d64eb4ad138d13

Observation ac1a2811-daad-4d69-bb65-a0193be5fac8 · outbound

This paper cites Reinforced Self-Training (ReST) for Language Modeling.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Reinforced Self-Training (ReST) for Language Modeling

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.940300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.940300Z digest=sha256:c2e27b1789ecd1a42b7c8613c8014cc1c503c668648f80a09dfb6c72b81e6911

Observation 25e9db84-767a-48ea-b70d-e03cd68fe5a9 · outbound

This paper cites Openchat: Advancing open-source language models with mixed-quality data.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Openchat: Advancing open-source language models with mixed-quality data

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:16:32.621160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T12:16:31.944005Z digest=sha256:5d90bb547178211726e4b5bde153998a6c2946715a3da639239725ae303795da

Observation 84b5bfc4-f444-45dd-ac05-f508bab73996 · outbound

This paper cites Fine-Tuning Language Models with Advantage-Induced Policy Alignment.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Fine-Tuning Language Models with Advantage-Induced Policy Alignment

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.947063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.947063Z digest=sha256:ed97b6c07c9344762b90fe17a3304728672e43dda460ac5808013b7705fb7945

Observation fa582487-b8bf-4e08-9c55-7a73db68d3b8 · outbound

This paper cites Machine learning: a probabilistic perspective.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Machine learning: a probabilistic perspective

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.949861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.949861Z digest=sha256:bb6bdf5c99cae9ad48a600f578e3ed6c462c27d415cf5ffe83666c6db6ce7f38

Observation 42d83879-5ccc-4995-84eb-62af187d97d4 · outbound

This paper cites Improving Policy Gradient by Exploring Under-appreciated Rewards.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Improving Policy Gradient by Exploring Under-appreciated Rewards

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.952713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.952713Z digest=sha256:5ca2c2313323065d7746d6c64d60dff020259ddbf1ef2c0c23788670ca3c185d

Observation 2a03ab81-e9a4-4ed8-a15e-156dba954253 · outbound

This paper cites Learning how to propagate messages in graph neural networks.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Learning how to propagate messages in graph neural networks

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:16:32.602814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T12:16:31.956300Z digest=sha256:056d1ea5fe473598dae35e07c9229afcb4f69d49509db93258a608535082ff39

Observation 00ecfb45-e50e-47c1-b8fa-c9e7af06eb87 · outbound

This paper cites Learning to generalize from sparse and underspecified rewards.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Learning to generalize from sparse and underspecified rewards

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:16:32.591645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T12:16:31.958843Z digest=sha256:368056627282a896d96ea38b68ed435040358b5dab2d5e6938519e77468c4cec

Observation f71633dc-4925-4aad-8ec1-3b59b4422e85 · outbound

This paper cites Decoupled self-supervised learning for graphs.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Decoupled self-supervised learning for graphs

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:16:32.581460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T12:16:31.961637Z digest=sha256:491c9efd8496c19242b626ce9144fcc73cf0561cd977cad37e5c8863b38d58b2

Observation f57ea6f7-d3af-4c71-b8aa-df3d3e2ac0ef · outbound

This paper cites UltraFeedback: Boosting Language Models with Scaled AI Feedback.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment UltraFeedback: Boosting Language Models with Scaled AI Feedback

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.964333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.964333Z digest=sha256:b02e7602b2b3f0cbd48a9715f40e749497a888363ebd26a5381c8c86e3572fc2

Observation 5a888593-3615-4760-a93b-621af2650ef9 · outbound

This paper cites Zephyr: Direct Distillation of LM Alignment.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Zephyr: Direct Distillation of LM Alignment

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.967147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.967147Z digest=sha256:7854caa46b7da0e4474a3ac9f22777c74e6b159e94ec3a2f1a2cefaee2fa85e9

Observation 8842ab19-1cfe-4140-910d-d92946430a09 · outbound

This paper cites MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.970515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.970515Z digest=sha256:d0680d856a7a2705a38168175f695d27223b5cdbd69c7dc29e88bd46d342432f

Observation 8a528a66-0fca-47d9-91f4-689ff083e0ae · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.973598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.973598Z digest=sha256:84aa508ca1755573a1bdd4a9c690f1b4e7fe54445a83d429a4f48692ff818d58

Observation 53e1e8af-cde6-4422-ac5d-4e030c0db960 · outbound

This paper cites Instruction-Following Evaluation for Large Language Models.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Instruction-Following Evaluation for Large Language Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.977059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.977059Z digest=sha256:cde336b99e2886ed1313a60808e3c0a3c3ab4f29b2f833acdf8a1833c96e381f

Observation 38d0ece0-5b9d-4acd-89ef-12140dc30d40 · outbound

This paper cites Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.980152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.980152Z digest=sha256:f28062a7b812684df58ce14854dd040b7d87c140cd88c0f71f2ad76e8c0aa8ee

Observation 9f242457-b2f7-4380-9e7b-3bcb877d3602 · outbound

This paper cites GPQA: A Graduate-Level Google-Proof Q&A Benchmark.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment GPQA: A Graduate-Level Google-Proof Q&A Benchmark

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.983210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.983210Z digest=sha256:4649c0c68196bc009fddfae6b5c9dbf6e7f2cf7fe5fef2db51672318c120bc68

Observation 085c13ee-d299-423b-b2c1-408a20437e10 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Training Verifiers to Solve Math Word Problems

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.986326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.986326Z digest=sha256:e4bcdb4f766c546d1a47b4287e4988e030c0afaaf37519147e84d9787ac89ae5

Observation f68d5328-0d1d-41ed-86d2-480ad6135bdf · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Measuring Mathematical Problem Solving With the MATH Dataset

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.989194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.989194Z digest=sha256:95822dad5dc05c23453f82cf2743a925afb2d887a25549e3f09e24b84a9910bb

Observation 959e2286-93db-4f1d-bf5b-cc7288c8181c · outbound

This paper cites Alpacaeval: An automatic evaluator of instruction-following models, 2023.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Alpacaeval: An automatic evaluator of instruction-following models, 2023

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.991750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.991750Z digest=sha256:612d05f81e838f27659b077fd8dd19d40727d4b2a8abc2b9401b7d65365a9870

Observation b5c75b4a-1225-4dd4-bbf8-58a1b9a2280f · outbound

This paper cites Pythia: A suite for analyzing large language models across training and scaling.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Pythia: A suite for analyzing large language models across training and scaling

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.994540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.994540Z digest=sha256:e7107361a1a9d5a63f95b07701dfd75e24e228a44699709cb0a7005df6a7c583

Observation 3eecb998-3173-402e-b8d0-ea74583ca919 · outbound

This paper cites Language models are unsupervised multitask learners.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Language models are unsupervised multitask learners

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:16:32.559450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T12:16:31.997113Z digest=sha256:3c3ecba170a11c3ea2fdbc636a70989a0731cbfee4e3ff0c41062c6de5b224b5

Observation d2f9f718-aa82-40fd-96c9-61af13d31330 · outbound

This paper cites Contrastive Preference Optimization: Pushing the Boundaries of LLM Performance in Machine Translation.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Contrastive Preference Optimization: Pushing the Boundaries of LLM Performance in Machine Translation

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.999493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.999493Z digest=sha256:52e036c1282d2cc1743410ef26b9ef045ef0b104f08af511e99c468249db8f1c

Observation dc65e0f0-fbbf-4f39-a91c-6616e4399414 · outbound

This paper cites Beyond reverse kl: Generalizing direct preference optimization with diverse divergence constraints.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Beyond reverse kl: Generalizing direct preference optimization with diverse divergence constraints

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:16:32.549991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T12:16:32.003524Z digest=sha256:724840f9ec59898626c80830d5ebe9f97a142ffe236ee6bb0f0050aef348e270

Observation 650830fe-ea77-4a4c-b4dd-a0dc74095e6d · outbound

This paper cites Iterative Reasoning Preference Optimization.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Iterative Reasoning Preference Optimization

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:32.006296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:32.006296Z digest=sha256:c79a963b088c41c1dfa8988679e46ef082599a5a3afc9b4deb232a11d6e540f3

Observation 33bd544e-d5b4-4168-b0a0-581e3d3d6058 · outbound

This paper cites Information, divergence and risk for binary experiments.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Information, divergence and risk for binary experiments

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:16:32.540757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T12:16:32.010038Z digest=sha256:d3b45ef51f61db485baa758b11fa35148662346ded8c5e5cea2d4444b080f47d

Observation da37e0fa-1e0f-437f-9935-12465a4d6df0 · outbound

This paper cites Reward augmented maximum likelihood for neural structured prediction.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Reward augmented maximum likelihood for neural structured prediction

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:16:32.530302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T12:16:32.012829Z digest=sha256:2118d87d51bb2f2f61473522371b1c948de6c7cbd1ac514d0ca756436faefb65

Observation 05b24dda-9261-44f0-9a92-abeb00905921 · outbound

This paper cites Clustering with bregman divergences.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Clustering with bregman divergences

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:16:32.520990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T12:16:32.016829Z digest=sha256:5b9ae46b35589256110cfd714fc423fe7883c60eb322074b92ade9aef93cc1ac

Observation 41168029-a0b7-4045-82c1-31903ff9efa5 · outbound

This paper cites Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:32.019439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:32.019439Z digest=sha256:334cbb1a30cbc4019c903e1a962a8177b5a07b5ca51525fcc0eed0b78d7e5f9d

Observation 484e4191-a869-4b78-bbf9-af172e8cac7c · outbound

This paper cites Understanding the Effects of RLHF on LLM Generalisation and Diversity.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:32.022520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:32.022520Z digest=sha256:05c60ea675ce29aa49ccc6be42b7e3bfcdbc7ae0bef912a885f52fa2d94d96a4

Observation 8210680e-89c7-418e-bff8-ce1a1fcc093d · outbound

This paper cites Guidelines: • The answer NA means that the abstract and introduction do not include the claims made in the paper.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Guidelines: • The answer NA means that the abstract and introduction do not include the claims made in the paper

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:16:32.511971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T12:16:32.027443Z digest=sha256:bcdd20bb4d5448497716b5d035f72085e7a09ff3aa8f9a3164bb8f253158a1c8

Observation 26608078-d248-42dd-abf9-23a4eb8d7b11 · outbound

This paper cites Limitations.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Limitations

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:16:32.501635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T12:16:32.030919Z digest=sha256:132f299ea5cd27dd98fc2ed8d0d09e5969f905fdfc4200553c7e47a9a8a22141

Observation 4884cd23-33a9-4595-8b34-3bae026efaf8 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not include theoretical results.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Guidelines: • The answer NA means that the paper does not include theoretical results

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:16:32.489123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T12:16:32.034227Z digest=sha256:3967f37b982a469be8f4e718f93c298b566fa5e8560e1327cf641440b3592e68

Observation 4826bd82-96f4-4751-875e-7cee43bda8b5 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not include experiments.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Guidelines: • The answer NA means that the paper does not include experiments

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:16:32.476708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T12:16:32.038330Z digest=sha256:6ff482fed53167900778558ceb0f8b8c28dbed4bbf03ad70c37185a607e82381

Observation c0db1cac-3a60-4de3-8d6a-1daad2ad7748 · outbound

This paper cites • Please see the NeurIPS code and data submission guidelines ( https://nips.cc/ public/guides/CodeSubmissionPolicy) for more details.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment • Please see the NeurIPS code and data submission guidelines ( https://nips.cc/ public/guides/CodeSubmissionPolicy) for more details

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:16:32.467290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T12:16:32.041960Z digest=sha256:b23a5ee7d203860bd54be390d00250d4848497bf5ebb56f97fe9eea5ff8a4b21

Observation aec86a50-2589-4464-aa65-e396144e2d65 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not include experiments.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Guidelines: • The answer NA means that the paper does not include experiments

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:16:32.455381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T12:16:32.044849Z digest=sha256:cdbdfdb791e41a3daa15242d1fbd93fd5852de09a4747209808a449fcc8dd81c

Observation 06642fd1-12b3-4dcd-bb92-f67df389f24a · outbound

This paper cites Guidelines: • The answer NA means that the paper does not include experiments.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Guidelines: • The answer NA means that the paper does not include experiments

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:16:32.446467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T12:16:32.047698Z digest=sha256:e5801bda21ba3d933a0d2371f10553e0ca498d1a71ff8be7f69d1c896f72994e

Observation 28738ff5-4694-48a8-8d58-87fd63ab9128 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not include experiments.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Guidelines: • The answer NA means that the paper does not include experiments

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:16:32.437364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T12:16:32.050305Z digest=sha256:826cfce2a03359c8c080d098d7d6a43c41629708195a96be5af1777af1b3e46b

Observation 30282f20-6d4c-4f28-aed4-c4af1e1fea3c · outbound

This paper cites Guidelines: • The answer NA means that the authors have not reviewed the NeurIPS Code of Ethics.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Guidelines: • The answer NA means that the authors have not reviewed the NeurIPS Code of Ethics

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:16:32.426109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T12:16:32.052865Z digest=sha256:20fcb779b940535d9cb099fd3acaa879580682f0213d258f26124fa9435ceb9f

Observation f9cea911-8a11-41e2-917e-369700efd3ee · outbound

This paper cites Guidelines: • The answer NA means that there is no societal impact of the work performed.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Guidelines: • The answer NA means that there is no societal impact of the work performed

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:16:32.414192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T12:16:32.055941Z digest=sha256:03334796f832e559782f067fc3e52530fdaa4c6eb18485af2ba3e3416b890e05

Observation 9bb2e117-02d4-49b7-b2ec-819966320e54 · outbound

This paper cites Guidelines: • The answer NA means that the paper poses no such risks.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Guidelines: • The answer NA means that the paper poses no such risks

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:32.058383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:32.058383Z digest=sha256:540009ade7cc5cd7486167057a968c60fc62f233c73c373df98a9818869ee880

Observation ea19d25d-37db-4e40-9eea-312ff696fdf3 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not use existing assets.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Guidelines: • The answer NA means that the paper does not use existing assets

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:16:32.397363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T12:16:32.060850Z digest=sha256:a76d1822b9e5bb26b19e116485c17b987a8f2befc24a14acb29a252fa715788d

Observation 76ebee73-8f0e-4919-9cc8-415d8e5c5bb6 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not release new assets.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Guidelines: • The answer NA means that the paper does not release new assets

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:32.064064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:32.064064Z digest=sha256:f124c2d9c20492ca2c070d47aee277ab4daba5f904dcfb607187575eb8449132

Observation 5df0804e-650e-45aa-91b6-34563bce2854 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:32.067064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:32.067064Z digest=sha256:42aae7485c7c3f2c9262b833e39dc3a9e031213fb4448de34eb383fdb53b250c

Observation ea1687f3-4a6e-4d5d-aeef-083d0bd918a7 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:16:32.375743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T12:16:32.070018Z digest=sha256:46a5179393e6523be452dbd7727cf0054d1f62c4f573d3c59cfdd5332e9157a5

Pith citing papers

Observation 5983e2f6-f1c3-49aa-999c-9598cfd3d96f · inbound

DPO-Shift: Shifting the Distribution of Direct Preference Optimization cites this paper.

DPO-Shift: Shifting the Distribution of Direct Preference Optimization Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-08T12:17:46.012360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:17:46.012360Z digest=sha256:ddc3b09065b5dde49b25737e27cd2a716302ae03b7d16d713af7a858d52f5e38

Observation b2ecad65-726b-42cc-a8d2-a16489530897 · inbound

Policy-labeled Preference Learning: Is Preference Enough for RLHF? cites this paper.

Policy-labeled Preference Learning: Is Preference Enough for RLHF? Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T23:58:38.854695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:58:38.854695Z digest=sha256:f4b54095cba6f54e3a29e3cbd2c4fd76a0aa52be3b888034c66a9fa62203b6e9

Observation 4b664141-2b19-4be9-bc82-8fb1de2a8e68 · inbound

Bridging Brains and Machines: A Unified Frontier in Neuroscience, Artificial Intelligence, and Neuromorphic Systems cites this paper.

Bridging Brains and Machines: A Unified Frontier in Neuroscience, Artificial Intelligence, and Neuromorphic Systems Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment

Reference 150

Resolution
verified exact
arxiv_id, observed 2026-05-19T04:42:04.901351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-19T04:37:33.928616Z digest=sha256:d4e50d69a8fb68ce6767a43b7369648310ec241e8b5e64ed9945a10b5b64eb1e

Observation 10e735fa-5805-42b3-a9da-57eff2c8f526 · inbound

Gradient-Gated DPO: Stabilizing Preference Optimization in Language Models cites this paper.

Gradient-Gated DPO: Stabilizing Preference Optimization in Language Models Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:20:41.691864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-08T18:35:13.659698Z digest=sha256:4d547631b1c383075a8e0db099ee3e7f66333327cd2e2ff4e80cf89c3807c8c3