Pith. sign in

Paper Citation Record · LEDGER

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment

As of 12 August 2026, this Paper Citation Record lists 87 of 87 outbound references and 3 inbound Pith citation observations for arXiv:2412.14516.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.14516 v1

Coverage vector

measured 87 of 87 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T12:16:32.070018Z

measured 90 of 90 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T12:17:46.012360Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-19T04:42:04.899091Z

Reference resolution

87 of 87 outbound references displayed

  • verified exact1
  • verified fuzzy30
  • unresolved56
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a3043f17-7c1a-4907-9627-c1f5b1960ec7 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.787908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.787908Z digest=sha256:ce58660c7010de5b89b902fd2b361c6c4ab493fa29d67860a6dfada146cc498f

Observation 7a53bf06-8323-4cfd-805a-41592a79ce75 · outbound

This paper cites Training language models to follow instructions with human feedback.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Training language models to follow instructions with human feedback

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.792312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.792312Z digest=sha256:775cab3a8569f84fcb61f48c429aee88358001dad2987ce0dedde55f892affc4

Observation f1c5331e-62a6-4b14-8f7e-9d06a1242baf · outbound

This paper cites Learning to summarize with human feedback.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Learning to summarize with human feedback

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.795485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.795485Z digest=sha256:4460e31df7aea06e5b11942d4689b437920bdcb335ca6bbbbc404c0c887d7632

Observation c21db7c7-e901-4ace-aa9d-f03d596f7616 · outbound

This paper cites Deep reinforcement learning from human preferences.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Deep reinforcement learning from human preferences

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.798253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.798253Z digest=sha256:d54d4a9d20668d476cdd14931c7c5d6f091cd6524c9e633f1663e92d0b8e30a8

Observation 4ab6c605-1107-4c47-ba23-174ba61bb750 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Proximal Policy Optimization Algorithms

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.800935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.800935Z digest=sha256:e6e17815fdd40dddcef4f85bbf44e27890e507cf9fc334aea6fd59f2fcb8aced

Observation c2e03700-8331-4b9a-9237-0832dace122f · outbound

This paper cites Implementation matters in deep policy gradients: A case study on ppo and trpo.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Implementation matters in deep policy gradients: A case study on ppo and trpo

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.805323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.805323Z digest=sha256:8184ce80469f1d6860397525d0d32f9cd799a18fc5686e146648f3642b643331

Observation c8364072-296c-4bfa-9f37-7b5b687dd2c0 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Direct preference optimization: Your language model is secretly a reward model

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.809910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.809910Z digest=sha256:e1c243f47eebee14935714c19700785af3d39765d6402dc61949686387da719e

Observation de430066-de2d-42a1-aab7-fe0e94031552 · outbound

This paper cites A general theoretical paradigm to understand learning from human preferences.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment A general theoretical paradigm to understand learning from human preferences

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.813699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.813699Z digest=sha256:f1f7950cc06fd9243af7096b2eff430ea7d896a78f52f56e19241fd71e0a1477

Observation 252bde81-682b-48cd-b034-cb66fe3413f1 · outbound

This paper cites SLiC-HF: Sequence Likelihood Calibration with Human Feedback.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment SLiC-HF: Sequence Likelihood Calibration with Human Feedback

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.816513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.816513Z digest=sha256:9abbb9ff6fac787adc252c4e8638b911056c7ca084411b747323c9dedec34160

Observation 1260564c-4074-4593-9df5-ec7ae10168ed · outbound

This paper cites Preference Fine-Tuning of LLMs Should Leverage Suboptimal, On-Policy Data.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Preference Fine-Tuning of LLMs Should Leverage Suboptimal, On-Policy Data

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.820225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.820225Z digest=sha256:8ea1f078ba4f48a063a395f7db5cd933cd0eb75cafe0111d7b048cb68f5e1365

Observation d69ef4d6-85ad-494f-a7e1-c9a4554f450f · outbound

This paper cites Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.823443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.823443Z digest=sha256:e977a357640d82c5b85d08a784bd478b0aa9436d8a7f6b6b0bab17f869ebce8f

Observation 99b3e0ae-f37d-4861-986f-cdd089439c71 · outbound

This paper cites Advancing LLM Reasoning Generalists with Preference Trees.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Advancing LLM Reasoning Generalists with Preference Trees

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.827786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.827786Z digest=sha256:57d82210f982cf963ded02dd6272c89a6d1e87dccad2546b03fc41a8157fed9b

Observation 6796dcc0-2ce7-4394-805f-483abf14982d · outbound

This paper cites Learning word vectors for sentiment analysis.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Learning word vectors for sentiment analysis

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.831179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.831179Z digest=sha256:a9ce6db9b274528a45d566292367ec7cf41590f676c9dc8fd1256de0bd40e4cc

Observation 4cd60717-3a3a-459c-8c01-687001bbea5c · outbound

This paper cites Tl; dr: Mining reddit to learn automatic summarization.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Tl; dr: Mining reddit to learn automatic summarization

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:16:32.711917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T12:16:31.834698Z digest=sha256:c2bb327d4f3b66d83721bfad33b99a7f51bd1a000411ba00289eb64c9e7b8d61

Observation edadd08e-1c6a-43ed-998e-18509afea875 · outbound

This paper cites A framework for few-shot language model evaluation, 12 2023.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment A framework for few-shot language model evaluation, 12 2023

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.838747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.838747Z digest=sha256:11fd1ec0da841ab99f0fd9cc12ceceaa28c1d87a0fea950db053cac65e597a94

Observation 65c40c16-36fe-473b-ad35-f090bd37ad5b · outbound

This paper cites Nash Learning from Human Feedback.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Nash Learning from Human Feedback

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.841413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.841413Z digest=sha256:86d9b5d9cf36c4194b66b229761a37ecefe1c706b53c67a22f6d01d9ff317476

Observation ec0e0994-acd2-4e42-a38b-9f8c176824b2 · outbound

This paper cites Statistical rejection sampling improves preference optimization.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Statistical rejection sampling improves preference optimization

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:16:32.700617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T12:16:31.844889Z digest=sha256:d4068223d53a5145e2458b0ef119f0e068b9e0c5deb319876320de66cf3e6a2c

Observation 70cce962-625c-40dd-ad7e-be59950a25ca · outbound

This paper cites Self-Rewarding Language Models.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Self-Rewarding Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.847944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.847944Z digest=sha256:d02e36a394f5ad0c06b18bf0114d852cf0e636aef764f30d5d6657775915c8a1

Observation 0762ebe0-d00d-443c-ba4f-58f67230247e · outbound

This paper cites Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.851281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.851281Z digest=sha256:3ef5eee19b8443411cca5e667f722e58431f4b5d786fca63aa287851d2b0405b

Observation 6819dafb-6643-4018-b926-430f6136ac17 · outbound

This paper cites Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Direct Nash Optimization: Teaching Language Models to Self-Improve with General Preferences

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.855210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.855210Z digest=sha256:70ae91c0efef8ab6da65ea699ff075f5063f7d10029b66a4ec2c345a1c05f468

Observation 087d7327-30ac-4fe0-9c68-53e750be321b · outbound

This paper cites Direct Language Model Alignment from Online AI Feedback.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Direct Language Model Alignment from Online AI Feedback

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.859337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.859337Z digest=sha256:39bedf5696ef871bd3462ffab12390c17254682f483d9efacfc131ade10f8611

Observation a7690f04-8146-4675-90d9-319bb5331e75 · outbound

This paper cites Simper: Simple preference fine-tuning without hyperparameters by perplexity optimization.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Simper: Simple preference fine-tuning without hyperparameters by perplexity optimization

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:16:32.693022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T12:16:31.863229Z digest=sha256:e18ae395298cf2172d6c4581887e16d98dd984d34a0d93d57567761d5f4b8328

Observation 07476bde-f31a-476d-8daa-12c33e203891 · outbound

This paper cites Policy Optimization in RLHF: The Impact of Out-of-preference Data.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Policy Optimization in RLHF: The Impact of Out-of-preference Data

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.865955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.865955Z digest=sha256:0661787d860b128487eeb66889f704c38597a4cc4e0d29b7a6adfbf145f9133a

Observation f3207bb2-2819-4a05-88bf-8338026d52c2 · outbound

This paper cites Is DPO Superior to PPO for LLM Alignment? A Comprehensive Study.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Is DPO Superior to PPO for LLM Alignment? A Comprehensive Study

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.868865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.868865Z digest=sha256:82a15f3fc29a6159b5efafbb25ea09f8928d1d134bf10b049dac47d7bc397081

Observation 00fe9418-3368-48a4-9310-8b5c3d35de57 · outbound

This paper cites Provably Robust DPO: Aligning Language Models with Noisy Feedback.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Provably Robust DPO: Aligning Language Models with Noisy Feedback

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.872076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.872076Z digest=sha256:b3d372137c4b7cfe1d8dd9a88e3107a16225154c8f32d1f8d3f09110bf283595

Observation c89ca9c3-a17a-4814-871d-11c438522a3a · outbound

This paper cites Noise Contrastive Alignment of Language Models with Explicit Rewards.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Noise Contrastive Alignment of Language Models with Explicit Rewards

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.875936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.875936Z digest=sha256:68cc79abbcca7bfec46c3880965fb2e05705f03f247d57477158ec56395a1178

Observation 4960576a-471c-4581-801a-9b1107e24467 · outbound

This paper cites COPR: Continual Human Preference Learning via Optimal Policy Regularization.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment COPR: Continual Human Preference Learning via Optimal Policy Regularization

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-08-11T12:16:32.256071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T12:16:31.879139Z digest=sha256:deb8899e91d8483af7a36bce3fe7e2c8ac17597019eecfddf325fcd310e81d90

Observation c4044b74-173e-4ab9-8a56-2590f91da535 · outbound

This paper cites Towards Efficient Exact Optimization of Language Model Alignment.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Towards Efficient Exact Optimization of Language Model Alignment

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.882201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.882201Z digest=sha256:9b54b2cc9f8e8b61108ca54280009d19bfeb929f1e429a6dff75899513b28a38

Observation 793ac3b3-1a83-4aea-9f02-c5230f473f25 · outbound

This paper cites Noise contrastive estimation and negative sampling for condi- tional models: Consistency and statistical efficiency.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Noise contrastive estimation and negative sampling for condi- tional models: Consistency and statistical efficiency

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:16:32.684069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T12:16:31.885458Z digest=sha256:8715c6e8ce7de5d5bd5b27e9994b2cbe3595b240bfe5abaeafb3bb42766bf033

Observation ea2b61fe-61ac-47a8-9026-a9ad5900318c · outbound

This paper cites Towards Analyzing and Understanding the Limitations of DPO: A Theoretical Perspective.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Towards Analyzing and Understanding the Limitations of DPO: A Theoretical Perspective

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.888937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.888937Z digest=sha256:a5c18ceffda8e9d8b396e0abeccf22680ef4833c4eadbf52365db6515f446429

Observation 53babd5a-45e5-4f9d-8b08-949cb10d1c75 · outbound

This paper cites From $r$ to $Q^*$: Your Language Model is Secretly a Q-Function.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment From $r$ to $Q^*$: Your Language Model is Secretly a Q-Function

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.891745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.891745Z digest=sha256:e7e296ec9a8df00d8ff19199508871de22607d55333bc988353d930fc282460e

Observation f4659949-2b58-4b5b-8532-ade787f3a012 · outbound

This paper cites DPO Meets PPO: Reinforced Token Optimization for RLHF.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment DPO Meets PPO: Reinforced Token Optimization for RLHF

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.895081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.895081Z digest=sha256:da648bb5c67ff79304f08568beb3fb7c39ce119bc373a965eb8235368d581fef

Observation df554fa1-4c4e-4e47-8839-3d33f48fea7a · outbound

This paper cites Token-level Direct Preference Optimization.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Token-level Direct Preference Optimization

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.898358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.898358Z digest=sha256:6de183445f6fc29f6a81baa07b47ff0c3c83198b34b3988d10fb360f9d5af0c3

Observation 1343b00a-f14a-4d03-ae9d-80f0692c6f1a · outbound

This paper cites A general offline reinforcement learning framework for interac- tive recommendation.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment A general offline reinforcement learning framework for interac- tive recommendation

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:16:32.674683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T12:16:31.903506Z digest=sha256:e91bdc54d9015603e47fb34c7e3649e51e1ffb191bba45400b30b5db4852fdb4

Observation 6f6fb95b-bc85-4058-9a5f-7997b2df7c71 · outbound

This paper cites On calibration of modern neural networks.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment On calibration of modern neural networks

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:16:32.664634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T12:16:31.906673Z digest=sha256:e47b3b086bc8ffe102fde5a74e34494dbfd223279f01409cce9f6ed3bf4cefb2

Observation 64781a79-dc7c-4d9e-82a9-962a674930b6 · outbound

This paper cites Scale calibration of deep ranking models.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Scale calibration of deep ranking models

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:16:32.655401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T12:16:31.910137Z digest=sha256:ada4f7c724ea0851a526fcbd7bb8c592152ef6c5c92d21479a302d3669340f67

Observation dfdcb30e-710b-4eb5-9e07-495b1f6bcc7d · outbound

This paper cites Calibrated model-based deep reinforcement learning.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Calibrated model-based deep reinforcement learning

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:16:32.646066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T12:16:31.912682Z digest=sha256:3ce350c10a06ef1255c5fefa15ff361c75200e17ea4bb97037994ea6f3d7dba6

Observation 5714eafd-90fb-415e-b599-e941e60eb9b0 · outbound

This paper cites Cal-ql: Calibrated offline rl pre-training for efficient online fine-tuning.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Cal-ql: Calibrated offline rl pre-training for efficient online fine-tuning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.916373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.916373Z digest=sha256:85181e95cd9290f68c3683e49bbd728f656ad1edd833442e56c653c81a1986ce

Observation 8437b2c8-1af5-4ddd-9ab1-d93fa3439987 · outbound

This paper cites On the Calibration of Large Language Models and Alignment.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment On the Calibration of Large Language Models and Alignment

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.919947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.919947Z digest=sha256:d8f4d12fb18653ea8940248d931b3e7cd27d19a7249e1e6e3a6651a14810e184

Observation 9312365d-b111-474d-b068-e16fa8b1bee6 · outbound

This paper cites Language Models (Mostly) Know What They Know.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Language Models (Mostly) Know What They Know

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.922990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.922990Z digest=sha256:8a4eeee0e0179acf911d5b060a630b5bc7933deadd6bbecf754fa3965a4a2696

Observation 56d76232-f930-4dd0-af5b-52442956b335 · outbound

This paper cites Rank analysis of incomplete block designs: I.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Rank analysis of incomplete block designs: I

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:16:32.631213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T12:16:31.927134Z digest=sha256:837f7fd1e5ae81307916a3f19bf6abd90c6f93164f6654a0748fa55141fdc490

Observation 83463c92-b4e6-4ccd-8540-2d4bb110ed37 · outbound

This paper cites Generalized Preference Optimization: A Unified Approach to Offline Alignment.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Generalized Preference Optimization: A Unified Approach to Offline Alignment

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.930130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.930130Z digest=sha256:b40ced8d670a2b9af5c75910dee04c598ae22ecd28dd5386c4fee491a180656a

Observation a4256de8-453d-45da-95c8-e554d4f05985 · outbound

This paper cites How to Leverage Demonstration Data in Alignment for Large Language Model? A Self-Imitation Learning Perspective.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment How to Leverage Demonstration Data in Alignment for Large Language Model? A Self-Imitation Learning Perspective

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.933188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.933188Z digest=sha256:8d7ceb4a9aac8a5139172dd3c4cec8b6b72376e6aab03a9a886a5fbfd66cab4f

Observation 0943c1fb-e0db-48ed-9cf5-9f167b4771ef · outbound

This paper cites SimPO: Simple Preference Optimization with a Reference-Free Reward.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.937213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.937213Z digest=sha256:f6ecda52dc5ec3d2380400d28835a95576951a82e4c79b4f1cf52182f06e9d76

Observation ac1a2811-daad-4d69-bb65-a0193be5fac8 · outbound

This paper cites Reinforced Self-Training (ReST) for Language Modeling.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Reinforced Self-Training (ReST) for Language Modeling

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.940300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.940300Z digest=sha256:7ffcf61f5778748f40c1049c4598af88b04f3aeef14fef28f9b1e0df7f6ee689

Observation 25e9db84-767a-48ea-b70d-e03cd68fe5a9 · outbound

This paper cites Openchat: Advancing open-source language models with mixed-quality data.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Openchat: Advancing open-source language models with mixed-quality data

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:16:32.621160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T12:16:31.944005Z digest=sha256:54e9769f3d2cc8a286bcec428d4e1c2f5d60bbe2d6df3e3d11270c2af152fd48

Observation 84b5bfc4-f444-45dd-ac05-f508bab73996 · outbound

This paper cites Fine-Tuning Language Models with Advantage-Induced Policy Alignment.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Fine-Tuning Language Models with Advantage-Induced Policy Alignment

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.947063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.947063Z digest=sha256:d772e99d1dbadd04ded65597edfc61ed1961a4f87356be9a07a730d44e0c6fdd

Observation fa582487-b8bf-4e08-9c55-7a73db68d3b8 · outbound

This paper cites Machine learning: a probabilistic perspective.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Machine learning: a probabilistic perspective

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.949861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.949861Z digest=sha256:a0a87535dd046bf88448ae82459056b0727a45f7f6f1523aa43cd84a45f73ba0

Observation 42d83879-5ccc-4995-84eb-62af187d97d4 · outbound

This paper cites Improving Policy Gradient by Exploring Under-appreciated Rewards.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Improving Policy Gradient by Exploring Under-appreciated Rewards

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.952713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.952713Z digest=sha256:1c4c597ff7979068c2069ab0e083adcaff97728abfcaef80bc07c02f43d45fe9

Observation 2a03ab81-e9a4-4ed8-a15e-156dba954253 · outbound

This paper cites Learning how to propagate messages in graph neural networks.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Learning how to propagate messages in graph neural networks

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:16:32.602814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T12:16:31.956300Z digest=sha256:e6ef6dc92e1239f6bff21e2f628f407e9be85cffd5430bd689df864cff6e5474

Observation 00ecfb45-e50e-47c1-b8fa-c9e7af06eb87 · outbound

This paper cites Learning to generalize from sparse and underspecified rewards.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Learning to generalize from sparse and underspecified rewards

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:16:32.591645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T12:16:31.958843Z digest=sha256:1cb0d6899a6cf2d672f4051cfbe0499b7fd607c7966d7b34c0cd110d5043f92b

Observation f71633dc-4925-4aad-8ec1-3b59b4422e85 · outbound

This paper cites Decoupled self-supervised learning for graphs.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Decoupled self-supervised learning for graphs

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:16:32.581460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T12:16:31.961637Z digest=sha256:5dd5296daed21b1a1f720930ca71c436794482cb5fa095f0aacaba6c0f02d9ff

Observation f57ea6f7-d3af-4c71-b8aa-df3d3e2ac0ef · outbound

This paper cites UltraFeedback: Boosting Language Models with Scaled AI Feedback.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment UltraFeedback: Boosting Language Models with Scaled AI Feedback

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.964333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.964333Z digest=sha256:df81af1297371d8454babe2dac479990aca0e866eb2bba751832806c62c1cef2

Observation 5a888593-3615-4760-a93b-621af2650ef9 · outbound

This paper cites Zephyr: Direct Distillation of LM Alignment.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Zephyr: Direct Distillation of LM Alignment

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.967147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.967147Z digest=sha256:26bab478491a75f34ecaba5532a97109df13cefc649d46d516e8272a1cb43e41

Observation 8842ab19-1cfe-4140-910d-d92946430a09 · outbound

This paper cites MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.970515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.970515Z digest=sha256:e1a6638ff7366bdb34643b36190a0be97ca0325899d0910cf6dd59b3c80682d2

Observation 8a528a66-0fca-47d9-91f4-689ff083e0ae · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.973598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.973598Z digest=sha256:134c779fbb51bcc53fd7dc4441cb688cb642ed0db02ca2f5f8cc240521bdb654

Observation 53e1e8af-cde6-4422-ac5d-4e030c0db960 · outbound

This paper cites Instruction-Following Evaluation for Large Language Models.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Instruction-Following Evaluation for Large Language Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.977059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.977059Z digest=sha256:ee10f5923a0e1c75a3d21c5028e6639096064aa483b95c88a7efc79e3322945e

Observation 38d0ece0-5b9d-4acd-89ef-12140dc30d40 · outbound

This paper cites Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.980152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.980152Z digest=sha256:7feb4abcc0c03a3ab68375c945a55b28405c3c2b2d5959b9a5ec02386b38e2cd

Observation 9f242457-b2f7-4380-9e7b-3bcb877d3602 · outbound

This paper cites GPQA: A Graduate-Level Google-Proof Q&A Benchmark.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment GPQA: A Graduate-Level Google-Proof Q&A Benchmark

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.983210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.983210Z digest=sha256:cd1c4074697aa474db577b52e5c9fc172a1cef7d8684c67ed3951be2cb68f550

Observation 085c13ee-d299-423b-b2c1-408a20437e10 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Training Verifiers to Solve Math Word Problems

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.986326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.986326Z digest=sha256:99e31c46977c98e5b60f15bb9dba296278fae3715a9d5b83789782283f735ffe

Observation f68d5328-0d1d-41ed-86d2-480ad6135bdf · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Measuring Mathematical Problem Solving With the MATH Dataset

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.989194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.989194Z digest=sha256:aac6ecf61bbddd0426f6205e4e79d050ba7fdd39195ed2ac7e80f275b730622b

Observation 959e2286-93db-4f1d-bf5b-cc7288c8181c · outbound

This paper cites Alpacaeval: An automatic evaluator of instruction-following models, 2023.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Alpacaeval: An automatic evaluator of instruction-following models, 2023

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.991750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.991750Z digest=sha256:1f9d7c9bb684067db73b1d21428122d6a7747eb96526b850d1c21c053148ec57

Observation b5c75b4a-1225-4dd4-bbf8-58a1b9a2280f · outbound

This paper cites Pythia: A suite for analyzing large language models across training and scaling.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Pythia: A suite for analyzing large language models across training and scaling

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.994540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.994540Z digest=sha256:981bfa9728ee87f1d50b0ca45f05bf8e22eff27804469fec862a3db55465b05b

Observation 3eecb998-3173-402e-b8d0-ea74583ca919 · outbound

This paper cites Language models are unsupervised multitask learners.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Language models are unsupervised multitask learners

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:16:32.559450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T12:16:31.997113Z digest=sha256:4de85cf50f94261623b83511fba7b221236d34a0e47abed2afd069dcc8cabaf4

Observation d2f9f718-aa82-40fd-96c9-61af13d31330 · outbound

This paper cites Contrastive Preference Optimization: Pushing the Boundaries of LLM Performance in Machine Translation.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Contrastive Preference Optimization: Pushing the Boundaries of LLM Performance in Machine Translation

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:31.999493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:31.999493Z digest=sha256:50b5390ab0b296dcaa6cd3f76b4cb2b124f77561f296d603ccbbc35d1acb0390

Observation dc65e0f0-fbbf-4f39-a91c-6616e4399414 · outbound

This paper cites Beyond reverse kl: Generalizing direct preference optimization with diverse divergence constraints.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Beyond reverse kl: Generalizing direct preference optimization with diverse divergence constraints

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:16:32.549991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T12:16:32.003524Z digest=sha256:ef52d2cd51622e3b87a913bcd056c0709933c49faa1e9b8259fc31da28ac4d37

Observation 650830fe-ea77-4a4c-b4dd-a0dc74095e6d · outbound

This paper cites Iterative Reasoning Preference Optimization.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Iterative Reasoning Preference Optimization

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:32.006296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:32.006296Z digest=sha256:bd42343c326c918e2edc86cd2aeb22b4727a1ecf361f5571dd4d27a6def11832

Observation 33bd544e-d5b4-4168-b0a0-581e3d3d6058 · outbound

This paper cites Information, divergence and risk for binary experiments.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Information, divergence and risk for binary experiments

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:16:32.540757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T12:16:32.010038Z digest=sha256:0e5480b18285ae36bb4ce8682b7baccbc9109e8c10049d69a26eedf34cf2853c

Observation da37e0fa-1e0f-437f-9935-12465a4d6df0 · outbound

This paper cites Reward augmented maximum likelihood for neural structured prediction.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Reward augmented maximum likelihood for neural structured prediction

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:16:32.530302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T12:16:32.012829Z digest=sha256:76fd1944001549a092301161cfdc5ec1b74ad4abd7d3637ee542f476e70344fb

Observation 05b24dda-9261-44f0-9a92-abeb00905921 · outbound

This paper cites Clustering with bregman divergences.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Clustering with bregman divergences

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:16:32.520990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T12:16:32.016829Z digest=sha256:60d0dab46d67beaaa60b7e45635932403302993b71f4850809cf708b9007d88a

Observation 41168029-a0b7-4045-82c1-31903ff9efa5 · outbound

This paper cites Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:32.019439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:32.019439Z digest=sha256:6cff528dd9ce46309fd5e6678f503331b5ebe20149811f7ae68fb9b9f081121f

Observation 484e4191-a869-4b78-bbf9-af172e8cac7c · outbound

This paper cites Understanding the Effects of RLHF on LLM Generalisation and Diversity.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:32.022520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:32.022520Z digest=sha256:9c15287346e70f5e4aee2a9a7d861f6d7e06f3385956a72023719499a0279275

Observation 8210680e-89c7-418e-bff8-ce1a1fcc093d · outbound

This paper cites Guidelines: • The answer NA means that the abstract and introduction do not include the claims made in the paper.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Guidelines: • The answer NA means that the abstract and introduction do not include the claims made in the paper

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:16:32.511971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T12:16:32.027443Z digest=sha256:ad004075c335117bbed76671877307e6056c0993e2dcccb07015b590af0a7417

Observation 26608078-d248-42dd-abf9-23a4eb8d7b11 · outbound

This paper cites Limitations.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Limitations

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:16:32.501635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T12:16:32.030919Z digest=sha256:ff5a53e9304635b454027b4944e869aa1be8859fefa6ae192774cc74600d1b31

Observation 4884cd23-33a9-4595-8b34-3bae026efaf8 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not include theoretical results.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Guidelines: • The answer NA means that the paper does not include theoretical results

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:16:32.489123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T12:16:32.034227Z digest=sha256:36ea4b0946d9cb5795bbf75dc65c4701cd9f5c94dda2984f0ada0498bb675443

Observation 4826bd82-96f4-4751-875e-7cee43bda8b5 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not include experiments.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Guidelines: • The answer NA means that the paper does not include experiments

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:16:32.476708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T12:16:32.038330Z digest=sha256:b1bcf09ba172e6cc20543d298a742ec4998f17f3f2715457d0226141a1f0d03c

Observation c0db1cac-3a60-4de3-8d6a-1daad2ad7748 · outbound

This paper cites • Please see the NeurIPS code and data submission guidelines ( https://nips.cc/ public/guides/CodeSubmissionPolicy) for more details.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment • Please see the NeurIPS code and data submission guidelines ( https://nips.cc/ public/guides/CodeSubmissionPolicy) for more details

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:16:32.467290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T12:16:32.041960Z digest=sha256:3b8012eec71e23dd119c6bdb00d3d65cfaea4492a0696b63b09c9932a02a8d61

Observation aec86a50-2589-4464-aa65-e396144e2d65 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not include experiments.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Guidelines: • The answer NA means that the paper does not include experiments

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:16:32.455381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T12:16:32.044849Z digest=sha256:e83a41a451389ac13ed063e6f7ee9415add2c94d8021f2586c17c9843fd5f849

Observation 06642fd1-12b3-4dcd-bb92-f67df389f24a · outbound

This paper cites Guidelines: • The answer NA means that the paper does not include experiments.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Guidelines: • The answer NA means that the paper does not include experiments

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:16:32.446467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T12:16:32.047698Z digest=sha256:e555bf2bc8f2994f207db4d2b563d22adbc00602b8f2f34ef4eff8f3d59ecbcf

Observation 28738ff5-4694-48a8-8d58-87fd63ab9128 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not include experiments.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Guidelines: • The answer NA means that the paper does not include experiments

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:16:32.437364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T12:16:32.050305Z digest=sha256:2e8cf53ee91b9a160c3c1bc7c38d881bcd82be5ff92e8c86371b997539365a90

Observation 30282f20-6d4c-4f28-aed4-c4af1e1fea3c · outbound

This paper cites Guidelines: • The answer NA means that the authors have not reviewed the NeurIPS Code of Ethics.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Guidelines: • The answer NA means that the authors have not reviewed the NeurIPS Code of Ethics

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:16:32.426109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T12:16:32.052865Z digest=sha256:4624ed0bfb05ae65b156374222fa1ca44d25bba84517348d4dacea31d6d83146

Observation f9cea911-8a11-41e2-917e-369700efd3ee · outbound

This paper cites Guidelines: • The answer NA means that there is no societal impact of the work performed.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Guidelines: • The answer NA means that there is no societal impact of the work performed

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:16:32.414192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T12:16:32.055941Z digest=sha256:1bad35882543cbe7b7a9964ada46e4e342bfe1cde32659cb2c30be18fe75f47f

Observation 9bb2e117-02d4-49b7-b2ec-819966320e54 · outbound

This paper cites Guidelines: • The answer NA means that the paper poses no such risks.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Guidelines: • The answer NA means that the paper poses no such risks

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:32.058383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:32.058383Z digest=sha256:2d6c6bdf10878de743080e6d3176678b4af85572c52d827a6a92ce6607ae9866

Observation ea19d25d-37db-4e40-9eea-312ff696fdf3 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not use existing assets.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Guidelines: • The answer NA means that the paper does not use existing assets

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:16:32.397363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T12:16:32.060850Z digest=sha256:ccd3f771babd6088eb67d90a20d4440ebc20f1fadf89501a73573143b834b418

Observation 76ebee73-8f0e-4919-9cc8-415d8e5c5bb6 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not release new assets.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Guidelines: • The answer NA means that the paper does not release new assets

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:32.064064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:32.064064Z digest=sha256:9a94a1cc3e58409a66f6cdc1a458dfc22360267572065f73de352f1f21e3d31a

Observation 5df0804e-650e-45aa-91b6-34563bce2854 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-11T12:16:32.067064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:16:32.067064Z digest=sha256:4a1e73eb53dd09c494bf786461dd0add4440a58f2cac467f158ba21b41342435

Observation ea1687f3-4a6e-4d5d-aeef-083d0bd918a7 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects.

Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:16:32.375743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T12:16:32.070018Z digest=sha256:a4aed3fdd0cf0a966af40c706a47c9da70c7f978f2d7e2fe6f35d7faefd2891a

Pith citing papers

Observation 5983e2f6-f1c3-49aa-999c-9598cfd3d96f · inbound

DPO-Shift: Shifting the Distribution of Direct Preference Optimization cites this paper.

DPO-Shift: Shifting the Distribution of Direct Preference Optimization Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-08T12:17:46.012360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:17:46.012360Z digest=sha256:3b06b3c7aafd9b25052f84daf4757b4060c78c2f91d40c1c515c77f7f0967181

Observation 4b664141-2b19-4be9-bc82-8fb1de2a8e68 · inbound

Bridging Brains and Machines: A Unified Frontier in Neuroscience, Artificial Intelligence, and Neuromorphic Systems cites this paper.

Bridging Brains and Machines: A Unified Frontier in Neuroscience, Artificial Intelligence, and Neuromorphic Systems Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment

Reference 150

Resolution
verified exact
arxiv_id, observed 2026-05-19T04:42:04.901351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-19T04:37:33.928616Z digest=sha256:ac6deb2e44ddd4d6d43a961d3e5f46d63267d5514751a8de92397561e60193e8

Observation 10e735fa-5805-42b3-a9da-57eff2c8f526 · inbound

Gradient-Gated DPO: Stabilizing Preference Optimization in Language Models cites this paper.

Gradient-Gated DPO: Stabilizing Preference Optimization in Language Models Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:20:41.691864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-08T18:35:13.659698Z digest=sha256:3aea889301f3b6fc641ea7395c682d13d584ffed0883b8188b2f601ea6a4ccee