Pith. sign in

Paper Citation Record · LEDGER

Energy-Based Preference Model Offers Better Offline Alignment than the Bradley-Terry Preference Model

As of 21 August 2026, this Paper Citation Record lists 54 of 54 outbound references and 1 inbound Pith citation observation for arXiv:2412.13862.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.13862 v1

Coverage vector

measured 54 of 54 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T12:48:59.097529Z

measured 55 of 55 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-10T19:43:51.965882Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-10T22:35:49.296372Z

Reference resolution

54 of 54 outbound references displayed

  • verified exact2
  • verified fuzzy18
  • unresolved34
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4e643310-f3c9-4ac2-9b82-7e2c1708c6fb · outbound

This paper cites Direct Preference Optimization with an Offset.

Energy-Based Preference Model Offers Better Offline Alignment than the Bradley-Terry Preference Model Direct Preference Optimization with an Offset

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T12:48:58.815716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:48:58.815716Z digest=sha256:ff3a2e1f8d3e5498e8672435e6142b1276455e4ac1495c23ed1593630fa29a6e

Observation 15dea84b-555b-4a49-98f8-50a8e061c7ca · outbound

This paper cites A General Theoretical Paradigm to Understand Learning from Human Preferences.

Energy-Based Preference Model Offers Better Offline Alignment than the Bradley-Terry Preference Model A General Theoretical Paradigm to Understand Learning from Human Preferences

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T12:48:58.822647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:48:58.822647Z digest=sha256:9cf5a45a6ea00b30b8ea77cd479896ff98d8c93b5e2a7bafa9b02dfb6458c2e2

Observation d45fb7b6-1d43-4ff6-93e1-9fbb14c543af · outbound

This paper cites and Rinaldo, A.

Energy-Based Preference Model Offers Better Offline Alignment than the Bradley-Terry Preference Model and Rinaldo, A

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:49:00.055997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T12:48:58.828616Z digest=sha256:c51eb5c3187cea3fdbf4f2b1f178bd0fd2921e2b97e6ad475ae3a79cb8225ac0

Observation ff5a3d7a-a53f-47e6-8719-9a5a608aa9e8 · outbound

This paper cites Noise Contrastive Alignment of Language Models with Explicit Rewards.

Energy-Based Preference Model Offers Better Offline Alignment than the Bradley-Terry Preference Model Noise Contrastive Alignment of Language Models with Explicit Rewards

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T12:48:58.834297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:48:58.834297Z digest=sha256:9a9f1a3f9bb0f58962e00ada46ee0fa36567744c768688741d227cb10f0d8835

Observation 0e03fce6-5637-485c-bebf-5d07130da182 · outbound

This paper cites Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models.

Energy-Based Preference Model Offers Better Offline Alignment than the Bradley-Terry Preference Model Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T12:48:58.839933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:48:58.839933Z digest=sha256:f4d72230a7f0428fd85963c55d44c7a7f91da34633ed1a2910ba23a87738162f

Observation 4a858c7f-d256-4759-a084-cacf9b754644 · outbound

This paper cites F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D.

Energy-Based Preference Model Offers Better Offline Alignment than the Bradley-Terry Preference Model F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T12:48:58.846024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:48:58.846024Z digest=sha256:a9005660836f925c1efb8f58613f548c1fdebf8ab5a261cd68d2ea9219438094

Observation 91b6e351-45dd-495b-bb75-784b1921c051 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Energy-Based Preference Model Offers Better Offline Alignment than the Bradley-Terry Preference Model Training Verifiers to Solve Math Word Problems

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T12:48:58.851831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:48:58.851831Z digest=sha256:6b6248244cda65c243d78f67c6198a7e2faa79c6cf8ccd71030ac35a8f644eba

Observation 1eaa6f73-9c6b-4db4-842a-de3ce7b7805d · outbound

This paper cites Ultrafeedback: Boosting language models with scaled ai feedback.

Energy-Based Preference Model Offers Better Offline Alignment than the Bradley-Terry Preference Model Ultrafeedback: Boosting language models with scaled ai feedback

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T12:48:58.857162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:48:58.857162Z digest=sha256:f7ec86d76b13d2200a0eaec36408d2a7d8348c3b8b744779803414e2224e09bc

Observation 86a88aec-301e-405e-88f4-e4568f17e852 · outbound

This paper cites Learning discrete energy-based models via auxiliary-variable local exploration.

Energy-Based Preference Model Offers Better Offline Alignment than the Bradley-Terry Preference Model Learning discrete energy-based models via auxiliary-variable local exploration

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T12:48:58.862036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:48:58.862036Z digest=sha256:5f0a961bea895999469b5bee428d0977beb55d9f19d1bfaddd89d6dc240b405f

Observation 449f708d-6ae1-4dbf-adcb-6ce0e9346f12 · outbound

This paper cites Residual energy-based models for text generation.

Energy-Based Preference Model Offers Better Offline Alignment than the Bradley-Terry Preference Model Residual energy-based models for text generation

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:48:59.998228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T12:48:58.866839Z digest=sha256:591d1bc0133055e95b48e57a4bafd824f8d4eba76a66dd18e275374bd143c34a

Observation 9f44cb8c-944d-4a1a-9a66-29859631b523 · outbound

This paper cites Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators.

Energy-Based Preference Model Offers Better Offline Alignment than the Bradley-Terry Preference Model Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T12:48:58.872968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:48:58.872968Z digest=sha256:cec57a339c2dca70620ed13171960e7c52549f4e8c253c162234dc9b002648fd

Observation d350913b-7986-4f91-888a-acf9c9761421 · outbound

This paper cites R., Elsahar, H., and Dymetman, M.

Energy-Based Preference Model Offers Better Offline Alignment than the Bradley-Terry Preference Model R., Elsahar, H., and Dymetman, M

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:48:59.981642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T12:48:58.878045Z digest=sha256:4b81c4fca837c1f92dc966ccf15cfaa7f771b1de8963fd874b7fb093dadaa9a1

Observation 828121dc-748d-4147-a966-1caf2d68abc2 · outbound

This paper cites Kto: Model alignment as prospect theoretic optimization, 2024.

Energy-Based Preference Model Offers Better Offline Alignment than the Bradley-Terry Preference Model Kto: Model alignment as prospect theoretic optimization, 2024

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:48:59.964578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T12:48:58.882894Z digest=sha256:3a425d908a74251d108649f7afdaaba5c7c697febd9e5cc6dc2d55cbb722ea8d

Observation 52dca5c1-f87d-4a03-ab07-7b7ca07815ca · outbound

This paper cites an unresolved cited work.

Energy-Based Preference Model Offers Better Offline Alignment than the Bradley-Terry Preference Model Unresolved cited work

Reference 14

Resolution
verified exact
raw_fallback, observed 2026-08-11T12:48:59.476360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T12:48:58.887771Z digest=sha256:e7227431c77c4eabef7ac6a917b86c23b6ac422af258707a9e1e616db0308b0a

Observation 2c811697-94e4-4243-8405-b80341263d10 · outbound

This paper cites Asymptotic theory of sparse Bradley–Terry model.

Energy-Based Preference Model Offers Better Offline Alignment than the Bradley-Terry Preference Model Asymptotic theory of sparse Bradley–Terry model

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T12:48:58.892294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:48:58.892294Z digest=sha256:448780c4bbdba19a058690120872d3af3b0adee5e4d480b8de60015a363c2501

Observation 4cbb1edf-b085-4815-be1f-0829d05f24e3 · outbound

This paper cites Minimax rate for learning from pairwise comparisons in the BTL model.

Energy-Based Preference Model Offers Better Offline Alignment than the Bradley-Terry Preference Model Minimax rate for learning from pairwise comparisons in the BTL model

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:48:59.945388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T12:48:58.897067Z digest=sha256:19ef0d2ae15a1ce1c0ad97c5c25c4d0864fb10d16433fe9636d04da05722d8e0

Observation 7eb9262a-763a-46d9-9b62-18ed5a70d129 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Energy-Based Preference Model Offers Better Offline Alignment than the Bradley-Terry Preference Model Measuring Massive Multitask Language Understanding

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T12:48:58.901587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:48:58.901587Z digest=sha256:e27f49b3b6d5b5f0ac1c01be30df0d193cfc9edf09b6cb3a86e04e96846a3f24

Observation 53924ce4-02b5-4848-a621-1a27c7fda0d4 · outbound

This paper cites ORPO: Monolithic Preference Optimization without Reference Model.

Energy-Based Preference Model Offers Better Offline Alignment than the Bradley-Terry Preference Model ORPO: Monolithic Preference Optimization without Reference Model

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T12:48:58.906625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:48:58.906625Z digest=sha256:0452dfae78ccc84539ebc587e18092b287654ba988c55c81f7f3b5eeced9d2df

Observation 116d37b5-abc7-4032-8e85-b3459118f259 · outbound

This paper cites Some extensions of score matching.

Energy-Based Preference Model Offers Better Offline Alignment than the Bradley-Terry Preference Model Some extensions of score matching

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:48:59.927093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T12:48:58.911994Z digest=sha256:38ef0c76f62f89278eb13ba775191f3d4a732f70813be665f36fc9c100edd0e3

Observation 9e6571af-398a-4cd2-9ca0-bbb2bd6782d2 · outbound

This paper cites Binary Classifier Optimization for Large Language Model Alignment.

Energy-Based Preference Model Offers Better Offline Alignment than the Bradley-Terry Preference Model Binary Classifier Optimization for Large Language Model Alignment

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T12:48:58.916634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:48:58.916634Z digest=sha256:9057305e5ed2ec022cf15347f974b060d48bed6433f9cb9829cbd64d9c2faf9a

Observation e0a39171-716c-4599-8fc1-3f1779848bcf · outbound

This paper cites and Langford, J.

Energy-Based Preference Model Offers Better Offline Alignment than the Bradley-Terry Preference Model and Langford, J

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T12:48:58.921656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:48:58.921656Z digest=sha256:575770a30161669f7d5c4ed638e3e34bbefbd70050f01460ec2bdf6d9a922d48

Observation 2bc77cfa-62e5-49d4-829c-a4b05839c4b9 · outbound

This paper cites A distributional approach to controlled text generation.

Energy-Based Preference Model Offers Better Offline Alignment than the Bradley-Terry Preference Model A distributional approach to controlled text generation

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:48:59.898378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T12:48:58.926599Z digest=sha256:723f45b4594af4379f836a800b0f0531a9ef15f5ab2586ce159cbc5b3810240a

Observation b94f1ae7-803b-44f3-b269-aba4a3ce4182 · outbound

This paper cites On reinforcement learning and distribution matching for fine-tuning language models with no catastrophic forgetting.

Energy-Based Preference Model Offers Better Offline Alignment than the Bradley-Terry Preference Model On reinforcement learning and distribution matching for fine-tuning language models with no catastrophic forgetting

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:48:59.881429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T12:48:58.931285Z digest=sha256:73ee467016f4df3184dea040406243fda962c48c96afd0749370707e46dbc3e2

Observation 58601b35-69b5-46ac-afd6-63283f1c592c · outbound

This paper cites Rl with kl penalties is better viewed as bayesian inference.

Energy-Based Preference Model Offers Better Offline Alignment than the Bradley-Terry Preference Model Rl with kl penalties is better viewed as bayesian inference

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:48:59.863106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T12:48:58.936142Z digest=sha256:eaffe13867a28353dd3b44d7e947445ba70934cccacc191eafd1272f303eeca3

Observation 2dbdaf73-3119-4d2a-bded-a33d6ce348aa · outbound

This paper cites Crafting papers on machine learning.

Energy-Based Preference Model Offers Better Offline Alignment than the Bradley-Terry Preference Model Crafting papers on machine learning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T12:48:58.941058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:48:58.941058Z digest=sha256:7480807dcae62fa8149897e3ec2f0e906a8618984ffc65028b9d2dee5db5c5ef

Observation 3df27b0e-663b-4639-9ee5-cc8616648eff · outbound

This paper cites Perturb-and-max-product: Sampling and learning in discrete energy-based models.

Energy-Based Preference Model Offers Better Offline Alignment than the Bradley-Terry Preference Model Perturb-and-max-product: Sampling and learning in discrete energy-based models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:48:59.835544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T12:48:58.946462Z digest=sha256:23eaecdb73766c6f35e9e64e2ad7e6e7941905fcbf354f6f3ad3ba116b27731c

Observation ab292aad-ea84-47e1-a8ed-acbfbb8fae1e · outbound

This paper cites A tutorial on energy-based learning.

Energy-Based Preference Model Offers Better Offline Alignment than the Bradley-Terry Preference Model A tutorial on energy-based learning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T12:48:58.951718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:48:58.951718Z digest=sha256:bd75cc18eca160a7dc5bbe8062a46353998fc4c59355348c2c8b1ddda9b7aef6

Observation 08f6e237-f9ba-4938-a3d4-ba5b03f99b89 · outbound

This paper cites Conditional strong law of large number.

Energy-Based Preference Model Offers Better Offline Alignment than the Bradley-Terry Preference Model Conditional strong law of large number

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:48:59.807814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T12:48:58.956555Z digest=sha256:4fd620d242bb8efbfc5de99127d932203ff2913dc8085c9dbcf202a09d1ba614

Observation e7ad1929-501a-4d4d-ba23-7a817d6d9d1f · outbound

This paper cites Concrete score matching: Generalized score matching for discrete data.

Energy-Based Preference Model Offers Better Offline Alignment than the Bradley-Terry Preference Model Concrete score matching: Generalized score matching for discrete data

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:48:59.791576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T12:48:58.961658Z digest=sha256:4e678784a77df1d4908386210a343bb5ae2eeb7e8090d1f3dce02633c6e2f1ba

Observation 5b26f855-819d-4343-aa5e-9c62415b335f · outbound

This paper cites SimPO: Simple Preference Optimization with a Reference-Free Reward.

Energy-Based Preference Model Offers Better Offline Alignment than the Bradley-Terry Preference Model SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T12:48:58.966786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:48:58.966786Z digest=sha256:8458b70d51a70b9c9f5fc9da2a019f07c853b4c6188572b383ec88df7066e78f

Observation 4ad4fb31-89a4-46d4-b5d0-9b9ec9d784c6 · outbound

This paper cites A note on dpo with noisy preferences & relationship to ipo, 2023.

Energy-Based Preference Model Offers Better Offline Alignment than the Bradley-Terry Preference Model A note on dpo with noisy preferences & relationship to ipo, 2023

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:48:59.774846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T12:48:58.972468Z digest=sha256:a5e800ab0c6aa1840bbec828f54425e879cba8ed2bc858c8279bb01c270cc865

Observation 15159972-dbd5-4056-8625-2a0397895fa9 · outbound

This paper cites and Szepesv \'a ri, C.

Energy-Based Preference Model Offers Better Offline Alignment than the Bradley-Terry Preference Model and Szepesv \'a ri, C

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T12:48:58.978098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:48:58.978098Z digest=sha256:77f93753c00ecca6615a05c6568cf42a1a83db89ae3420a8d2d0545d93ade4e5

Observation 005ebca7-3c69-4b12-b83b-9af5fd11e260 · outbound

This paper cites Training language models to follow instructions with human feedback.

Energy-Based Preference Model Offers Better Offline Alignment than the Bradley-Terry Preference Model Training language models to follow instructions with human feedback

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T12:48:58.983342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:48:58.983342Z digest=sha256:5c5f65f6525ba02e8e03892245e470e8b3c442bbed65b5d3945f13dad343bdd8

Observation 8f291d2b-6762-4ace-80c8-727e3371753d · outbound

This paper cites Disentangling Length from Quality in Direct Preference Optimization.

Energy-Based Preference Model Offers Better Offline Alignment than the Bradley-Terry Preference Model Disentangling Length from Quality in Direct Preference Optimization

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T12:48:58.989023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:48:58.989023Z digest=sha256:afeb5759cdd113aced1cefad3f57374c7beba35fe91e62d906ef8e1a5c2318f0

Observation dce41a7c-42ea-430f-b70b-1f3ecee24051 · outbound

This paper cites Distributional reinforcement learning for energy-based sequential models.

Energy-Based Preference Model Offers Better Offline Alignment than the Bradley-Terry Preference Model Distributional reinforcement learning for energy-based sequential models

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:48:59.736327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T12:48:58.994572Z digest=sha256:01c0a5f2fdcbedfa228782ac6ffbe218db655567cb12ddb22bc83e98788c3ad4

Observation b82470cd-8163-41d7-991a-4e2e314da0a4 · outbound

This paper cites Red teaming language models with language models.

Energy-Based Preference Model Offers Better Offline Alignment than the Bradley-Terry Preference Model Red teaming language models with language models

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:48:59.720009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T12:48:58.999638Z digest=sha256:5739af1597a798e38e2062b06c7167a4632f1675d04ea566825b93e831683fa0

Observation dcd78151-8bd9-4e8c-a7d4-3ac44e66082a · outbound

This paper cites D., Ermon, S., and Finn, C.

Energy-Based Preference Model Offers Better Offline Alignment than the Bradley-Terry Preference Model D., Ermon, S., and Finn, C

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T12:48:59.006303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:48:59.006303Z digest=sha256:80119b732d3489d6ab5e25e8769aa9b58a1e0dd742081ee76d681908893baad0

Observation ffa53833-34fc-4044-b7bc-e5ddbf7c1d55 · outbound

This paper cites an unresolved cited work.

Energy-Based Preference Model Offers Better Offline Alignment than the Bradley-Terry Preference Model Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-11T12:48:59.694059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T12:48:59.011961Z digest=sha256:7211ce46bf75176db46688fd830c6e6ec5ea4cc51b15a2ba373e27fcb337b1c2

Observation 037f79df-dffe-4495-b58c-a2ae781b06ed · outbound

This paper cites and Yao, Y.-C.

Energy-Based Preference Model Offers Better Offline Alignment than the Bradley-Terry Preference Model and Yao, Y.-C

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:48:59.678282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T12:48:59.017603Z digest=sha256:5ef48b5ef534503675e5a2a81a6368c4d92350f76565bee41949d1ccfae68c00

Observation b38a4a33-1473-433a-ac39-ebf1467aa183 · outbound

This paper cites Preference ranking optimization for human alignment.

Energy-Based Preference Model Offers Better Offline Alignment than the Bradley-Terry Preference Model Preference ranking optimization for human alignment

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:48:59.662482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T12:48:59.023076Z digest=sha256:af84f8a04718bf7471cdaf73c6b5961168d8bef0e6eae59675fdd650ef956de3

Observation fdc8a931-4580-4b3f-954a-bbd70659ff33 · outbound

This paper cites How to Train Your Energy-Based Models.

Energy-Based Preference Model Offers Better Offline Alignment than the Bradley-Terry Preference Model How to Train Your Energy-Based Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T12:48:59.028265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:48:59.028265Z digest=sha256:4605e00e19b5c84b3b0371e32836f3250ffe42aef0fe2d99eb48b6de67e3d311

Observation fd091cf0-8cdd-4fed-8e57-d72ee52be94b · outbound

This paper cites an unresolved cited work.

Energy-Based Preference Model Offers Better Offline Alignment than the Bradley-Terry Preference Model Unresolved cited work

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T12:48:59.034333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:48:59.034333Z digest=sha256:3f2346996771b9c0a7e908d1798ec12b7e5859bdbf4dbc868d5e28487d2427cc

Observation 77207307-2b79-4c21-928d-6557bafce82e · outbound

This paper cites Generalized Preference Optimization: A Unified Approach to Offline Alignment.

Energy-Based Preference Model Offers Better Offline Alignment than the Bradley-Terry Preference Model Generalized Preference Optimization: A Unified Approach to Offline Alignment

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T12:48:59.040041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:48:59.040041Z digest=sha256:b37258908fca41bdc303627a2dd1b111015a2967e105b4429188c44e840b119a

Observation ee60719a-a4d6-430a-af2a-150cba8896c4 · outbound

This paper cites It's All in the Heads: Using Attention Heads as a Baseline for Cross-Lingual Transfer in Commonsense Reasoning.

Energy-Based Preference Model Offers Better Offline Alignment than the Bradley-Terry Preference Model It's All in the Heads: Using Attention Heads as a Baseline for Cross-Lingual Transfer in Commonsense Reasoning

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T12:48:59.045427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:48:59.045427Z digest=sha256:b369db0a32b23d408175ce94d76ee069ce3d3755e1aae817232b884c1f317526

Observation 4e4f9b66-1c89-4f85-b756-cda3b00f5b2f · outbound

This paper cites Zephyr: Direct Distillation of LM Alignment.

Energy-Based Preference Model Offers Better Offline Alignment than the Bradley-Terry Preference Model Zephyr: Direct Distillation of LM Alignment

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T12:48:59.050458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:48:59.050458Z digest=sha256:06828e73800c015f8ddaa0d59c4a761438dfb80f9bdbff5e6602bbf5b307dac8

Observation 01e071f4-380c-4e77-bf70-8461579cce64 · outbound

This paper cites Asymptotic comparison of identifying constraints for Bradley-Terry models.

Energy-Based Preference Model Offers Better Offline Alignment than the Bradley-Terry Preference Model Asymptotic comparison of identifying constraints for Bradley-Terry models

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-08-11T12:48:59.205867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T12:48:59.055874Z digest=sha256:ac0791109075bb19004fdfe2d20cb627cd4e71eb1f780365b23d6bd0d01d6548

Observation fe111916-a9b7-41d7-aa9b-a64bb9ef0bf7 · outbound

This paper cites Contrastive Preference Optimization: Pushing the Boundaries of LLM Performance in Machine Translation.

Energy-Based Preference Model Offers Better Offline Alignment than the Bradley-Terry Preference Model Contrastive Preference Optimization: Pushing the Boundaries of LLM Performance in Machine Translation

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T12:48:59.061553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:48:59.061553Z digest=sha256:a206a463607e85568452fb01557a6bf8295c92b415da72f3028fc68a96383f33

Observation fbe02fec-a988-4395-8111-dad8e2699249 · outbound

This paper cites Q., Salamatian, S., Sun, Z., Suresh, A.

Energy-Based Preference Model Offers Better Offline Alignment than the Bradley-Terry Preference Model Q., Salamatian, S., Sun, Z., Suresh, A

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:48:59.635861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T12:48:59.067084Z digest=sha256:c5e366a5a7afc4c26f75950aa45c1a523513e73d203e7dfe4482e24a17f5fb5f

Observation 5d502206-73bc-4899-93c3-2aa4911776e8 · outbound

This paper cites RRHF : Rank responses to align language models with human feedback.

Energy-Based Preference Model Offers Better Offline Alignment than the Bradley-Terry Preference Model RRHF : Rank responses to align language models with human feedback

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T12:48:59.071970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:48:59.071970Z digest=sha256:064cf26a3795e4ab4de0d36050face8ea81ece3fba25a4f72c15efd022d1fc93

Observation 0c8bd438-bdf2-4a90-b12a-3584b9180ecf · outbound

This paper cites Offline reinforcement learning with realizability and single-policy concentrability.

Energy-Based Preference Model Offers Better Offline Alignment than the Bradley-Terry Preference Model Offline reinforcement learning with realizability and single-policy concentrability

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T12:48:59.077015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:48:59.077015Z digest=sha256:cfe76f73a760d9f8cd338f7460a332a8234b5469b89bd4bb06e644dc731d8393

Observation 6a98a1c6-2865-4970-b9b2-d61bd2128927 · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena.

Energy-Based Preference Model Offers Better Offline Alignment than the Bradley-Terry Preference Model Judging llm-as-a-judge with mt-bench and chatbot arena

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T12:48:59.082195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:48:59.082195Z digest=sha256:69d0485578d4bf23a03404ef1359840bafba1fcdbc780b12165cd67ac91c9910

Observation 5051a299-b8c4-4b80-b238-4c089ae2725a · outbound

This paper cites WPO: Enhancing RLHF with Weighted Preference Optimization.

Energy-Based Preference Model Offers Better Offline Alignment than the Bradley-Terry Preference Model WPO: Enhancing RLHF with Weighted Preference Optimization

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T12:48:59.087392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:48:59.087392Z digest=sha256:c3bb8d1239100a0d42482b71cd0714e3a0f529a2be166fc3d372862fb285ae7c

Observation 9e76b640-4a81-4fd8-b7cd-d66c14432c06 · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

Energy-Based Preference Model Offers Better Offline Alignment than the Bradley-Terry Preference Model Fine-Tuning Language Models from Human Preferences

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T12:48:59.092277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:48:59.092277Z digest=sha256:0947c94a33e857559586441722655ea64f55e11d66bfbbddbe89a58bbcf89b3b

Observation 0d4675f9-bdff-4f27-a786-747aa81bd8a7 · outbound

This paper cites write newline.

Energy-Based Preference Model Offers Better Offline Alignment than the Bradley-Terry Preference Model write newline

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T12:48:59.097529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:48:59.097529Z digest=sha256:3168c48664133222446e43148897892d4a3b9a38a3682a653ec847aed9aef2cd

Pith citing papers

Observation 7bc0c716-0e7d-4d43-9191-4dcd7c866d1e · inbound

Relative Density Ratio Optimization for Stable and Statistically Consistent Model Alignment cites this paper.

Relative Density Ratio Optimization for Stable and Statistically Consistent Model Alignment Energy-Based Preference Model Offers Better Offline Alignment than the Bradley-Terry Preference Model

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:35:49.303293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T19:43:51.965882Z digest=sha256:51e205dc617ff27ffc7dfddfe631b54726cb1356836c4e2223a235f049a9f873