Pith. sign in

Paper Citation Record · LEDGER

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization

As of 13 August 2026, this Paper Citation Record lists 61 of 61 outbound references and 0 inbound Pith citation observations for arXiv:2501.03271.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.03271 v3

Coverage vector

measured 61 of 61 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T22:19:52.822369Z

measured 61 of 61 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

61 of 61 outbound references displayed

  • verified exact0
  • verified fuzzy35
  • unresolved24
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2b8df58e-5994-41f7-8a91-f8e5c8f778b3 · outbound

This paper cites It provides the core preference alignment signal commonly used in reinforcement learning from human feedback (RLHF) (Christiano et al., 2017).

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization It provides the core preference alignment signal commonly used in reinforcement learning from human feedback (RLHF) (Christiano et al., 2017)

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:53.423899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:19:52.660007Z digest=sha256:2b6710d8b0e92ca48582bf70cc0f5943f0397acb93ceb7d1bdad87f9d82d1561

Observation fbe7a981-3107-4b70-90b2-cae9a7b39673 · outbound

This paper cites The factor γ determines how much the model should focus on aligning responses semantically.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization The factor γ determines how much the model should focus on aligning responses semantically

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:53.413752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:19:52.663766Z digest=sha256:f4bcbf49d31934b7d946a0619c7a23e2b74fd1c7314e98548ea086c36c532db4

Observation faaf65be-ec15-492b-963a-a6dab4c232d7 · outbound

This paper cites semantic margin.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization semantic margin

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:53.384130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:19:52.674734Z digest=sha256:92cc45559d010d031b1151cf4e6561673346f915f333581242c9fe4a0e1d1ee7

Observation 39713119-a846-4b51-b765-999a12a5242d · outbound

This paper cites Larger devi- ations in NAG suggest the suitability of RBF and Spectral kernels to handle the increased separation.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Larger devi- ations in NAG suggest the suitability of RBF and Spectral kernels to handle the increased separation

Reference 4

Resolution
malformed identifier
raw_fallback, observed 2026-08-10T22:19:53.320279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:19:52.696282Z digest=sha256:684621239445dd3a9a2f25ad59f14558e61070f0293cfc64adbcaf53f58dc9f2

Observation 98083ef1-ae6d-43d4-b938-84e025f9787e · outbound

This paper cites Advances in Neural Informa- tion Processing Systems.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Advances in Neural Informa- tion Processing Systems

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:53.443777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:19:52.649154Z digest=sha256:108222fbf04974c0b53bc90c996bd21923c464fa939b040c04e0da296af758cc

Observation 1a6a632e-4802-44c0-86a7-9cb8366928f6 · outbound

This paper cites SAFREE: Training-Free and Adaptive Guard for Safe Text-to-Image And Video Generation.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization SAFREE: Training-Free and Adaptive Guard for Safe Text-to-Image And Video Generation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:52.652196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:52.652196Z digest=sha256:f038aeae869de0fdfc9007d6b5108fc69ccd5d8ac1b3c2ca4f490811a2fef56e

Observation d0d48f6e-9af0-4a58-ad8b-d144d1b7fe4d · outbound

This paper cites So the answer is,.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization So the answer is,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:53.433940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:19:52.656680Z digest=sha256:d1cd9d229e26190c3a633f15c0b235a631284e146e13645c846f4ba7de342899

Observation a2251d4b-5fa0-41f6-96c6-f4bddf191a58 · outbound

This paper cites • γ >0: Embedding-based alignment is included, encouraging the model to consider semantic co- herence alongside probability alignment.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization • γ >0: Embedding-based alignment is included, encouraging the model to consider semantic co- herence alongside probability alignment

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:53.404798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:19:52.667668Z digest=sha256:2be930885443b644aae9c45853b51b3e5eea525ce008dd6eca17f4d43799df13

Observation e8dd7623-0101-4704-b1e8-3310ce513e25 · outbound

This paper cites This helps the model avoid reinforc- ing incorrect preferences when probability-based signals are uncertain.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization This helps the model avoid reinforc- ing incorrect preferences when probability-based signals are uncertain

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:53.394905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:19:52.671293Z digest=sha256:9d68a32b80644831db65af133e8b5f3fa108674d7be888aa163a687dafe77b8f

Observation 9688a092-39bd-45ca-9740-a25b6a7d8856 · outbound

This paper cites an unresolved cited work.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:19:53.373753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:19:52.677865Z digest=sha256:a78392c5b8b4cc60603c70126d9bdf2d54fcd7cdf4b5f91a2b7cdfa180713b6a

Observation 85cbe184-9c23-47ea-bbed-ee6722f31e0f · outbound

This paper cites the reward model serves as a learned proxy for human judgment, guiding the policy to generate more desirable out- puts.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization the reward model serves as a learned proxy for human judgment, guiding the policy to generate more desirable out- puts

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:53.363449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:19:52.681179Z digest=sha256:7a4f5ecedbf66ebedf9f88aff2e2200ea6cebbbc102015cae8e3b030d32d7421

Observation 5e404d99-e947-4833-8ecb-a24e47cb5d7c · outbound

This paper cites It is defined as: PND = d(x, y+) − d(x, y−) where d(x, y+) and d(x, y−) denote the distances from x to the positive and negative responses, re- spectively.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization It is defined as: PND = d(x, y+) − d(x, y−) where d(x, y+) and d(x, y−) denote the distances from x to the positive and negative responses, re- spectively

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:53.351871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:19:52.685421Z digest=sha256:a55dfcc5e8a4ecfbd6fcc0c4789812a7f40bbbf350a41419038c9b93cc8248b6

Observation 483f8a42-42ba-44e7-9546-f1f423ca072d · outbound

This paper cites Conversely, low PNA V values imply stable alignment, favoring simpler kernels such as Mahalanobis or Spectral.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Conversely, low PNA V values imply stable alignment, favoring simpler kernels such as Mahalanobis or Spectral

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:53.341110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:19:52.688857Z digest=sha256:111a7a8277c49a3d47adf59e227c4212284fc55360cae7a11bc31d2216a89848

Observation 04900dec-2822-4706-a463-90d00936f45f · outbound

This paper cites an unresolved cited work.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:19:53.330517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:19:52.692931Z digest=sha256:7edd1ae612dc5f3de74242cf61009ba81afc045be3686e7e18873c79c5ca7c25

Observation c151bc98-d2e7-4869-bef6-cab0287aef28 · outbound

This paper cites tailedness.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization tailedness

Reference 26

Resolution
malformed identifier
raw_fallback, observed 2026-08-10T22:19:53.310912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:19:52.699586Z digest=sha256:2fc5c60012815245f63e4d85d1f2725884c0101b60d9506aa7a414489b5c0725

Observation d223f0bb-a233-4dbc-b864-7f5968d2e625 · outbound

This paper cites an unresolved cited work.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:19:53.299967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:19:52.703029Z digest=sha256:1be7c6c23d6d88f966a68a32917a2f930f028e538855d48044b1160c4a8b7635

Observation f8c00141-ce8b-4703-8ae9-fef6093ab31b · outbound

This paper cites tailedness.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization tailedness

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:53.290903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:19:52.706207Z digest=sha256:b88b4822c8aaf1d1012c9fdaccc8b21a5ff90c2e7caf7351aaa26598302733bf

Observation 2eb72395-f5a7-4f3a-9dcc-e5abccee9398 · outbound

This paper cites tailedness.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization tailedness

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:53.280277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:19:52.709623Z digest=sha256:71074cbbda376eae5d4b03c08ab80d60c00ee75c2d03710b8975837f7072e4fa

Observation f21edbfe-7652-4769-b057-5f057ada81b5 · outbound

This paper cites The RBF kernel exhibits isotropic influence (circular), while the Polynomial kernel allows nonlinear, bounded in- fluence.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization The RBF kernel exhibits isotropic influence (circular), while the Polynomial kernel allows nonlinear, bounded in- fluence

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:53.268810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:19:52.712628Z digest=sha256:eb14a789b1193d9d41776e0f28f8be6ee3cf297fcfd959d6c9ef6b14ab470db5

Observation 5aa28a01-606a-437e-a347-f81020c21f8c · outbound

This paper cites local" kernels. In contrast, the Mahalanobis and Spectral kernels show a slower decay, reflecting their role as.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization local" kernels. In contrast, the Mahalanobis and Spectral kernels show a slower decay, reflecting their role as

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:53.257548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:19:52.715537Z digest=sha256:fd3abad7e00c022fb30e4b7973835c558e28220f96d656fb3cfe712b11239651

Observation c6e680dc-33de-4d25-bd4c-fe538c67ad57 · outbound

This paper cites • Computing the logarithm of the ratio between the positive and negative class probabilities.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization • Computing the logarithm of the ratio between the positive and negative class probabilities

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:53.237843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:19:52.722134Z digest=sha256:56bdbbd18b889c520c497ccf5101f94de3bba7aa5993bf7dc54f8172ee533ba5

Observation 5b876d01-a45e-43d9-bdc3-6d3c931ec274 · outbound

This paper cites (e⊤y−ex+c)∇θ(e⊤y+ex)−(e⊤y+ex+c)∇θ(e⊤y−ex) (e⊤y−ex+c)2 # =γd e⊤y+ex+c e⊤y−ex+c !d−1 ·.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization (e⊤y−ex+c)∇θ(e⊤y+ex)−(e⊤y+ex+c)∇θ(e⊤y−ex) (e⊤y−ex+c)2 # =γd e⊤y+ex+c e⊤y−ex+c !d−1 ·

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:53.228407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:19:52.725251Z digest=sha256:655336140d458d64e33db73a77591d2015b5a535fa523cfcb689ef89fa1451d6

Observation 654f3029-6129-4f9a-b50d-13332c6bdf9a · outbound

This paper cites • Softmax Calculation: Compute the exponential efθ(x,y) for each class and normalize by the sum over all classes.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization • Softmax Calculation: Compute the exponential efθ(x,y) for each class and normalize by the sum over all classes

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:53.216828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:19:52.728455Z digest=sha256:4f1c9320acc731d0e98d6a38532f5dd21586eeecc956df8d35f40683f87d9c02

Observation a8b48bb8-0205-4ac5-8c9c-887316c40714 · outbound

This paper cites − 1 σ2 logπ(y+| x) π(y−| x) ·exp  − logπ(y+|x) π(y−|x) 2 2σ2   ·∇θlogπ(y+| x)− ∇θlogπ(y−| x) − γ σ2 · e⊤xey+ e⊤xey− ·exp  − e⊤xey+ e⊤xey− 2 2σ2   ·.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization − 1 σ2 logπ(y+| x) π(y−| x) ·exp  − logπ(y+|x) π(y−|x) 2 2σ2   ·∇θlogπ(y+| x)− ∇θlogπ(y−| x) − γ σ2 · e⊤xey+ e⊤xey− ·exp  − e⊤xey+ e⊤xey− 2 2σ2   ·

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:53.207555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:19:52.732117Z digest=sha256:384559a2f08f60db476c33c7c6f6fa50b994a810859d7f8679335ddf0e1a6422

Observation 769a11e7-3947-472e-b436-a357e4b65815 · outbound

This paper cites where πθ(y | x) is modeled using a softmax func- tion: πθ(y | x) = efθ(x,y) P y′ efθ(x,y′).

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization where πθ(y | x) is modeled using a softmax func- tion: πθ(y | x) = efθ(x,y) P y′ efθ(x,y′)

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:53.197893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:19:52.735984Z digest=sha256:df1664168b7ddfa98891fd7b892575ba9fc619ddb51c18ef43fd89efec58c621

Observation 6f955f84-b230-4149-ae58-a95c08e763b5 · outbound

This paper cites • Ratio Calculation: Compute the ratio e⊤ x ey+ e⊤x ey−.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization • Ratio Calculation: Compute the ratio e⊤ x ey+ e⊤x ey−

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:53.187809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:19:52.739234Z digest=sha256:b37d33527d48c8dbb8c8a7895d90cfe8a4f8000c241a8e4222bd2e6959b4180a

Observation ccfdada3-6bd9-45ed-a461-af5c88418575 · outbound

This paper cites an unresolved cited work.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:19:53.178490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:19:52.742997Z digest=sha256:c2477fb347fe2e8193ca6accc1e6461ec00c4a201902bb6c4a12dd9b1ea2d660

Observation 58cdebd1-44ab-4239-8527-29214910d430 · outbound

This paper cites an unresolved cited work.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:19:53.168497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:19:52.746872Z digest=sha256:c90dbb200db48ae3ae38ad020f6e7cd4841c4f17e76d41a55f8965f9226744ff

Observation e4db9c6b-6be1-451f-a2f2-f5a1c6209108 · outbound

This paper cites an unresolved cited work.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:19:53.159314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:19:52.750944Z digest=sha256:0986fbdd6dbf4c54266088be7fc1669add50a7148dce9958988b9e4bbdad7107

Observation 5c4b5597-7c17-4547-af11-348b02da768e · outbound

This paper cites an unresolved cited work.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:19:53.147808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:19:52.755229Z digest=sha256:734f3ef7c8de8d9e41fa4fbe1f3acaac238188c7a337668f11d3e9189d10c285

Observation 31a46170-d363-4e13-b13b-0d3e0638337e · outbound

This paper cites an unresolved cited work.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:19:53.138212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:19:52.758966Z digest=sha256:a6f24a0ec7296bff23a4d2fd4ad139217c79c6ea9106f68762a0714e60f89e1b

Observation 128b0f3d-9da8-4dfd-8f30-9863717d3453 · outbound

This paper cites an unresolved cited work.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:19:53.129088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:19:52.762517Z digest=sha256:0f89725fedc59a175806e436ebdcf6cb43fcbc39a4caa11e2f9665a891da1119

Observation 5eea5219-43ca-4e1d-a53f-2ded0c5a691c · outbound

This paper cites Steps Involved: • Score Computation: Calculate fθ(x, y) for each class y, which involves a dot product between input features and model parameters.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Steps Involved: • Score Computation: Calculate fθ(x, y) for each class y, which involves a dot product between input features and model parameters

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:53.119709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:19:52.765952Z digest=sha256:1b3edb9c2ac20943f7b8dcae97a343f1d341e425eb5e41f97a917e5d09589b4a

Observation eb5ad9ae-ef82-4688-8a1f-2496f3fde62a · outbound

This paper cites Steps Involved: • Dot Product Computation : Calculate the dot products e⊤ x ey+ and e⊤ x ey−, where ex, ey+, ey− ∈ Rd.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Steps Involved: • Dot Product Computation : Calculate the dot products e⊤ x ey+ and e⊤ x ey−, where ex, ey+, ey− ∈ Rd

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:53.110711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:19:52.769188Z digest=sha256:19a3e43a53df33fb5e7e2391a48d1507796a164cb1bf49ba00efc5e6a4209bd5

Observation 860061fe-95bb-4ff3-bbca-a4e2824bad87 · outbound

This paper cites Steps Involved: • Score Computation: Calculate fθ(x, y) for each class y, which involves a dot product between input features and model parameters.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Steps Involved: • Score Computation: Calculate fθ(x, y) for each class y, which involves a dot product between input features and model parameters

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:53.100371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:19:52.772893Z digest=sha256:0731e8d39d54ccfab5caf4b3b9f6cc94ab9f9a4d6fdf75454ebce8d8051a845c

Observation 7ce2653e-f1c6-4cee-95a5-3148b580bac6 · outbound

This paper cites Steps Involved: • Dot Product Computation : Calculate the dot products e⊤ x ey+ and e⊤ x ey−, where ex, ey+, ey− ∈ Rd.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Steps Involved: • Dot Product Computation : Calculate the dot products e⊤ x ey+ and e⊤ x ey−, where ex, ey+, ey− ∈ Rd

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:53.090521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:19:52.777240Z digest=sha256:217cae596c9667e158bb1b9736346575dd1ebef60069e625b12472682c3662e6

Observation 46f1883d-dea7-4ef1-9e4f-1611435d13ac · outbound

This paper cites an unresolved cited work.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Unresolved cited work

Reference 49

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:19:53.080425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:19:52.780784Z digest=sha256:e18cfe51f1d0615cdd3a1beb726de413e25c2120b4ee6c9de9d6475bd3ee27c9

Observation 4202698c-b24e-4ee2-8a50-34fcfa2c643e · outbound

This paper cites an unresolved cited work.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:19:53.070692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:19:52.784682Z digest=sha256:77c633a5f768ef59edc49e983cddc92fc1751f61bd8b0b7e5c0ca2e67de8b0e0

Observation 90bbd48b-8d29-464b-b884-a3e9dc6ffc07 · outbound

This paper cites an unresolved cited work.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:19:53.060044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:19:52.788403Z digest=sha256:caf63702872357febff2ba771955ff876061190af2f97a6271761ccf28ccd61b

Observation c80db629-ddcc-42c7-b1d9-02ef01cf52f5 · outbound

This paper cites an unresolved cited work.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:19:53.048614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:19:52.791496Z digest=sha256:00ee5120d55d68802cf1e076ded3350ac35be25e6e9d743cf1db3862652ac044

Observation 84a73a3e-fa9d-4902-b819-a7bed2d439ba · outbound

This paper cites RBF Kernel KRBF(x, x′) = exp − ∥x − x′∥2 2σ2 Steps Involved: • Compute the Euclidean distance ∥x − x′∥, which involves O(d) operations, where d is the dimen- sion of the input.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization RBF Kernel KRBF(x, x′) = exp − ∥x − x′∥2 2σ2 Steps Involved: • Compute the Euclidean distance ∥x − x′∥, which involves O(d) operations, where d is the dimen- sion of the input

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:53.039153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:19:52.794313Z digest=sha256:99fc47170d7ff14498a9b1d1e3bf6e812416435f8b7eb12c0cbd3853a4469864

Observation d58bd280-5e59-4bc2-bb9c-7b545da80ca6 · outbound

This paper cites Spectral Kernel KSpectral(x, x′) = pX i=1 exp −λiz2 i ϕi(zi), where zi = log π(y+|x) π(y−|x).

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Spectral Kernel KSpectral(x, x′) = pX i=1 exp −λiz2 i ϕi(zi), where zi = log π(y+|x) π(y−|x)

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:53.029184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:19:52.797742Z digest=sha256:6993440f13323762e3ccc2c5a6fc5a93b6cbe19469b34100b4f0d28b62a2435b

Observation 199e6983-76aa-4b0f-9b1c-0edcdc30b325 · outbound

This paper cites • Lipschitz Continuity: The gradient of the RBF kernel is Lipschitz continuous due to its exponen- tial decay property.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization • Lipschitz Continuity: The gradient of the RBF kernel is Lipschitz continuous due to its exponen- tial decay property

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:53.019218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:19:52.801100Z digest=sha256:7221ec173963790e61ead0895adfc4052c6fc183a1df43b3e6203d98e9f00af2

Observation 96c4d69c-0fb8-4d61-945e-57d561595d32 · outbound

This paper cites Higher degrees in- troduce non-convexity, resulting in a more rugged loss landscape with multiple local minima and saddle points.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Higher degrees in- troduce non-convexity, resulting in a more rugged loss landscape with multiple local minima and saddle points

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:53.003717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:19:52.804028Z digest=sha256:ec2927d89d75335bf6871c6a8b819f2675adeda2921fa6f13bb173718e4c0986

Observation 2b6085b4-e054-4701-ba60-6a1f7ead3b5b · outbound

This paper cites Orthonormal basis functions, such as wavelets, can introduce oscillatory behavior in the loss land- scape (Ng et al., 2001).

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Orthonormal basis functions, such as wavelets, can introduce oscillatory behavior in the loss land- scape (Ng et al., 2001)

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:52.991567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:19:52.807196Z digest=sha256:8478ec829fa2bdb97b3f6c82e7e33559e729f1dab539a8c34e69e3b3a5704f5d

Observation cc4920d5-3a07-414c-bd63-6218ca5b34fa · outbound

This paper cites distance.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization distance

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:52.980602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:19:52.810821Z digest=sha256:b62a58949a11ca845b6763b1e66169e52aace3abf71a1e77b8ef941b482b5ee6

Observation b85bd9b0-1645-44b2-8d4d-ab4e3fa171b8 · outbound

This paper cites HT-SR theory posits that ρ(λ) often follows a truncated power law: ρ(λ) ∝ λ−α, for λmin ≤ λ ≤ λmax.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization HT-SR theory posits that ρ(λ) often follows a truncated power law: ρ(λ) ∝ λ−α, for λmin ≤ λ ≤ λmax

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:52.970574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:19:52.814975Z digest=sha256:836bacc3c83a2be969f8f760f64914ddedf77936f71e2c4d91055519e1d74bc1

Observation 06cd8f15-93eb-4c3d-9565-45e359970603 · outbound

This paper cites an unresolved cited work.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Unresolved cited work

Reference 60

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:19:52.958609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:19:52.818646Z digest=sha256:c41dcda314de90d46f811a975f28eb74c44bc84cb92b21610052ce205f276bd6

Observation dff8693b-f982-4353-bb16-a2fb432f6df7 · outbound

This paper cites Correlation Flow,.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Correlation Flow,

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:52.948103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:19:52.822369Z digest=sha256:eb0aff6ce83d2084f4312b42fd7a974bd592ef3cf6aeb88f940b5e4c97ba4040

Observation 627141cc-c26f-4354-8f0b-063b89ed57f2 · outbound

This paper cites Towards A Rigorous Science of Interpretable Machine Learning.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Towards A Rigorous Science of Interpretable Machine Learning

Reference 465

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:52.616572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:52.616572Z digest=sha256:32bb74fb265192aa56f0f009bd92d4ea21c2365cbed82c2a3d8024ced765a97e

Observation 4d2e9a84-5d0d-4993-a114-52a94c6f4b7c · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Representation Learning with Contrastive Predictive Coding

Reference 2001

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:52.636195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:52.636195Z digest=sha256:3c4972d48313edcd04adfcf3eaf398c5ea8a57c5ab50b4cf34399a653e5737a3

Observation e8ab9796-9968-4d86-b365-93bab8e9ef69 · outbound

This paper cites stop execution if X is true.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization stop execution if X is true

Reference 2004

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:53.246861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:19:52.718765Z digest=sha256:6f57f70f2fa015599b16bb7f47a40a9e332e99062c3982cae1eb148d8cf28dab

Observation 735ef1a2-1143-4332-8459-25218799a086 · outbound

This paper cites In IEEE Conference on Com- puter Vision and Pattern Recognition (CVPR), pages 1735–1742.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization In IEEE Conference on Com- puter Vision and Pattern Recognition (CVPR), pages 1735–1742

Reference 2006

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:53.463425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:19:52.620948Z digest=sha256:4ce16bccd4cc4faa7665b6f38a543cb5ff277156cc57f278c97727a1e3b5aaa7

Observation 161117fc-6c47-4808-bf29-27efd01ab84c · outbound

This paper cites Unpacking DPO and PPO: Disentangling Best Practices for Learning from Preference Feedback.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Unpacking DPO and PPO: Disentangling Best Practices for Learning from Preference Feedback

Reference 2007

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:52.624338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:52.624338Z digest=sha256:42c43c86452dbaa296bd90221b53fd3c40289b9ffacc85b23e8df78598e8a762

Observation 8efc0ae7-d759-4fe7-b41a-a83990e3f85f · outbound

This paper cites Diffusion Model Alignment Using Direct Preference Optimization.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Diffusion Model Alignment Using Direct Preference Optimization

Reference 2008

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:52.646022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:52.646022Z digest=sha256:ff638a358c83545a2ff40f06c03c4b4d7a3c6128b5a53566bb15beac3586bdfe

Observation c7c61bca-92d7-45f7-a2de-b5a16724f374 · outbound

This paper cites Direct Preference Optimization: Your Language Model is Secretly a Reward Model.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Direct Preference Optimization: Your Language Model is Secretly a Reward Model

Reference 2009

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:52.643057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:52.643057Z digest=sha256:e1226136fdf554e2e6f1ffbb2ac70e1d27e90072a8fe9c72f271074d2504240a

Observation 36d39408-377b-4298-ab49-eb471f7f3c6e · outbound

This paper cites In International Conference on Learning Representations (ICLR).

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization In International Conference on Learning Representations (ICLR)

Reference 2018

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:53.453574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:19:52.632670Z digest=sha256:019afd87420cabbc5ba63ffc7d2dd406723c9932ff148d6ea5c7ef11ee004581

Observation c685b90a-17d4-48ae-a7c7-1b2dda82f10a · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Training Verifiers to Solve Math Word Problems

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:52.610934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:52.610934Z digest=sha256:59ef8d62f208c014df59dd1a297cd1623bbde1925316f4a23f808d99fe3122a2

Observation f71bc17a-6336-4ed3-bc88-ea28b62fdd03 · outbound

This paper cites Training language models to follow instructions with human feedback.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Training language models to follow instructions with human feedback

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:52.639456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:52.639456Z digest=sha256:33b29a0c145116a4234e2e46871a3240d18069922f3a1a432ab5d38f10a41743

Observation 3e278364-c54a-4a64-b026-deea65b7dd89 · outbound

This paper cites Let's Verify Step by Step.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Let's Verify Step by Step

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:52.628333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:52.628333Z digest=sha256:6e2808ef12dff906528f17784596ae2bed1e953142950c89ff2566def0c5a82f

Observation 32c14187-a3d3-4346-9292-90415eed753b · outbound

This paper cites Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:52.606670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:52.606670Z digest=sha256:fb5a670930b71511118b8aa92ca8484765d4b3ac078ed456009bddeb31cfd079

Pith citing papers

No inbound Pith citation observations are available.