Pith. sign in

Paper Citation Record · LEDGER

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization

As of 23 August 2026, this Paper Citation Record lists 61 of 61 outbound references and 0 inbound Pith citation observations for arXiv:2501.03271.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.03271 v3

Coverage vector

measured 61 of 61 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T22:19:52.822369Z

measured 61 of 61 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

61 of 61 outbound references displayed

  • verified exact0
  • verified fuzzy35
  • unresolved24
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2b8df58e-5994-41f7-8a91-f8e5c8f778b3 · outbound

This paper cites It provides the core preference alignment signal commonly used in reinforcement learning from human feedback (RLHF) (Christiano et al., 2017).

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization It provides the core preference alignment signal commonly used in reinforcement learning from human feedback (RLHF) (Christiano et al., 2017)

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:53.423899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:19:52.660007Z digest=sha256:4f3fcbcc22d73ae050e50b7e82b7274196066ed073159f4b93029df263280c22

Observation fbe7a981-3107-4b70-90b2-cae9a7b39673 · outbound

This paper cites The factor γ determines how much the model should focus on aligning responses semantically.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization The factor γ determines how much the model should focus on aligning responses semantically

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:53.413752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:19:52.663766Z digest=sha256:3ada93465532cb6ebb47b1f426c6e7c35ab74b5de2f3202ea049d76b3e827a9e

Observation faaf65be-ec15-492b-963a-a6dab4c232d7 · outbound

This paper cites semantic margin.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization semantic margin

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:53.384130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:19:52.674734Z digest=sha256:777bcbee7b28b1d5bfaa8b3857a57b2870730c54542d5afb049312c599ee9f63

Observation 39713119-a846-4b51-b765-999a12a5242d · outbound

This paper cites Larger devi- ations in NAG suggest the suitability of RBF and Spectral kernels to handle the increased separation.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Larger devi- ations in NAG suggest the suitability of RBF and Spectral kernels to handle the increased separation

Reference 4

Resolution
malformed identifier
raw_fallback, observed 2026-08-10T22:19:53.320279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:19:52.696282Z digest=sha256:6599f7a302e44abcc7f1be0877a333d6006ed5a033f7c7d21d97ee98a4469a94

Observation 98083ef1-ae6d-43d4-b938-84e025f9787e · outbound

This paper cites Advances in Neural Informa- tion Processing Systems.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Advances in Neural Informa- tion Processing Systems

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:53.443777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:19:52.649154Z digest=sha256:1f84143fffbdaf82d68f3593cbad2836689b058ad30f61a4df5cfa6124fad9af

Observation 1a6a632e-4802-44c0-86a7-9cb8366928f6 · outbound

This paper cites SAFREE: Training-Free and Adaptive Guard for Safe Text-to-Image And Video Generation.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization SAFREE: Training-Free and Adaptive Guard for Safe Text-to-Image And Video Generation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:52.652196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:52.652196Z digest=sha256:12938ef16b7e1b674aacc7522aa7118a1ff78eaec39b0ea8b58b1e523001de83

Observation d0d48f6e-9af0-4a58-ad8b-d144d1b7fe4d · outbound

This paper cites So the answer is,.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization So the answer is,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:53.433940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:19:52.656680Z digest=sha256:27538d40399bd7ef7013dae2ecae309dea227c0c55d251449da804b6cca5643d

Observation a2251d4b-5fa0-41f6-96c6-f4bddf191a58 · outbound

This paper cites • γ >0: Embedding-based alignment is included, encouraging the model to consider semantic co- herence alongside probability alignment.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization • γ >0: Embedding-based alignment is included, encouraging the model to consider semantic co- herence alongside probability alignment

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:53.404798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:19:52.667668Z digest=sha256:431f0d26f24df8fc2a8c5eafa47c2ee6fa0a8b6fe7f4dcc791d8a0dd98783af4

Observation e8dd7623-0101-4704-b1e8-3310ce513e25 · outbound

This paper cites This helps the model avoid reinforc- ing incorrect preferences when probability-based signals are uncertain.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization This helps the model avoid reinforc- ing incorrect preferences when probability-based signals are uncertain

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:53.394905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:19:52.671293Z digest=sha256:9e18042733a904df1084db9e74ae0f7a248aa1f4dd0e09ba8bfb21ce1ca4ddab

Observation 9688a092-39bd-45ca-9740-a25b6a7d8856 · outbound

This paper cites an unresolved cited work.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:19:53.373753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:19:52.677865Z digest=sha256:d8a4540fe5eb30106eec464c0d163df7785d1301ceae07fd17a0d532783224d6

Observation 85cbe184-9c23-47ea-bbed-ee6722f31e0f · outbound

This paper cites the reward model serves as a learned proxy for human judgment, guiding the policy to generate more desirable out- puts.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization the reward model serves as a learned proxy for human judgment, guiding the policy to generate more desirable out- puts

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:53.363449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:19:52.681179Z digest=sha256:c3f43398d78f484cfd687aa4419de1b51ccf0ed0e8f9f1380633bd763d642d76

Observation 5e404d99-e947-4833-8ecb-a24e47cb5d7c · outbound

This paper cites It is defined as: PND = d(x, y+) − d(x, y−) where d(x, y+) and d(x, y−) denote the distances from x to the positive and negative responses, re- spectively.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization It is defined as: PND = d(x, y+) − d(x, y−) where d(x, y+) and d(x, y−) denote the distances from x to the positive and negative responses, re- spectively

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:53.351871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:19:52.685421Z digest=sha256:4f2171ca828a665711e6c80005eb0f3ecce65dfa9170b740505f52469c81fb03

Observation 483f8a42-42ba-44e7-9546-f1f423ca072d · outbound

This paper cites Conversely, low PNA V values imply stable alignment, favoring simpler kernels such as Mahalanobis or Spectral.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Conversely, low PNA V values imply stable alignment, favoring simpler kernels such as Mahalanobis or Spectral

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:53.341110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:19:52.688857Z digest=sha256:e345dcd04bb15169549231dc1554132bebce6a1410a9499e91bf05a898ef8fc7

Observation 04900dec-2822-4706-a463-90d00936f45f · outbound

This paper cites an unresolved cited work.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:19:53.330517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:19:52.692931Z digest=sha256:4ebb4d98775c5ec40c0c842c8d405307a7a09a14d2425970eff18803b5651547

Observation c151bc98-d2e7-4869-bef6-cab0287aef28 · outbound

This paper cites tailedness.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization tailedness

Reference 26

Resolution
malformed identifier
raw_fallback, observed 2026-08-10T22:19:53.310912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:19:52.699586Z digest=sha256:2f246fd829b1f8b39942168669e3f814b5104ad1fe8b86ef00ecfca37102c39d

Observation d223f0bb-a233-4dbc-b864-7f5968d2e625 · outbound

This paper cites an unresolved cited work.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:19:53.299967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:19:52.703029Z digest=sha256:c51e2fd616163f681200dc9568eabf379f476e2241208f0c63eef0030e270453

Observation f8c00141-ce8b-4703-8ae9-fef6093ab31b · outbound

This paper cites tailedness.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization tailedness

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:53.290903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:19:52.706207Z digest=sha256:c837fe826d5bb3b3e9821151dcb55147f92be54f278f47210781d3eca4c84c47

Observation 2eb72395-f5a7-4f3a-9dcc-e5abccee9398 · outbound

This paper cites tailedness.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization tailedness

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:53.280277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:19:52.709623Z digest=sha256:593eaf457329db0102f4f91287009650c80a628a6bba96db6eb2a38c316ab2d2

Observation f21edbfe-7652-4769-b057-5f057ada81b5 · outbound

This paper cites The RBF kernel exhibits isotropic influence (circular), while the Polynomial kernel allows nonlinear, bounded in- fluence.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization The RBF kernel exhibits isotropic influence (circular), while the Polynomial kernel allows nonlinear, bounded in- fluence

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:53.268810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:19:52.712628Z digest=sha256:640d5066707cd88f22b34a5118baed20f8c13d551bcd071cf2ed545ddca1990b

Observation 5aa28a01-606a-437e-a347-f81020c21f8c · outbound

This paper cites local" kernels. In contrast, the Mahalanobis and Spectral kernels show a slower decay, reflecting their role as.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization local" kernels. In contrast, the Mahalanobis and Spectral kernels show a slower decay, reflecting their role as

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:53.257548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:19:52.715537Z digest=sha256:5aa53db97086870d5fc5e44ae7370a11a13d6dac7af1170deecc4e35c8d95b2d

Observation c6e680dc-33de-4d25-bd4c-fe538c67ad57 · outbound

This paper cites • Computing the logarithm of the ratio between the positive and negative class probabilities.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization • Computing the logarithm of the ratio between the positive and negative class probabilities

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:53.237843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:19:52.722134Z digest=sha256:3d2d18ed3ca453c625704f1c9c2c641c00b2944f6698c51de00c1c52dc0d7415

Observation 5b876d01-a45e-43d9-bdc3-6d3c931ec274 · outbound

This paper cites (e⊤y−ex+c)∇θ(e⊤y+ex)−(e⊤y+ex+c)∇θ(e⊤y−ex) (e⊤y−ex+c)2 # =γd e⊤y+ex+c e⊤y−ex+c !d−1 ·.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization (e⊤y−ex+c)∇θ(e⊤y+ex)−(e⊤y+ex+c)∇θ(e⊤y−ex) (e⊤y−ex+c)2 # =γd e⊤y+ex+c e⊤y−ex+c !d−1 ·

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:53.228407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:19:52.725251Z digest=sha256:712fe57fabe3ed890452202e3a15a01bf89e49dcd7a4aae5b85578280ec3d7db

Observation 654f3029-6129-4f9a-b50d-13332c6bdf9a · outbound

This paper cites • Softmax Calculation: Compute the exponential efθ(x,y) for each class and normalize by the sum over all classes.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization • Softmax Calculation: Compute the exponential efθ(x,y) for each class and normalize by the sum over all classes

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:53.216828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:19:52.728455Z digest=sha256:995eb2bacc38e744deccb21bee4d6ca58c4efb884d81e26e583d521873a3a05d

Observation a8b48bb8-0205-4ac5-8c9c-887316c40714 · outbound

This paper cites − 1 σ2 logπ(y+| x) π(y−| x) ·exp  − logπ(y+|x) π(y−|x) 2 2σ2   ·∇θlogπ(y+| x)− ∇θlogπ(y−| x) − γ σ2 · e⊤xey+ e⊤xey− ·exp  − e⊤xey+ e⊤xey− 2 2σ2   ·.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization − 1 σ2 logπ(y+| x) π(y−| x) ·exp  − logπ(y+|x) π(y−|x) 2 2σ2   ·∇θlogπ(y+| x)− ∇θlogπ(y−| x) − γ σ2 · e⊤xey+ e⊤xey− ·exp  − e⊤xey+ e⊤xey− 2 2σ2   ·

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:53.207555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:19:52.732117Z digest=sha256:031840a36ebc938b558e106bb8d36d38cb2eb94505d883e9ee8bafb4477f9a42

Observation 769a11e7-3947-472e-b436-a357e4b65815 · outbound

This paper cites where πθ(y | x) is modeled using a softmax func- tion: πθ(y | x) = efθ(x,y) P y′ efθ(x,y′).

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization where πθ(y | x) is modeled using a softmax func- tion: πθ(y | x) = efθ(x,y) P y′ efθ(x,y′)

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:53.197893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:19:52.735984Z digest=sha256:2a127f82174b918171b21c95e1e90cec738303eb4203ccfbb1053c9544cc990a

Observation 6f955f84-b230-4149-ae58-a95c08e763b5 · outbound

This paper cites • Ratio Calculation: Compute the ratio e⊤ x ey+ e⊤x ey−.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization • Ratio Calculation: Compute the ratio e⊤ x ey+ e⊤x ey−

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:53.187809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:19:52.739234Z digest=sha256:da6d3bf57d0928a6a4e497c7369137d425f9128b55a392b5230ac1e9c699971d

Observation ccfdada3-6bd9-45ed-a461-af5c88418575 · outbound

This paper cites an unresolved cited work.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:19:53.178490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:19:52.742997Z digest=sha256:d86349ab6e67d11b0ce2d95f49c4970dabfe8aef37598b5715629e10d1f3e18f

Observation 58cdebd1-44ab-4239-8527-29214910d430 · outbound

This paper cites an unresolved cited work.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:19:53.168497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:19:52.746872Z digest=sha256:e70a510d28623d8fad498bd9c4d93979dc182c47ec3c3ba8802dcbb2eaf8c65a

Observation e4db9c6b-6be1-451f-a2f2-f5a1c6209108 · outbound

This paper cites an unresolved cited work.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:19:53.159314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:19:52.750944Z digest=sha256:ff717c722e267235868661f58ccfc7458894df3111cad0abe0790cdf1222d916

Observation 5c4b5597-7c17-4547-af11-348b02da768e · outbound

This paper cites an unresolved cited work.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:19:53.147808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:19:52.755229Z digest=sha256:f69d534874a3d4c2f170147d8998dbf456286bdaa80619fd2512eb886643aec2

Observation 31a46170-d363-4e13-b13b-0d3e0638337e · outbound

This paper cites an unresolved cited work.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:19:53.138212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:19:52.758966Z digest=sha256:0750ef055a8518740561479c18cc5a612745c0d920c0e33f9108b1653d78194e

Observation 128b0f3d-9da8-4dfd-8f30-9863717d3453 · outbound

This paper cites an unresolved cited work.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:19:53.129088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:19:52.762517Z digest=sha256:c9863b0f825d2ccf393518f4f96c29eafb41ba4d341f1a3277160ac8b930a14b

Observation 5eea5219-43ca-4e1d-a53f-2ded0c5a691c · outbound

This paper cites Steps Involved: • Score Computation: Calculate fθ(x, y) for each class y, which involves a dot product between input features and model parameters.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Steps Involved: • Score Computation: Calculate fθ(x, y) for each class y, which involves a dot product between input features and model parameters

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:53.119709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:19:52.765952Z digest=sha256:f3603d3127641e7ca34620054f71344f449ce4969e6015570ac1db8195d1281c

Observation eb5ad9ae-ef82-4688-8a1f-2496f3fde62a · outbound

This paper cites Steps Involved: • Dot Product Computation : Calculate the dot products e⊤ x ey+ and e⊤ x ey−, where ex, ey+, ey− ∈ Rd.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Steps Involved: • Dot Product Computation : Calculate the dot products e⊤ x ey+ and e⊤ x ey−, where ex, ey+, ey− ∈ Rd

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:53.110711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:19:52.769188Z digest=sha256:6294a1f7ab5d359a109d889ba0611acb163ee0c3d6de43f8de7f2ca23749fa5d

Observation 860061fe-95bb-4ff3-bbca-a4e2824bad87 · outbound

This paper cites Steps Involved: • Score Computation: Calculate fθ(x, y) for each class y, which involves a dot product between input features and model parameters.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Steps Involved: • Score Computation: Calculate fθ(x, y) for each class y, which involves a dot product between input features and model parameters

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:53.100371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:19:52.772893Z digest=sha256:aa8c3e9ceaeba244e79d998228398aa0d6db8563c6b940e7087508a10c4bdc68

Observation 7ce2653e-f1c6-4cee-95a5-3148b580bac6 · outbound

This paper cites Steps Involved: • Dot Product Computation : Calculate the dot products e⊤ x ey+ and e⊤ x ey−, where ex, ey+, ey− ∈ Rd.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Steps Involved: • Dot Product Computation : Calculate the dot products e⊤ x ey+ and e⊤ x ey−, where ex, ey+, ey− ∈ Rd

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:53.090521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:19:52.777240Z digest=sha256:751c7e393321e262835341a88f907f5902b394c3af343a69c96359f4bee014d4

Observation 46f1883d-dea7-4ef1-9e4f-1611435d13ac · outbound

This paper cites an unresolved cited work.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Unresolved cited work

Reference 49

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:19:53.080425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:19:52.780784Z digest=sha256:f39a42b059bd5cb0bc71f016e2835162ff0f74adc14f77e80457fe7f6ac42a18

Observation 4202698c-b24e-4ee2-8a50-34fcfa2c643e · outbound

This paper cites an unresolved cited work.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:19:53.070692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:19:52.784682Z digest=sha256:197d1eefe3ce84c96555e6e56311a51668edbe7bd072a961da18993e15960a63

Observation 90bbd48b-8d29-464b-b884-a3e9dc6ffc07 · outbound

This paper cites an unresolved cited work.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:19:53.060044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:19:52.788403Z digest=sha256:706bb777d0fcfadf22d950efc6e0259ea0cd047cfebbb703cda8a8e38994d296

Observation c80db629-ddcc-42c7-b1d9-02ef01cf52f5 · outbound

This paper cites an unresolved cited work.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:19:53.048614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:19:52.791496Z digest=sha256:9d03fd64d8699265de063c9cb332ea7844d759c24f2c0f26fa6b656f6480ce44

Observation 84a73a3e-fa9d-4902-b819-a7bed2d439ba · outbound

This paper cites RBF Kernel KRBF(x, x′) = exp − ∥x − x′∥2 2σ2 Steps Involved: • Compute the Euclidean distance ∥x − x′∥, which involves O(d) operations, where d is the dimen- sion of the input.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization RBF Kernel KRBF(x, x′) = exp − ∥x − x′∥2 2σ2 Steps Involved: • Compute the Euclidean distance ∥x − x′∥, which involves O(d) operations, where d is the dimen- sion of the input

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:53.039153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:19:52.794313Z digest=sha256:5b2e9fb9bdaaacf54b8f5ef070855bb04662d19a0c294d6fa1b117f8d03c25cc

Observation d58bd280-5e59-4bc2-bb9c-7b545da80ca6 · outbound

This paper cites Spectral Kernel KSpectral(x, x′) = pX i=1 exp −λiz2 i ϕi(zi), where zi = log π(y+|x) π(y−|x).

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Spectral Kernel KSpectral(x, x′) = pX i=1 exp −λiz2 i ϕi(zi), where zi = log π(y+|x) π(y−|x)

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:53.029184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:19:52.797742Z digest=sha256:529b84ab9a4765e878168b2db0fb3c4d90d4c7ed60c0641d27ae3888157e5b29

Observation 199e6983-76aa-4b0f-9b1c-0edcdc30b325 · outbound

This paper cites • Lipschitz Continuity: The gradient of the RBF kernel is Lipschitz continuous due to its exponen- tial decay property.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization • Lipschitz Continuity: The gradient of the RBF kernel is Lipschitz continuous due to its exponen- tial decay property

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:53.019218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:19:52.801100Z digest=sha256:31a9b8ff9d87fe19bf640cbc578dd7b1550b021a0f2a699f6e78e4ee9967b38f

Observation 96c4d69c-0fb8-4d61-945e-57d561595d32 · outbound

This paper cites Higher degrees in- troduce non-convexity, resulting in a more rugged loss landscape with multiple local minima and saddle points.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Higher degrees in- troduce non-convexity, resulting in a more rugged loss landscape with multiple local minima and saddle points

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:53.003717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:19:52.804028Z digest=sha256:6e81ff586b7b2b06e63fb7200517c8cbd15432d7294849a05a813b1d0ea8c84f

Observation 2b6085b4-e054-4701-ba60-6a1f7ead3b5b · outbound

This paper cites Orthonormal basis functions, such as wavelets, can introduce oscillatory behavior in the loss land- scape (Ng et al., 2001).

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Orthonormal basis functions, such as wavelets, can introduce oscillatory behavior in the loss land- scape (Ng et al., 2001)

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:52.991567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:19:52.807196Z digest=sha256:3abda9077888ae07747ed5a0091ddff7a0ceebdf79cb32fe358b7255713d032b

Observation cc4920d5-3a07-414c-bd63-6218ca5b34fa · outbound

This paper cites distance.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization distance

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:52.980602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:19:52.810821Z digest=sha256:99238810f70d4800c280150509d10e4638475297e72a23a0ec2de04459720d42

Observation b85bd9b0-1645-44b2-8d4d-ab4e3fa171b8 · outbound

This paper cites HT-SR theory posits that ρ(λ) often follows a truncated power law: ρ(λ) ∝ λ−α, for λmin ≤ λ ≤ λmax.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization HT-SR theory posits that ρ(λ) often follows a truncated power law: ρ(λ) ∝ λ−α, for λmin ≤ λ ≤ λmax

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:52.970574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:19:52.814975Z digest=sha256:ab661f0fe4aa521d9546d03430dff2c521855715cd4b8e3ce87e405aa1d9030f

Observation 06cd8f15-93eb-4c3d-9565-45e359970603 · outbound

This paper cites an unresolved cited work.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Unresolved cited work

Reference 60

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:19:52.958609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:19:52.818646Z digest=sha256:bb289a8ce620797ee6d5f571b68123db9b1728a1f7481a99a833fe0ba4894a24

Observation dff8693b-f982-4353-bb16-a2fb432f6df7 · outbound

This paper cites Correlation Flow,.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Correlation Flow,

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:52.948103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:19:52.822369Z digest=sha256:5954e17403ba1a82f27527346cdcb40a9a45ef36d732a8c02f49148c729095cc

Observation 627141cc-c26f-4354-8f0b-063b89ed57f2 · outbound

This paper cites Towards A Rigorous Science of Interpretable Machine Learning.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Towards A Rigorous Science of Interpretable Machine Learning

Reference 465

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:52.616572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:52.616572Z digest=sha256:a7826c6ea786e92af1ec332e05cdff8552cd72b76c4c47b957569a8098c8ac49

Observation 4d2e9a84-5d0d-4993-a114-52a94c6f4b7c · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Representation Learning with Contrastive Predictive Coding

Reference 2001

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:52.636195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:52.636195Z digest=sha256:6335479b0a62db1487e334b33e709a1f0d6c752a8319e68eae16b7a77474d83b

Observation e8ab9796-9968-4d86-b365-93bab8e9ef69 · outbound

This paper cites stop execution if X is true.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization stop execution if X is true

Reference 2004

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:53.246861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:19:52.718765Z digest=sha256:2c54ff9ddded1b540f749bffda0bad4e5cb9581ae796e899594c0c66f0e8ddfe

Observation 735ef1a2-1143-4332-8459-25218799a086 · outbound

This paper cites In IEEE Conference on Com- puter Vision and Pattern Recognition (CVPR), pages 1735–1742.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization In IEEE Conference on Com- puter Vision and Pattern Recognition (CVPR), pages 1735–1742

Reference 2006

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:53.463425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:19:52.620948Z digest=sha256:31a47d0040a6ccd95f5e1e97e345db89036ed68e84d2524e4d83e04759db2905

Observation 161117fc-6c47-4808-bf29-27efd01ab84c · outbound

This paper cites Unpacking DPO and PPO: Disentangling Best Practices for Learning from Preference Feedback.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Unpacking DPO and PPO: Disentangling Best Practices for Learning from Preference Feedback

Reference 2007

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:52.624338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:52.624338Z digest=sha256:814698fb2afcc5437949a7d8a2db0892175e7015c5ef49f1e64192f96e144d9e

Observation 8efc0ae7-d759-4fe7-b41a-a83990e3f85f · outbound

This paper cites Diffusion Model Alignment Using Direct Preference Optimization.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Diffusion Model Alignment Using Direct Preference Optimization

Reference 2008

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:52.646022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:52.646022Z digest=sha256:2b3b4dd434cfca8d1e854165b94f9c9e3d5b06d8cee2c88f1ae1ff9bbaf3d785

Observation c7c61bca-92d7-45f7-a2de-b5a16724f374 · outbound

This paper cites Direct Preference Optimization: Your Language Model is Secretly a Reward Model.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Direct Preference Optimization: Your Language Model is Secretly a Reward Model

Reference 2009

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:52.643057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:52.643057Z digest=sha256:1c76cc6287f5164ec9c6b1125d625104487626a599b3c2591d19c7233f428f99

Observation 36d39408-377b-4298-ab49-eb471f7f3c6e · outbound

This paper cites In International Conference on Learning Representations (ICLR).

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization In International Conference on Learning Representations (ICLR)

Reference 2018

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:19:53.453574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T22:19:52.632670Z digest=sha256:13e946be1d9350ea915e30a918b63ba13db928dca15d03410b533aa2f81e55bc

Observation c685b90a-17d4-48ae-a7c7-1b2dda82f10a · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Training Verifiers to Solve Math Word Problems

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:52.610934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:52.610934Z digest=sha256:38439b0857c28c9437c509e314a31247fe4df475ff6e0e146078a177cc91600f

Observation f71bc17a-6336-4ed3-bc88-ea28b62fdd03 · outbound

This paper cites Training language models to follow instructions with human feedback.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Training language models to follow instructions with human feedback

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:52.639456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:52.639456Z digest=sha256:ea167abdd07a1c13375aaf3531cc428a032c0de3eb13dd4c06d3342020926789

Observation 3e278364-c54a-4a64-b026-deea65b7dd89 · outbound

This paper cites Let's Verify Step by Step.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Let's Verify Step by Step

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:52.628333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:52.628333Z digest=sha256:7dc2eee9aab8bfe4af9b2d252482b429d0ea973d3a448636806f741a110b5c23

Observation 32c14187-a3d3-4346-9292-90415eed753b · outbound

This paper cites Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference.

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-10T22:19:52.606670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:19:52.606670Z digest=sha256:e9212953138ebf5ddc56dbba78304dd35590335ea7587812d3974cbc0affb9a1

Pith citing papers

No inbound Pith citation observations are available.