Pith. sign in

Paper Citation Record · LEDGER

Se-DPO: Self-Evolving Token Credit for Direct Preference Optimization

As of 14 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 0 inbound Pith citation observations for arXiv:2608.09568.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.09568 v1

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T14:59:12.228176Z

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

29 of 29 outbound references displayed

  • verified exact2
  • verified fuzzy6
  • unresolved18
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c6fa8945-4983-4e15-b8ee-fb64a46a421f · outbound

This paper cites The Softplus activation ensures non-negative output.

Se-DPO: Self-Evolving Token Credit for Direct Preference Optimization The Softplus activation ensures non-negative output

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:59:13.099952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T14:59:12.228176Z digest=sha256:4cdc3022644c44f4f6e908b32b284a4eb87493fef0d265ad1c3b33b3bedc60ed

Observation 8a8df1ca-40a5-474e-8c68-72426770d7b4 · outbound

This paper cites TAB-PO: Preference Optimization with a Token-Level Adaptive Barrier for Token-Critical Structured Generation.

Se-DPO: Self-Evolving Token Credit for Direct Preference Optimization TAB-PO: Preference Optimization with a Token-Level Adaptive Barrier for Token-Critical Structured Generation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:12.089794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:12.089794Z digest=sha256:71a484ac32d6981f2d849a965414b1f038ebada86a561e036e4aa4548008dc5c

Observation 33b45022-8242-4122-9b16-6b43985a1c3a · outbound

This paper cites The Llama 3 Herd of Models.

Se-DPO: Self-Evolving Token Credit for Direct Preference Optimization The Llama 3 Herd of Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:12.102438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:12.102438Z digest=sha256:910c950afce7dd67570cdc0957dfb9f27d45afea98b409fe4b6c117d8694b175

Observation 70d24357-23bb-4e09-8381-07e1d75c7ae3 · outbound

This paper cites Adaptive batch-wise sample scheduling for direct preference optimization.arXiv preprint arXiv:2506.17252,.

Se-DPO: Self-Evolving Token Credit for Direct Preference Optimization Adaptive batch-wise sample scheduling for direct preference optimization.arXiv preprint arXiv:2506.17252,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:12.108854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:12.108854Z digest=sha256:34197baac442e4b6260ef8dba1fafad1c6cca94b006e85974c08ab8c93d4740f

Observation 13b60a3a-61fa-46b4-872d-bc34ccacddfb · outbound

This paper cites Kl penalty control via perturbation for direct preference optimization.arXiv preprint arXiv:2502.13177,.

Se-DPO: Self-Evolving Token Credit for Direct Preference Optimization Kl penalty control via perturbation for direct preference optimization.arXiv preprint arXiv:2502.13177,

Reference 10

Resolution
verified exact
raw_fallback, observed 2026-08-11T14:59:12.785116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T14:59:12.120007Z digest=sha256:4eabca466c7bf12ff27415a65ad609309b14cde17decaa3a641dfcc6ee432b59

Observation 9b1089d8-0b76-4953-8b22-cc2c7b53c939 · outbound

This paper cites AutoSupervision: Closing the Feedback Loop in Scientific Workflows with Grounded Revision Verification.

Se-DPO: Self-Evolving Token Credit for Direct Preference Optimization AutoSupervision: Closing the Feedback Loop in Scientific Workflows with Grounded Revision Verification

Reference 11

Resolution
metadata mismatch
local_arxiv, observed 2026-08-11T14:59:12.690701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T14:59:12.125194Z digest=sha256:3711fddfa7b89690cac3f15c94b1ad83bb504efc673f86038620e30b04f6fa2c

Observation 121b761d-a040-4c1f-8fb9-9f2acc0b40d4 · outbound

This paper cites From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline.

Se-DPO: Self-Evolving Token Credit for Direct Preference Optimization From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:12.131103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:12.131103Z digest=sha256:ad9066e7585ce8f485f54abfc3918a4db6a1ea2a6a1851a38500b39d072389ce

Observation 49016e83-99d3-461e-9614-038d2905fce0 · outbound

This paper cites TIS-DPO: Token-level Importance Sampling for Direct Preference Optimization With Estimated Weights.

Se-DPO: Self-Evolving Token Credit for Direct Preference Optimization TIS-DPO: Token-level Importance Sampling for Direct Preference Optimization With Estimated Weights

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:12.137307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:12.137307Z digest=sha256:34e0abc5b9658b0af8b880c9c31630e225b6ea4777e4fc97984dc9475fbbcbad

Observation fa33753e-55f5-45cd-9414-fc31134649ba · outbound

This paper cites Autoregressive Direct Preference Optimization.

Se-DPO: Self-Evolving Token Credit for Direct Preference Optimization Autoregressive Direct Preference Optimization

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-08-11T14:59:12.616020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T14:59:12.142774Z digest=sha256:f6973d4f95a34b4b24f19dd6c922b954125cb4d18ca52b05ee7ccd2550a16e5b

Observation 56850233-3fec-4c8e-8ff8-4b7147d966d8 · outbound

This paper cites Small-margin preferences still matter—if you train them right.arXiv preprint arXiv:2602.00954,.

Se-DPO: Self-Evolving Token Credit for Direct Preference Optimization Small-margin preferences still matter—if you train them right.arXiv preprint arXiv:2602.00954,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:12.148038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:12.148038Z digest=sha256:f930dbab26d90fda402309d1c8a910014f4efdbcb3590f150712cb8bedb693e6

Observation 45ea035a-bcc7-47b1-a214-9cf9a7f8fc28 · outbound

This paper cites From $r$ to $Q^*$: Your Language Model is Secretly a Q-Function.

Se-DPO: Self-Evolving Token Credit for Direct Preference Optimization From $r$ to $Q^*$: Your Language Model is Secretly a Q-Function

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:12.152668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:12.152668Z digest=sha256:a7384be3654e369072110d12906d9161749f348896fb28e553156e1cac2fcb62

Observation a99d5f1a-a66f-4978-839a-6f23484af853 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Se-DPO: Self-Evolving Token Credit for Direct Preference Optimization Proximal Policy Optimization Algorithms

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:12.157429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:12.157429Z digest=sha256:b861dc82d325180682599dce815df2918689b684103920570f9be59f73f23cdf

Observation 71bbb7a6-f4c8-4a2c-bf4c-ed864b7bcea5 · outbound

This paper cites Interpretable preferences via multi-objective reward modeling and mixture-of-experts.

Se-DPO: Self-Evolving Token Credit for Direct Preference Optimization Interpretable preferences via multi-objective reward modeling and mixture-of-experts

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:59:13.194401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T14:59:12.166894Z digest=sha256:a814dfdbcd6bb548e5464caff0c50851420ff97059aa85ed5648d780c2837b50

Observation 63b16c52-ce48-4e97-8a1e-7bbafd9a170b · outbound

This paper cites Explore the reasoning capability of LLMs in the chess testbed.

Se-DPO: Self-Evolving Token Credit for Direct Preference Optimization Explore the reasoning capability of LLMs in the chess testbed

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:59:13.173052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T14:59:12.172751Z digest=sha256:8ac2def389bc341bcb5bcfb6455bab7abd522b308de41b842080c23b8bcc606d

Observation 065bbcbc-0ea8-4de3-9c63-5b860f4f90ab · outbound

This paper cites URLhttps://aclanthology.org/2025.naacl-short.52/.

Se-DPO: Self-Evolving Token Credit for Direct Preference Optimization URLhttps://aclanthology.org/2025.naacl-short.52/

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:59:13.154557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T14:59:12.180368Z digest=sha256:af2fb7eca90548f380d8ecda3239e5d42689448e09157bc204b950123c928177

Observation a0d28849-a59f-498a-bcf0-7315c6904fc0 · outbound

This paper cites Se- lective preference optimization via token-level reward function estimation.

Se-DPO: Self-Evolving Token Credit for Direct Preference Optimization Se- lective preference optimization via token-level reward function estimation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:12.186010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:12.186010Z digest=sha256:c16b6de254b8aa9f5d4b67a58b84e49e70d04036749e444abee38d49da28421f

Observation c911f9fa-1abc-46b8-9a46-443e0f913cc4 · outbound

This paper cites RUBRIC: Realism--Utility Balanced Ranking for Imbalanced Classification.

Se-DPO: Self-Evolving Token Credit for Direct Preference Optimization RUBRIC: Realism--Utility Balanced Ranking for Imbalanced Classification

Reference 23

Resolution
metadata mismatch
local_arxiv, observed 2026-08-11T14:59:12.357820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T14:59:12.192104Z digest=sha256:3023a21c1ade68980229184f8271f0aac2b3ebf459d475ddb37d8f47223aec67

Observation 90cbbf2e-91f9-40bf-9088-e69e9fdbd730 · outbound

This paper cites A-SR: Self-Evolving Agentic LLMs for Symbolic Regression via Hierarchical Coordination.

Se-DPO: Self-Evolving Token Credit for Direct Preference Optimization A-SR: Self-Evolving Agentic LLMs for Symbolic Regression via Hierarchical Coordination

Reference 24

Resolution
metadata mismatch
local_arxiv, observed 2026-08-11T14:59:12.329770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T14:59:12.197591Z digest=sha256:fc66b336d43800f4d15a814c7bf5cca96a8ccb761b473d58a9010db0fde2cd82

Observation 3ae436d4-1233-42b5-b5d3-5b999e4641c1 · outbound

This paper cites Wpo: Enhancing rlhf with weighted preference optimization.

Se-DPO: Self-Evolving Token Credit for Direct Preference Optimization Wpo: Enhancing rlhf with weighted preference optimization

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:59:13.136581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T14:59:12.202844Z digest=sha256:f63ff6808109748723471f1d37ce4d6840665d222929c1ba2473d793e22567ef

Observation 053fcd64-447e-461e-b378-fd25c27137ca · outbound

This paper cites TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization.

Se-DPO: Self-Evolving Token Credit for Direct Preference Optimization TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:12.210012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:12.210012Z digest=sha256:a9e3ddd1a348d68863913d6eb786b37286f1200466fbdb6528c9750f9b591fde

Observation d324eaf6-ebbf-4695-8e22-4b4855768640 · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

Se-DPO: Self-Evolving Token Credit for Direct Preference Optimization Fine-Tuning Language Models from Human Preferences

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:12.216222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:12.216222Z digest=sha256:ab35397aff9f9819914be0291f681ca3b270804e513cd69b5c167024d160d76f

Observation fee99bea-86ff-4e65-b828-eb08c74dd85a · outbound

This paper cites Sparsepo: Controlling preference alignment of llms via sparse token masks.

Se-DPO: Self-Evolving Token Credit for Direct Preference Optimization Sparsepo: Controlling preference alignment of llms via sparse token masks

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:12.076346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:12.076346Z digest=sha256:79aa0fcce4ac516c1c9ccb642eac18a7211e533c3104ef2f96d0100b3ee86e6c

Observation 4a452e08-4ea5-4a91-80dc-99e6096b74e7 · outbound

This paper cites Under independent noise, Var[ˆ∆] =∑t c2 t σ2 t.

Se-DPO: Self-Evolving Token Credit for Direct Preference Optimization Under independent noise, Var[ˆ∆] =∑t c2 t σ2 t

Reference 2019

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:59:13.117690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T14:59:12.222521Z digest=sha256:d21098428c10628319e8a3d51a3b44ded6e55b85cabd8b931e052f7cf51786e5

Observation 564c92a3-b4ae-40ee-ae6c-fa07150f4530 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Se-DPO: Self-Evolving Token Credit for Direct Preference Optimization Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:12.161810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:12.161810Z digest=sha256:128568dad3bb936614990e6a7c3e17ec048f6db028760d53b647b126c309ff4e

Observation 959f098a-23bb-411c-9ad1-170ba632a6f2 · outbound

This paper cites Discriminative Policy Optimization for Token-Level Reward Models.

Se-DPO: Self-Evolving Token Credit for Direct Preference Optimization Discriminative Policy Optimization for Token-Level Reward Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:12.070055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:12.070055Z digest=sha256:df8cfb1416799536df4129ff6df1891d463225c564058d0c5af0574d6a5c2080

Observation e050620a-34d5-46d5-b4d6-852f9bc765c1 · outbound

This paper cites Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs.

Se-DPO: Self-Evolving Token Credit for Direct Preference Optimization Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:12.114561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:12.114561Z digest=sha256:2c115cd429174526eea8f0acbb2dedfc52e2a80cea8303aad01c0b644004902a

Observation aa42e25b-f3b6-4a78-bc79-2cee2585bf2e · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Se-DPO: Self-Evolving Token Credit for Direct Preference Optimization Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:12.063432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:12.063432Z digest=sha256:4f934d2164de89f37da3dafbba2e43005e6c7ae03e5f10870a3536a4e574dba2

Observation f4e325dd-45d5-4f10-88c2-f9f7f3e02579 · outbound

This paper cites Process Reinforcement through Implicit Rewards.

Se-DPO: Self-Evolving Token Credit for Direct Preference Optimization Process Reinforcement through Implicit Rewards

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:12.083599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:12.083599Z digest=sha256:f5d9d5384d721dd3077ad9205dbd01c40d5acf0330c444e62119a4942927d598

Observation de9a80bd-8553-44e9-b2f7-41399341a89c · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

Se-DPO: Self-Evolving Token Credit for Direct Preference Optimization Gemma 2: Improving Open Language Models at a Practical Size

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:12.096578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:12.096578Z digest=sha256:75929072565380b14976bb8165624b58f196f4c6fff4126884f8cf5de7a25b5d

Pith citing papers

No inbound Pith citation observations are available.