Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T14:59:12.228176Z
Paper Citation Record · LEDGER
As of 14 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 0 inbound Pith citation observations for arXiv:2608.09568.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T14:59:12.228176Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
29 of 29 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation c6fa8945-4983-4e15-b8ee-fb64a46a421f · outbound
Se-DPO: Self-Evolving Token Credit for Direct Preference Optimization The Softplus activation ensures non-negative output
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 8a8df1ca-40a5-474e-8c68-72426770d7b4 · outbound
Se-DPO: Self-Evolving Token Credit for Direct Preference Optimization TAB-PO: Preference Optimization with a Token-Level Adaptive Barrier for Token-Critical Structured Generation
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33b45022-8242-4122-9b16-6b43985a1c3a · outbound
Se-DPO: Self-Evolving Token Credit for Direct Preference Optimization The Llama 3 Herd of Models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70d24357-23bb-4e09-8381-07e1d75c7ae3 · outbound
Se-DPO: Self-Evolving Token Credit for Direct Preference Optimization Adaptive batch-wise sample scheduling for direct preference optimization.arXiv preprint arXiv:2506.17252,
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13b60a3a-61fa-46b4-872d-bc34ccacddfb · outbound
Se-DPO: Self-Evolving Token Credit for Direct Preference Optimization Kl penalty control via perturbation for direct preference optimization.arXiv preprint arXiv:2502.13177,
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 9b1089d8-0b76-4953-8b22-cc2c7b53c939 · outbound
Se-DPO: Self-Evolving Token Credit for Direct Preference Optimization AutoSupervision: Closing the Feedback Loop in Scientific Workflows with Grounded Revision Verification
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 121b761d-a040-4c1f-8fb9-9f2acc0b40d4 · outbound
Se-DPO: Self-Evolving Token Credit for Direct Preference Optimization From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49016e83-99d3-461e-9614-038d2905fce0 · outbound
Se-DPO: Self-Evolving Token Credit for Direct Preference Optimization TIS-DPO: Token-level Importance Sampling for Direct Preference Optimization With Estimated Weights
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa33753e-55f5-45cd-9414-fc31134649ba · outbound
Se-DPO: Self-Evolving Token Credit for Direct Preference Optimization Autoregressive Direct Preference Optimization
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 56850233-3fec-4c8e-8ff8-4b7147d966d8 · outbound
Se-DPO: Self-Evolving Token Credit for Direct Preference Optimization Small-margin preferences still matter—if you train them right.arXiv preprint arXiv:2602.00954,
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45ea035a-bcc7-47b1-a214-9cf9a7f8fc28 · outbound
Se-DPO: Self-Evolving Token Credit for Direct Preference Optimization From $r$ to $Q^*$: Your Language Model is Secretly a Q-Function
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a99d5f1a-a66f-4978-839a-6f23484af853 · outbound
Se-DPO: Self-Evolving Token Credit for Direct Preference Optimization Proximal Policy Optimization Algorithms
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 71bbb7a6-f4c8-4a2c-bf4c-ed864b7bcea5 · outbound
Se-DPO: Self-Evolving Token Credit for Direct Preference Optimization Interpretable preferences via multi-objective reward modeling and mixture-of-experts
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 63b16c52-ce48-4e97-8a1e-7bbafd9a170b · outbound
Se-DPO: Self-Evolving Token Credit for Direct Preference Optimization Explore the reasoning capability of LLMs in the chess testbed
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 065bbcbc-0ea8-4de3-9c63-5b860f4f90ab · outbound
Se-DPO: Self-Evolving Token Credit for Direct Preference Optimization URLhttps://aclanthology.org/2025.naacl-short.52/
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation a0d28849-a59f-498a-bcf0-7315c6904fc0 · outbound
Se-DPO: Self-Evolving Token Credit for Direct Preference Optimization Se- lective preference optimization via token-level reward function estimation
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c911f9fa-1abc-46b8-9a46-443e0f913cc4 · outbound
Se-DPO: Self-Evolving Token Credit for Direct Preference Optimization RUBRIC: Realism--Utility Balanced Ranking for Imbalanced Classification
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 90cbbf2e-91f9-40bf-9088-e69e9fdbd730 · outbound
Se-DPO: Self-Evolving Token Credit for Direct Preference Optimization A-SR: Self-Evolving Agentic LLMs for Symbolic Regression via Hierarchical Coordination
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 3ae436d4-1233-42b5-b5d3-5b999e4641c1 · outbound
Se-DPO: Self-Evolving Token Credit for Direct Preference Optimization Wpo: Enhancing rlhf with weighted preference optimization
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 053fcd64-447e-461e-b378-fd25c27137ca · outbound
Se-DPO: Self-Evolving Token Credit for Direct Preference Optimization TGDPO: Harnessing Token-Level Reward Guidance for Enhancing Direct Preference Optimization
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d324eaf6-ebbf-4695-8e22-4b4855768640 · outbound
Se-DPO: Self-Evolving Token Credit for Direct Preference Optimization Fine-Tuning Language Models from Human Preferences
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fee99bea-86ff-4e65-b828-eb08c74dd85a · outbound
Se-DPO: Self-Evolving Token Credit for Direct Preference Optimization Sparsepo: Controlling preference alignment of llms via sparse token masks
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a452e08-4ea5-4a91-80dc-99e6096b74e7 · outbound
Se-DPO: Self-Evolving Token Credit for Direct Preference Optimization Under independent noise, Var[ˆ∆] =∑t c2 t σ2 t
Reference 2019
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 564c92a3-b4ae-40ee-ae6c-fa07150f4530 · outbound
Se-DPO: Self-Evolving Token Credit for Direct Preference Optimization Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 959f098a-23bb-411c-9ad1-170ba632a6f2 · outbound
Se-DPO: Self-Evolving Token Credit for Direct Preference Optimization Discriminative Policy Optimization for Token-Level Reward Models
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e050620a-34d5-46d5-b4d6-852f9bc765c1 · outbound
Se-DPO: Self-Evolving Token Credit for Direct Preference Optimization Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa42e25b-f3b6-4a78-bc79-2cee2585bf2e · outbound
Se-DPO: Self-Evolving Token Credit for Direct Preference Optimization Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f4e325dd-45d5-4f10-88c2-f9f7f3e02579 · outbound
Se-DPO: Self-Evolving Token Credit for Direct Preference Optimization Process Reinforcement through Implicit Rewards
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de9a80bd-8553-44e9-b2f7-41399341a89c · outbound
Se-DPO: Self-Evolving Token Credit for Direct Preference Optimization Gemma 2: Improving Open Language Models at a Practical Size
Reference 2026
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.