Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T17:54:26.542958Z
Paper Citation Record · LEDGER
As of 20 August 2026, this Paper Citation Record lists 12 of 12 outbound references and 0 inbound Pith citation observations for arXiv:2507.20133.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T17:54:26.542958Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
12 of 12 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 95e95346-6727-40db-90e9-842654018759 · outbound
Sem-DPO: Mitigating Semantic Inconsistency in Preference Optimization for Prompt Engineering AI Alignment through Reinforcement Learning from Human Feedback? Contradictions and Limitations
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b249688-81cc-4792-99d5-303bfac6a3a9 · outbound
Sem-DPO: Mitigating Semantic Inconsistency in Preference Optimization for Prompt Engineering Entropy Controllable Direct Preference Optimization
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ade77db5-f931-4a63-8a69-74c51116f0bb · outbound
Sem-DPO: Mitigating Semantic Inconsistency in Preference Optimization for Prompt Engineering Disentangling Length from Quality in Direct Preference Optimization
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7699e76f-f7c0-423e-9396-ebf1a485146a · outbound
Sem-DPO: Mitigating Semantic Inconsistency in Preference Optimization for Prompt Engineering Direct Preference Optimization: Your Language Model is Secretly a Reward Model
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7a8ee54-fc26-48af-97f1-dba559aa856a · outbound
Sem-DPO: Mitigating Semantic Inconsistency in Preference Optimization for Prompt Engineering arXiv preprint arXiv:2411.04712
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 822f6461-fac5-43b8-a6c7-0771abc94bf4 · outbound
Sem-DPO: Mitigating Semantic Inconsistency in Preference Optimization for Prompt Engineering RSPO: Regularized Self-Play Alignment of Large Language Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8934c5be-c620-4ef0-b3b4-b4cf842d666f · outbound
Sem-DPO: Mitigating Semantic Inconsistency in Preference Optimization for Prompt Engineering Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 165200ea-805e-4b0e-82ac-54fe1e5aeda6 · outbound
Sem-DPO: Mitigating Semantic Inconsistency in Preference Optimization for Prompt Engineering Orthogonal Finetuning for Direct Preference Optimization
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation a80da313-aa7a-4318-ac82-292c3282deec · outbound
Sem-DPO: Mitigating Semantic Inconsistency in Preference Optimization for Prompt Engineering Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc8ecfaa-e4ca-4fde-9767-b4d0cdf7e45c · outbound
Sem-DPO: Mitigating Semantic Inconsistency in Preference Optimization for Prompt Engineering RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c11746e-2098-4200-bd06-e9bbc5f8ce73 · outbound
Sem-DPO: Mitigating Semantic Inconsistency in Preference Optimization for Prompt Engineering Self-play with Execution Feedback: Improving Instruction-following Capabilities of Large Language Models
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef8f7377-8197-47fa-945b-04a680345a5e · outbound
Sem-DPO: Mitigating Semantic Inconsistency in Preference Optimization for Prompt Engineering DiaTool-DPO: Multi-Turn Direct Preference Optimization for Tool-Augmented Large Language Models
Reference 2025
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
No inbound Pith citation observations are available.