Pith. sign in

Paper Citation Record · LEDGER

Sem-DPO: Mitigating Semantic Inconsistency in Preference Optimization for Prompt Engineering

As of 20 August 2026, this Paper Citation Record lists 12 of 12 outbound references and 0 inbound Pith citation observations for arXiv:2507.20133.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.20133 v2

Coverage vector

measured 12 of 12 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T17:54:26.542958Z

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

12 of 12 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 95e95346-6727-40db-90e9-842654018759 · outbound

This paper cites AI Alignment through Reinforcement Learning from Human Feedback? Contradictions and Limitations.

Sem-DPO: Mitigating Semantic Inconsistency in Preference Optimization for Prompt Engineering AI Alignment through Reinforcement Learning from Human Feedback? Contradictions and Limitations

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T17:54:26.500695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:54:26.500695Z digest=sha256:5598a05f44fab6d70acb7a90454bc9da685d8f16e72043d9e549957f8cfaf687

Observation 6b249688-81cc-4792-99d5-303bfac6a3a9 · outbound

This paper cites Entropy Controllable Direct Preference Optimization.

Sem-DPO: Mitigating Semantic Inconsistency in Preference Optimization for Prompt Engineering Entropy Controllable Direct Preference Optimization

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T17:54:26.507754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:54:26.507754Z digest=sha256:6792ee2bf8b7dafebe1a2e3e16876877e91b836fe18995b35f84ed8ce3c276b4

Observation ade77db5-f931-4a63-8a69-74c51116f0bb · outbound

This paper cites Disentangling Length from Quality in Direct Preference Optimization.

Sem-DPO: Mitigating Semantic Inconsistency in Preference Optimization for Prompt Engineering Disentangling Length from Quality in Direct Preference Optimization

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T17:54:26.513629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:54:26.513629Z digest=sha256:ccb315a52e00cd20efdd321eaa28262b621837ce680d75a3d6a541117de4ec46

Observation 7699e76f-f7c0-423e-9396-ebf1a485146a · outbound

This paper cites Direct Preference Optimization: Your Language Model is Secretly a Reward Model.

Sem-DPO: Mitigating Semantic Inconsistency in Preference Optimization for Prompt Engineering Direct Preference Optimization: Your Language Model is Secretly a Reward Model

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T17:54:26.518764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:54:26.518764Z digest=sha256:fb300bb6c838aef8abd883b0444280fc5298e7c22a500cbb32405cb2ac70e8f8

Observation a7a8ee54-fc26-48af-97f1-dba559aa856a · outbound

This paper cites arXiv preprint arXiv:2411.04712.

Sem-DPO: Mitigating Semantic Inconsistency in Preference Optimization for Prompt Engineering arXiv preprint arXiv:2411.04712

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T17:54:26.523900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:54:26.523900Z digest=sha256:7fa94d28aada95edbb73a8296ca131d4a7372d802121d421477a2082347082fd

Observation 822f6461-fac5-43b8-a6c7-0771abc94bf4 · outbound

This paper cites RSPO: Regularized Self-Play Alignment of Large Language Models.

Sem-DPO: Mitigating Semantic Inconsistency in Preference Optimization for Prompt Engineering RSPO: Regularized Self-Play Alignment of Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T17:54:26.531019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:54:26.531019Z digest=sha256:1a1fae6a13b2f0f3ef67f259c7bcff96822d83caef0dd49f7da3fbe2ac548c87

Observation 8934c5be-c620-4ef0-b3b4-b4cf842d666f · outbound

This paper cites Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints.

Sem-DPO: Mitigating Semantic Inconsistency in Preference Optimization for Prompt Engineering Beyond Reverse KL: Generalizing Direct Preference Optimization with Diverse Divergence Constraints

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T17:54:26.537395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:54:26.537395Z digest=sha256:f947fb758916bcdd1049926e61e3e21dd0fb42ba66123510ff8ae19208406bf8

Observation 165200ea-805e-4b0e-82ac-54fe1e5aeda6 · outbound

This paper cites Orthogonal Finetuning for Direct Preference Optimization.

Sem-DPO: Mitigating Semantic Inconsistency in Preference Optimization for Prompt Engineering Orthogonal Finetuning for Direct Preference Optimization

Reference 12

Resolution
metadata mismatch
local_arxiv, observed 2026-08-15T17:54:26.601911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T17:54:26.542958Z digest=sha256:36bfa452959bcbb108621789cfadd182724baa8fc8bf18c8c65b1b5803276609

Observation a80da313-aa7a-4318-ac82-292c3282deec · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Sem-DPO: Mitigating Semantic Inconsistency in Preference Optimization for Prompt Engineering Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-15T17:54:26.471746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:54:26.471746Z digest=sha256:5da19ea263cc6dd915cfc6dfa11c18d57af08527dcb422ce7b32285e9090c95b

Observation dc8ecfaa-e4ca-4fde-9767-b4d0cdf7e45c · outbound

This paper cites RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback.

Sem-DPO: Mitigating Semantic Inconsistency in Preference Optimization for Prompt Engineering RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-15T17:54:26.493999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:54:26.493999Z digest=sha256:590f90aef8451b0947218ed8ff0dd1f65682d1715b5ca77c215af42fd416b233

Observation 2c11746e-2098-4200-bd06-e9bbc5f8ce73 · outbound

This paper cites Self-play with Execution Feedback: Improving Instruction-following Capabilities of Large Language Models.

Sem-DPO: Mitigating Semantic Inconsistency in Preference Optimization for Prompt Engineering Self-play with Execution Feedback: Improving Instruction-following Capabilities of Large Language Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-15T17:54:26.479849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:54:26.479849Z digest=sha256:c7378d6e4f9102f07a191373e9869c5ca0fe682fb1c6698d3bf16252177dd262

Observation ef8f7377-8197-47fa-945b-04a680345a5e · outbound

This paper cites DiaTool-DPO: Multi-Turn Direct Preference Optimization for Tool-Augmented Large Language Models.

Sem-DPO: Mitigating Semantic Inconsistency in Preference Optimization for Prompt Engineering DiaTool-DPO: Multi-Turn Direct Preference Optimization for Tool-Augmented Large Language Models

Reference 2025

Resolution
metadata mismatch
local_arxiv, observed 2026-08-15T17:54:26.877988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T17:54:26.488087Z digest=sha256:0790a9fe24082e172b928584462fd2821ce4aa32649ac50361ee1de515bf5b8d

Pith citing papers

No inbound Pith citation observations are available.