Pith. sign in

Paper Citation Record · LEDGER

DuPO: Enabling Reliable LLM Self-Verification via Dual Preference Optimization

As of 5 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 5 inbound Pith citation observations for arXiv:2508.14460.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.14460 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 5 of 5 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-09T03:36:57.168246Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-09T03:45:55.769708Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 2c089767-b6be-4159-b829-90b0d5308803 · inbound

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey cites this paper.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey DuPO: Enabling Reliable LLM Self-Verification via Dual Preference Optimization

Reference 168

Resolution
verified exact
arxiv_id, observed 2026-05-18T19:21:48.822252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:1c01bca4a8393f22b9dfbe17f7bbaf38e30ce5a72a3b2f8b2c134a2fa49d6bdb

Observation f50b5d55-6c77-4489-8a46-2d1762a29fa6 · inbound

HEALing Entropy Collapse: Enhancing Exploration in Few-Shot RLVR via Hybrid-Domain Entropy Dynamics Alignment cites this paper.

HEALing Entropy Collapse: Enhancing Exploration in Few-Shot RLVR via Hybrid-Domain Entropy Dynamics Alignment DuPO: Enabling Reliable LLM Self-Verification via Dual Preference Optimization

Reference 38

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T09:23:37.538719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-10T05:23:08.478393Z digest=sha256:64deedf1d0a28ec415355a00a18e1731e2e050cd88b9869e42ffa0941c5d039a

Observation 5f83a916-69e9-4b21-9e86-0743939841f6 · inbound

Reinforced Collaboration in Multi-Agent Flow Networks cites this paper.

Reinforced Collaboration in Multi-Agent Flow Networks DuPO: Enabling Reliable LLM Self-Verification via Dual Preference Optimization

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:42:57.152042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T20:42:48.438057Z digest=sha256:e81fa284e1d759d4b4f0bcd04a79158d8bf610d377a2adcbef4ec7dbbece78e2

Observation 6be0487b-94c1-4ea9-b66b-0180bdc3eeac · inbound

Be Your Own Teacher: Steering Protein Language Models via Unsupervised Reward Optimization cites this paper.

Be Your Own Teacher: Steering Protein Language Models via Unsupervised Reward Optimization DuPO: Enabling Reliable LLM Self-Verification via Dual Preference Optimization

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-07-03T23:49:02.982884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-26T21:40:55.946788Z digest=sha256:488adf629c5bc3b9cb84c9c28deeae34d93556cfc85dadcbebe7a2f096489d96

Observation 54bbfa46-f077-471c-8e24-920f431e2757 · inbound

Recursive Self-Improvement in AI: From Bounded Self-Refinement to Autonomous Research Loops cites this paper.

Recursive Self-Improvement in AI: From Bounded Self-Refinement to Autonomous Research Loops DuPO: Enabling Reliable LLM Self-Verification via Dual Preference Optimization

Reference 147

Resolution
verified exact
local_arxiv, observed 2026-07-09T03:45:55.770922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-09T03:36:57.168246Z digest=sha256:d8012564265fc2adb34898543d24a39b11ce1448865a952f13146bcb0eacdfb3