Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 5 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 5 inbound Pith citation observations for arXiv:2508.14460.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-09T03:36:57.168246Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-09T03:45:55.769708Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 2c089767-b6be-4159-b829-90b0d5308803 · inbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey DuPO: Enabling Reliable LLM Self-Verification via Dual Preference Optimization
Reference 168
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation f50b5d55-6c77-4489-8a46-2d1762a29fa6 · inbound
HEALing Entropy Collapse: Enhancing Exploration in Few-Shot RLVR via Hybrid-Domain Entropy Dynamics Alignment DuPO: Enabling Reliable LLM Self-Verification via Dual Preference Optimization
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 5f83a916-69e9-4b21-9e86-0743939841f6 · inbound
Reinforced Collaboration in Multi-Agent Flow Networks DuPO: Enabling Reliable LLM Self-Verification via Dual Preference Optimization
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 6be0487b-94c1-4ea9-b66b-0180bdc3eeac · inbound
Be Your Own Teacher: Steering Protein Language Models via Unsupervised Reward Optimization DuPO: Enabling Reliable LLM Self-Verification via Dual Preference Optimization
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 54bbfa46-f077-471c-8e24-920f431e2757 · inbound
Recursive Self-Improvement in AI: From Bounded Self-Refinement to Autonomous Research Loops DuPO: Enabling Reliable LLM Self-Verification via Dual Preference Optimization
Reference 147
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.