Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 6 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 5 inbound Pith citation observations for arXiv:1909.12238.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-06-29T19:44:02.317303Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
0 of 0 outbound references displayed
External citation measurements
39
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
No outbound reference observations are available for this paper version.
Observation 4debb08a-1c94-4462-82f0-c91fa6b6f80d · inbound
Improving alignment of dialogue agents via targeted human judgements V-MPO: On-Policy Maximum a Posteriori Policy Optimization for Discrete and Continuous Control
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f30a9262-2c2d-4d6c-b07d-5a5eaaed3143 · inbound
Promptbreeder: Self-Referential Self-Improvement Via Prompt Evolution V-MPO: On-Policy Maximum a Posteriori Policy Optimization for Discrete and Continuous Control
Reference 299
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c6cc8c3f-608c-48a1-b990-c82fdc6ef923 · inbound
Listwise Policy Optimization: Group-based RLVR as Target-Projection on the LLM Response Simplex V-MPO: On-Policy Maximum a Posteriori Policy Optimization for Discrete and Continuous Control
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d1786848-5e9b-4e83-89f9-627c49c605bc · inbound
Listwise Policy Optimization: Group-based RLVR as Target-Projection on the LLM Response Simplex V-MPO: On-Policy Maximum a Posteriori Policy Optimization for Discrete and Continuous Control
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f108616f-79d7-43ac-adf5-b3c5c05d03dc · inbound
Ratio-Variance Regularized Policy Optimization V-MPO: On-Policy Maximum a Posteriori Policy Optimization for Discrete and Continuous Control
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.