Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-02T14:39:41.941604Z
Paper Citation Record · LEDGER
As of 11 August 2026, this Paper Citation Record lists 9 of 9 outbound references and 0 inbound Pith citation observations for arXiv:2607.16206.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-02T14:39:41.941604Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
9 of 9 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 4e5d2703-1aa4-4b6c-bef6-852ff8dd2f3b · outbound
PPO-HSC: An Exploratory Reinforcement Learning Framework Based on Wide-Area Policy Coverage Optimization Proximal Policy Optimization Algorithms
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 78db689e-a9b8-4d48-8a96-3b0e42ec475e · outbound
PPO-HSC: An Exploratory Reinforcement Learning Framework Based on Wide-Area Policy Coverage Optimization Advances in Neural Information Processing Systems, 36 (2023) PPO-HSC: Exploratory RL via Policy Coverage Optimization 13
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f0e6076-2c99-483a-921b-6606026c0a5e · outbound
PPO-HSC: An Exploratory Reinforcement Learning Framework Based on Wide-Area Policy Coverage Optimization Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e3d26c6a-edaf-419d-8803-3a1e3e894c4e · outbound
PPO-HSC: An Exploratory Reinforcement Learning Framework Based on Wide-Area Policy Coverage Optimization Advances in Neural Information Processing Systems, 29 (2016)
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 72246951-7bf9-4f5e-9c37-4e4017687e27 · outbound
PPO-HSC: An Exploratory Reinforcement Learning Framework Based on Wide-Area Policy Coverage Optimization In: International Conference on Machine Learning, pp
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae3c72eb-87a7-43e9-95e9-bae7cce5a9cb · outbound
PPO-HSC: An Exploratory Reinforcement Learning Framework Based on Wide-Area Policy Coverage Optimization Connection Science, 3(3), 241-268 (1991)
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 079e982c-2674-4b67-8236-be916da7b3f8 · outbound
PPO-HSC: An Exploratory Reinforcement Learning Framework Based on Wide-Area Policy Coverage Optimization Frontiers in Robotics and AI, 3, 40 (2016)
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e19e80e0-c24c-4057-b940-d802b63f77a1 · outbound
PPO-HSC: An Exploratory Reinforcement Learning Framework Based on Wide-Area Policy Coverage Optimization O.: Abandoning Objectives: Evolution Through the Search for Novelty Alone
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69e4a543-aa9d-45b7-9829-c43d79582777 · outbound
PPO-HSC: An Exploratory Reinforcement Learning Framework Based on Wide-Area Policy Coverage Optimization Illuminating search spaces by mapping elites
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.