Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-29T09:01:03.618498Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 23 of 23 outbound references and 0 inbound Pith citation observations for arXiv:2605.30201.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-29T09:01:03.618498Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
23 of 23 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation e39a9985-dd35-4723-b520-3d5480392985 · outbound
HPO: Hysteretic Policy Optimization for Stable and Efficient Training under Sparse-Reward Regime Enhancing reinforcement learning with dense rewards from language model critic
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69caea79-9fe4-46ea-9d42-fd4eff773694 · outbound
HPO: Hysteretic Policy Optimization for Stable and Efficient Training under Sparse-Reward Regime Unresolved cited work
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 611dcfd6-8825-4c0b-97a5-d13de040fd84 · outbound
HPO: Hysteretic Policy Optimization for Stable and Efficient Training under Sparse-Reward Regime Zhang, Han Bao, Hanwei Xu, Haocheng Wang, Haowei Zhang, Honghui Ding, Huajian Xin, Huazuo Gao, Hui Li, Hui Qu, J
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d5d61ed-08a6-4266-9fb1-27a59694138a · outbound
HPO: Hysteretic Policy Optimization for Stable and Efficient Training under Sparse-Reward Regime DenseGRPO: From sparse to dense reward for flow matching model alignment
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 39fc10bb-581c-4451-91b6-d2f5bce7668c · outbound
HPO: Hysteretic Policy Optimization for Stable and Efficient Training under Sparse-Reward Regime RAFT: Reward ranked finetuning for generative foundation model alignment.Transactions on Machine Learning Research, 2023
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bab408fa-a595-4120-b294-c0ef882a71b7 · outbound
HPO: Hysteretic Policy Optimization for Stable and Efficient Training under Sparse-Reward Regime Re- wardmap: Tackling sparse rewards in fine-grained visual reasoning via multi-stage reinforce- ment learning
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aacf218f-d7d6-47a5-929d-39514e07cbd5 · outbound
HPO: Hysteretic Policy Optimization for Stable and Efficient Training under Sparse-Reward Regime Unresolved cited work
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 09a5a2da-c1e2-455c-84b2-a5f5cd768f49 · outbound
HPO: Hysteretic Policy Optimization for Stable and Efficient Training under Sparse-Reward Regime Soft adaptive policy optimization, 2025
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bdf835ff-720b-4f43-ac3d-6abf490b13b4 · outbound
HPO: Hysteretic Policy Optimization for Stable and Efficient Training under Sparse-Reward Regime Rewarding the unlikely: Lifting grpo beyond distribution sharpening, 2025
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f0f7608-ec8c-4634-959c-e421be09efc5 · outbound
HPO: Hysteretic Policy Optimization for Stable and Efficient Training under Sparse-Reward Regime Reinforcement Learning via Self-Distillation
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a82861da-32fc-4f52-9bb3-d870baa9158a · outbound
HPO: Hysteretic Policy Optimization for Stable and Efficient Training under Sparse-Reward Regime Understanding r1-zero-like training: A critical perspective
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90001c7c-6a19-4a2c-9d22-7af35690bd14 · outbound
HPO: Hysteretic Policy Optimization for Stable and Efficient Training under Sparse-Reward Regime Hysteretic Q-Learning: An Algorithm for Decentralized Reinforcement Learning in Cooperative Multi-agent Teams
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee0fb757-c0d5-4099-b7ea-692ddbb5cc69 · outbound
HPO: Hysteretic Policy Optimization for Stable and Efficient Training under Sparse-Reward Regime Unresolved cited work
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f9c2ff86-5807-4dd2-867f-535ac000ceef · outbound
HPO: Hysteretic Policy Optimization for Stable and Efficient Training under Sparse-Reward Regime Tinyzero
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79636ff0-2e79-4f0f-a5f1-4bf6fb3644df · outbound
HPO: Hysteretic Policy Optimization for Stable and Efficient Training under Sparse-Reward Regime Direct preference optimization: Your language model is secretly a reward model
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 17f6492f-5404-4f03-9f86-2988ac74c19c · outbound
HPO: Hysteretic Policy Optimization for Stable and Efficient Training under Sparse-Reward Regime Reasoning Language Models for Root Cause Analysis in 5G Wireless Networks
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 88713596-c662-4a69-bfc0-466991631b9e · outbound
HPO: Hysteretic Policy Optimization for Stable and Efficient Training under Sparse-Reward Regime Proximal Policy Optimization Algorithms
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b576e0ec-3515-4803-9bf4-6852d2adfd61 · outbound
HPO: Hysteretic Policy Optimization for Stable and Efficient Training under Sparse-Reward Regime Unresolved cited work
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4839da91-6396-4c61-a9ed-26222be91207 · outbound
HPO: Hysteretic Policy Optimization for Stable and Efficient Training under Sparse-Reward Regime Hybridflow: A flexible and efficient rlhf framework
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f3b0af17-308b-461f-88e6-68848bae52c0 · outbound
HPO: Hysteretic Policy Optimization for Stable and Efficient Training under Sparse-Reward Regime A minimalist approach to llm reasoning: from rejection sampling to reinforce, 2025
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ad9e361-91c0-4369-aa80-4692c14bd898 · outbound
HPO: Hysteretic Policy Optimization for Stable and Efficient Training under Sparse-Reward Regime Qwen3 technical report, 2025
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 29e2724e-02d8-4372-9a41-891c4beb1d01 · outbound
HPO: Hysteretic Policy Optimization for Stable and Efficient Training under Sparse-Reward Regime Dapo: An open-source llm reinforcement learning system at scale, 2025
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff42b07c-5c05-402d-a176-6895a276780d · outbound
HPO: Hysteretic Policy Optimization for Stable and Efficient Training under Sparse-Reward Regime Group sequence policy optimization, 2025
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.