Pith. sign in

Paper Citation Record · LEDGER

HPO: Hysteretic Policy Optimization for Stable and Efficient Training under Sparse-Reward Regime

As of 8 August 2026, this Paper Citation Record lists 23 of 23 outbound references and 0 inbound Pith citation observations for arXiv:2605.30201.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.30201 v1

Coverage vector

measured 23 of 23 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-29T09:01:03.618498Z

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

23 of 23 outbound references displayed

  • verified exact3
  • verified fuzzy0
  • unresolved20
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e39a9985-dd35-4723-b520-3d5480392985 · outbound

This paper cites Enhancing reinforcement learning with dense rewards from language model critic.

HPO: Hysteretic Policy Optimization for Stable and Efficient Training under Sparse-Reward Regime Enhancing reinforcement learning with dense rewards from language model critic

Reference 1

Resolution
unresolved
no resolver link, observed 2026-06-29T09:01:03.618498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T09:01:03.618498Z digest=sha256:542668e7bbddf1ab044336068d604ba117f736f3f8c65b7b954a699f1c972e9a

Observation 69caea79-9fe4-46ea-9d42-fd4eff773694 · outbound

This paper cites an unresolved cited work.

HPO: Hysteretic Policy Optimization for Stable and Efficient Training under Sparse-Reward Regime Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-06-29T09:01:03.618498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T09:01:03.618498Z digest=sha256:9a6a031868832184e599ea0791d60bf72255034dd1cdd2b4861cfc5976b65a31

Observation 611dcfd6-8825-4c0b-97a5-d13de040fd84 · outbound

This paper cites Zhang, Han Bao, Hanwei Xu, Haocheng Wang, Haowei Zhang, Honghui Ding, Huajian Xin, Huazuo Gao, Hui Li, Hui Qu, J.

HPO: Hysteretic Policy Optimization for Stable and Efficient Training under Sparse-Reward Regime Zhang, Han Bao, Hanwei Xu, Haocheng Wang, Haowei Zhang, Honghui Ding, Huajian Xin, Huazuo Gao, Hui Li, Hui Qu, J

Reference 3

Resolution
unresolved
no resolver link, observed 2026-06-29T09:01:03.618498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T09:01:03.618498Z digest=sha256:7c9f25177cab671d450147f9aacd11c1db8281567cbe0cce14ea600d810f3f04

Observation 7d5d61ed-08a6-4266-9fb1-27a59694138a · outbound

This paper cites DenseGRPO: From sparse to dense reward for flow matching model alignment.

HPO: Hysteretic Policy Optimization for Stable and Efficient Training under Sparse-Reward Regime DenseGRPO: From sparse to dense reward for flow matching model alignment

Reference 4

Resolution
unresolved
no resolver link, observed 2026-06-29T09:01:03.618498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T09:01:03.618498Z digest=sha256:be4120f3da8123cf39288a43e9d0a6ceebff6cc9dfbe8a4b6817f80ec512c1a5

Observation 39fc10bb-581c-4451-91b6-d2f5bce7668c · outbound

This paper cites RAFT: Reward ranked finetuning for generative foundation model alignment.Transactions on Machine Learning Research, 2023.

HPO: Hysteretic Policy Optimization for Stable and Efficient Training under Sparse-Reward Regime RAFT: Reward ranked finetuning for generative foundation model alignment.Transactions on Machine Learning Research, 2023

Reference 5

Resolution
unresolved
no resolver link, observed 2026-06-29T09:01:03.618498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T09:01:03.618498Z digest=sha256:5dc37b6d272187b95cb1666de9cc4e85ce23fab51f6180482b8c9db9b83ddeea

Observation bab408fa-a595-4120-b294-c0ef882a71b7 · outbound

This paper cites Re- wardmap: Tackling sparse rewards in fine-grained visual reasoning via multi-stage reinforce- ment learning.

HPO: Hysteretic Policy Optimization for Stable and Efficient Training under Sparse-Reward Regime Re- wardmap: Tackling sparse rewards in fine-grained visual reasoning via multi-stage reinforce- ment learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-06-29T09:01:03.618498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T09:01:03.618498Z digest=sha256:96b576a3b3bd26c7ce5631aa355da3fd336b45662b49b97c10e5e5e4eb34dd6f

Observation aacf218f-d7d6-47a5-929d-39514e07cbd5 · outbound

This paper cites an unresolved cited work.

HPO: Hysteretic Policy Optimization for Stable and Efficient Training under Sparse-Reward Regime Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-06-29T09:01:03.618498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T09:01:03.618498Z digest=sha256:4c8a9c4b0e1fb01c0d19c1d433d0c23ef2f2667121e39042f04ece5bbb798dd4

Observation 09a5a2da-c1e2-455c-84b2-a5f5cd768f49 · outbound

This paper cites Soft adaptive policy optimization, 2025.

HPO: Hysteretic Policy Optimization for Stable and Efficient Training under Sparse-Reward Regime Soft adaptive policy optimization, 2025

Reference 8

Resolution
unresolved
no resolver link, observed 2026-06-29T09:01:03.618498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T09:01:03.618498Z digest=sha256:9ffa1e09c7f6fbb12a652497d67e3dd4a00bcef7d436703f12c5bd1e8ccd0a5d

Observation bdf835ff-720b-4f43-ac3d-6abf490b13b4 · outbound

This paper cites Rewarding the unlikely: Lifting grpo beyond distribution sharpening, 2025.

HPO: Hysteretic Policy Optimization for Stable and Efficient Training under Sparse-Reward Regime Rewarding the unlikely: Lifting grpo beyond distribution sharpening, 2025

Reference 9

Resolution
unresolved
no resolver link, observed 2026-06-29T09:01:03.618498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T09:01:03.618498Z digest=sha256:944f9940bc53209dd8af05c9cdfbc1123a589e9f1a4678808975a77c203ec570

Observation 1f0f7608-ec8c-4634-959c-e421be09efc5 · outbound

This paper cites Reinforcement Learning via Self-Distillation.

HPO: Hysteretic Policy Optimization for Stable and Efficient Training under Sparse-Reward Regime Reinforcement Learning via Self-Distillation

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-06-29T09:03:15.687060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T09:01:03.618498Z digest=sha256:133509f2fa62d98d15396127b0c47ab9cfbce431db2d4abf7b2255c51f215847

Observation a82861da-32fc-4f52-9bb3-d870baa9158a · outbound

This paper cites Understanding r1-zero-like training: A critical perspective.

HPO: Hysteretic Policy Optimization for Stable and Efficient Training under Sparse-Reward Regime Understanding r1-zero-like training: A critical perspective

Reference 11

Resolution
unresolved
no resolver link, observed 2026-06-29T09:01:03.618498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T09:01:03.618498Z digest=sha256:76d58651ca417f93fde558d1edf229b4dce4c9d9299d385e607a4e1a69ad5a52

Observation 90001c7c-6a19-4a2c-9d22-7af35690bd14 · outbound

This paper cites Hysteretic Q-Learning: An Algorithm for Decentralized Reinforcement Learning in Cooperative Multi-agent Teams.

HPO: Hysteretic Policy Optimization for Stable and Efficient Training under Sparse-Reward Regime Hysteretic Q-Learning: An Algorithm for Decentralized Reinforcement Learning in Cooperative Multi-agent Teams

Reference 12

Resolution
unresolved
no resolver link, observed 2026-06-29T09:01:03.618498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T09:01:03.618498Z digest=sha256:2732159ac7251868314491cd25903568aac9e3facfca759fb178f686d2543239

Observation ee0fb757-c0d5-4099-b7ea-692ddbb5cc69 · outbound

This paper cites an unresolved cited work.

HPO: Hysteretic Policy Optimization for Stable and Efficient Training under Sparse-Reward Regime Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-06-29T09:01:03.618498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T09:01:03.618498Z digest=sha256:23fb4923145650e4e89bf3267968e2db795fdc32357c98fa41212c71d06dfd36

Observation f9c2ff86-5807-4dd2-867f-535ac000ceef · outbound

This paper cites Tinyzero.

HPO: Hysteretic Policy Optimization for Stable and Efficient Training under Sparse-Reward Regime Tinyzero

Reference 14

Resolution
unresolved
no resolver link, observed 2026-06-29T09:01:03.618498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T09:01:03.618498Z digest=sha256:d203042072eb3b84b50168f772e3cd0a02abbd0f859e43d69b658b669c73096a

Observation 79636ff0-2e79-4f0f-a5f1-4bf6fb3644df · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

HPO: Hysteretic Policy Optimization for Stable and Efficient Training under Sparse-Reward Regime Direct preference optimization: Your language model is secretly a reward model

Reference 15

Resolution
unresolved
no resolver link, observed 2026-06-29T09:01:03.618498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T09:01:03.618498Z digest=sha256:2998a190e32005c96cc2fc773c34bddf319da14ffa45c31f8b653b9868014bd7

Observation 17f6492f-5404-4f03-9f86-2988ac74c19c · outbound

This paper cites Reasoning Language Models for Root Cause Analysis in 5G Wireless Networks.

HPO: Hysteretic Policy Optimization for Stable and Efficient Training under Sparse-Reward Regime Reasoning Language Models for Root Cause Analysis in 5G Wireless Networks

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-06-29T09:03:15.687531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T09:01:03.618498Z digest=sha256:66d6b373a997474674dd6f7b02a4735bc01c1d7d5da9f015b85b74d46cfa1863

Observation 88713596-c662-4a69-bfc0-466991631b9e · outbound

This paper cites Proximal Policy Optimization Algorithms.

HPO: Hysteretic Policy Optimization for Stable and Efficient Training under Sparse-Reward Regime Proximal Policy Optimization Algorithms

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-06-29T09:03:15.690088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T09:01:03.618498Z digest=sha256:0c2d94ca2de2ffb8d86c81665a7d52408e3276b01cdd2ba5c71a42eb64ab7116

Observation b576e0ec-3515-4803-9bf4-6852d2adfd61 · outbound

This paper cites an unresolved cited work.

HPO: Hysteretic Policy Optimization for Stable and Efficient Training under Sparse-Reward Regime Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-06-29T09:01:03.618498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T09:01:03.618498Z digest=sha256:bde5973f98f4222612a2df05cdcb0a23815c205310a80e0663c2f48cdeda48ed

Observation 4839da91-6396-4c61-a9ed-26222be91207 · outbound

This paper cites Hybridflow: A flexible and efficient rlhf framework.

HPO: Hysteretic Policy Optimization for Stable and Efficient Training under Sparse-Reward Regime Hybridflow: A flexible and efficient rlhf framework

Reference 19

Resolution
unresolved
no resolver link, observed 2026-06-29T09:01:03.618498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T09:01:03.618498Z digest=sha256:afeaae9b8f86a87649e256e5a18c680ca8a1d63454e3c020e5fda141f05c7ba1

Observation f3b0af17-308b-461f-88e6-68848bae52c0 · outbound

This paper cites A minimalist approach to llm reasoning: from rejection sampling to reinforce, 2025.

HPO: Hysteretic Policy Optimization for Stable and Efficient Training under Sparse-Reward Regime A minimalist approach to llm reasoning: from rejection sampling to reinforce, 2025

Reference 20

Resolution
unresolved
no resolver link, observed 2026-06-29T09:01:03.618498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T09:01:03.618498Z digest=sha256:2b94a4489eae068a60848346067ead600ff2b8e6826563a9692a4c09386528e3

Observation 3ad9e361-91c0-4369-aa80-4692c14bd898 · outbound

This paper cites Qwen3 technical report, 2025.

HPO: Hysteretic Policy Optimization for Stable and Efficient Training under Sparse-Reward Regime Qwen3 technical report, 2025

Reference 21

Resolution
unresolved
no resolver link, observed 2026-06-29T09:01:03.618498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T09:01:03.618498Z digest=sha256:3543daf72071d5b33f2f44f3efce4682c35b0bafda134b10b2a1d282e1011488

Observation 29e2724e-02d8-4372-9a41-891c4beb1d01 · outbound

This paper cites Dapo: An open-source llm reinforcement learning system at scale, 2025.

HPO: Hysteretic Policy Optimization for Stable and Efficient Training under Sparse-Reward Regime Dapo: An open-source llm reinforcement learning system at scale, 2025

Reference 22

Resolution
unresolved
no resolver link, observed 2026-06-29T09:01:03.618498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T09:01:03.618498Z digest=sha256:80a2155eb1f88c40c73c9c680d7fce99d1c8fb6a149984731d62d75b22503124

Observation ff42b07c-5c05-402d-a176-6895a276780d · outbound

This paper cites Group sequence policy optimization, 2025.

HPO: Hysteretic Policy Optimization for Stable and Efficient Training under Sparse-Reward Regime Group sequence policy optimization, 2025

Reference 23

Resolution
unresolved
no resolver link, observed 2026-06-29T09:01:03.618498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T09:01:03.618498Z digest=sha256:7828f83b8d26e5974a39df67e95663e9a5ec5e55bcfb50c03312cfa7ae566256

Pith citing papers

No inbound Pith citation observations are available.