Pith. sign in

Paper Citation Record · LEDGER

Retry Policy Gradients in Continuous Action Spaces

As of 18 August 2026, this Paper Citation Record lists 12 of 12 outbound references and 1 inbound Pith citation observation for arXiv:2606.05888.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.05888 v1

Coverage vector

measured 12 of 12 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-28T01:44:37.492572Z

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T00:41:12.545963Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

12 of 12 outbound references displayed

  • verified exact3
  • verified fuzzy0
  • unresolved8
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 19363c93-9a7d-400c-9e42-7115da93686e · outbound

This paper cites Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models.

Retry Policy Gradients in Continuous Action Spaces Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:56:57.024097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-28T01:44:37.492572Z digest=sha256:ec6c8fc23fbc57af92cc9b8bee89306482bf34027aed5cbe61b0432b898cb31d

Observation 4ce9cad1-9684-465b-a774-bee3b9e4f106 · outbound

This paper cites Soft Actor-Critic for Discrete Action Settings.

Retry Policy Gradients in Continuous Action Spaces Soft Actor-Critic for Discrete Action Settings

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T12:56:57.023856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-28T01:44:37.492572Z digest=sha256:a7fa42fd92161b6a551d61600d2bd4e547b16a3a18d739bd9d2bcb611ec1a4cf

Observation 79e9363a-b13f-405e-8daa-6d80fd84adbf · outbound

This paper cites Formal Theory of Creativity, Fun, and Intrinsic Motivation (1990–2010).IEEE Transactions on Autonomous Mental Development, 2(3):230–247,.

Retry Policy Gradients in Continuous Action Spaces Formal Theory of Creativity, Fun, and Intrinsic Motivation (1990–2010).IEEE Transactions on Autonomous Mental Development, 2(3):230–247,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-06-28T01:44:37.492572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T01:44:37.492572Z digest=sha256:e0482b7cd6b6bdadcb120e503092b74f1a3d5a2550b8496f58a04ecfd93d1356

Observation 84630d76-08e1-4318-a6c9-567e436357e1 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Retry Policy Gradients in Continuous Action Spaces Proximal Policy Optimization Algorithms

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-07-02T12:56:57.021274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-28T01:44:37.492572Z digest=sha256:78ec4eb9e57aec33ccf970f055d7f5e8878e1c516284295b2ab2eaa4d95ff87a

Observation c57fcdf4-207c-4831-89dd-8ba69012d046 · outbound

This paper cites Finite-Time Regret Analysis of Retry-Aware Bandits.

Retry Policy Gradients in Continuous Action Spaces Finite-Time Regret Analysis of Retry-Aware Bandits

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-07-02T12:56:57.026584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-28T01:44:37.492572Z digest=sha256:559b628b3227f3216a317be0b2cc2b776d6eb8aabfd576c30584250e60ee69f3

Observation 51fabc0b-fc4e-4553-9a4b-4ec75f2d21a6 · outbound

This paper cites We include the code in the supplementary material.

Retry Policy Gradients in Continuous Action Spaces We include the code in the supplementary material

Reference 6

Resolution
unresolved
no resolver link, observed 2026-06-28T01:44:37.492572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T01:44:37.492572Z digest=sha256:f719556cf5191dff376a86f8c616b8d602e0ad016a343a2af3888bb7231ce078

Observation e32aa80c-fdb8-4386-890c-08b762c4d3e9 · outbound

This paper cites an unresolved cited work.

Retry Policy Gradients in Continuous Action Spaces Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-06-28T01:44:37.492572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T01:44:37.492572Z digest=sha256:1a85b88fa9074b6bfbb15e726226835048862736bb9f1dc9661b9c4fac651aa4

Observation 4c4bf057-044b-443a-8e96-d388c9dafaa1 · outbound

This paper cites an unresolved cited work.

Retry Policy Gradients in Continuous Action Spaces Unresolved cited work

Reference 8

Resolution
unresolved
no resolver link, observed 2026-06-28T01:44:37.492572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T01:44:37.492572Z digest=sha256:7ce7254b08ee54722dc039379a1c7b5a03fc65e7d593f7c2a51462fe28fee1fb

Observation 3c66705e-b816-4883-8c1e-141309130430 · outbound

This paper cites an unresolved cited work.

Retry Policy Gradients in Continuous Action Spaces Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-06-28T01:44:37.492572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T01:44:37.492572Z digest=sha256:0eec8fc56e2dc209998956fe74e144a882c948083e11d07ee422372ac795eaa3

Observation 8f4d7fd7-1c3a-4697-9a36-cfffca2b93b1 · outbound

This paper cites an unresolved cited work.

Retry Policy Gradients in Continuous Action Spaces Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-06-28T01:44:37.492572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T01:44:37.492572Z digest=sha256:4d5028fca1a5f8c2ed1ae463e11a3bc04a23d708d90c0d65fd476363f78e8f6a

Observation 1340e0ee-731f-4367-9d16-932ee51e3069 · outbound

This paper cites an unresolved cited work.

Retry Policy Gradients in Continuous Action Spaces Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-06-28T01:44:37.492572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T01:44:37.492572Z digest=sha256:3c21ecdb8ef9cc8c2a9114b3e666e2797f74b09a2c449181aea11b09c102c298

Observation a25c0331-12a3-4eb4-8d5b-8c315a10d5e5 · outbound

This paper cites an unresolved cited work.

Retry Policy Gradients in Continuous Action Spaces Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-06-28T01:44:37.492572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T01:44:37.492572Z digest=sha256:6f8c1dd41af02d75642b75b38d9e8abf3c2f51a83661c83781d5ba64e1f26344

Pith citing papers

Observation 98fbfd37-543b-4703-916f-ca483b3488e8 · inbound

Multi-Axis Max@K Reinforcement Learning for Representative Diversity in Text-to-Image Generation cites this paper.

Multi-Axis Max@K Reinforcement Learning for Representative Diversity in Text-to-Image Generation Retry Policy Gradients in Continuous Action Spaces

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-02T00:41:12.545963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:41:12.545963Z digest=sha256:1eab5dd3b1f4e4fc5dcfe37386027e1fab453c49b1e2f15f2df71649e2dcc2be