Pith. sign in

Paper Citation Record · LEDGER

A Survey of Exploration Methods in Reinforcement Learning

As of 6 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 10 inbound Pith citation observations for arXiv:2109.00157.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2109.00157 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 10 of 10 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T20:21:04.419668Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

5
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 62f72c5a-6235-4654-bc26-fe632c7dce51 · inbound

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models cites this paper.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models A Survey of Exploration Methods in Reinforcement Learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T20:21:04.419668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:21:04.419668Z digest=sha256:9836bb1eddc86c9b4add8babeb8a097c0e44838c8efb8a86859dda30c55bf2fd

Observation 7c8d4574-d36e-4a66-9513-d4dc796d50d7 · inbound

CDE: Curiosity-Driven Exploration for Efficient Reinforcement Learning in Large Language Models cites this paper.

CDE: Curiosity-Driven Exploration for Efficient Reinforcement Learning in Large Language Models A Survey of Exploration Methods in Reinforcement Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T18:53:02.597599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:53:02.597599Z digest=sha256:fc2b6826dfe18b593e2fa9049323ae008e5dfdc8231d8b98bff63183f3f4bd20

Observation 877220da-97a2-4565-9c7a-a2f10d61f3a2 · inbound

Smart Walkers in Discrete Space cites this paper.

Smart Walkers in Discrete Space A Survey of Exploration Methods in Reinforcement Learning

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-16T09:20:47.682510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T09:19:47.050591Z digest=sha256:9e34eb6102ac50e0203a717c2f946f56c13e2a43293efe92326b8cdf2ae57976

Observation 0ae1acbb-c158-48d9-bb24-f7539e748d95 · inbound

SpecRL: Reinforcement Learning with Test-Based Completeness Rewards for Formal Specification Synthesis cites this paper.

SpecRL: Reinforcement Learning with Test-Based Completeness Rewards for Formal Specification Synthesis A Survey of Exploration Methods in Reinforcement Learning

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T23:30:50.844784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T19:06:10.718673Z digest=sha256:c5f7f8a2c144257d8848101baa8d7b45cf14bbff25b89065b46f5608efd89f33

Observation 3413dc2e-46f0-4972-9d79-0f866fc922a9 · inbound

Contextual Multi-Task Reinforcement Learning for Autonomous Reef Monitoring cites this paper.

Contextual Multi-Task Reinforcement Learning for Autonomous Reef Monitoring A Survey of Exploration Methods in Reinforcement Learning

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-10T15:15:31.460106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T15:13:16.406160Z digest=sha256:28d129a9ae6c1485feb7e54c69ed094d37fd5719ca2a1979a655e2c543c33e64

Observation 2864721b-faa5-461b-a298-3e582951cbdf · inbound

Contextual Multi-Task Reinforcement Learning for Autonomous Reef Monitoring cites this paper.

Contextual Multi-Task Reinforcement Learning for Autonomous Reef Monitoring A Survey of Exploration Methods in Reinforcement Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-12T21:15:13.179473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T21:15:13.179473Z digest=sha256:6b7fc48fc1345d50b87b49489e237fe15d540a1a75f4110c2ed0f994f6ec4572

Observation 6fbaf575-13ae-449d-af90-d33c5b0a82b7 · inbound

Flexible Empowerment at Reasoning with Extended Best-of-N Sampling cites this paper.

Flexible Empowerment at Reasoning with Extended Best-of-N Sampling A Survey of Exploration Methods in Reinforcement Learning

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:43:49.278744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T09:39:29.875977Z digest=sha256:ace4d182dcea6a13b32fa93da7f2e4207d9830524cb18d1642266220f54585d3

Observation c62a7275-2b45-4015-8da9-98638dbab921 · inbound

Optimal Semiparametric Dynamic Pricing with Feature Diversity cites this paper.

Optimal Semiparametric Dynamic Pricing with Feature Diversity A Survey of Exploration Methods in Reinforcement Learning

Reference 55

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T17:26:06.944659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-08T17:34:49.811590Z digest=sha256:770bad345328516df37e2cb7a63a16e88e78fc4d1174f93f88423f2dbc72f6cb

Observation 768a09fc-41e6-4135-8ef3-2afd43b35674 · inbound

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning cites this paper.

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning A Survey of Exploration Methods in Reinforcement Learning

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-04T07:59:40.653383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T12:15:08.304150Z digest=sha256:981efa294f3d02e17e0c84908b61c71e6736e6d2433dd8ca197c6593b00e6793

Observation 5e0b398b-60e6-4157-aa44-89cf8ebba0e2 · inbound

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback cites this paper.

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback A Survey of Exploration Methods in Reinforcement Learning

Reference 234

Resolution
unresolved
no resolver link, observed 2026-08-03T04:39:32.089880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:39:32.089880Z digest=sha256:b52f469392fab22371f5cc93b31c6d67170faa37e32ffe641fee24d5c12177d5