Pith. sign in

Paper Citation Record · LEDGER

A Survey of Exploration Methods in Reinforcement Learning

As of 23 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 18 inbound Pith citation observations for arXiv:2109.00157.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2109.00157 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 18 of 18 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T21:59:58.574784Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

5
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 028e9a40-e4cf-43b0-87e4-ed52803edc9a · inbound

ELEMENT: Episodic and Lifelong Exploration via Maximum Entropy cites this paper.

ELEMENT: Episodic and Lifelong Exploration via Maximum Entropy A Survey of Exploration Methods in Reinforcement Learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T22:12:10.452960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:12:10.452960Z digest=sha256:d67a5e0ab499b70fd037c03e4285fe0071d53968c42e32d8cd7385b0dee84e10

Observation 595fc39b-f931-4b4d-abbb-d9a20d1d694f · inbound

Effective Reward Specification in Deep Reinforcement Learning cites this paper.

Effective Reward Specification in Deep Reinforcement Learning A Survey of Exploration Methods in Reinforcement Learning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T19:09:53.307309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:09:53.307309Z digest=sha256:b039de0b2f822924da34b1af76e68e532bb8cb167ba55405b480d8b1706eb25c

Observation 52181ded-a820-4ff4-a78a-693a6396b17b · inbound

Active Inference and Human--Computer Interaction cites this paper.

Active Inference and Human--Computer Interaction A Survey of Exploration Methods in Reinforcement Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T11:59:02.571787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:59:02.571787Z digest=sha256:bbe25d7c83e5f4056d7f0b403acbbb9fa5570a5613d4ca1dccc3b1ca6cf0c9f8

Observation 4c4361b5-ff42-4174-8b5a-f2b15909f195 · inbound

The impact of intrinsic rewards on exploration in Reinforcement Learning cites this paper.

The impact of intrinsic rewards on exploration in Reinforcement Learning A Survey of Exploration Methods in Reinforcement Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T18:11:48.173645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:11:48.173645Z digest=sha256:54b20a1166512557de31b05aaf6eae81b5cf76c9dca1739df0cfd77fda3c1f20

Observation 43a209fe-3b9d-41c8-957b-fbb07b00cbc4 · inbound

Parameter Estimation using Reinforcement Learning Causal Curiosity: Limits and Challenges cites this paper.

Parameter Estimation using Reinforcement Learning Causal Curiosity: Limits and Challenges A Survey of Exploration Methods in Reinforcement Learning

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T21:59:58.574784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:59:58.574784Z digest=sha256:005d007b351ecc3f7eed62217b8c510cfc8936e810cb41468d43864f09223225

Observation 7e36c451-4f03-4bec-af1c-54f73c8859ec · inbound

Exploring Large Action Sets with Hyperspherical Embeddings using von Mises-Fisher Sampling cites this paper.

Exploring Large Action Sets with Hyperspherical Embeddings using von Mises-Fisher Sampling A Survey of Exploration Methods in Reinforcement Learning

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T21:21:54.730532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:21:54.730532Z digest=sha256:3d5de7f323751ed467b60e09149ee32f0dad7d6f47ab0c1ada13b0cf94a08b91

Observation 3cd80d02-d104-44df-a2c3-ee39a4cb426c · inbound

Towards Human-level Dexterity via Robot Learning cites this paper.

Towards Human-level Dexterity via Robot Learning A Survey of Exploration Methods in Reinforcement Learning

Reference 102

Resolution
unresolved
no resolver link, observed 2026-08-06T18:14:16.231424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:14:16.231424Z digest=sha256:80bbf4f79f726e959566b658f311e72877060adef33c1e5296b2d0c55a926abe

Observation 62f72c5a-6235-4654-bc26-fe632c7dce51 · inbound

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models cites this paper.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models A Survey of Exploration Methods in Reinforcement Learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T20:21:04.419668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:21:04.419668Z digest=sha256:b094fbad9fa773b052d2e71aa2547555fe5b20d0d8558827f3b29196d541ed77

Observation 7c8d4574-d36e-4a66-9513-d4dc796d50d7 · inbound

CDE: Curiosity-Driven Exploration for Efficient Reinforcement Learning in Large Language Models cites this paper.

CDE: Curiosity-Driven Exploration for Efficient Reinforcement Learning in Large Language Models A Survey of Exploration Methods in Reinforcement Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T18:53:02.597599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:53:02.597599Z digest=sha256:f6efc0991a99c8c184f6992c2e43add2780aa57d332b5b6be41a8c0c173857ce

Observation 877220da-97a2-4565-9c7a-a2f10d61f3a2 · inbound

Smart Walkers in Discrete Space cites this paper.

Smart Walkers in Discrete Space A Survey of Exploration Methods in Reinforcement Learning

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-16T09:20:47.682510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-16T09:19:47.050591Z digest=sha256:8645e8e09091803e793bf3f7589947e2c12ea97c1245eae120ed6c99afc52abf

Observation 0ae1acbb-c158-48d9-bb24-f7539e748d95 · inbound

SpecRL: Reinforcement Learning with Test-Based Completeness Rewards for Formal Specification Synthesis cites this paper.

SpecRL: Reinforcement Learning with Test-Based Completeness Rewards for Formal Specification Synthesis A Survey of Exploration Methods in Reinforcement Learning

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T23:30:50.844784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T19:06:10.718673Z digest=sha256:e21d22eba5155d77ad98679b462f2e599f0f9cdb46b2c8736970f798d4279ad6

Observation 3413dc2e-46f0-4972-9d79-0f866fc922a9 · inbound

Contextual Multi-Task Reinforcement Learning for Autonomous Reef Monitoring cites this paper.

Contextual Multi-Task Reinforcement Learning for Autonomous Reef Monitoring A Survey of Exploration Methods in Reinforcement Learning

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-10T15:15:31.460106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T15:13:16.406160Z digest=sha256:29430576e978e4cb04ea346362011eae2f1bce3d677067f8dadfe8072cf0ac0b

Observation 2864721b-faa5-461b-a298-3e582951cbdf · inbound

Contextual Multi-Task Reinforcement Learning for Autonomous Reef Monitoring cites this paper.

Contextual Multi-Task Reinforcement Learning for Autonomous Reef Monitoring A Survey of Exploration Methods in Reinforcement Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-12T21:15:13.179473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T21:15:13.179473Z digest=sha256:dd4f845eba6893cfbbb9cc9668eb592d063458cc9c6f01c4ac0b055078ba9db3

Observation 6fbaf575-13ae-449d-af90-d33c5b0a82b7 · inbound

Flexible Empowerment at Reasoning with Extended Best-of-N Sampling cites this paper.

Flexible Empowerment at Reasoning with Extended Best-of-N Sampling A Survey of Exploration Methods in Reinforcement Learning

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:43:49.278744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T09:39:29.875977Z digest=sha256:f9553b649d6b9087e0da9b66b58699409d0fe8cb513a58667d181085c4a1b22d

Observation c62a7275-2b45-4015-8da9-98638dbab921 · inbound

Optimal Semiparametric Dynamic Pricing with Feature Diversity cites this paper.

Optimal Semiparametric Dynamic Pricing with Feature Diversity A Survey of Exploration Methods in Reinforcement Learning

Reference 55

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T17:26:06.944659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-05-08T17:34:49.811590Z digest=sha256:51259bb69bd254b9172b2f2eafffb94deedebcdb02f00f803eaf371369dc9d1b

Observation 768a09fc-41e6-4135-8ef3-2afd43b35674 · inbound

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning cites this paper.

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning A Survey of Exploration Methods in Reinforcement Learning

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-04T07:59:40.653383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-26T12:15:08.304150Z digest=sha256:eb61640e445b2525477c01ffe12b48ce7b47057048d4765d00832dacd87b6422

Observation 5e0b398b-60e6-4157-aa44-89cf8ebba0e2 · inbound

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback cites this paper.

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback A Survey of Exploration Methods in Reinforcement Learning

Reference 234

Resolution
unresolved
no resolver link, observed 2026-08-03T04:39:32.089880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T04:39:32.089880Z digest=sha256:bac5222691b96dce002e995e6e9859dc894f4318a14db403dea37550bf314b10

Observation 463ddaed-3d44-4da6-9ce5-80f49d67d6bd · inbound

Parameter Exploration for RLVR via Variational Learning cites this paper.

Parameter Exploration for RLVR via Variational Learning A Survey of Exploration Methods in Reinforcement Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T10:32:12.730067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:32:12.730067Z digest=sha256:876af5b155f26a3f012e4f43245b0ac05e651a9952ebf1858bfdf7b21d44d570