Pith. sign in

Paper Citation Record · LEDGER

Exploration Behavior of Untrained Policies

As of 11 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 0 inbound Pith citation observations for arXiv:2506.22566.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.22566 v3

Coverage vector

measured 21 of 21 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:10:29.525354Z

measured 21 of 21 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

21 of 21 outbound references displayed

  • verified exact2
  • verified fuzzy13
  • unresolved6
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4f742340-d001-4ecb-ad71-3213a896ec27 · outbound

This paper cites Lipbab: Computing exact lipschitz constant of relu networks.

Exploration Behavior of Untrained Policies Lipbab: Computing exact lipschitz constant of relu networks

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:10:32.943686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T22:10:27.275619Z digest=sha256:32bf4902b3b676b48a39f21bfbbff0efa236666f45c1058e62c883d9e68568dc

Observation 5e3a4c18-c91a-436b-a530-2607ec7bbadc · outbound

This paper cites Radial basis functions.

Exploration Behavior of Untrained Policies Radial basis functions

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:10:32.810908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T22:10:27.379296Z digest=sha256:5867e1eea1e3a8416509a3d8efe6347a439736c397245f09f6d76feb627e553a

Observation c93c7a78-3e4a-4457-b0fe-9f056c9059e1 · outbound

This paper cites Exploration by Random Network Distillation.

Exploration Behavior of Untrained Policies Exploration by Random Network Distillation

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T22:10:27.481099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:10:27.481099Z digest=sha256:eff4b13c3103affc4e8b8b2702ed00213d96e7ffeaf4348d61b9b7e9d37e25a5

Observation aea7d39d-5684-45d1-a88a-db50d36e9520 · outbound

This paper cites Rainbow: Combining im- provements in deep reinforcement learning.

Exploration Behavior of Untrained Policies Rainbow: Combining im- provements in deep reinforcement learning

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:10:32.593138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T22:10:27.621983Z digest=sha256:8c06f92492b249baf074da1502be7568341a4d3d28c2c6dfb6f427b48d7a8faa

Observation 921bc39c-40bb-4719-a0bf-387044182b2c · outbound

This paper cites Neural tangent kernel: Convergence and generalization in neural networks.

Exploration Behavior of Untrained Policies Neural tangent kernel: Convergence and generalization in neural networks

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:10:32.248330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T22:10:27.797985Z digest=sha256:2c69016a1fc4305f7a350bc4546c80b2e30b8d639ddd4bb37118fe96b49d7f05

Observation 4a01a4cc-b16d-411d-83f0-78a7da33c1a8 · outbound

This paper cites Exploration in deep reinforce- ment learning: A survey.

Exploration Behavior of Untrained Policies Exploration in deep reinforce- ment learning: A survey

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:10:32.009910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T22:10:27.956796Z digest=sha256:36c37810a0d23b0b6f28687551ae2e0cdd574c0ab7e36f6c90602f5df81e671a

Observation 648b2eca-f65a-4e42-aa9b-c8e86a180c32 · outbound

This paper cites Lipschitz constant estimation of Neural Networks via sparse polynomial optimization.

Exploration Behavior of Untrained Policies Lipschitz constant estimation of Neural Networks via sparse polynomial optimization

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:10:29.933889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T22:10:28.129862Z digest=sha256:5b7e949b6bb07c1cac77776badaa0d27912a5d6707dc2a8c1619e01fbd20d2c8

Observation 1a691914-3af9-435c-8dab-fd76864640ee · outbound

This paper cites Flipping coins to estimate pseudocounts for exploration in reinforcement learning.

Exploration Behavior of Untrained Policies Flipping coins to estimate pseudocounts for exploration in reinforcement learning

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:10:31.740520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T22:10:28.283250Z digest=sha256:e63612bf5f5dd1c0914f5676216f5383fea8a5b6e0c61481f10f56f07660f2dd

Observation 24c8ee25-e0a5-4cff-86a8-0d3e5110a3b3 · outbound

This paper cites Periodic activation functions induce stationar- ity.

Exploration Behavior of Untrained Policies Periodic activation functions induce stationar- ity

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:10:31.470042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T22:10:28.463799Z digest=sha256:a6d10f695956a208b7539b9290635fca74b18b5fd8ac6f2bb8ff9eefeb394e24

Observation a0542598-25d5-440f-9006-12808fac717e · outbound

This paper cites Human-level control through deep reinforcement learning.

Exploration Behavior of Untrained Policies Human-level control through deep reinforcement learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T22:10:28.624563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:10:28.624563Z digest=sha256:214600cc8fda0cac2c595bf0671357ec7f6061295237f4a1d61fedfb6d09ab48

Observation e3830a1c-008c-4c0e-ae77-39efec3bb782 · outbound

This paper cites Bayesian learning for neural networks , volume 118.

Exploration Behavior of Untrained Policies Bayesian learning for neural networks , volume 118

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T22:10:28.721081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:10:28.721081Z digest=sha256:60b5fe9193d70d984e5e5a9075a29602e698d509bedb5852b3d0ff81c9653b8f

Observation ff014368-7c2d-4e38-a154-f1cf0c515db3 · outbound

This paper cites The primacy bias in deep reinforcement learning.

Exploration Behavior of Untrained Policies The primacy bias in deep reinforcement learning

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:10:31.187261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T22:10:28.801986Z digest=sha256:20d76690f4f0179c014e5c3acf4eff48ed1f577fc8c6e3ca01a8f152a004c015

Observation 744ac8ad-46de-4499-bc53-31ac34d2eb37 · outbound

This paper cites Deep reinforcement learning with plasticity injection.

Exploration Behavior of Untrained Policies Deep reinforcement learning with plasticity injection

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:10:30.923444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T22:10:28.888210Z digest=sha256:668215f3f7ed5a8495e79034c5cceca5fdde75681808adfa4790da9cf2837c85

Observation e483f8f7-c5b7-41e9-b252-789164ae8490 · outbound

This paper cites Deep exploration via bootstrapped dqn.

Exploration Behavior of Untrained Policies Deep exploration via bootstrapped dqn

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:10:30.670591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T22:10:28.961018Z digest=sha256:8dd9440a955581ea0f9e1a0af771f16c11eb0b4fade5f32c02484d37b4284308

Observation 3c452bac-3594-4190-bf34-11c4375a7721 · outbound

This paper cites Trust region policy optimization.

Exploration Behavior of Untrained Policies Trust region policy optimization

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:10:30.438864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T22:10:29.041931Z digest=sha256:d3faab84a5a4f69bbfaf3f869a4c278e0be50659d173b763495345343446fec2

Observation 1cf297be-0c97-490d-ab4d-efcd470d9e22 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Exploration Behavior of Untrained Policies Proximal Policy Optimization Algorithms

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T22:10:29.122205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:10:29.122205Z digest=sha256:0232eea7e4351d2c59d946078283e2f53853d3d8d60b8b1e2c757e20863cea87

Observation 419d81e5-aa47-4426-b564-2a250826ae5e · outbound

This paper cites On Bonus-Based Exploration Methods in the Arcade Learning Environment.

Exploration Behavior of Untrained Policies On Bonus-Based Exploration Methods in the Arcade Learning Environment

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:10:29.743035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T22:10:29.210486Z digest=sha256:dc65321c007e4ae9aa3d6da9985198ea0bf4c4c7b1145f54aee414caa74798c1

Observation 22c9fc1d-d33a-44f0-a281-3d3b3fef7bf1 · outbound

This paper cites Lipschitz regularity of deep neural networks: analysis and efficient estimation.

Exploration Behavior of Untrained Policies Lipschitz regularity of deep neural networks: analysis and efficient estimation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T22:10:29.278455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:10:29.278455Z digest=sha256:d90bbaaa4f34230ef1b22ce9aecf8964362c6ce36d68e98adcc7d9280c91157c

Observation aec46387-79f1-45fe-b2f3-ce3a690111e0 · outbound

This paper cites Q-learning.

Exploration Behavior of Untrained Policies Q-learning

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:10:30.220688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T22:10:29.345229Z digest=sha256:4508d8bfe749fcefc3097ca188b46131567ac928717c08b2dc2e6f6b5b52e04d

Observation c5b8af5f-2bc5-42f1-8abf-0e501571d7c2 · outbound

This paper cites Simple statistical gradient-following algorithms for connectionist rein- forcement learning.

Exploration Behavior of Untrained Policies Simple statistical gradient-following algorithms for connectionist rein- forcement learning

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:10:30.062294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T22:10:29.435222Z digest=sha256:852c96a6afb9b90db6f757ef890b965ca34074cb57173ef46b980d22764c1352

Observation fbe3f23d-dab5-49d1-b34c-e2efb76b71c6 · outbound

This paper cites Neural Architecture Search with Reinforcement Learning.

Exploration Behavior of Untrained Policies Neural Architecture Search with Reinforcement Learning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T22:10:29.525354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:10:29.525354Z digest=sha256:cee65db95af1911353b5a3eb9832fb38791bf50a34106edb8cd5f6688a7659f2

Pith citing papers

No inbound Pith citation observations are available.