Pith. sign in

Paper Citation Record · LEDGER

A Method for Evaluating Hyperparameter Sensitivity in Reinforcement Learning

As of 21 August 2026, this Paper Citation Record lists 23 of 23 outbound references and 3 inbound Pith citation observations for arXiv:2412.07165.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.07165 v2

Coverage vector

measured 23 of 23 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T19:07:39.750969Z

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:24:57.402374Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

23 of 23 outbound references displayed

  • verified exact1
  • verified fuzzy18
  • unresolved4
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

2
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 5fe30a1c-3e83-42db-9d9d-5f334ddbb3b2 · outbound

This paper cites What Matters in On - Policy Reinforcement Learning ? A Large - Scale Empirical Study.

A Method for Evaluating Hyperparameter Sensitivity in Reinforcement Learning What Matters in On - Policy Reinforcement Learning ? A Large - Scale Empirical Study

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:07:40.340418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-11T19:07:39.613300Z digest=sha256:9d142daec2a9d6958d69e3a1f7d582212325d2245db650c7cebb238b6e9fef71

Observation 7eaf0051-ba98-4ef0-9559-9eddf46d8cfd · outbound

This paper cites Hyperparameters in Contextual RL are Highly Situational.

A Method for Evaluating Hyperparameter Sensitivity in Reinforcement Learning Hyperparameters in Contextual RL are Highly Situational

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-08-11T19:07:39.909523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-11T19:07:39.619016Z digest=sha256:020a3a9f1f261fd7ba5681351ec16cfbae147c0885df9759cdc42e4f2b2a4f17

Observation 5e66c21d-4a0a-4e45-b3da-601ab2eedd79 · outbound

This paper cites Hyperparameters in Reinforcement Learning and How To Tune Them.

A Method for Evaluating Hyperparameter Sensitivity in Reinforcement Learning Hyperparameters in Reinforcement Learning and How To Tune Them

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:07:40.321521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-11T19:07:39.624658Z digest=sha256:4d94c3f3ba463ba35b83a334f1c51a19a501d08c77bd5ad4af3f0c223e600de8

Observation 93eacf2b-6b48-43ed-b064-78ce63a6ec3a · outbound

This paper cites o rg KH Franke, Gregor K \.

A Method for Evaluating Hyperparameter Sensitivity in Reinforcement Learning o rg KH Franke, Gregor K \

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:07:40.295995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-11T19:07:39.630778Z digest=sha256:bf528a7baa226caed9efb407ffbabea24024ad7ad2098b2caa3d194ebc3b44d2

Observation cd48d35c-bf21-473e-aaf3-b009d8f9cc95 · outbound

This paper cites Daniel Freeman, Erik Frey, Anton Raichuk, Sertan Girgin, Igor Mordatch, and Olivier Bachem.

A Method for Evaluating Hyperparameter Sensitivity in Reinforcement Learning Daniel Freeman, Erik Frey, Anton Raichuk, Sertan Girgin, Igor Mordatch, and Olivier Bachem

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:07:40.273567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-11T19:07:39.640442Z digest=sha256:05c84aa0fd3a8343c62d934c45f4e056e097c46ee0f819bea543b9a7202d6ce1

Observation 045cc84f-7936-4f53-891c-eb8fa1b401d1 · outbound

This paper cites Mastering Diverse Domains through World Models.

A Method for Evaluating Hyperparameter Sensitivity in Reinforcement Learning Mastering Diverse Domains through World Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T19:07:39.646721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:07:39.646721Z digest=sha256:4aa28d67e11f298f679cbddc9e71979bb3749683d7e96f653d4b215fdf75f42f

Observation 92c051d2-df4c-4061-acbc-1ab33e92625b · outbound

This paper cites Rainbow: Combining Improvements in Deep Reinforcement Learning.

A Method for Evaluating Hyperparameter Sensitivity in Reinforcement Learning Rainbow: Combining Improvements in Deep Reinforcement Learning

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:07:40.232213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-11T19:07:39.653259Z digest=sha256:378e049bf93f8df67789389755d6db374537b8093ec4e582e63516d67ae12995

Observation 93424955-b8ae-47e0-b65f-1b634411e059 · outbound

This paper cites The 37 Implementation Details of Proximal Policy Optimization.

A Method for Evaluating Hyperparameter Sensitivity in Reinforcement Learning The 37 Implementation Details of Proximal Policy Optimization

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:07:40.208497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-11T19:07:39.659061Z digest=sha256:9481feb53e05f5007869da245aa381890707598d309c477d61b1b20c37fe97eb

Observation 5e81d34b-3333-489d-9e18-58d988198e09 · outbound

This paper cites Jordan, Yash Chandak, Daniel Cohen, Mengxue Zhang, and Philip S.

A Method for Evaluating Hyperparameter Sensitivity in Reinforcement Learning Jordan, Yash Chandak, Daniel Cohen, Mengxue Zhang, and Philip S

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:07:40.184249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-11T19:07:39.665748Z digest=sha256:55b44dcc5e23c3f8a774f62cc4e6b98dcb8e21b00f8c9169ea0a688eb0683967

Observation 1ff7292e-6741-444a-a8a2-a8a7b866e0ab · outbound

This paper cites Kingma and Jimmy Ba.

A Method for Evaluating Hyperparameter Sensitivity in Reinforcement Learning Kingma and Jimmy Ba

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:07:40.159913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-11T19:07:39.671587Z digest=sha256:3a61d6f9354268cad362c92156dd0eb59627a75dbc3f36756a2698e59cb30d14

Observation 67e71ef6-5f64-4201-92a9-b400e97f5acc · outbound

This paper cites Lagoudakis and Ronald Parr.

A Method for Evaluating Hyperparameter Sensitivity in Reinforcement Learning Lagoudakis and Ronald Parr

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:07:40.139322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-11T19:07:39.678055Z digest=sha256:07d76bb98e0ea0279473703936c9b7b47f741189b9b70afa08167370b0d56ec9

Observation b630384c-4699-4e03-a77d-cd1d2543a8f0 · outbound

This paper cites Discovered Policy Optimisation.

A Method for Evaluating Hyperparameter Sensitivity in Reinforcement Learning Discovered Policy Optimisation

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:07:40.122246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-11T19:07:39.684002Z digest=sha256:ba485c37b5cc52d142e2e7606637e3d5eaad6ed096fa46102caeb21374d0c40c

Observation de7177af-accf-433c-ba2c-f130844956bc · outbound

This paper cites Rusu, Joel Veness, Marc G.

A Method for Evaluating Hyperparameter Sensitivity in Reinforcement Learning Rusu, Joel Veness, Marc G

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:07:40.105243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-11T19:07:39.691014Z digest=sha256:4abbcba1b7df2618869029898e9967956baad0a6b4d27d5315a75efbfbdb5f73

Observation e1851ac6-ef43-4eb4-9fec-131d91b15020 · outbound

This paper cites Empirical Design in Reinforcement Learning.

A Method for Evaluating Hyperparameter Sensitivity in Reinforcement Learning Empirical Design in Reinforcement Learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T19:07:39.698889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:07:39.698889Z digest=sha256:f44fb8dbd713bac620c14462a2c0b9fdffc713be7aaf0b5ebdf80acb7c21524e

Observation 80449247-d683-48e7-8d9b-a2bdda6df319 · outbound

This paper cites The Cross - Environment Hyperparameter Setting Benchmark for Reinforcement Learning.

A Method for Evaluating Hyperparameter Sensitivity in Reinforcement Learning The Cross - Environment Hyperparameter Setting Benchmark for Reinforcement Learning

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:07:40.086941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-11T19:07:39.705482Z digest=sha256:ace44f07d057b97c689ee090b0227da903112e0e4905e3526d695bbfb58c5eb3

Observation 25b6592a-5a47-44d7-ad64-afc19bcc79f9 · outbound

This paper cites Neural fitted q iteration - first experiences with a data efficient neural reinforcement learning method.

A Method for Evaluating Hyperparameter Sensitivity in Reinforcement Learning Neural fitted q iteration - first experiences with a data efficient neural reinforcement learning method

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:07:40.063461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-11T19:07:39.710916Z digest=sha256:fbe4b84123b6a4b419f303cd5dbb7c75799638744d2b86201a65fed34808ca88

Observation b19e3072-29de-4c58-aec9-063952b2d161 · outbound

This paper cites Proximal Policy Optimization Algorithms.

A Method for Evaluating Hyperparameter Sensitivity in Reinforcement Learning Proximal Policy Optimization Algorithms

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T19:07:39.716115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:07:39.716115Z digest=sha256:c83ebea61f1f6d21d24ddc09bb0118e845557aa0f4959f8a503976c3c34892c8

Observation cf16789d-9655-40e6-9b5a-e4b8169bce63 · outbound

This paper cites A reinforcement learning method for maximizing undiscounted rewards.

A Method for Evaluating Hyperparameter Sensitivity in Reinforcement Learning A reinforcement learning method for maximizing undiscounted rewards

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:07:40.040645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-11T19:07:39.722669Z digest=sha256:7e4448cfa023e39da90405c2111cd87b65f25dff333e648c65c858bdf076e951

Observation 695a8875-defd-46de-9593-b905024d9cca · outbound

This paper cites Dickerson, and Joseph Suarez.

A Method for Evaluating Hyperparameter Sensitivity in Reinforcement Learning Dickerson, and Joseph Suarez

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:07:40.022256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-11T19:07:39.727969Z digest=sha256:8cd6ff1578b615584d11590eb13f3793575cb414a8b9e441e7454011057f096b

Observation b7836f55-462b-4123-aa07-e744122403ae · outbound

This paper cites Reinforcement Learning : An Introduction.

A Method for Evaluating Hyperparameter Sensitivity in Reinforcement Learning Reinforcement Learning : An Introduction

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:07:40.002884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-11T19:07:39.733238Z digest=sha256:ffcbbfc6ee275fbe7e9020871886172e19c4ad225090b245269da19d5f3e13fa

Observation 30e5cbdd-c6ee-4b28-8726-5778092e61b1 · outbound

This paper cites Learning Values Across Many Orders of Magnitude.

A Method for Evaluating Hyperparameter Sensitivity in Reinforcement Learning Learning Values Across Many Orders of Magnitude

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:07:39.983700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-11T19:07:39.738929Z digest=sha256:1613b5beee9765969f29af8f6f1e3de07bb3d70562b373e02e3724170a132402

Observation 71501654-42b8-408f-93d8-80ab33280237 · outbound

This paper cites Learning from Delayed Rewards.

A Method for Evaluating Hyperparameter Sensitivity in Reinforcement Learning Learning from Delayed Rewards

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:07:39.957633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-11T19:07:39.744232Z digest=sha256:4e1e4033e0c6e671ae89b291952242d2c571f51a6e54c9ddc43582b874720869

Observation 436d46a4-1008-4f7f-97b4-cdac143ab0f5 · outbound

This paper cites write newline.

A Method for Evaluating Hyperparameter Sensitivity in Reinforcement Learning write newline

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T19:07:39.750969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:07:39.750969Z digest=sha256:1e131386607c269f20949240168980136614bf3ca47b8175f952a99d50ece7a6

Pith citing papers

Observation a86bb074-2a9b-47c2-90ba-95720d7c2bec · inbound

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One cites this paper.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One A Method for Evaluating Hyperparameter Sensitivity in Reinforcement Learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T15:24:57.402374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:24:57.402374Z digest=sha256:712d1b90e092478476320703635f52dc22fe165f46a5bae01c1bfe9fb64fc969

Observation f48faf5f-8f10-4095-b923-e3fe5bf3f8f4 · inbound

How Should We Meta-Learn Reinforcement Learning Algorithms? cites this paper.

How Should We Meta-Learn Reinforcement Learning Algorithms? A Method for Evaluating Hyperparameter Sensitivity in Reinforcement Learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:44.755232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:44.755232Z digest=sha256:a1a58ac511d4bead80d2ffba8104c8da8c2e436fc670f3d902b494ffc746f131

Observation ebba4398-d90e-42c5-9f66-de8e2bf794a0 · inbound

Feedback-Normalized Developer Memory for Reinforcement-Learning Coding Agents: A Safety-Gated MCP Architecture cites this paper.

Feedback-Normalized Developer Memory for Reinforcement-Learning Coding Agents: A Safety-Gated MCP Architecture A Method for Evaluating Hyperparameter Sensitivity in Reinforcement Learning

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-09T22:29:06.380913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-09T13:59:40.638085Z digest=sha256:d46f55831977e52f2a7767a572dcc1bc02a1bde49544a0374a7514519771338f