Pith. sign in

Paper Citation Record · LEDGER

A Method for Evaluating Hyperparameter Sensitivity in Reinforcement Learning

As of 20 August 2026, this Paper Citation Record lists 23 of 23 outbound references and 3 inbound Pith citation observations for arXiv:2412.07165.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.07165 v2

Coverage vector

measured 23 of 23 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T19:07:39.750969Z

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:24:57.402374Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

23 of 23 outbound references displayed

  • verified exact1
  • verified fuzzy18
  • unresolved4
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

2
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 5fe30a1c-3e83-42db-9d9d-5f334ddbb3b2 · outbound

This paper cites What Matters in On - Policy Reinforcement Learning ? A Large - Scale Empirical Study.

A Method for Evaluating Hyperparameter Sensitivity in Reinforcement Learning What Matters in On - Policy Reinforcement Learning ? A Large - Scale Empirical Study

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:07:40.340418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T19:07:39.613300Z digest=sha256:f05294f6d2dcf271148ace3e3b9966bd5b7c00df9f768ec1f18034603dab4022

Observation 7eaf0051-ba98-4ef0-9559-9eddf46d8cfd · outbound

This paper cites Hyperparameters in Contextual RL are Highly Situational.

A Method for Evaluating Hyperparameter Sensitivity in Reinforcement Learning Hyperparameters in Contextual RL are Highly Situational

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-08-11T19:07:39.909523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T19:07:39.619016Z digest=sha256:cbbea08fa5c8172698b3dd8e9adf1e19ef8a05abe2dbe4ea3c06fd45adeb4b71

Observation 5e66c21d-4a0a-4e45-b3da-601ab2eedd79 · outbound

This paper cites Hyperparameters in Reinforcement Learning and How To Tune Them.

A Method for Evaluating Hyperparameter Sensitivity in Reinforcement Learning Hyperparameters in Reinforcement Learning and How To Tune Them

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:07:40.321521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T19:07:39.624658Z digest=sha256:22306b796136f4b0ea6163e7c292d07fabfcabc18081138314e0ad3e6ad3824d

Observation 93eacf2b-6b48-43ed-b064-78ce63a6ec3a · outbound

This paper cites o rg KH Franke, Gregor K \.

A Method for Evaluating Hyperparameter Sensitivity in Reinforcement Learning o rg KH Franke, Gregor K \

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:07:40.295995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T19:07:39.630778Z digest=sha256:b4e75e8fc1adc177657de1773bb4a34f43f4b290ded4f5f8272991223052a79e

Observation cd48d35c-bf21-473e-aaf3-b009d8f9cc95 · outbound

This paper cites Daniel Freeman, Erik Frey, Anton Raichuk, Sertan Girgin, Igor Mordatch, and Olivier Bachem.

A Method for Evaluating Hyperparameter Sensitivity in Reinforcement Learning Daniel Freeman, Erik Frey, Anton Raichuk, Sertan Girgin, Igor Mordatch, and Olivier Bachem

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:07:40.273567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T19:07:39.640442Z digest=sha256:542e5376966f575a8021c37027ed7859e4a8d445efd4b193df6e1d8a1ce8cacc

Observation 045cc84f-7936-4f53-891c-eb8fa1b401d1 · outbound

This paper cites Mastering Diverse Domains through World Models.

A Method for Evaluating Hyperparameter Sensitivity in Reinforcement Learning Mastering Diverse Domains through World Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T19:07:39.646721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:07:39.646721Z digest=sha256:4aa28d67e11f298f679cbddc9e71979bb3749683d7e96f653d4b215fdf75f42f

Observation 92c051d2-df4c-4061-acbc-1ab33e92625b · outbound

This paper cites Rainbow: Combining Improvements in Deep Reinforcement Learning.

A Method for Evaluating Hyperparameter Sensitivity in Reinforcement Learning Rainbow: Combining Improvements in Deep Reinforcement Learning

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:07:40.232213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T19:07:39.653259Z digest=sha256:3e7e01ff8d177bb5ba8cf31fb9741b79e79097f1e48534d67ca67d91f0c0b9fb

Observation 93424955-b8ae-47e0-b65f-1b634411e059 · outbound

This paper cites The 37 Implementation Details of Proximal Policy Optimization.

A Method for Evaluating Hyperparameter Sensitivity in Reinforcement Learning The 37 Implementation Details of Proximal Policy Optimization

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:07:40.208497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T19:07:39.659061Z digest=sha256:d2687778eaa348d9f09040aaa9c5ba1b370e1a7d149387f11c9e4f2267b9436a

Observation 5e81d34b-3333-489d-9e18-58d988198e09 · outbound

This paper cites Jordan, Yash Chandak, Daniel Cohen, Mengxue Zhang, and Philip S.

A Method for Evaluating Hyperparameter Sensitivity in Reinforcement Learning Jordan, Yash Chandak, Daniel Cohen, Mengxue Zhang, and Philip S

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:07:40.184249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T19:07:39.665748Z digest=sha256:2c2cb58a2a35d44fafa701824b8040887f7c1e17fc5a56ead69cbd0b45284210

Observation 1ff7292e-6741-444a-a8a2-a8a7b866e0ab · outbound

This paper cites Kingma and Jimmy Ba.

A Method for Evaluating Hyperparameter Sensitivity in Reinforcement Learning Kingma and Jimmy Ba

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:07:40.159913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T19:07:39.671587Z digest=sha256:37c4c3481dd444346446c618b21a5f99bda2fbdb9c9b053a2a36258e9715cf1d

Observation 67e71ef6-5f64-4201-92a9-b400e97f5acc · outbound

This paper cites Lagoudakis and Ronald Parr.

A Method for Evaluating Hyperparameter Sensitivity in Reinforcement Learning Lagoudakis and Ronald Parr

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:07:40.139322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T19:07:39.678055Z digest=sha256:97f9aeedef1076d88bc9741845e831765dba2dc279f793f14fa97bd4b66a6e7d

Observation b630384c-4699-4e03-a77d-cd1d2543a8f0 · outbound

This paper cites Discovered Policy Optimisation.

A Method for Evaluating Hyperparameter Sensitivity in Reinforcement Learning Discovered Policy Optimisation

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:07:40.122246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T19:07:39.684002Z digest=sha256:4aaac0ab381abfde430fe5192679d12b2aad442abbcf43aa4d3dc215cd7533f6

Observation de7177af-accf-433c-ba2c-f130844956bc · outbound

This paper cites Rusu, Joel Veness, Marc G.

A Method for Evaluating Hyperparameter Sensitivity in Reinforcement Learning Rusu, Joel Veness, Marc G

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:07:40.105243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T19:07:39.691014Z digest=sha256:c75d1887e01ba8f340a6d9cbc14934ca6ae1202761872e2955bf04c714f8a803

Observation e1851ac6-ef43-4eb4-9fec-131d91b15020 · outbound

This paper cites Empirical Design in Reinforcement Learning.

A Method for Evaluating Hyperparameter Sensitivity in Reinforcement Learning Empirical Design in Reinforcement Learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T19:07:39.698889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:07:39.698889Z digest=sha256:f44fb8dbd713bac620c14462a2c0b9fdffc713be7aaf0b5ebdf80acb7c21524e

Observation 80449247-d683-48e7-8d9b-a2bdda6df319 · outbound

This paper cites The Cross - Environment Hyperparameter Setting Benchmark for Reinforcement Learning.

A Method for Evaluating Hyperparameter Sensitivity in Reinforcement Learning The Cross - Environment Hyperparameter Setting Benchmark for Reinforcement Learning

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:07:40.086941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T19:07:39.705482Z digest=sha256:dbd986dfe9a9d2d9026fbb7ef36d5805989fa75053c000b53db7df6d5eb20b2d

Observation 25b6592a-5a47-44d7-ad64-afc19bcc79f9 · outbound

This paper cites Neural fitted q iteration - first experiences with a data efficient neural reinforcement learning method.

A Method for Evaluating Hyperparameter Sensitivity in Reinforcement Learning Neural fitted q iteration - first experiences with a data efficient neural reinforcement learning method

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:07:40.063461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T19:07:39.710916Z digest=sha256:0da1ddcfdf59a2ec369017cb9f87200b7943fe768b8c3bb5b68ff58d04bc6138

Observation b19e3072-29de-4c58-aec9-063952b2d161 · outbound

This paper cites Proximal Policy Optimization Algorithms.

A Method for Evaluating Hyperparameter Sensitivity in Reinforcement Learning Proximal Policy Optimization Algorithms

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T19:07:39.716115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:07:39.716115Z digest=sha256:c83ebea61f1f6d21d24ddc09bb0118e845557aa0f4959f8a503976c3c34892c8

Observation cf16789d-9655-40e6-9b5a-e4b8169bce63 · outbound

This paper cites A reinforcement learning method for maximizing undiscounted rewards.

A Method for Evaluating Hyperparameter Sensitivity in Reinforcement Learning A reinforcement learning method for maximizing undiscounted rewards

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:07:40.040645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T19:07:39.722669Z digest=sha256:57c8e143c2d09e468f3bb4e51a62b788c5da228e6ab46a979f023d7ad8920c17

Observation 695a8875-defd-46de-9593-b905024d9cca · outbound

This paper cites Dickerson, and Joseph Suarez.

A Method for Evaluating Hyperparameter Sensitivity in Reinforcement Learning Dickerson, and Joseph Suarez

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:07:40.022256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T19:07:39.727969Z digest=sha256:bc67d2c44cd904639a4b923c5889f39f1ea4e6a945b4fef9cf15d0b96f497e15

Observation b7836f55-462b-4123-aa07-e744122403ae · outbound

This paper cites Reinforcement Learning : An Introduction.

A Method for Evaluating Hyperparameter Sensitivity in Reinforcement Learning Reinforcement Learning : An Introduction

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:07:40.002884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T19:07:39.733238Z digest=sha256:ddb73cd07f29c928668682a7201e5e5ea6a74045b705feebfb4e68cd18c39d5b

Observation 30e5cbdd-c6ee-4b28-8726-5778092e61b1 · outbound

This paper cites Learning Values Across Many Orders of Magnitude.

A Method for Evaluating Hyperparameter Sensitivity in Reinforcement Learning Learning Values Across Many Orders of Magnitude

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:07:39.983700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T19:07:39.738929Z digest=sha256:efd60b7ebbb2ceb6f59fe8d1ed6eba09e6dfb52faf89681a7151b54ed5112079

Observation 71501654-42b8-408f-93d8-80ab33280237 · outbound

This paper cites Learning from Delayed Rewards.

A Method for Evaluating Hyperparameter Sensitivity in Reinforcement Learning Learning from Delayed Rewards

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:07:39.957633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-11T19:07:39.744232Z digest=sha256:5e27c91e3ba9f08b9ab01597422685f889bd7766518d28fc1ea08f28454e3041

Observation 436d46a4-1008-4f7f-97b4-cdac143ab0f5 · outbound

This paper cites write newline.

A Method for Evaluating Hyperparameter Sensitivity in Reinforcement Learning write newline

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T19:07:39.750969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:07:39.750969Z digest=sha256:1e131386607c269f20949240168980136614bf3ca47b8175f952a99d50ece7a6

Pith citing papers

Observation a86bb074-2a9b-47c2-90ba-95720d7c2bec · inbound

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One cites this paper.

Multiple Weaks Win Single Strong: Large Language Models Ensemble Weak Reinforcement Learning Agents into a Supreme One A Method for Evaluating Hyperparameter Sensitivity in Reinforcement Learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T15:24:57.402374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:24:57.402374Z digest=sha256:712d1b90e092478476320703635f52dc22fe165f46a5bae01c1bfe9fb64fc969

Observation f48faf5f-8f10-4095-b923-e3fe5bf3f8f4 · inbound

How Should We Meta-Learn Reinforcement Learning Algorithms? cites this paper.

How Should We Meta-Learn Reinforcement Learning Algorithms? A Method for Evaluating Hyperparameter Sensitivity in Reinforcement Learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:44.755232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:44.755232Z digest=sha256:a1a58ac511d4bead80d2ffba8104c8da8c2e436fc670f3d902b494ffc746f131

Observation ebba4398-d90e-42c5-9f66-de8e2bf794a0 · inbound

Feedback-Normalized Developer Memory for Reinforcement-Learning Coding Agents: A Safety-Gated MCP Architecture cites this paper.

Feedback-Normalized Developer Memory for Reinforcement-Learning Coding Agents: A Safety-Gated MCP Architecture A Method for Evaluating Hyperparameter Sensitivity in Reinforcement Learning

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-09T22:29:06.380913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-09T13:59:40.638085Z digest=sha256:4f1166ade9564888e02b78f15b810ec736e7d4424c41b8f576b5eefdfa69d105