Pith. sign in

Paper Citation Record · LEDGER

Hyperparameter Selection for Offline Reinforcement Learning

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 14 inbound Pith citation observations for arXiv:2007.09055.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2007.09055 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 14 of 14 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T04:36:12.787317Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-08T22:35:40.699101Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 97df3bbe-b358-41e8-9e0f-eba154358607 · inbound

What Matters in Learning from Offline Human Demonstrations for Robot Manipulation cites this paper.

What Matters in Learning from Offline Human Demonstrations for Robot Manipulation Hyperparameter Selection for Offline Reinforcement Learning

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-13T08:51:55.909662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T08:51:55.826747Z digest=sha256:14eae25fe2d2005d67f2f73ea488340d039b1e1e561cdd0c6a1b3426223de905

Observation 503d20eb-01b0-463b-a778-cf34efdc8a83 · inbound

Simulating Errors in Touchscreen Typing cites this paper.

Simulating Errors in Touchscreen Typing Hyperparameter Selection for Offline Reinforcement Learning

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-09T04:36:12.787317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T04:36:12.787317Z digest=sha256:ac162ca49eab054c74296f10b58aaf8934fe9caef6f68bffa716d089b47f5a21

Observation 4f965292-2506-4043-9daa-6e2b00cb5c13 · inbound

Learning to Trust Bellman Updates: Selective State-Adaptive Regularization for Offline RL cites this paper.

Learning to Trust Bellman Updates: Selective State-Adaptive Regularization for Offline RL Hyperparameter Selection for Offline Reinforcement Learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:46.200212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:14:46.200212Z digest=sha256:b3d491dd14c9f0d8ba1d071ab262fee2604a12a1bf88bf4aff71d9e6ebfed536

Observation 3f69e0b2-8fea-4d05-92fa-65b796d5717b · inbound

Fully Offline Reinforcement Learning cites this paper.

Fully Offline Reinforcement Learning Hyperparameter Selection for Offline Reinforcement Learning

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:02.204750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:16:02.204750Z digest=sha256:38b031938b2f19370734ccd2b789a9093009183224d508fcf0ad1b8e2f01515f

Observation 61c75cfd-8bf8-4d42-94d8-4f1ad775cd03 · inbound

Learning to Evaluate Autonomous Behaviour in Human-Robot Interaction cites this paper.

Learning to Evaluate Autonomous Behaviour in Human-Robot Interaction Hyperparameter Selection for Offline Reinforcement Learning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T19:11:35.228663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:11:35.228663Z digest=sha256:af9ee0c8e889ff6aa2ea21365e35fef40eea8f7b0cd6b83a244ec4e2cb06b29a

Observation ff3857be-c107-487d-924d-b1acc5806e36 · inbound

Accelerating Detailed Routing Convergence through Offline Reinforcement Learning cites this paper.

Accelerating Detailed Routing Convergence through Offline Reinforcement Learning Hyperparameter Selection for Offline Reinforcement Learning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-03T18:47:48.077033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:47:48.077033Z digest=sha256:efcb44972016cfea1b33629814ed64bbf8731faa42b076e1f3a4266885b6bf52

Observation 95d05ae8-6b0e-49c3-add4-668a5af5f3e9 · inbound

Adaptive Policy Selection and Fine-Tuning under Interaction Budgets for Offline-to-Online Reinforcement Learning cites this paper.

Adaptive Policy Selection and Fine-Tuning under Interaction Budgets for Offline-to-Online Reinforcement Learning Hyperparameter Selection for Offline Reinforcement Learning

Reference 33

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T17:41:09.214704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-08T17:17:14.300877Z digest=sha256:362e430afaa3bd7a5c217c7bd7c395b27034038fa65ae5957fa43af50912fa16

Observation 89e47c35-418a-4152-b106-d5f27aa3e2b0 · inbound

Sample-efficient inductive matrix completion with noise and inexact side-information cites this paper.

Sample-efficient inductive matrix completion with noise and inexact side-information Hyperparameter Selection for Offline Reinforcement Learning

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-20T14:08:21.004025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-20T14:04:44.364824Z digest=sha256:97fd41c177d274fa80982b55eb1ed60d36597c11f0dd453cc95d7f51f871fa54

Observation ce6e5a52-5af1-444e-9dc0-b7455dfa9314 · inbound

Autoregressive Diffusion World Models for Off-Policy Evaluation of LLM Agents cites this paper.

Autoregressive Diffusion World Models for Off-Policy Evaluation of LLM Agents Hyperparameter Selection for Offline Reinforcement Learning

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-07-02T11:46:56.118548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T02:53:54.942603Z digest=sha256:b64c20c07433601cd83fb121302da7edec1b7ba09311affa9c5da8e6da2cf843

Observation 61327d7e-f6e5-4e51-b616-4ddf7b986675 · inbound

Autoregressive Diffusion World Models for Off-Policy Evaluation of LLM Agents cites this paper.

Autoregressive Diffusion World Models for Off-Policy Evaluation of LLM Agents Hyperparameter Selection for Offline Reinforcement Learning

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-02T12:24:48.795303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:24:48.795303Z digest=sha256:cecd3905aded162683db988d7c62a0e1cf8d46cb0b50a1dcc68b4fe1646156e3

Observation 9c22ebbc-b5a0-47a6-9c31-fe65c5bccc80 · inbound

Some Essential Constructive Foundations for Systems and Control cites this paper.

Some Essential Constructive Foundations for Systems and Control Hyperparameter Selection for Offline Reinforcement Learning

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T23:47:28.426328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-27T17:48:35.402903Z digest=sha256:d053d5019ccd39f51e73375cc6fe881505babf86e7a810ba1b1b8c23995c6d1d

Observation 68bb7b42-0d07-4fb7-b15d-0f90912b8842 · inbound

$\text{DT}^2$: Decision-Targeted Digital Twins cites this paper.

$\text{DT}^2$: Decision-Targeted Digital Twins Hyperparameter Selection for Offline Reinforcement Learning

Reference 129

Resolution
verified exact
arxiv_id, observed 2026-07-04T20:30:07.312491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-25T20:08:13.039445Z digest=sha256:7e4a61f4c9c7f913354d2893a1f7e2de2bdd2a627d7349060f7b30f3a455330e

Observation 3b6f9a96-ec65-4ca4-923b-9cc60bf59f8b · inbound

Strategic Bargaining in Multi-Buyer Markets: Reinforcement Learning from Verifiable Rewards for LLM Negotiations cites this paper.

Strategic Bargaining in Multi-Buyer Markets: Reinforcement Learning from Verifiable Rewards for LLM Negotiations Hyperparameter Selection for Offline Reinforcement Learning

Reference 175

Resolution
verified exact
local_arxiv, observed 2026-07-08T22:35:40.700354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-08T22:26:44.052574Z digest=sha256:80f364306549ee5263a8097bdc64fefdbe9cbf8bc20869c1a26a10908fe43a7e

Observation e951cdb5-a131-4551-ba6f-3b7b55ee904c · inbound

Active Offline-to-Online Reinforcement Learning cites this paper.

Active Offline-to-Online Reinforcement Learning Hyperparameter Selection for Offline Reinforcement Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-14T03:38:59.132093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T03:38:59.132093Z digest=sha256:584d154e2aac8e6507e6bbe19679413bc4ba2c8629f826793f15a7d2e85b637c