Pith. sign in

Paper Citation Record · LEDGER

Empirical Study of Off-Policy Policy Evaluation for Reinforcement Learning

As of 14 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 8 inbound Pith citation observations for arXiv:1911.06854.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1911.06854 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 8 of 8 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:59:43.198273Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-07T14:33:54.441902Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation daf5ec59-b9ec-45d2-9b0b-1355b7dab123 · inbound

Off-Policy Evaluation Under Nonignorable Missing Data cites this paper.

Off-Policy Evaluation Under Nonignorable Missing Data Empirical Study of Off-Policy Policy Evaluation for Reinforcement Learning

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-06T18:59:43.198273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:59:43.198273Z digest=sha256:5b18f398bdfae0c6dec30ca4e9e7caf26410548724f03cc9d8c32939652fb8eb

Observation 5f03b118-fe6d-4584-bc95-413a32e2d38a · inbound

The Three Regimes of Offline-to-Online Reinforcement Learning cites this paper.

The Three Regimes of Offline-to-Online Reinforcement Learning Empirical Study of Off-Policy Policy Evaluation for Reinforcement Learning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T13:00:09.551445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:00:09.551445Z digest=sha256:e30f702a283063d821d52c4edd6e5416f7ce0355444b1d06c0b5e784f7039aa7

Observation 17fb03ed-4e93-4b4e-b7cb-4295feb810d8 · inbound

MedGym:A Unified Continuous-Time Benchmark for Dynamic Medical Treatment Reinforcement Learning cites this paper.

MedGym:A Unified Continuous-Time Benchmark for Dynamic Medical Treatment Reinforcement Learning Empirical Study of Off-Policy Policy Evaluation for Reinforcement Learning

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T21:06:13.695187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-28T17:29:37.376343Z digest=sha256:b44f85c37d3e95725c5ef756665097faecbbc97997ae2f863293e01facc162ee

Observation a77fb6c2-c596-4c43-b9b6-49d309c4b0c2 · inbound

Autoregressive Diffusion World Models for Off-Policy Evaluation of LLM Agents cites this paper.

Autoregressive Diffusion World Models for Off-Policy Evaluation of LLM Agents Empirical Study of Off-Policy Policy Evaluation for Reinforcement Learning

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-02T11:46:56.123242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-28T02:53:54.942603Z digest=sha256:daa5a89926f7a1c6e679eb5c79042b8d2d3da0d5311006a015d99dabc7ba3dec

Observation 09354f68-833e-42fd-953b-a641e070e521 · inbound

Autoregressive Diffusion World Models for Off-Policy Evaluation of LLM Agents cites this paper.

Autoregressive Diffusion World Models for Off-Policy Evaluation of LLM Agents Empirical Study of Off-Policy Policy Evaluation for Reinforcement Learning

Reference 1996

Resolution
unresolved
no resolver link, observed 2026-08-02T12:24:48.810739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:24:48.810739Z digest=sha256:2117091e923561a34cc83995380e191781dc67e9a6271aa512ee78b6033a2a82

Observation 456e3113-319c-4082-9cf0-80b2ab594428 · inbound

Off-Policy Evaluation for Missingness-Aware Policies in MDPs with Rewards Missing Not at Random cites this paper.

Off-Policy Evaluation for Missingness-Aware Policies in MDPs with Rewards Missing Not at Random Empirical Study of Off-Policy Policy Evaluation for Reinforcement Learning

Reference 43

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T05:39:40.718347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-26T15:38:16.806415Z digest=sha256:54fc875e2d9a50b4461c00bd83039a5ff3c7bdfee5a71507a5276236cf872a66

Observation acdee659-2dcf-4622-804e-1e19289bd25e · inbound

Fitted Occupancy-Ratio Evaluation without Bellman Completeness cites this paper.

Fitted Occupancy-Ratio Evaluation without Bellman Completeness Empirical Study of Off-Policy Policy Evaluation for Reinforcement Learning

Reference 59

Resolution
verified exact
local_arxiv, observed 2026-07-07T14:33:54.443180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-07-07T14:25:33.921937Z digest=sha256:6cb5e1e5060b501b1870e31c19b3e3d2e2390e5f4d7ef5b684129a0b3578ce65

Observation c5f9debe-2b57-41a2-914f-90e7352b1446 · inbound

Fitted Occupancy-Ratio Evaluation without Bellman Completeness cites this paper.

Fitted Occupancy-Ratio Evaluation without Bellman Completeness Empirical Study of Off-Policy Policy Evaluation for Reinforcement Learning

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-02T08:33:40.502503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:33:40.502503Z digest=sha256:f5f782703a98245e384c31b14c2e3bcb6a34498982a9cf3f50be67e27fab6e2c