Pith. sign in

Paper Citation Record · LEDGER

Semiparametric Off-Policy Inference for Optimal Policy Values under Possible Non-Uniqueness

As of 16 August 2026, this Paper Citation Record lists 12 of 12 outbound references and 0 inbound Pith citation observations for arXiv:2505.13809.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.13809 v5

Coverage vector

measured 12 of 12 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:14:55.422016Z

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

12 of 12 outbound references displayed

  • verified exact0
  • verified fuzzy7
  • unresolved5
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8afd3d04-3793-45d1-9288-f83cdee74b35 · outbound

This paper cites & Abbeel, P.

Semiparametric Off-Policy Inference for Optimal Policy Values under Possible Non-Uniqueness & Abbeel, P

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:14:55.655675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T20:14:55.353000Z digest=sha256:958535b34e7c37be229cf02bf2c52f3e257684bee02e4f0ec94b7b183d84925b

Observation 92053903-744a-47ac-bd72-203bd4179011 · outbound

This paper cites Ψ˚pPq´E P “ QpPqpA,S;πqπpA|Sq ‰ “EP “ QpPqpA,S;π ˚pPqqπ˚pPqpA|Sq ‰ ´EP “ QpPqpA,S;πqπpA|Sq ‰ “EP “ QpPq ` A,S;π ˚pPq ˘` π˚pPqpA|Sq´πpA|Sq ˘‰ `EP.

Semiparametric Off-Policy Inference for Optimal Policy Values under Possible Non-Uniqueness Ψ˚pPq´E P “ QpPqpA,S;πqπpA|Sq ‰ “EP “ QpPqpA,S;π ˚pPqqπ˚pPqpA|Sq ‰ ´EP “ QpPqpA,S;πqπpA|Sq ‰ “EP “ QpPq ` A,S;π ˚pPq ˘` π˚pPqpA|Sq´πpA|Sq ˘‰ `EP

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:14:55.589897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T20:14:55.405526Z digest=sha256:bb0a2010f4cf6eb5fff0e36eb942dd016f1495eb0f2c21df4fc0634a4283f9fa

Observation 1c9dd6f1-f501-4084-bcd0-dc0ade227e07 · outbound

This paper cites By Theorem 25.32 of Van Der Vaart (2000), there exists no estimator sequence for Ψ ˚ that is regular atP ϵ, and hence there exists no RAL estimator for Ψ ˚ atP.

Semiparametric Off-Policy Inference for Optimal Policy Values under Possible Non-Uniqueness By Theorem 25.32 of Van Der Vaart (2000), there exists no estimator sequence for Ψ ˚ that is regular atP ϵ, and hence there exists no RAL estimator for Ψ ˚ atP

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:14:55.555576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T20:14:55.416702Z digest=sha256:8a482706b56e659932d377350ffc14ce6c7fc4e16a69641e0b78c695e26c0948

Observation 58f163da-b520-4dc4-b13b-98d2f4e2163f · outbound

This paper cites pηNSAVE´ηwppπpQqq “.

Semiparametric Off-Policy Inference for Optimal Policy Values under Possible Non-Uniqueness pηNSAVE´ηwppπpQqq “

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:14:55.538269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T20:14:55.422016Z digest=sha256:ecbcc3cf1c39e453f5a769730bff7749d0a180e83baa930380e1503d82e360c0

Observation e3ea3dba-1c28-43c9-8b96-cb5c678c4cea · outbound

This paper cites The Role of Environment Access in Agnostic Reinforcement Learning.

Semiparametric Off-Policy Inference for Optimal Policy Values under Possible Non-Uniqueness The Role of Environment Access in Agnostic Reinforcement Learning

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-15T20:14:55.367289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:14:55.367289Z digest=sha256:e0d656a17090542bcb263e4eb987b119a4dc35745d90eb110018a881fc5dfe16

Observation 2717d0b3-b903-49de-a6ea-1f562a96d27b · outbound

This paper cites an unresolved cited work.

Semiparametric Off-Policy Inference for Optimal Policy Values under Possible Non-Uniqueness Unresolved cited work

Reference 688

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:14:55.640304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T20:14:55.360336Z digest=sha256:f87280ddc033460bfc9bcad55b61ee79d227b5181abea62d2cc88a987237275f

Observation 528d8283-46c7-44e1-a8dc-bf98ed3b662c · outbound

This paper cites A unified view of entropy-regularized Markov decision processes.

Semiparametric Off-Policy Inference for Optimal Policy Values under Possible Non-Uniqueness A unified view of entropy-regularized Markov decision processes

Reference 713

Resolution
unresolved
no resolver link, observed 2026-08-15T20:14:55.379705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:14:55.379705Z digest=sha256:9ccd9519fb82a51b2a013229c8dd9cc205a619c1ba80ba9195cbc16d31b42861

Observation d881e405-e303-4f3a-983a-cfbe19340b71 · outbound

This paper cites Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems.

Semiparametric Off-Policy Inference for Optimal Policy Values under Possible Non-Uniqueness Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems

Reference 1225

Resolution
unresolved
no resolver link, observed 2026-08-15T20:14:55.373331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:14:55.373331Z digest=sha256:8816a76befd31ba6379e6be198ed8fad72607fe4e605825fa9d54fcf188009aa

Observation 7acdadff-1da1-49e1-95a3-fba1adf803fc · outbound

This paper cites 1 p1´γq 2 EP0 “ V ` S;π˚pPϵq ˘ ´V ` S;π˚pP0q ˘‰ ` 1 1´γ EP0,S„ωp¨¨¨;π˚pP0qq „ EA„π˚pPϵq.

Semiparametric Off-Policy Inference for Optimal Policy Values under Possible Non-Uniqueness 1 p1´γq 2 EP0 “ V ` S;π˚pPϵq ˘ ´V ` S;π˚pP0q ˘‰ ` 1 1´γ EP0,S„ωp¨¨¨;π˚pP0qq „ EA„π˚pPϵq

Reference 2000

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:14:55.571994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T20:14:55.411001Z digest=sha256:c2f0cdaed08a7c164ba6a0a6b636fc42dd8990c8087ddff42ea5ce77728e1776

Observation 0f76220f-9f6a-4113-ba6a-38c1652077df · outbound

This paper cites (13) The second tool for obtaining the lower bound is the reverse Cauchy-Schwarz inequality (see the result for integrals in Corollary 6.1 of Aldaz et al.

Semiparametric Off-Policy Inference for Optimal Policy Values under Possible Non-Uniqueness (13) The second tool for obtaining the lower bound is the reverse Cauchy-Schwarz inequality (see the result for integrals in Corollary 6.1 of Aldaz et al

Reference 2009

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:14:55.623647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T20:14:55.393729Z digest=sha256:74c30bc0811c46ecfeda7aafda06fd84699d2855a740d4c54388250bf34a902f

Observation 7e554d34-5e25-4908-80db-d8d8af7d830d · outbound

This paper cites It remains to upper bound the first term on the right-hand side of (14).

Semiparametric Off-Policy Inference for Optimal Policy Values under Possible Non-Uniqueness It remains to upper bound the first term on the right-hand side of (14)

Reference 2015

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:14:55.606406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T20:14:55.399029Z digest=sha256:a52f1f55b7b936f5c0ec24655dfa21696da24f5fae061976faeb6d156e11ecd3

Observation 6ebb3824-b34c-4ddb-a2ba-ceddff9ba563 · outbound

This paper cites A Review of Off-Policy Evaluation in Reinforcement Learning.

Semiparametric Off-Policy Inference for Optimal Policy Values under Possible Non-Uniqueness A Review of Off-Policy Evaluation in Reinforcement Learning

Reference 9668

Resolution
unresolved
no resolver link, observed 2026-08-15T20:14:55.385777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:14:55.385777Z digest=sha256:284aef8ec1067d9cb380c0a0e069268999d02b6fcb743b5edfe45b4fb83292aa

Pith citing papers

No inbound Pith citation observations are available.