Pith. sign in

Paper Citation Record · LEDGER

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation

As of 8 August 2026, this Paper Citation Record lists 95 of 95 outbound references and 0 inbound Pith citation observations for arXiv:2505.22492.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.22492 v1

Coverage vector

measured 95 of 95 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:17:01.124208Z

measured 95 of 95 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

95 of 95 outbound references displayed

  • verified exact7
  • verified fuzzy50
  • unresolved37
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation fe970c9f-0548-4569-aaa3-4eab2ff15df9 · outbound

This paper cites an unresolved cited work.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:52.903422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:16:52.903422Z digest=sha256:4f01921d7973b985ffca1299f8cf05dec65828002b9662b74f0e667fa47930dc

Observation b62aa258-e78b-4938-bb97-799ce524f0ef · outbound

This paper cites and Kallus, N.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation and Kallus, N

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:53.021013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:16:53.021013Z digest=sha256:6994f64dfc198a5e834215a6f88c15df3d86ece5b6ae8a65a3673dbbd24834ed

Observation 3b741a11-b4dc-4569-b70d-945c5839b408 · outbound

This paper cites Off-policy evaluation in doubly inhomogeneous environments.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Off-policy evaluation in doubly inhomogeneous environments

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:53.103572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:16:53.103572Z digest=sha256:fe5f40459e533aa1a681c5eb41077732b0d8234cee6d8d44d7952cd5f3227639

Observation 440371c0-981b-4670-bec6-f0aacf55fdae · outbound

This paper cites More efficient off-policy evaluation through regularized targeted learning.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation More efficient off-policy evaluation through regularized targeted learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:53.186263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:16:53.186263Z digest=sha256:5d7043bdba72c61884cf56435ffa78c3b8811ce1484a1f8daf3c1ac216508e14

Observation 342ab1ad-17f1-455a-9f63-4a19cc565efe · outbound

This paper cites an unresolved cited work.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:53.264406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:16:53.264406Z digest=sha256:a944520213aa34045cf35f6999d2168ad80b330ac715c0fa93a8da0db9596b7e

Observation 552f657a-6eda-4337-a6d5-a218da5e2d98 · outbound

This paper cites OpenAI Gym.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation OpenAI Gym

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:53.388701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:16:53.388701Z digest=sha256:85d94554011bba37a57df21e56201840dbf2bcb2dfb5e6d10a3abeab20cc7f70

Observation c6c5e743-b5b3-4cbb-aebf-ad81d8d95b29 · outbound

This paper cites and Zhou, A.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation and Zhou, A

Reference 7

Resolution
verified exact
raw_fallback, observed 2026-08-07T13:17:04.709883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:16:53.496289Z digest=sha256:dc462a443a3ef03632a4d4dbff6ec49ee7bacd9eb6b6be3ded38ee157c136251

Observation 8a06d0b8-b9d7-40c3-916f-86b10aa7c50a · outbound

This paper cites Structured Difference-of-Q via Orthogonal Learning.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Structured Difference-of-Q via Orthogonal Learning

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:17:04.447238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:16:53.610687Z digest=sha256:491baedf73289ea9d7c6f3ef57b7269164ed8c9a954187a4127baf8a3f0835dd

Observation 67a52a5d-4608-4582-a6c2-bd0f88c6a259 · outbound

This paper cites and Berger, R.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation and Berger, R

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:53.686598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:16:53.686598Z digest=sha256:a97e32bb07fe5da0bd492da77ee5204c724b6bcf0b7399fb74f947fc466d54dd

Observation bc2ae633-0ad7-44ca-a0db-07a0658b0efa · outbound

This paper cites and Li, L.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation and Li, L

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:53.794136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:16:53.794136Z digest=sha256:cc6f1dd39b597b02c80cd667047b89435b2993665c9f5b642da10577cf6da32d

Observation 49de3aaf-d1f8-4dbd-a4a4-759ad9b84b90 · outbound

This paper cites and Jiang, N.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation and Jiang, N

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:53.924369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:16:53.924369Z digest=sha256:67aa23b8070d7dd88995cd003f529279aecee89284db82fc0d3f0918557670c6

Observation b8a8b783-825e-4208-95fc-5d209ea87241 · outbound

This paper cites and Qi, Z.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation and Qi, Z

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:54.023087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:16:54.023087Z digest=sha256:cf3afccba990c3075d7cc195d2604ee26e8f72ac18d3827e3bed915cf2b46c8f

Observation fca476a0-cfa0-4d92-ab2b-9bf5fb2ca2f7 · outbound

This paper cites Gaussian approximation of suprema of empirical processes.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Gaussian approximation of suprema of empirical processes

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:54.133862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:16:54.133862Z digest=sha256:3fa08a9601870a8104a78d13ee33384553abbc0002f3dad748c024c4b447d422

Observation 52b4bb87-f570-4feb-84e1-1403a8ab136a · outbound

This paper cites Double/debiased machine learning for treatment and structural parameters.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Double/debiased machine learning for treatment and structural parameters

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:54.246149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:16:54.246149Z digest=sha256:a60e760763e45569d72029ca0da0ea46ecf3325d370c060bac355af87ff92944

Observation db63a9a8-f309-47aa-acac-bd6bb2d1c379 · outbound

This paper cites Coindice: Off-policy confidence interval estimation.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Coindice: Off-policy confidence interval estimation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:54.339217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:16:54.339217Z digest=sha256:42fc866ba7e91dcb9d337ed85450f3a0e4092c0ec924ed6177e4a8911f61c151

Observation 0a9c8274-3776-490b-a0ed-558411d6bf08 · outbound

This paper cites Doubly Robust Policy Evaluation and Optimization.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Doubly Robust Policy Evaluation and Optimization

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:54.432077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:16:54.432077Z digest=sha256:291b02862af1306287396fed31e73d252cfe50a44c121a53edc438bf2c21a4de

Observation 8446019f-3832-4df8-9d51-6b1368b142b9 · outbound

This paper cites A theoretical analysis of deep q-learning.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation A theoretical analysis of deep q-learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:54.524675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:16:54.524675Z digest=sha256:e9110743239081791f05ee56520c2610415e7817e563bddc6a44f7386d836a8c

Observation 25caa724-aca7-4a07-b9f0-dda5c47cd091 · outbound

This paper cites More Robust Doubly Robust Off-policy Evaluation.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation More Robust Doubly Robust Off-policy Evaluation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:54.617714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:16:54.617714Z digest=sha256:fbb5781c01ffeaf37099700bf91e133a76bd2542b08b7791b1375dd54b2195c8

Observation 69d160b1-511b-4984-a4d4-ec8006edc122 · outbound

This paper cites Accountable off-policy evaluation with kernel B ellman statistics.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Accountable off-policy evaluation with kernel B ellman statistics

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:54.732514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:16:54.732514Z digest=sha256:b0fb21de4826041135ce2f8a7753de4e4145628eabcecb7e50f2cbd4152f312b

Observation 1d2f3fc2-4499-45b8-ae5c-0061bd1e4779 · outbound

This paper cites Combining parametric and nonparametric models for off-policy evaluation.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Combining parametric and nonparametric models for off-policy evaluation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:54.828722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:16:54.828722Z digest=sha256:e7be742119d1714b68735cb204e743026101b8dbfa6e60a4025ed5e22519784f

Observation 3e674bf6-aede-4609-a97b-8498caf0fb30 · outbound

This paper cites D., Thomas, P.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation D., Thomas, P

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:15.124519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:16:54.910774Z digest=sha256:15122fe1da0369dba12b1c6ddd124c21990c3bea6f64f720682b412120bda13b

Observation 238d6f10-29ef-4d5b-b3bf-b06331b78393 · outbound

This paper cites Importance sampling policy evaluation with an estimated behavior policy.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Importance sampling policy evaluation with an estimated behavior policy

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:14.980638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:16:54.995797Z digest=sha256:7c78137afd9f3d5bd009e67d5a02354b3f1d5ef804ce96957911cf770308882b

Observation 715be3ee-ce98-4c84-ac12-c5babdbdda54 · outbound

This paper cites P., Thomas, P.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation P., Thomas, P

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:14.829766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:16:55.071207Z digest=sha256:e46ba36508a21d0857e4a59058ae6ffbe7fcf1fa28acbe616e84b2762f702993

Observation 5db9c47a-84a6-4898-9dc1-e9c7a05cdcd1 · outbound

This paper cites P., Niekum, S., and Stone, P.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation P., Niekum, S., and Stone, P

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:14.679690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:16:55.196586Z digest=sha256:a6be58f70ada30149e9ef9aba6ccc8f0f5c8696093a450d2f0d1ce1bb01fdf51

Observation 9c90fb45-6232-4e37-b518-1c03af849c8c · outbound

This paper cites Bootstrapping fitted q-evaluation for off-policy inference.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Bootstrapping fitted q-evaluation for off-policy inference

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:14.525737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:16:55.317248Z digest=sha256:e16d763cf834c0c03e0fb2b7b7ae1d04e1c2775dc5551c77bbdcb84b7420d312

Observation 891b15a5-8bf1-45dc-abba-88a38f30a1b1 · outbound

This paper cites Importance sampling via the estimated sampler.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Importance sampling via the estimated sampler

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:14.384125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:16:55.387170Z digest=sha256:008d3786741f9a5b8288424e79b0f1ed7da215bccd687fa754c6c47e2ae2c1e0

Observation 2ad6d719-9b29-446e-a5d2-f9e1a16fc9ca · outbound

This paper cites W., and Ridder, G.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation W., and Ridder, G

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:14.222790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:16:55.459117Z digest=sha256:cbdd43393c5715af938c318d5d6134cd61cc13e8edcd1b1b8fee59cbdb267617

Observation db7f3592-4268-4450-ba42-a0663dfe4af2 · outbound

This paper cites and Wager, S.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation and Wager, S

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:55.534857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:16:55.534857Z digest=sha256:d31bc9ef69dbf5cdd0569a7df6b941ffd38e455c7cfdca3707a7c9f08151bfe3

Observation ff3d687f-0b80-4dd4-a4d7-db5b0d93866d · outbound

This paper cites and Li, L.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation and Li, L

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:14.029859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:16:55.616369Z digest=sha256:6511080929709951dd2d3d63241907bfc0818375d7ad2365f11513f48c251805

Observation a42c0dd8-5648-4354-b424-84d1a4e77faa · outbound

This paper cites and Uehara, M.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation and Uehara, M

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:13.838362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:16:55.716241Z digest=sha256:67c4848ae8a3429b26b44ae461476ef5000250c4a33cc237dd5fadeb4b5161f7

Observation d404d600-7de2-4150-946b-5307bac95f8c · outbound

This paper cites and Uehara, M.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation and Uehara, M

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:13.626828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:16:55.793014Z digest=sha256:7eee8a4e179e331031ac52e98291c2ba6218d853f11ea29835794210944e5be9

Observation 6411d0f2-b200-46ad-b612-6dd15b17e52c · outbound

This paper cites and Zhou, A.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation and Zhou, A

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:13.420849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:16:55.872168Z digest=sha256:0549b46deac70bae9bd0221d7ed1292fccb5a35b34d899d53a257d956c711423

Observation 63221bff-3c62-45ab-9178-79969858e05b · outbound

This paper cites an unresolved cited work.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:17:13.160576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:16:55.932696Z digest=sha256:eaa6fc2de034946e2b24bc7b7215e1bcb4d600e3c5e482e75dd39db29daab20d

Observation 5a9cc8b5-daf7-428d-9acd-b4d241a9daf0 · outbound

This paper cites Batch policy learning under constraints.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Batch policy learning under constraints

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:12.937701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:16:56.017006Z digest=sha256:1eb61877e5ca3621504d1b6f65fc608125bfda7f8734b14b862ac725447a393e

Observation 1924c1ed-bc7b-4f1b-98f6-816e702217d1 · outbound

This paper cites Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:56.144780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:16:56.144780Z digest=sha256:77e332472c7d44b82454df0dc2c2c8f0102c4c5b9d9c9b344a04b3dda4c72b99

Observation 977884f7-69c5-4e80-aa98-3b38ff7c87f1 · outbound

This paper cites High-probability sample complexities for policy evaluation with linear function approximation.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation High-probability sample complexities for policy evaluation with linear function approximation

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:17:04.262424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:16:56.248494Z digest=sha256:45c62e4f5344a272dd180ab72a4b14982d74cf88b22532e29ab2a09a35b23889

Observation 38ccf486-2bf4-47ce-91d8-e2cb5ef3fff3 · outbound

This paper cites Off-policy estimation of long-term average outcomes with applications to mobile health.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Off-policy estimation of long-term average outcomes with applications to mobile health

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:12.773972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:16:56.359148Z digest=sha256:6b512464d0e2512c646c8ec3d702cbf8203c810a3dd56b0717fcb78adabd81cd

Observation 67f38c91-6c96-4360-9f51-f15cbc3b435a · outbound

This paper cites an unresolved cited work.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:17:12.624851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:16:56.466433Z digest=sha256:1efe02ada5bedcbb89d2aa9d215b558e538ea7197a91a64f4768bb31c496d2c9

Observation 03b93b2d-2ad1-411f-9a43-8b252edd6bf6 · outbound

This paper cites Breaking the curse of horizon: infinite-horizon off-policy estimation.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Breaking the curse of horizon: infinite-horizon off-policy estimation

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:12.478362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:16:56.565112Z digest=sha256:ddb1d52a8260c6d6b3313349daaf9a5febbe105dcefa2bea0207055242fadff3

Observation b1b84187-62f8-49de-bd51-6289711aa20c · outbound

This paper cites and Zhang, S.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation and Zhang, S

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:12.254767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:16:56.671738Z digest=sha256:347b8d8492b6a900576fe84f9317584ac31b843c959439837e884ea5c96f926b

Observation 20212785-c506-4558-a002-57494379b523 · outbound

This paper cites Doubly Optimal Policy Evaluation for Reinforcement Learning.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Doubly Optimal Policy Evaluation for Reinforcement Learning

Reference 41

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T13:17:03.912586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:16:56.740396Z digest=sha256:02d29b5725a0ae67eb988b772eb7bdebbd1aa453cdbed65216f793f789586c06

Observation b6029ff5-5b9f-48b3-9947-4d1d444ba384 · outbound

This paper cites Online Estimation and Inference for Robust Policy Evaluation in Reinforcement Learning.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Online Estimation and Inference for Robust Policy Evaluation in Reinforcement Learning

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:17:03.551774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:16:56.834090Z digest=sha256:0d73b1dd0d017a28ceb400fa7692d0a894d6ea887d414e2ecee1b63da81fa676

Observation 1ed53d7c-3231-44c0-8f6b-0feeedbf69dc · outbound

This paper cites J., Laber, E.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation J., Laber, E

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:56.946901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:16:56.946901Z digest=sha256:96c0cf0233bfcedac07aa69ece51251d9cf8ada9d4af4af8c2051e9074aba3f3

Observation 548001c1-8ac1-463d-adbe-e8cb79d621a4 · outbound

This paper cites P., and Nowak, R.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation P., and Nowak, R

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:12.077014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:16:57.007129Z digest=sha256:efe959033eb796eba0878a99ce63ccd3d8153afc5dcd9fdbcdbce50bba6db5ea

Observation e038af15-6e2e-40e6-8d48-d63427dd26d3 · outbound

This paper cites A., van der Laan, M.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation A., van der Laan, M

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:11.923588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:16:57.074159Z digest=sha256:a91b858e544bc6441bb48d8bbd97027505088c8e53e67bb213cef130b5f4bf13

Observation cb654943-6911-4027-b699-509c2785a743 · outbound

This paper cites Dualdice: Behavior-agnostic estimation of discounted stationary distribution corrections.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Dualdice: Behavior-agnostic estimation of discounted stationary distribution corrections

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:11.691191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:16:57.135333Z digest=sha256:8407d94673efa5ff484ec54b0e2ece6a965cd6f27dddbd3d36a855e2e6d8bf0d

Observation 3bbf4a0c-7e88-437c-ba2e-7be99f60802d · outbound

This paper cites A Spectral Approach to Off-Policy Evaluation for POMDPs.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation A Spectral Approach to Off-Policy Evaluation for POMDPs

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:57.195687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:16:57.195687Z digest=sha256:63025498f9d7db49c01c613e0dc8b980eb6839b4d8e3aadbe13f8083f03bcd89

Observation 877ba37c-7916-41a1-bbda-3021e7f63f2b · outbound

This paper cites Off-policy policy evaluation for sequential decisions under unobserved confounding.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Off-policy policy evaluation for sequential decisions under unobserved confounding

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:11.540247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:16:57.246652Z digest=sha256:860673f6168b39d0c3d35c520d3e0ff785c73e053363b3ddb455223c2ef9f7ba

Observation fe049d9e-3dad-43a7-876b-39568be90b97 · outbound

This paper cites K., Hsieh, F., and Robins, J.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation K., Hsieh, F., and Robins, J

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:11.289965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:16:57.316362Z digest=sha256:27b3270d43e37f56505497271ab4472847c14b8477672fbe924b097c24bd95db

Observation 220033ba-5dc0-445d-b8e8-855b2dd8aecf · outbound

This paper cites Training language models to follow instructions with human feedback.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Training language models to follow instructions with human feedback

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:57.362862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:16:57.362862Z digest=sha256:8fb200d5e6073fe8b7a63f83303c63cada558b8db9f55c4494a591ec70ed0dee

Observation 7e730401-bf8b-41ff-a17e-a81d69418f46 · outbound

This paper cites S., and Singh, S.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation S., and Singh, S

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:11.104301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:16:57.449604Z digest=sha256:7deefcd8a38847bded347a94877ed665262406133ee9b91ea303afaafbddbfc6

Observation c5644559-bd83-4111-bc92-4c2224a768b3 · outbound

This paper cites an unresolved cited work.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Unresolved cited work

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:57.519872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:16:57.519872Z digest=sha256:01245aa44f966e62cb35a5b6c1406a90cc3c1c8cdc92257627047a2f4e174627

Observation eb28c104-9602-4e69-8561-9b0715d8a30c · outbound

This paper cites an unresolved cited work.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:17:10.808173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:16:57.582139Z digest=sha256:665565426570eb6f862a96000e07369525c795f3732ad275dad6de7670805ef0

Observation 4695e73b-5847-4e3f-82b9-e682b16dc01b · outbound

This paper cites Conditional importance sampling for off-policy learning.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Conditional importance sampling for off-policy learning

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:10.692678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:16:57.632221Z digest=sha256:7554523a19026699c73f1e9b0863b2ba7f110900a68b4cccffd2c0ad74f4d4f4

Observation a5078da9-a7fa-4985-85d6-b91f6d93ef28 · outbound

This paper cites G., and Dabney, W.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation G., and Dabney, W

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:10.457993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:16:57.718028Z digest=sha256:aad8bcdf620793202fe29064b59dbb379911793b6e68e88b6b2e09a0794f9b1f

Observation b9e6c4a0-196f-4998-8a2a-b23b71887564 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Proximal Policy Optimization Algorithms

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:57.801762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:16:57.801762Z digest=sha256:cc1551480b3a2e296d7007af213560a54b771f2d10738a6d30bc4dc9a99fa93e

Observation 7e7a7233-70a3-4e87-b622-49670b846ed1 · outbound

This paper cites Estimating the dimension of a model.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Estimating the dimension of a model

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:10.220070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:16:57.866226Z digest=sha256:28bf9312c9cda64ee60f2cc71b2b613026d1c82336d111ce6dc367d89bd6542c

Observation 850e99cd-29f6-4007-a1c9-33882eea3cf1 · outbound

This paper cites and Ben-David, S.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation and Ben-David, S

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:57.934021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:16:57.934021Z digest=sha256:8904937c13663a2aa05d63065b2691dcb00669fc87c033a0f78dca50cc5daa8a

Observation 9afe5e2f-6267-448a-b52c-d34ce187313a · outbound

This paper cites Mathematical Statistics.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Mathematical Statistics

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:58.022585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:16:58.022585Z digest=sha256:75472ce459db573b06ad6b92a24f3a49ed167ecb12a1c6ff4ec2fdd3760a0006

Observation a5a578f7-f424-47bd-9594-7bb5521b9d41 · outbound

This paper cites On methods of sieves and penalization.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation On methods of sieves and penalization

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:10.002895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:16:58.081016Z digest=sha256:067e1305bc55e7bf4e6114c8005189281d71b57dd6204aaf8f10ace92dff1254

Observation 067892b6-6dcf-460f-93c9-2e9dfe291a08 · outbound

This paper cites Deeply-debiased off-policy interval estimation.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Deeply-debiased off-policy interval estimation

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:09.855972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:16:58.143606Z digest=sha256:94d68d051d2c79dc4f04fd6a23eb9ab7025ed9a0430bd346ff8ec0efcfbbe16e

Observation 3587ffd9-3294-4bd0-b23e-85ae9726279a · outbound

This paper cites A minimax learning approach to off-policy evaluation in confounded partially observable markov decision processes.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation A minimax learning approach to off-policy evaluation in confounded partially observable markov decision processes

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:09.668572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:16:58.211350Z digest=sha256:0cec5d801443b464a416ca8f8d4b1db2768578ad3fb7b407ce6eb471a401fe22

Observation 337a06a4-da2c-4071-a7f1-ce0b1005ff5f · outbound

This paper cites Statistical inference of the value function for reinforcement learning in infinite-horizon settings.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Statistical inference of the value function for reinforcement learning in infinite-horizon settings

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:09.489721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:16:58.281999Z digest=sha256:81337cfbd98b103cc21bd42b86e082910ddee26f31eb6ebc75835c9e664c887a

Observation d688c1de-2597-498c-a6cf-596f4ba2351b · outbound

This paper cites Dynamic causal effects evaluation in a/b testing with a reinforcement learning framework.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Dynamic causal effects evaluation in a/b testing with a reinforcement learning framework

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:09.222749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:16:58.378112Z digest=sha256:27f5cb419e5e0eb53d1fca0ae98fa2b370f321f577e0416d72e090ffbefcc88a

Observation d1a16f9e-e59f-46b5-a50d-514496496e92 · outbound

This paper cites Off-policy confidence interval estimation with confounded markov decision process.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Off-policy confidence interval estimation with confounded markov decision process

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:08.987934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:16:58.470043Z digest=sha256:2a4f8aa6f5d4bda7013d7a8cd5cd3ce05310a1297ade71604f211038d572a21c

Observation 61258582-d81d-4811-8a87-20401a18103d · outbound

This paper cites ARMA-Design: Optimal Treatment Allocation Strategies for A/B Testing in Partially Observable Time Series Experiments.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation ARMA-Design: Optimal Treatment Allocation Strategies for A/B Testing in Partially Observable Time Series Experiments

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:58.476262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:16:58.476262Z digest=sha256:fbc22e9b243da405f14e93f5d69ef0b3e03de8dbbe8b83685713ee6c3d819d35

Observation 02bd134f-5140-496e-867d-2a6dc4b7aea0 · outbound

This paper cites S., Szepesv \'a ri, C., and Maei, H.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation S., Szepesv \'a ri, C., and Maei, H

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:08.811401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:16:58.553538Z digest=sha256:9ca0fab437690b2c54e3e40edbaa6176d2fe7ad0dee0620c6b202e2d50e5a01d

Observation a266e348-4d60-4a95-a6c3-91f20e10358b · outbound

This paper cites Doubly robust bias reduction in infinite horizon off-policy estimation.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Doubly robust bias reduction in infinite horizon off-policy estimation

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:08.583530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:16:58.648961Z digest=sha256:74a29d0fe36c49641f575d01cf1f1e67c82b94a979644cd620f04f0c9bfeba17

Observation bd7a0f40-a240-49ca-9b93-4c4ed5bba7b4 · outbound

This paper cites Off-policy evaluation in partially observable environments.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Off-policy evaluation in partially observable environments

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:08.354914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:16:58.740428Z digest=sha256:d8a4415645a43457d43e88569c18ea7fc678b3f4a27baa35eefdb74caef9dd7f

Observation 6b8460eb-5193-457b-a1b5-abfd2a007cdb · outbound

This paper cites an unresolved cited work.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Unresolved cited work

Reference 70

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:17:08.090302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:16:58.839634Z digest=sha256:382c379a1b6d57391ec57818db190f3a57e03e1ae37672646ff037db7dd83484

Observation 01ac0eea-2e90-47dc-b9e5-108682734923 · outbound

This paper cites S., Theocharous, G., and Ghavamzadeh, M.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation S., Theocharous, G., and Ghavamzadeh, M

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:07.914057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:16:58.954745Z digest=sha256:6cc1e8a90299808e4957ad979cc6f2f692db73bd4f710487e489c257cb4cc13e

Observation 1319aa44-ad66-41b0-a89e-f24c72174927 · outbound

This paper cites an unresolved cited work.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Unresolved cited work

Reference 72

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:17:07.764164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:16:59.070422Z digest=sha256:ca7e3223468b8d72859560626f9eab5df1cf4a30ad7d6eb9f6767aa6c0091b74

Observation 95427da5-953e-4828-a1a3-9afece42be71 · outbound

This paper cites Minimax weight and q-function learning for off-policy evaluation.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Minimax weight and q-function learning for off-policy evaluation

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:07.582206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:16:59.160195Z digest=sha256:93b355a6477456efc6bd1199102831523b3f1567f88b5857a513d6bddf2cf17e

Observation 545f29b9-a6c7-4efa-b8dc-a8ed35540f44 · outbound

This paper cites A Review of Off-Policy Evaluation in Reinforcement Learning.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation A Review of Off-Policy Evaluation in Reinforcement Learning

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:59.244313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:16:59.244313Z digest=sha256:f8f5689f3bbfa553baeeca9ca124333363fd32991079e8c61dc33c29b6271ad4

Observation 4faaad34-d831-47bf-9316-24161d54223a · outbound

This paper cites Future-dependent value-based off-policy evaluation in pomdps.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Future-dependent value-based off-policy evaluation in pomdps

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:07.402734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:16:59.334783Z digest=sha256:7fff9cfc6032ada9293cf157924b606ab4a9041e61ee0b24732cde1bb29669fc

Observation dbca4510-59cb-4e97-9bcb-3d4bd89e501e · outbound

This paper cites W., Wellner, J.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation W., Wellner, J

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:07.227633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:16:59.448099Z digest=sha256:c494906e881861eec4e120c6444b124d129a411051e99fd0f90e8b864075a30c

Observation bce7b07e-cbc8-4d16-b354-076b5e4193d5 · outbound

This paper cites Safe exploration for efficient policy evaluation and comparison.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Safe exploration for efficient policy evaluation and comparison

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:07.029178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:16:59.535998Z digest=sha256:5b12cb310c9e36148f515ca4dd82b8184979a9090eccc8c652c5a1983a78bc66

Observation a245a698-85c7-4393-ba08-1b02e7f94aea · outbound

This paper cites Blessing from Human-AI Interaction: Super Reinforcement Learning in Confounded Environments.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Blessing from Human-AI Interaction: Super Reinforcement Learning in Confounded Environments

Reference 78

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:17:02.769580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:16:59.636730Z digest=sha256:081bbaca1bc2af89ccaedd58ed3c720d467c8fa7273a2907ec0102908171416f

Observation 323967c1-c2f2-4265-87d9-3bc98a27f1c8 · outbound

This paper cites an unresolved cited work.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Unresolved cited work

Reference 79

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:17:06.847343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:16:59.714007Z digest=sha256:19cf3973e0a93eae656317629bc5d32913cde125844ba1e2a6bc55868986ba7c

Observation e4fc61c2-0fef-43a3-9468-12ebcf38ec60 · outbound

This paper cites Off-policy evaluation for tabular reinforcement learning with synthetic trajectories.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Off-policy evaluation for tabular reinforcement learning with synthetic trajectories

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:06.682959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:16:59.793756Z digest=sha256:67cf64080d548fdbee8bd061f8307debfe58bfa5e8791e5e1ed2805ab1865abe

Observation 563ceaa3-3eff-4226-94b7-4f6ddc600d0b · outbound

This paper cites Unraveling the interplay between carryover effects and reward autocorrelations in switchback experiments.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Unraveling the interplay between carryover effects and reward autocorrelations in switchback experiments

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:06.487321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:16:59.899396Z digest=sha256:0bea5c0a6b132b35c33a845c188c849a2658786f4c9e3188dbcf9424856b29aa

Observation babb5910-858b-4ed3-8b50-cc5542ed781b · outbound

This paper cites Semiparametrically efficient off-policy evaluation in linear markov decision processes.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Semiparametrically efficient off-policy evaluation in linear markov decision processes

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:06.298936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:16:59.966690Z digest=sha256:9003506f42f7548d9ed5eeb12452a5434fcd99f5d82ad4337c1e2f552756d80b

Observation ffa3eb27-0ac8-42a5-a7a2-7949cdab0a73 · outbound

This paper cites Towards optimal off-policy evaluation for reinforcement learning with marginalized importance sampling.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Towards optimal off-policy evaluation for reinforcement learning with marginalized importance sampling

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:06.117651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:17:00.089605Z digest=sha256:14d0adf668057765dda193b1cb25fdbd235f7633e6bb24c30a0ceb0763aa988c

Observation 748abfb2-46a1-4789-b90b-a28ee93d0e2f · outbound

This paper cites Towards optimal off-policy evaluation for reinforcement learning with marginalized importance sampling.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Towards optimal off-policy evaluation for reinforcement learning with marginalized importance sampling

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:05.956955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:17:00.188458Z digest=sha256:0e323aef2eed0d855947fe34f9541089f85a6feb546015867b0b4a2c7585e8d4

Observation 70ddb5a8-7bf2-4b41-83e2-ba540f4ed7b1 · outbound

This paper cites Quantile Off-Policy Evaluation via Deep Conditional Generative Learning.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Quantile Off-Policy Evaluation via Deep Conditional Generative Learning

Reference 85

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:17:02.401023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:17:00.283152Z digest=sha256:e2cd9512cb821b7e20db3c2b569f7f4bf362ce84ea370086fc35e773c3a410f2

Observation 6de15f68-f83f-4b87-b49d-e48db67e2cf9 · outbound

This paper cites An instrumental variable approach to confounded off-policy evaluation.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation An instrumental variable approach to confounded off-policy evaluation

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-07T13:17:00.380335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:17:00.380335Z digest=sha256:ce30923d4950d54ce6c585ca218a9f6fba3db3916664bf6429825beb1c244cf4

Observation 84acc115-2bbd-4e54-8447-7ca1fb3ddd45 · outbound

This paper cites and Wang, Y.-X.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation and Wang, Y.-X

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:05.792318Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:17:00.494044Z digest=sha256:f96b3d003f70a08f4099ce67c2c5270d4863cbbfa16b6700d95112db47037838

Observation e0498785-3fb3-4552-984a-f41cacf3261e · outbound

This paper cites Two-way deconfounder for off-policy evaluation in causal reinforcement learning.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Two-way deconfounder for off-policy evaluation in causal reinforcement learning

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:05.635897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:17:00.581039Z digest=sha256:c3ac36cbd64dadafd99911c3db65b1201d885a9133ec0b96a343cf3dc8b1d589

Observation 9856e75b-2403-45c1-9963-e03bb168f867 · outbound

This paper cites A., Laber, E.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation A., Laber, E

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:05.492543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:17:00.642581Z digest=sha256:bf40801a398da0cdf25aad04fce73781103fbe97e8ae105ea1ae8d094f6274d2

Observation aff6e40f-7662-485f-bf4f-6aecd7d41bad · outbound

This paper cites A., Laber, E.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation A., Laber, E

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:05.360563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:17:00.743625Z digest=sha256:a29cfafa6417bd002ba76b2ad73da1e06ee78916e552454c2ec49744bcf4829b

Observation 091369b4-6ed4-49e7-9e1b-bbe140f8acdd · outbound

This paper cites and Zhang, Y.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation and Zhang, Y

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:05.217784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:17:00.837982Z digest=sha256:98674d5383f2ec7d1452584411bf27a64aea9a714345d1dee822a3b9499867a5

Observation 817f2570-b3f1-411c-8418-e5bb063fee49 · outbound

This paper cites B., and Kosorok, M.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation B., and Kosorok, M

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:05.066644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:17:00.918164Z digest=sha256:c9e78350011afa90949d46703543d99b203c904b066ab33414a14ba337be18b0

Observation de491975-afd8-4969-9253-61c05ba1bdab · outbound

This paper cites Distributional Shift-Aware Off-Policy Interval Estimation: A Unified Error Quantification Framework.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Distributional Shift-Aware Off-Policy Interval Estimation: A Unified Error Quantification Framework

Reference 93

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:17:01.728469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:17:00.981742Z digest=sha256:360fa89bfedd4a57298d03b870cbe87ac953feaad2fa511c26220804085e74c6

Observation 8c1e2da9-074b-421d-9278-2c4e3909cf93 · outbound

This paper cites Robust offline reinforcement learning with heavy-tailed rewards.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Robust offline reinforcement learning with heavy-tailed rewards

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:04.924461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T13:17:01.051859Z digest=sha256:f60dab6fb76382d6cf35ec6426d1ddd4bc3615c7d2a5f6520ab26d106000b07a

Observation 34152c82-94af-4f4e-a037-fa18debd6725 · outbound

This paper cites write newline.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation write newline

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-07T13:17:01.124208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:17:01.124208Z digest=sha256:66d6112dd490c52b26e9d28b0e5ef06832f0d707a274592b646b33a7996a7da4

Pith citing papers

No inbound Pith citation observations are available.