Pith. sign in

Paper Citation Record · LEDGER

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation

As of 8 August 2026, this Paper Citation Record lists 95 of 95 outbound references and 0 inbound Pith citation observations for arXiv:2505.22492.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.22492 v1

Coverage vector

measured 95 of 95 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:17:01.124208Z

measured 95 of 95 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

95 of 95 outbound references displayed

  • verified exact7
  • verified fuzzy50
  • unresolved37
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation fe970c9f-0548-4569-aaa3-4eab2ff15df9 · outbound

This paper cites an unresolved cited work.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:52.903422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:16:52.903422Z digest=sha256:47c720ecc55330ac68e2456f9388b4bdc2ff55270857de5e2870eb138c9f42ce

Observation b62aa258-e78b-4938-bb97-799ce524f0ef · outbound

This paper cites and Kallus, N.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation and Kallus, N

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:53.021013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:16:53.021013Z digest=sha256:8ac96c560c55d43686aefb7583e1af7b4eba16fed83f14de5bee2049b6fc2d8e

Observation 3b741a11-b4dc-4569-b70d-945c5839b408 · outbound

This paper cites Off-policy evaluation in doubly inhomogeneous environments.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Off-policy evaluation in doubly inhomogeneous environments

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:53.103572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:16:53.103572Z digest=sha256:2ba9c0fcf2794ee9c989f3e162284b5f6b4ac658b9fb990b7d2f3c4c4a40176f

Observation 440371c0-981b-4670-bec6-f0aacf55fdae · outbound

This paper cites More efficient off-policy evaluation through regularized targeted learning.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation More efficient off-policy evaluation through regularized targeted learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:53.186263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:16:53.186263Z digest=sha256:78a46936a10ce144606de11834f2dc22d49c576f34cfb63f5a9b503a423323f8

Observation 342ab1ad-17f1-455a-9f63-4a19cc565efe · outbound

This paper cites an unresolved cited work.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:53.264406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:16:53.264406Z digest=sha256:06758a6b95e7e3d33a1407e9e30e49ad81b9be92317789e1fefa5cff7418d80b

Observation 552f657a-6eda-4337-a6d5-a218da5e2d98 · outbound

This paper cites OpenAI Gym.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation OpenAI Gym

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:53.388701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:16:53.388701Z digest=sha256:6c9fc9d78f5626553c17c4cb899e39b257133e952cadb4c8d34f963f18d85334

Observation c6c5e743-b5b3-4cbb-aebf-ad81d8d95b29 · outbound

This paper cites and Zhou, A.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation and Zhou, A

Reference 7

Resolution
verified exact
raw_fallback, observed 2026-08-07T13:17:04.709883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:16:53.496289Z digest=sha256:3c871a398271d3b0e23a6f79d16a3c8e768b7aeadc87d946e452af96d7ccde69

Observation 8a06d0b8-b9d7-40c3-916f-86b10aa7c50a · outbound

This paper cites Structured Difference-of-Q via Orthogonal Learning.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Structured Difference-of-Q via Orthogonal Learning

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:17:04.447238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:16:53.610687Z digest=sha256:58074c18d20be23503c0a99ce4a146d45236e88effa3543e38263de97187c246

Observation 67a52a5d-4608-4582-a6c2-bd0f88c6a259 · outbound

This paper cites and Berger, R.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation and Berger, R

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:53.686598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:16:53.686598Z digest=sha256:a355fc3244eaf036f3146dbf1d9f195a9ff93d2587159900d2f70f6e05f33faa

Observation bc2ae633-0ad7-44ca-a0db-07a0658b0efa · outbound

This paper cites and Li, L.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation and Li, L

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:53.794136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:16:53.794136Z digest=sha256:d24f804783482353372edb51dd4f5eadf82c405b01734971e07a05497aebd455

Observation 49de3aaf-d1f8-4dbd-a4a4-759ad9b84b90 · outbound

This paper cites and Jiang, N.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation and Jiang, N

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:53.924369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:16:53.924369Z digest=sha256:54b23948b70db5bcce592e5e481f8b3d086bfe3a86ef2206a6cd3ebb3417f6c9

Observation b8a8b783-825e-4208-95fc-5d209ea87241 · outbound

This paper cites and Qi, Z.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation and Qi, Z

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:54.023087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:16:54.023087Z digest=sha256:412068e6a24e419e1c2ee9a3c6de872556f723569912c8f65ab8423e0b2e5697

Observation fca476a0-cfa0-4d92-ab2b-9bf5fb2ca2f7 · outbound

This paper cites Gaussian approximation of suprema of empirical processes.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Gaussian approximation of suprema of empirical processes

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:54.133862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:16:54.133862Z digest=sha256:d1fbe59b36e216a5787e01e500553f4bd5ced2a44cc72b9a1f364b9686bac408

Observation 52b4bb87-f570-4feb-84e1-1403a8ab136a · outbound

This paper cites Double/debiased machine learning for treatment and structural parameters.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Double/debiased machine learning for treatment and structural parameters

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:54.246149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:16:54.246149Z digest=sha256:bc53cd180b8f982fd6b5a3a63146d5c919c4fe3d58d5f6388fb99226cd8d48a9

Observation db63a9a8-f309-47aa-acac-bd6bb2d1c379 · outbound

This paper cites Coindice: Off-policy confidence interval estimation.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Coindice: Off-policy confidence interval estimation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:54.339217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:16:54.339217Z digest=sha256:6beefae3106697cbbff8435a24a479a76cde540646231d45c0c23847bf931b6f

Observation 0a9c8274-3776-490b-a0ed-558411d6bf08 · outbound

This paper cites Doubly Robust Policy Evaluation and Optimization.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Doubly Robust Policy Evaluation and Optimization

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:54.432077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:16:54.432077Z digest=sha256:a34aadd1422b89a4b9b389f9b4d1f0f2fb3834af2fa0e6599b76c7a4172ca59e

Observation 8446019f-3832-4df8-9d51-6b1368b142b9 · outbound

This paper cites A theoretical analysis of deep q-learning.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation A theoretical analysis of deep q-learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:54.524675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:16:54.524675Z digest=sha256:6aa486b6aa24ffa9d2ddab353cf14331ad83363258d4eaaa92245ab6c7562625

Observation 25caa724-aca7-4a07-b9f0-dda5c47cd091 · outbound

This paper cites More Robust Doubly Robust Off-policy Evaluation.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation More Robust Doubly Robust Off-policy Evaluation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:54.617714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:16:54.617714Z digest=sha256:534e37ce8dfdf02a5d614c9bd49aad12f1f660a94820e4c086f9f9ce44904962

Observation 69d160b1-511b-4984-a4d4-ec8006edc122 · outbound

This paper cites Accountable off-policy evaluation with kernel B ellman statistics.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Accountable off-policy evaluation with kernel B ellman statistics

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:54.732514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:16:54.732514Z digest=sha256:64208d278b78012a87423fc5a7798bb36e076b941b7328d678027af82c22c547

Observation 1d2f3fc2-4499-45b8-ae5c-0061bd1e4779 · outbound

This paper cites Combining parametric and nonparametric models for off-policy evaluation.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Combining parametric and nonparametric models for off-policy evaluation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:54.828722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:16:54.828722Z digest=sha256:e1c1f05176dd649b250cd84d1274ecd70cb3c8bf2d79037188a2858a8b33294c

Observation 3e674bf6-aede-4609-a97b-8498caf0fb30 · outbound

This paper cites D., Thomas, P.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation D., Thomas, P

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:15.124519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:16:54.910774Z digest=sha256:37adb378e577589790f287ee6eaca0f93f7fc294e8a75698219f96f37a534408

Observation 238d6f10-29ef-4d5b-b3bf-b06331b78393 · outbound

This paper cites Importance sampling policy evaluation with an estimated behavior policy.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Importance sampling policy evaluation with an estimated behavior policy

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:14.980638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:16:54.995797Z digest=sha256:f969b13cf4cd4b63a481faa71588a3c85661fb4557499242d9e8e495089f11cd

Observation 715be3ee-ce98-4c84-ac12-c5babdbdda54 · outbound

This paper cites P., Thomas, P.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation P., Thomas, P

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:14.829766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:16:55.071207Z digest=sha256:22d1e177de00e58421fcb9366f1a0693eb3b846765bb0cc102d6753f849e52cb

Observation 5db9c47a-84a6-4898-9dc1-e9c7a05cdcd1 · outbound

This paper cites P., Niekum, S., and Stone, P.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation P., Niekum, S., and Stone, P

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:14.679690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:16:55.196586Z digest=sha256:adb982cb163000d851e07ca38a19db0eaab2a9aac8289d8e1a2d14c5bf5d71fb

Observation 9c90fb45-6232-4e37-b518-1c03af849c8c · outbound

This paper cites Bootstrapping fitted q-evaluation for off-policy inference.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Bootstrapping fitted q-evaluation for off-policy inference

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:14.525737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:16:55.317248Z digest=sha256:d1c435bfd369056091191c350f04d76e4af4a2f4921f3df2a3c7a96d1b5bc65c

Observation 891b15a5-8bf1-45dc-abba-88a38f30a1b1 · outbound

This paper cites Importance sampling via the estimated sampler.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Importance sampling via the estimated sampler

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:14.384125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:16:55.387170Z digest=sha256:6965083618133048f67744f209d0b35353908d91ab6f64d246b11b294f9b9031

Observation 2ad6d719-9b29-446e-a5d2-f9e1a16fc9ca · outbound

This paper cites W., and Ridder, G.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation W., and Ridder, G

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:14.222790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:16:55.459117Z digest=sha256:39d7d10c2f73449923499330565acb0f61d7ee2dd2cd63505c4d34ea61f44d6c

Observation db7f3592-4268-4450-ba42-a0663dfe4af2 · outbound

This paper cites and Wager, S.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation and Wager, S

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:55.534857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:16:55.534857Z digest=sha256:a8a2c983e8ad6a50f872477b7d798371bc24ceded61bd32583c7f365c7d7a885

Observation ff3d687f-0b80-4dd4-a4d7-db5b0d93866d · outbound

This paper cites and Li, L.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation and Li, L

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:14.029859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:16:55.616369Z digest=sha256:12aed742fae2dbbea46c0bac751c25e77fed8c1005b9b833bbe9fde7518c4806

Observation a42c0dd8-5648-4354-b424-84d1a4e77faa · outbound

This paper cites and Uehara, M.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation and Uehara, M

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:13.838362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:16:55.716241Z digest=sha256:d688863e9b8de57d047242ce0ea23769ded6805d120e8213df1c57b2b83747a6

Observation d404d600-7de2-4150-946b-5307bac95f8c · outbound

This paper cites and Uehara, M.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation and Uehara, M

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:13.626828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:16:55.793014Z digest=sha256:9262a92fdfdf3ceb7012bdb0400f0b4f698f8753aea36a5efd44aae317077428

Observation 6411d0f2-b200-46ad-b612-6dd15b17e52c · outbound

This paper cites and Zhou, A.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation and Zhou, A

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:13.420849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:16:55.872168Z digest=sha256:d3247a9176af56549ad4566e719b68d1da7239dafc6fdfacaa45f813f38c9c56

Observation 63221bff-3c62-45ab-9178-79969858e05b · outbound

This paper cites an unresolved cited work.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:17:13.160576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:16:55.932696Z digest=sha256:3eedf2f80ae7323d70e461b81d8cda8848f7d0da732a917d19b7ab4734f1994f

Observation 5a9cc8b5-daf7-428d-9acd-b4d241a9daf0 · outbound

This paper cites Batch policy learning under constraints.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Batch policy learning under constraints

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:12.937701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:16:56.017006Z digest=sha256:949a4380312f39487416d9e5835d377015fc0df4f2e8cdd249a5fed8c48e6f6c

Observation 1924c1ed-bc7b-4f1b-98f6-816e702217d1 · outbound

This paper cites Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:56.144780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:16:56.144780Z digest=sha256:4042fd36874170257c4df27c6e58e96c76acaa5c8b21199008b9074b01d92be4

Observation 977884f7-69c5-4e80-aa98-3b38ff7c87f1 · outbound

This paper cites High-probability sample complexities for policy evaluation with linear function approximation.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation High-probability sample complexities for policy evaluation with linear function approximation

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:17:04.262424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:16:56.248494Z digest=sha256:1c11e0cc829db5b16d0dcd96f4d1926f3497ae63db192482c72c75d434c68dc1

Observation 38ccf486-2bf4-47ce-91d8-e2cb5ef3fff3 · outbound

This paper cites Off-policy estimation of long-term average outcomes with applications to mobile health.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Off-policy estimation of long-term average outcomes with applications to mobile health

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:12.773972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:16:56.359148Z digest=sha256:7ec78386077bd787941738943cc2f998d51621f27b9f802373d8fb326848df7a

Observation 67f38c91-6c96-4360-9f51-f15cbc3b435a · outbound

This paper cites an unresolved cited work.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:17:12.624851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:16:56.466433Z digest=sha256:e27172e1185c9e6a51011e587e32d85bc67aa0cb86ce333128f676abca542656

Observation 03b93b2d-2ad1-411f-9a43-8b252edd6bf6 · outbound

This paper cites Breaking the curse of horizon: infinite-horizon off-policy estimation.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Breaking the curse of horizon: infinite-horizon off-policy estimation

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:12.478362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:16:56.565112Z digest=sha256:068f983dfce37c2fe30c3eef6783453a4306f7287cf7a7897bfe755db880a39b

Observation b1b84187-62f8-49de-bd51-6289711aa20c · outbound

This paper cites and Zhang, S.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation and Zhang, S

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:12.254767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:16:56.671738Z digest=sha256:95a018e567add30051685266615735410ce8586f6168184db62aed78ff1de6df

Observation 20212785-c506-4558-a002-57494379b523 · outbound

This paper cites Doubly Optimal Policy Evaluation for Reinforcement Learning.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Doubly Optimal Policy Evaluation for Reinforcement Learning

Reference 41

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T13:17:03.912586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:16:56.740396Z digest=sha256:b262a66d4dec0b3ff75b00834b01f7343a56ea5532552a28d028079dc714db85

Observation b6029ff5-5b9f-48b3-9947-4d1d444ba384 · outbound

This paper cites Online Estimation and Inference for Robust Policy Evaluation in Reinforcement Learning.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Online Estimation and Inference for Robust Policy Evaluation in Reinforcement Learning

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:17:03.551774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:16:56.834090Z digest=sha256:0393dcbd384254d4773dd7e87cbdcf92429f41ed0da3a7b93ede19601d9fbcbf

Observation 1ed53d7c-3231-44c0-8f6b-0feeedbf69dc · outbound

This paper cites J., Laber, E.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation J., Laber, E

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:56.946901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:16:56.946901Z digest=sha256:f43fa3df893e096c37d0f29f5ca806850aaec52122947f84cf0da7c9906bd416

Observation 548001c1-8ac1-463d-adbe-e8cb79d621a4 · outbound

This paper cites P., and Nowak, R.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation P., and Nowak, R

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:12.077014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:16:57.007129Z digest=sha256:b856e191123bc5c67b248d59d28e307eae15e3d9d3eef668a34ec97678afaecf

Observation e038af15-6e2e-40e6-8d48-d63427dd26d3 · outbound

This paper cites A., van der Laan, M.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation A., van der Laan, M

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:11.923588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:16:57.074159Z digest=sha256:08d3a583eb2e081425ce81d6aaeab5a6b8a74b37d3aa032143fe8bd051e40538

Observation cb654943-6911-4027-b699-509c2785a743 · outbound

This paper cites Dualdice: Behavior-agnostic estimation of discounted stationary distribution corrections.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Dualdice: Behavior-agnostic estimation of discounted stationary distribution corrections

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:11.691191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:16:57.135333Z digest=sha256:41b0467ea9c6e88a1dd5218b1f5abdf428d347ca4ff8136f8eaf335fc20ec2bf

Observation 3bbf4a0c-7e88-437c-ba2e-7be99f60802d · outbound

This paper cites A Spectral Approach to Off-Policy Evaluation for POMDPs.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation A Spectral Approach to Off-Policy Evaluation for POMDPs

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:57.195687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:16:57.195687Z digest=sha256:8d5b09062c3e7136c104bc5a36ff448693fe6a9e6be910f3f66497ba5b492a82

Observation 877ba37c-7916-41a1-bbda-3021e7f63f2b · outbound

This paper cites Off-policy policy evaluation for sequential decisions under unobserved confounding.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Off-policy policy evaluation for sequential decisions under unobserved confounding

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:11.540247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:16:57.246652Z digest=sha256:ffd7ac484eb1d69fb7b3d90a6e7bb287f280cf72af99f1e67e4a0f2b7d01bf7b

Observation fe049d9e-3dad-43a7-876b-39568be90b97 · outbound

This paper cites K., Hsieh, F., and Robins, J.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation K., Hsieh, F., and Robins, J

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:11.289965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:16:57.316362Z digest=sha256:9f295c7dcaa27b5f59c65a03b2e8b97ea1cddade988866e86f1a4b464b689cb7

Observation 220033ba-5dc0-445d-b8e8-855b2dd8aecf · outbound

This paper cites Training language models to follow instructions with human feedback.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Training language models to follow instructions with human feedback

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:57.362862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:16:57.362862Z digest=sha256:66597c4cc4019c3768e03d0802e1aab3944f673ffd72ad1ee3a3be45af7a37ff

Observation 7e730401-bf8b-41ff-a17e-a81d69418f46 · outbound

This paper cites S., and Singh, S.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation S., and Singh, S

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:11.104301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:16:57.449604Z digest=sha256:c1220577f81f1bcb2163d5a0ef8140436b1bdd0f34396eab4a9119123d7e53d7

Observation c5644559-bd83-4111-bc92-4c2224a768b3 · outbound

This paper cites an unresolved cited work.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Unresolved cited work

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:57.519872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:16:57.519872Z digest=sha256:fa8638a1ac86caf9ce411f7a79188e9f9d2040a6e0f5e621b64ccbc5b69fdbd6

Observation eb28c104-9602-4e69-8561-9b0715d8a30c · outbound

This paper cites an unresolved cited work.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:17:10.808173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:16:57.582139Z digest=sha256:9b4d2a62507872cef777682813411b68d3f453be40b43e9f807bdaa278e45af9

Observation 4695e73b-5847-4e3f-82b9-e682b16dc01b · outbound

This paper cites Conditional importance sampling for off-policy learning.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Conditional importance sampling for off-policy learning

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:10.692678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:16:57.632221Z digest=sha256:201e930a721fc15ec2e579c5c8241924aa56f26f9471a7d82d4bbf133822a327

Observation a5078da9-a7fa-4985-85d6-b91f6d93ef28 · outbound

This paper cites G., and Dabney, W.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation G., and Dabney, W

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:10.457993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:16:57.718028Z digest=sha256:7e64ad6fcce8fbc24836e5ed1bdb07fd3de02761705bceef647c262cb09f0b08

Observation b9e6c4a0-196f-4998-8a2a-b23b71887564 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Proximal Policy Optimization Algorithms

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:57.801762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:16:57.801762Z digest=sha256:f8e086464eff8f2cf895f3dc04aefa8e9e32966b74acb93c46472ff9e549c271

Observation 7e7a7233-70a3-4e87-b622-49670b846ed1 · outbound

This paper cites Estimating the dimension of a model.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Estimating the dimension of a model

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:10.220070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:16:57.866226Z digest=sha256:555a00efe3980781bac0041ee7dd2d557e316191b09fea7305fe1c39b87cdd50

Observation 850e99cd-29f6-4007-a1c9-33882eea3cf1 · outbound

This paper cites and Ben-David, S.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation and Ben-David, S

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:57.934021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:16:57.934021Z digest=sha256:4c8a62c225236125cda9c29c916420e2cc284dceb056a080b38e0647e2555855

Observation 9afe5e2f-6267-448a-b52c-d34ce187313a · outbound

This paper cites Mathematical Statistics.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Mathematical Statistics

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:58.022585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:16:58.022585Z digest=sha256:a694557fdd8aef078efde0e899ad5f990d727521bb03e4ccf689260731b655f0

Observation a5a578f7-f424-47bd-9594-7bb5521b9d41 · outbound

This paper cites On methods of sieves and penalization.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation On methods of sieves and penalization

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:10.002895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:16:58.081016Z digest=sha256:94305bb34e3fb1ed4669200227f7d473b4e33eaeda33296b13a82d842bbc8c4c

Observation 067892b6-6dcf-460f-93c9-2e9dfe291a08 · outbound

This paper cites Deeply-debiased off-policy interval estimation.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Deeply-debiased off-policy interval estimation

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:09.855972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:16:58.143606Z digest=sha256:6fdb01b80df95fab2f83025b686c2cfefcdc8adb0c4e3e26064ce68fc0544786

Observation 3587ffd9-3294-4bd0-b23e-85ae9726279a · outbound

This paper cites A minimax learning approach to off-policy evaluation in confounded partially observable markov decision processes.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation A minimax learning approach to off-policy evaluation in confounded partially observable markov decision processes

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:09.668572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:16:58.211350Z digest=sha256:1c46397337a7df7ccba2cb2a77963264b900675e0c173d7738de76985091220d

Observation 337a06a4-da2c-4071-a7f1-ce0b1005ff5f · outbound

This paper cites Statistical inference of the value function for reinforcement learning in infinite-horizon settings.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Statistical inference of the value function for reinforcement learning in infinite-horizon settings

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:09.489721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:16:58.281999Z digest=sha256:53d9485f3572d9040758316ce99a4e31ad2665d529ce0fb113cc0872406e8716

Observation d688c1de-2597-498c-a6cf-596f4ba2351b · outbound

This paper cites Dynamic causal effects evaluation in a/b testing with a reinforcement learning framework.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Dynamic causal effects evaluation in a/b testing with a reinforcement learning framework

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:09.222749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:16:58.378112Z digest=sha256:fa0bbce62fede61b5985d2bbd8d62cf8b71ab667d053e4ca9d9d1d4b27dbd55c

Observation d1a16f9e-e59f-46b5-a50d-514496496e92 · outbound

This paper cites Off-policy confidence interval estimation with confounded markov decision process.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Off-policy confidence interval estimation with confounded markov decision process

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:08.987934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:16:58.470043Z digest=sha256:ca642a4d64e691f600ebc0cde5e91b1371821576fc4cd251af962a7e83dba04d

Observation 61258582-d81d-4811-8a87-20401a18103d · outbound

This paper cites ARMA-Design: Optimal Treatment Allocation Strategies for A/B Testing in Partially Observable Time Series Experiments.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation ARMA-Design: Optimal Treatment Allocation Strategies for A/B Testing in Partially Observable Time Series Experiments

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:58.476262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:16:58.476262Z digest=sha256:4d8c088c12f99fb0b7e0f8e97307857b2e6ceabac4ca183e02a19e4667a53cff

Observation 02bd134f-5140-496e-867d-2a6dc4b7aea0 · outbound

This paper cites S., Szepesv \'a ri, C., and Maei, H.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation S., Szepesv \'a ri, C., and Maei, H

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:08.811401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:16:58.553538Z digest=sha256:9d00e3b73dd12aa4b7865cd35bb14f11286b7851748124c8cb33f8ba8ea02df3

Observation a266e348-4d60-4a95-a6c3-91f20e10358b · outbound

This paper cites Doubly robust bias reduction in infinite horizon off-policy estimation.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Doubly robust bias reduction in infinite horizon off-policy estimation

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:08.583530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:16:58.648961Z digest=sha256:b4875d874e52883d9857f444d90dfb09f3b011fe2837b0c4e702257426e12151

Observation bd7a0f40-a240-49ca-9b93-4c4ed5bba7b4 · outbound

This paper cites Off-policy evaluation in partially observable environments.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Off-policy evaluation in partially observable environments

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:08.354914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:16:58.740428Z digest=sha256:8af6c04889b4c932d681ba744f3768b27005d726374117ed8598b069952012c9

Observation 6b8460eb-5193-457b-a1b5-abfd2a007cdb · outbound

This paper cites an unresolved cited work.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Unresolved cited work

Reference 70

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:17:08.090302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:16:58.839634Z digest=sha256:91efebdb45b5c94c0f9adcf3ba72e7417d7e2781a423936a0aaae9676f023a3f

Observation 01ac0eea-2e90-47dc-b9e5-108682734923 · outbound

This paper cites S., Theocharous, G., and Ghavamzadeh, M.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation S., Theocharous, G., and Ghavamzadeh, M

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:07.914057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:16:58.954745Z digest=sha256:2451e89428d4ef471db1dc02ea91cb1ed281babd00edcc8da7ab5f3a20f96183

Observation 1319aa44-ad66-41b0-a89e-f24c72174927 · outbound

This paper cites an unresolved cited work.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Unresolved cited work

Reference 72

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:17:07.764164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:16:59.070422Z digest=sha256:9c1cfb80b8d732bcc0b21483e5b0fef7470e51b37770aa244a6590d74496066f

Observation 95427da5-953e-4828-a1a3-9afece42be71 · outbound

This paper cites Minimax weight and q-function learning for off-policy evaluation.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Minimax weight and q-function learning for off-policy evaluation

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:07.582206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:16:59.160195Z digest=sha256:9967429ea391a8a4e7344bf4bfc1aa95e833c0e2cf0f082aaba3eedf90a6f53b

Observation 545f29b9-a6c7-4efa-b8dc-a8ed35540f44 · outbound

This paper cites A Review of Off-Policy Evaluation in Reinforcement Learning.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation A Review of Off-Policy Evaluation in Reinforcement Learning

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T13:16:59.244313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:16:59.244313Z digest=sha256:5e09e2c05b63a1354928043a2668deae62522b874b78dc42169e1aa5270fc613

Observation 4faaad34-d831-47bf-9316-24161d54223a · outbound

This paper cites Future-dependent value-based off-policy evaluation in pomdps.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Future-dependent value-based off-policy evaluation in pomdps

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:07.402734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:16:59.334783Z digest=sha256:e0ee5242b6c0c27b1df611382e25d282f485333494c3106b83b54827ab8c0c92

Observation dbca4510-59cb-4e97-9bcb-3d4bd89e501e · outbound

This paper cites W., Wellner, J.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation W., Wellner, J

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:07.227633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:16:59.448099Z digest=sha256:8932962ddf2c99c8e42d40bf9478be76884204c3b4a2f71280030fc098106f28

Observation bce7b07e-cbc8-4d16-b354-076b5e4193d5 · outbound

This paper cites Safe exploration for efficient policy evaluation and comparison.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Safe exploration for efficient policy evaluation and comparison

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:07.029178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:16:59.535998Z digest=sha256:bf157fd6857fc24630f68869598069baacaa3bde8c71f831ac437a0119a848ba

Observation a245a698-85c7-4393-ba08-1b02e7f94aea · outbound

This paper cites Blessing from Human-AI Interaction: Super Reinforcement Learning in Confounded Environments.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Blessing from Human-AI Interaction: Super Reinforcement Learning in Confounded Environments

Reference 78

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:17:02.769580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:16:59.636730Z digest=sha256:c2b446e688afa42a93305eacdaaa1725e13b329572c1d710d8d68c772fc4b1e3

Observation 323967c1-c2f2-4265-87d9-3bc98a27f1c8 · outbound

This paper cites an unresolved cited work.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Unresolved cited work

Reference 79

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:17:06.847343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:16:59.714007Z digest=sha256:ebdae853995c069c5f92d2f7682d0d9a7076723c221754dd8d4db17221b2c125

Observation e4fc61c2-0fef-43a3-9468-12ebcf38ec60 · outbound

This paper cites Off-policy evaluation for tabular reinforcement learning with synthetic trajectories.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Off-policy evaluation for tabular reinforcement learning with synthetic trajectories

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:06.682959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:16:59.793756Z digest=sha256:5b7a1a58557c631ad0d74c22e347f707e2c217eb722b7b4677257f320d4b1bd0

Observation 563ceaa3-3eff-4226-94b7-4f6ddc600d0b · outbound

This paper cites Unraveling the interplay between carryover effects and reward autocorrelations in switchback experiments.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Unraveling the interplay between carryover effects and reward autocorrelations in switchback experiments

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:06.487321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:16:59.899396Z digest=sha256:a85bfa4bb3e89c0e314ad5ac76d7e058c6332c94f80b9e4cf948f77293f9806d

Observation babb5910-858b-4ed3-8b50-cc5542ed781b · outbound

This paper cites Semiparametrically efficient off-policy evaluation in linear markov decision processes.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Semiparametrically efficient off-policy evaluation in linear markov decision processes

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:06.298936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:16:59.966690Z digest=sha256:434201a6e49adcdbc8d1599bedee9d2091446253cddfe76927019f6d0cf4a3fb

Observation ffa3eb27-0ac8-42a5-a7a2-7949cdab0a73 · outbound

This paper cites Towards optimal off-policy evaluation for reinforcement learning with marginalized importance sampling.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Towards optimal off-policy evaluation for reinforcement learning with marginalized importance sampling

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:06.117651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:17:00.089605Z digest=sha256:631e802be92ce3e869a0f5574c00d44a45234fd7f605a9f93a109ed20ac82e72

Observation 748abfb2-46a1-4789-b90b-a28ee93d0e2f · outbound

This paper cites Towards optimal off-policy evaluation for reinforcement learning with marginalized importance sampling.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Towards optimal off-policy evaluation for reinforcement learning with marginalized importance sampling

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:05.956955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:17:00.188458Z digest=sha256:6c81a796ccb3ac940dcd77fddbc3661ab2532dc499507d972577ab0f9b21b59b

Observation 70ddb5a8-7bf2-4b41-83e2-ba540f4ed7b1 · outbound

This paper cites Quantile Off-Policy Evaluation via Deep Conditional Generative Learning.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Quantile Off-Policy Evaluation via Deep Conditional Generative Learning

Reference 85

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:17:02.401023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:17:00.283152Z digest=sha256:4c3ed3817261f71b241f0fb16e758e3772a832001f4a5ce7c31f346e074c89b9

Observation 6de15f68-f83f-4b87-b49d-e48db67e2cf9 · outbound

This paper cites An instrumental variable approach to confounded off-policy evaluation.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation An instrumental variable approach to confounded off-policy evaluation

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-07T13:17:00.380335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:17:00.380335Z digest=sha256:130e5f030801e7c0da1f7a28b2b6dce10778af04558d141476e6e12df04e6a5d

Observation 84acc115-2bbd-4e54-8447-7ca1fb3ddd45 · outbound

This paper cites and Wang, Y.-X.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation and Wang, Y.-X

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:05.792318Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:17:00.494044Z digest=sha256:d15a1dec78145e5d2a270ead7a32bd86b567264f42981615f13975a3e3d41177

Observation e0498785-3fb3-4552-984a-f41cacf3261e · outbound

This paper cites Two-way deconfounder for off-policy evaluation in causal reinforcement learning.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Two-way deconfounder for off-policy evaluation in causal reinforcement learning

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:05.635897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:17:00.581039Z digest=sha256:23cd1e8ded2482a9a86d15cda1ee3363ed674092cc6cd068dceb625c2f881d8b

Observation 9856e75b-2403-45c1-9963-e03bb168f867 · outbound

This paper cites A., Laber, E.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation A., Laber, E

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:05.492543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:17:00.642581Z digest=sha256:679b89ccebc03d4713da4bafbdf4618dba69475b4eddcaeaafa7bd04ca893c9b

Observation aff6e40f-7662-485f-bf4f-6aecd7d41bad · outbound

This paper cites A., Laber, E.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation A., Laber, E

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:05.360563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:17:00.743625Z digest=sha256:32ba821a3dbb2bc3b77302889ec3c4983f5dc7680737654d907f607f5b715379

Observation 091369b4-6ed4-49e7-9e1b-bbe140f8acdd · outbound

This paper cites and Zhang, Y.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation and Zhang, Y

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:05.217784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:17:00.837982Z digest=sha256:d44776a25aa1539d8801b3abf0c1fce04bc9476818c22f188fc627f2d534f4d1

Observation 817f2570-b3f1-411c-8418-e5bb063fee49 · outbound

This paper cites B., and Kosorok, M.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation B., and Kosorok, M

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:05.066644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:17:00.918164Z digest=sha256:a77f5eca06997ff2cb4df82c7b5693b7c5a3b56fb5be16952abe4ed5d3da23aa

Observation de491975-afd8-4969-9253-61c05ba1bdab · outbound

This paper cites Distributional Shift-Aware Off-Policy Interval Estimation: A Unified Error Quantification Framework.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Distributional Shift-Aware Off-Policy Interval Estimation: A Unified Error Quantification Framework

Reference 93

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:17:01.728469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:17:00.981742Z digest=sha256:813a741915cfe723ebc518acccb82d75df53f503c99e4ed4f8d68e715d372f2f

Observation 8c1e2da9-074b-421d-9278-2c4e3909cf93 · outbound

This paper cites Robust offline reinforcement learning with heavy-tailed rewards.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation Robust offline reinforcement learning with heavy-tailed rewards

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:17:04.924461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:17:01.051859Z digest=sha256:9f1209db611513ee89232e52de6e13cb6be131622a47dd09ff1c4143e86e0dde

Observation 34152c82-94af-4f4e-a037-fa18debd6725 · outbound

This paper cites write newline.

Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy Evaluation write newline

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-07T13:17:01.124208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:17:01.124208Z digest=sha256:477a1e6ccda85edb2c73d34d9e9759152594836b479a7a37e9660458139e1de8

Pith citing papers

No inbound Pith citation observations are available.