Pith. sign in

Paper Citation Record · LEDGER

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes

As of 20 August 2026, this Paper Citation Record lists 71 of 71 outbound references and 0 inbound Pith citation observations for arXiv:1908.08526.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1908.08526 v3

Coverage vector

measured 71 of 71 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-14T11:46:30.168252Z

measured 71 of 71 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

71 of 71 outbound references displayed

  • verified exact2
  • verified fuzzy33
  • unresolved36
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2f7b9a31-3e26-4e61-97aa-6095eeb23fe4 · outbound

This paper cites Ai and X.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Ai and X

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:46:32.310499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-14T11:46:29.634294Z digest=sha256:1747f85617ee5fe52cbeb9bfb7feabc28ec561af8d3146b57b624fb7dfedb944

Observation 6a09b9d4-0a22-46dc-b54f-f4599ea5e1dc · outbound

This paper cites Ai and X.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Ai and X

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:46:32.273051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-14T11:46:29.641706Z digest=sha256:772aef756904f509a737de03b64172b689ecf6e6fd8ed6ccd5bd8468ca1f7a13

Observation 9fa8affe-7a93-470e-9298-b2c917d49aa1 · outbound

This paper cites Antos, C.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Antos, C

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:46:32.232092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-14T11:46:29.648353Z digest=sha256:63d2faebfda0a270cbb0be4f55de84ebe393ff3bbed6821ec34e59c8b5bd8223

Observation d818f678-21e1-4dc0-b942-8fb649b593cd · outbound

This paper cites an unresolved cited work.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-14T11:46:32.190388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-14T11:46:29.656662Z digest=sha256:3562f44b3364493a7a5b9f4047699039475bda9ef9ce4e5fb28e609481ad994d

Observation d6edfed1-f4db-4313-b3e2-5fcfc0cbef8f · outbound

This paper cites Benkeser and M.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Benkeser and M

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:46:32.154243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-14T11:46:29.662248Z digest=sha256:993d33e0027e5c5af1ec919de7521fd411fc2b8e2a698b9c6115f03a1edfe04c

Observation 09e5b3f6-384d-40a6-96fc-8ccab2344dc3 · outbound

This paper cites an unresolved cited work.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-14T11:46:32.120086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-14T11:46:29.668735Z digest=sha256:8ac5ca2aad030b1bbd137ff56afd0e23465a3a7e92174506c7acfddb9750eaf9

Observation de1af119-467d-4b1a-a29c-254f140b8e54 · outbound

This paper cites Bibaut, I.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Bibaut, I

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:46:32.088567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-14T11:46:29.682231Z digest=sha256:ec40b3d186020ea9dcc2224813db2ec385b8c72bd3116d7f748a5765577cb21e

Observation 1e23951c-9ce7-4762-8e03-a682dda75c17 · outbound

This paper cites Fast rates for empirical risk minimization over c\`adl\`ag functions with bounded sectional variation norm.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Fast rates for empirical risk minimization over c\`adl\`ag functions with bounded sectional variation norm

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-14T11:46:29.690268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T11:46:29.690268Z digest=sha256:80e85ed31ca99aca8f80febc373054142e6fc1a5671281652146f7a9f814376f

Observation a43cee4a-eafd-4766-9a3c-8e4461c1c3c6 · outbound

This paper cites an unresolved cited work.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-14T11:46:29.700225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T11:46:29.700225Z digest=sha256:4930b84861eb9e005c1276658d2413183d51100c92fbec512219b759fc096e56

Observation 4dceaf47-f150-4468-b2bd-37098574e674 · outbound

This paper cites an unresolved cited work.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-14T11:46:32.030080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-14T11:46:29.708272Z digest=sha256:b45c35f3c120e3d2eee634cd7f6b0812ffc49d04b1e3dc68cfe9b9e5b6491600

Observation 4e9e4388-96a0-43dd-98b3-2b0404f82f21 · outbound

This paper cites OpenAI Gym.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes OpenAI Gym

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-14T11:46:29.713760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T11:46:29.713760Z digest=sha256:a4d8cbf9be5a48f5480971be9ff55120a52fdbcd6f38318e532909de64363708

Observation 6c60ff96-8cce-4aa1-b15f-e891c1d266fa · outbound

This paper cites Chakraborty and E.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Chakraborty and E

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:46:31.993407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-14T11:46:29.720851Z digest=sha256:334fc0e9c007a55dcc84c6aaff2cb07054af2cb9471976787ba34235db08dfc4

Observation 9a64125f-fc43-4087-a8be-7c5722234d21 · outbound

This paper cites Chamberlain.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Chamberlain

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:46:31.967525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-14T11:46:29.727579Z digest=sha256:93c776c5cb22f6767d21ff4b9a3b941391abd0b60337e224e6392465bea217ce

Observation 9fab8462-5abe-4776-979a-edd1d53d53cb · outbound

This paper cites an unresolved cited work.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-14T11:46:31.946969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-14T11:46:29.733332Z digest=sha256:7c99494f1c6e7120a0eae90c65ff614da86c5426e54e53fc1aa29fff2763f2eb

Observation dc187219-a481-4ab6-8918-8d4fe65e1d18 · outbound

This paper cites Chernozhukov, D.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Chernozhukov, D

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:46:31.922449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-14T11:46:29.739432Z digest=sha256:f21e31407477eeeed92c0e669cbd4d89e605e63e26b0eae4732813c1a3757275

Observation f1477ac4-7b3e-43c7-9e51-78000f8b228f · outbound

This paper cites an unresolved cited work.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-14T11:46:31.874087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-14T11:46:29.744066Z digest=sha256:731bef54a1f39f34d11c0927fa515b75b1f8fe8a46efdd69b887b09b1fcedcdf

Observation f048094e-e66a-4e83-9a46-670416b3cf26 · outbound

This paper cites Dudik, D.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Dudik, D

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:46:31.843720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-14T11:46:29.751133Z digest=sha256:2a6015410ab66bc5c3cb72e15debc2e1b54e0e7657b54129a39751c4906dd2c8

Observation d4341ded-05a6-40df-9260-59191525b859 · outbound

This paper cites Ertefaie and R.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Ertefaie and R

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:46:31.818799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-14T11:46:29.756664Z digest=sha256:a25ecefdd2667c4e3c325ed817e15f3a42a1de37410580bf1f254a8228cc0bfc

Observation e8741784-62bd-4a8e-ab96-26a3abe57435 · outbound

This paper cites Farajtabar, Y.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Farajtabar, Y

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:46:31.790363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-14T11:46:29.763061Z digest=sha256:7f2950b1894eb04f2a8ba6db54e31b6400b2fc50c8b24292af31fa3f156d0cac

Observation 187b1056-ac8d-418a-8ffe-06d61e441410 · outbound

This paper cites Gottesman, F.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Gottesman, F

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:46:31.771074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-14T11:46:29.768468Z digest=sha256:c53352fab53ddddc58241fe98d3ab775fccbae3577dbdc3224e04a91b1cab053

Observation 7188bc26-2d3a-4d1c-aee4-7ba21dd13e0c · outbound

This paper cites Gy \"o rfi, M.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Gy \"o rfi, M

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:46:31.747914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-14T11:46:29.779224Z digest=sha256:90000940d7df8dcf43b6f8171c86f38da0084658079b0cffd3d530ee06ed7cfb

Observation 27e3deb6-450b-44d1-be84-81f952661e85 · outbound

This paper cites an unresolved cited work.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-14T11:46:31.724609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-14T11:46:29.784699Z digest=sha256:ce059f42d5b8ff79cbdbb61a2105ca09d974a9b024d0f10219afb4e904d5436f

Observation 96fac50b-3a5c-45d8-83ac-e8a784e66047 · outbound

This paper cites an unresolved cited work.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-14T11:46:31.701751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-14T11:46:29.791151Z digest=sha256:d99d3043997f13b19501e53182202fa28e573748a2d165c37b119dbbee1ef1dd

Observation bc4b7f22-3f27-4757-b432-6e2ae0f6a4f5 · outbound

This paper cites an unresolved cited work.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-14T11:46:31.681744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-14T11:46:29.797140Z digest=sha256:6a8b0a9c8cc2f4a8ce466d0f7029807c94377ef3f181d298d660799735f4e28a

Observation 40afc250-0db6-4f50-8d13-411bf036bc69 · outbound

This paper cites an unresolved cited work.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-14T11:46:31.658624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-14T11:46:29.805065Z digest=sha256:29616456f71a9af8292416350a4cbce679b26a7233236178b448cca7f74d1cba

Observation daa26224-7df5-44e2-a5a5-a44414a2197e · outbound

This paper cites Hernan and J.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Hernan and J

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:46:31.637722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-14T11:46:29.812436Z digest=sha256:f19652b1cc6018b3b066aed58e86a3e73f23ec1ccec15272d604500ca4adf65a

Observation 582a21ef-731e-4fcd-855a-f837ff193c4e · outbound

This paper cites Hirano, G.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Hirano, G

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:46:31.588084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-14T11:46:29.818987Z digest=sha256:feb83280c7836ce42761a99a0a6d605424773fa4d32a5e939bacba7431dd05fb

Observation e5cfc05c-9586-4051-ba4f-7268017bb553 · outbound

This paper cites Deep Neural Networks Learn Non-Smooth Functions Effectively.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Deep Neural Networks Learn Non-Smooth Functions Effectively

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-08-14T11:46:30.378997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-14T11:46:29.826640Z digest=sha256:db26a3302a8e97593b63ba3e257027bbe2e1e00ce5d10e3a435ddfc81975e9bc

Observation e833b080-1600-4f94-a90c-6b76babb06d3 · outbound

This paper cites Jiang and L.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Jiang and L

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:46:31.560180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-14T11:46:29.832031Z digest=sha256:47e511d5769ef21f8e05c832084d486b93f26a70c0d347a5d1f05ae8fe829bad

Observation 4130b4d5-b87b-4fce-a4c4-8a8211510a98 · outbound

This paper cites Kallus and M.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Kallus and M

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:46:31.534839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-14T11:46:29.839456Z digest=sha256:19a0a1b3d497d116536389c43b0e369877d27d665bd1fd9af69bc971fb5ae090

Observation 2784b18d-a65b-4730-9624-81580d3ef43e · outbound

This paper cites Non-Parametric Inference Adaptive to Intrinsic Dimension.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Non-Parametric Inference Adaptive to Intrinsic Dimension

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-08-14T11:46:30.339063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-14T11:46:29.845940Z digest=sha256:4c9b5861795bacb9759aa781fc14d10c8993dbc05597a973d966da6bb603a3f1

Observation d974909d-c01f-419c-a4c3-b1691fdb1f99 · outbound

This paper cites an unresolved cited work.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-14T11:46:31.501943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-14T11:46:29.852119Z digest=sha256:04d84e88b72291989d28c2361a1f72f13c0df02cbd871e27d7792220c6828c91

Observation 7a009ad2-ec69-4248-8704-ab35dee85572 · outbound

This paper cites an unresolved cited work.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-14T11:46:31.465329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-14T11:46:29.857083Z digest=sha256:5cd17dc6c3256f2c85da19ffb195ed373edf690e1e37dae5e276d11aa5ba317c

Observation d225f321-f385-411e-a58b-87a7943f1e5c · outbound

This paper cites Lagoudakis and R.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Lagoudakis and R

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:46:31.430843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-14T11:46:29.862257Z digest=sha256:490fdadf31b79eb5edbcfdc0e0ab571f65420777cadb9bc7b2e3d4f13ca67940

Observation dd1a6b8b-b1df-4d79-9ef3-40a075870779 · outbound

This paper cites an unresolved cited work.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-14T11:46:31.397232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-14T11:46:29.873067Z digest=sha256:f0cfcdb1c11e91bf2705945e74aab8ec224c5ddfb05e58113cfbaa917535f887

Observation c809529c-5f0f-4cc5-9162-a5d15a9a7ab7 · outbound

This paper cites an unresolved cited work.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-14T11:46:31.375862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-14T11:46:29.878243Z digest=sha256:c6d072d45f804c0a310521625462b5ccf117e08db146aef865874c66bb9c32e7

Observation 2b630a9a-167b-4935-acd9-17134ff9891c · outbound

This paper cites Li and J.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Li and J

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:46:31.354132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-14T11:46:29.887722Z digest=sha256:5d942ae8fd44083a56fc00ede1ea7edb53df34e04c0a83e553885557648aaac8

Observation bfb89cd7-e8a2-4360-bac8-2c7669c1f636 · outbound

This paper cites an unresolved cited work.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-14T11:46:31.326902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-14T11:46:29.894740Z digest=sha256:e21ce2adf1ce7047e3ec722b0920d6980a51cdd86afd14dda9215b1659e3c6f2

Observation 6d863ec6-1291-4917-a0a0-a21422839f2b · outbound

This paper cites an unresolved cited work.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-08-14T11:46:31.291614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-14T11:46:29.903268Z digest=sha256:267f401b481965efa040923c04604f3c8807f13d83cbb487a958ea4d6c0ced62

Observation f89ec68f-efb6-4d90-a9ee-d020a1f21e7e · outbound

This paper cites an unresolved cited work.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-14T11:46:31.265746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-14T11:46:29.911261Z digest=sha256:83c19cbbde81dd4135d3bb33c1f3d9a2e279766ed5c4ad48698cd495421943f6

Observation a33cf5bf-3eee-4da9-adc9-09fc9f8b6d5a · outbound

This paper cites Mandel, Y.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Mandel, Y

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:46:31.240629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-14T11:46:29.916643Z digest=sha256:24dcd99c642b1318e1434c795d92687599e3a79e25c66a943e06aedbd1151e8e

Observation e2f2b559-f3fc-4543-b051-9461d4f8e323 · outbound

This paper cites Mannor, D.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Mannor, D

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:46:31.214132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-14T11:46:29.927349Z digest=sha256:29d74a47f720003909dcabe69f46904d8817ed2cdcb52424f0a256d79898af8f

Observation d8b0a424-78a6-47d9-91d7-319aa22a519c · outbound

This paper cites Munos, T.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Munos, T

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:46:31.192089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-14T11:46:29.933559Z digest=sha256:012c3895b02faf9f689ba3ad0bfd99ff64f469e7d2bb578b4f7674fc6321f894

Observation 822b46da-14c6-4ba8-bfef-6d0673854268 · outbound

This paper cites an unresolved cited work.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-08-14T11:46:31.166768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-14T11:46:29.940144Z digest=sha256:8e1328bc4e9783b3a9d244b6ba3f2d97a92d02fa9958e44ee8b7c36b52a58425

Observation 3c527b1f-0f8d-4bc6-9832-a66756345532 · outbound

This paper cites an unresolved cited work.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-14T11:46:31.129731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-14T11:46:29.950509Z digest=sha256:365547c8abd598c930e2733b420b9fe7a9c1b988460cd377f7858504a3a2c0cb

Observation 20279ca8-88b9-4e63-b109-6c0971aecc98 · outbound

This paper cites an unresolved cited work.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Unresolved cited work

Reference 47

Resolution
unresolved
raw_fallback, observed 2026-08-14T11:46:31.095052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-14T11:46:29.958145Z digest=sha256:e07c6ee61644f366f2ac191dc42d1352662d6e61ae52c83c6f454b0de77fdd1b

Observation 51abad15-06c2-498f-9d52-25c699b96e12 · outbound

This paper cites Precup, R.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Precup, R

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:46:31.069128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-14T11:46:29.965527Z digest=sha256:807131d4948ad23b6ebcbd22eebf583d4833a49f75935cfb00a53c11719491ae

Observation 4d4a39b5-8936-4af7-8593-97bd6094a46c · outbound

This paper cites Rahimi and B.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Rahimi and B

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:46:31.040737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-14T11:46:29.973681Z digest=sha256:334a9a119d9a271861b2e265058e8e6f57865fbb0c4a0599b30ec67071030187

Observation a29adcc7-471b-4a9a-bd40-02b089af0ab7 · outbound

This paper cites an unresolved cited work.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Unresolved cited work

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-14T11:46:29.978792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T11:46:29.978792Z digest=sha256:20f701733753ee0e44f082c8e0f7f20841fddb447838a49d0d955845089e70ef

Observation c5474ff6-ed48-4ff7-9f19-a2047e5088a0 · outbound

This paper cites an unresolved cited work.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-08-14T11:46:30.988725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-14T11:46:29.984207Z digest=sha256:ceea623050703b5e8bef5d5f0fa4fc2b21ed5b58831d55221994d4bedaf24ade

Observation 3891080d-8386-47e1-9ee7-9440383b4ffa · outbound

This paper cites an unresolved cited work.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-08-14T11:46:30.961020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-14T11:46:29.991655Z digest=sha256:b271b6c22f1f4cf79a3a256ea4c021b1162f61dc26c7917e6d760682938c4865

Observation f35b1f7a-ee20-4320-a4bf-ebfc7c5a3e12 · outbound

This paper cites Rotnitzky and S.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Rotnitzky and S

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:46:30.932614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-14T11:46:30.007990Z digest=sha256:66dd638c6b1da8e8f4ec1e58afd40d3c2e86e2af4031f8c055697bba37ec2a1e

Observation 93237f4b-a33a-4e24-a832-e67a5bd7ced2 · outbound

This paper cites Identification, Doubly Robust Estimation, and Semiparametric Efficiency Theory of Nonignorable Missing Data With a Shadow Variable.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Identification, Doubly Robust Estimation, and Semiparametric Efficiency Theory of Nonignorable Missing Data With a Shadow Variable

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-14T11:46:30.017260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T11:46:30.017260Z digest=sha256:d06e13280f014ac9a33e97e5cedf6120ada9973cfbdf420143b0ad1c1451a975

Observation 57dc18f1-35d7-4f56-992e-64068604e7f5 · outbound

This paper cites Scharfstein, A.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Scharfstein, A

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:46:30.908750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-14T11:46:30.027133Z digest=sha256:a106d3174aa079033e92b3316e3a6eb83464b47ef59c70ac0d3db141e12afb58

Observation c5646464-6853-4f32-929c-2a9e204ae2e7 · outbound

This paper cites an unresolved cited work.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Unresolved cited work

Reference 56

Resolution
unresolved
raw_fallback, observed 2026-08-14T11:46:30.886535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-14T11:46:30.036300Z digest=sha256:24877c3cfc0aede452fc8e86ab67c20eb0ccdd3a63e4d09c0dc4f44f76dd20b1

Observation fdf71f61-1d43-4320-ba9b-e839a7b4f1d0 · outbound

This paper cites an unresolved cited work.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Unresolved cited work

Reference 57

Resolution
unresolved
raw_fallback, observed 2026-08-14T11:46:30.863362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-14T11:46:30.043767Z digest=sha256:8153f1e5513b6bebd0e0263453110fa6e60739cff7873fd726e4bca3222e2b7e

Observation f07166fe-6c58-4096-a7d9-e03af9c18acb · outbound

This paper cites an unresolved cited work.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Unresolved cited work

Reference 58

Resolution
unresolved
raw_fallback, observed 2026-08-14T11:46:30.840593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-14T11:46:30.052525Z digest=sha256:150e5641f749d5b93f7a50a8810f97eb5ff86ff9b21f6b61cebdb869deb15f0c

Observation a851e44f-fed0-4bbb-8988-eee8049152d0 · outbound

This paper cites Swaminathan and T.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Swaminathan and T

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:46:30.822877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-14T11:46:30.059703Z digest=sha256:d66c9c4513e62fe878a9b60559bf0b1cc66d6320bc2016c284134d6f96fcab6a

Observation 1abeb396-058d-45fe-abe3-bc2e6bc0f24f · outbound

This paper cites Thomas and E.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Thomas and E

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:46:30.800713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-14T11:46:30.067720Z digest=sha256:1179bf68765fabf7b322a800b00fa2529a81199c8d76f0ffc1cb409641ae30d8

Observation ea5317c0-baf0-426e-9596-c2fd244d828a · outbound

This paper cites Tripathi.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Tripathi

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:46:30.775764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-14T11:46:30.074811Z digest=sha256:1176fd8c0d2b81b4df7644f0df37bf210b54ec5142686aad497794a40a912d82

Observation b58a750b-1315-4ed8-b028-2bef40a88917 · outbound

This paper cites an unresolved cited work.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Unresolved cited work

Reference 62

Resolution
unresolved
raw_fallback, observed 2026-08-14T11:46:30.747367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-14T11:46:30.080935Z digest=sha256:41f7456b1c35432c43a2e4739074dcfdb1a0e304767454606a8ca2635a12a071

Observation 50296aed-8de8-4c85-976e-022a1cdb4864 · outbound

This paper cites an unresolved cited work.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Unresolved cited work

Reference 63

Resolution
unresolved
raw_fallback, observed 2026-08-14T11:46:30.711610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-14T11:46:30.086598Z digest=sha256:7c6bbb6dffe9084ea7c090a68984473a0fbce3ea9376aa30266227a36302f77b

Observation c7eb9fa1-801b-49f8-9920-8861b86df38e · outbound

This paper cites an unresolved cited work.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Unresolved cited work

Reference 64

Resolution
unresolved
raw_fallback, observed 2026-08-14T11:46:30.683598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-14T11:46:30.097210Z digest=sha256:db6e6c72fa06e8adc49cf064d73aed5816237aaf4b7c0d0cee0058952e715f2b

Observation f816fdf0-9662-48fd-9c08-fb500a359bec · outbound

This paper cites an unresolved cited work.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Unresolved cited work

Reference 65

Resolution
unresolved
raw_fallback, observed 2026-08-14T11:46:30.664588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-14T11:46:30.104286Z digest=sha256:70be1f0f3434b4b489b81b19d8985edabd4109c2304962c2d003ccc027cca98e

Observation 8a76e8bd-8928-4624-add7-95d5f7079e89 · outbound

This paper cites an unresolved cited work.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Unresolved cited work

Reference 66

Resolution
unresolved
raw_fallback, observed 2026-08-14T11:46:30.646811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-14T11:46:30.112904Z digest=sha256:2d8a8cbee3e6039175afdac46c2e07796c972f432142004b8a619cd94ee94f49

Observation 686ae564-f31c-4a78-85a4-c6b9f51b6a91 · outbound

This paper cites Vermeulen.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Vermeulen

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:46:30.622835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-14T11:46:30.119558Z digest=sha256:8c8f92ced281ab7c29b36ec2966fe0cfa60c56fcf51ac1ed89a6c697a214dfa5

Observation 1cbdc28f-9c5a-49e4-8aa7-5769a7bb7978 · outbound

This paper cites Adaptive Concentration of Regression Trees, with Application to Random Forests.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Adaptive Concentration of Regression Trees, with Application to Random Forests

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-14T11:46:30.130845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T11:46:30.130845Z digest=sha256:9c4799f7fc010aee17e4172cfe428c7325c1019cd627827fbdf8c4740373a675

Observation 9341a610-d4e4-4f3a-8465-1a1099db1c87 · outbound

This paper cites an unresolved cited work.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Unresolved cited work

Reference 69

Resolution
unresolved
raw_fallback, observed 2026-08-14T11:46:30.601862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-14T11:46:30.136827Z digest=sha256:a7411f29bf575e333d2a3c086b21e1738e0bc6ad5ee52bdf644bfb197af3c907

Observation ff344948-ef2b-4eac-a8dd-2c89ebf7e840 · outbound

This paper cites Yin and Y.-X.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Yin and Y.-X

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:46:30.569546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-14T11:46:30.144544Z digest=sha256:21abd1e6587fbcece6d21cf457822fcdf1b060cff5aed8967c274ff4a247048b

Observation a34c62ce-fd75-4651-bb54-fd8f7ef4f919 · outbound

This paper cites Zhang, A.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Zhang, A

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:46:30.542262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-14T11:46:30.161025Z digest=sha256:d23aee0ee00cb232112aa1c1a64fe97e4f8a5e30e6c651e99c0cbd68777922f4

Observation 36e41956-8ef9-4a85-84a7-a6e2bdd2c41c · outbound

This paper cites Zheng and M.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Zheng and M

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:46:30.492426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-14T11:46:30.168252Z digest=sha256:31461e934e4afb7978e3153c9004bcfd2db338d85305499ebfc306c6f395f11a

Pith citing papers

No inbound Pith citation observations are available.