Pith. sign in

Paper Citation Record · LEDGER

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes

As of 17 August 2026, this Paper Citation Record lists 71 of 71 outbound references and 0 inbound Pith citation observations for arXiv:1908.08526.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1908.08526 v3

Coverage vector

measured 71 of 71 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-14T11:46:30.168252Z

measured 71 of 71 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

71 of 71 outbound references displayed

  • verified exact2
  • verified fuzzy33
  • unresolved36
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2f7b9a31-3e26-4e61-97aa-6095eeb23fe4 · outbound

This paper cites Ai and X.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Ai and X

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:46:32.310499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-14T11:46:29.634294Z digest=sha256:b7187938560e4c2b0162d58014fce3d2e44f5fa3025f4051dcb83af3a7ac254b

Observation 6a09b9d4-0a22-46dc-b54f-f4599ea5e1dc · outbound

This paper cites Ai and X.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Ai and X

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:46:32.273051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-14T11:46:29.641706Z digest=sha256:04a531e7a97eb73130cbd58d017681f3bf188451dd701c9e9ef011e09dd81142

Observation 9fa8affe-7a93-470e-9298-b2c917d49aa1 · outbound

This paper cites Antos, C.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Antos, C

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:46:32.232092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-14T11:46:29.648353Z digest=sha256:614003fadafb82a895cf353d2963d1e458c21d775af9efc26f931809f8f884b1

Observation d818f678-21e1-4dc0-b942-8fb649b593cd · outbound

This paper cites an unresolved cited work.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-14T11:46:32.190388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-14T11:46:29.656662Z digest=sha256:0d1e50613b5eae1aaf1c9b26e4c4b261d18afe90620b80a15fc624bce3f76a9b

Observation d6edfed1-f4db-4313-b3e2-5fcfc0cbef8f · outbound

This paper cites Benkeser and M.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Benkeser and M

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:46:32.154243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-14T11:46:29.662248Z digest=sha256:5ef26f1aa85e21a54f438b4ad90ca1ee92c2760e7ed1d0a5a67b04f2d5b93422

Observation 09e5b3f6-384d-40a6-96fc-8ccab2344dc3 · outbound

This paper cites an unresolved cited work.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-14T11:46:32.120086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-14T11:46:29.668735Z digest=sha256:c2a3c4b1502d548ad26c2462b610cb275339cd232e5e3354bbbb5a19e0ad12c9

Observation de1af119-467d-4b1a-a29c-254f140b8e54 · outbound

This paper cites Bibaut, I.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Bibaut, I

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:46:32.088567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-14T11:46:29.682231Z digest=sha256:55f488235d6ea9dc414cf3b966df82ecc5088250b388d071b0abbfa80708e3fc

Observation 1e23951c-9ce7-4762-8e03-a682dda75c17 · outbound

This paper cites Fast rates for empirical risk minimization over c\`adl\`ag functions with bounded sectional variation norm.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Fast rates for empirical risk minimization over c\`adl\`ag functions with bounded sectional variation norm

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-14T11:46:29.690268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T11:46:29.690268Z digest=sha256:7a6b21c6782df687c89a9c13208827a99ee5404df85daf1cea2f107d1128797b

Observation a43cee4a-eafd-4766-9a3c-8e4461c1c3c6 · outbound

This paper cites an unresolved cited work.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-14T11:46:29.700225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T11:46:29.700225Z digest=sha256:4cc23e13e5514f144e3f822f0c4872bb394101dc5db065a5890ee2dd9bdd137f

Observation 4dceaf47-f150-4468-b2bd-37098574e674 · outbound

This paper cites an unresolved cited work.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-14T11:46:32.030080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-14T11:46:29.708272Z digest=sha256:7c46f31335bcb2750e4b5f17089e7920c063ba33ea48035aef32242905ae6488

Observation 4e9e4388-96a0-43dd-98b3-2b0404f82f21 · outbound

This paper cites OpenAI Gym.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes OpenAI Gym

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-14T11:46:29.713760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T11:46:29.713760Z digest=sha256:697728bff1aa2a39a87400f11a350738bfd02ca724ec014b913551c7aa166d8d

Observation 6c60ff96-8cce-4aa1-b15f-e891c1d266fa · outbound

This paper cites Chakraborty and E.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Chakraborty and E

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:46:31.993407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-14T11:46:29.720851Z digest=sha256:87e288f146ff0158d72a604f4540a55f26ca6587a28ff420d2eeb156a71854e2

Observation 9a64125f-fc43-4087-a8be-7c5722234d21 · outbound

This paper cites Chamberlain.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Chamberlain

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:46:31.967525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-14T11:46:29.727579Z digest=sha256:4f9bcdafcc9a16c8346f976301afadfd8a83a5a1564258ffee8da2ba67acea08

Observation 9fab8462-5abe-4776-979a-edd1d53d53cb · outbound

This paper cites an unresolved cited work.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-14T11:46:31.946969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-14T11:46:29.733332Z digest=sha256:14e866c029fc448bb4c613742ad4542049d7e951a9b81ad24c38c9c330189cb8

Observation dc187219-a481-4ab6-8918-8d4fe65e1d18 · outbound

This paper cites Chernozhukov, D.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Chernozhukov, D

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:46:31.922449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-14T11:46:29.739432Z digest=sha256:4ff19e16571b83a5145df3f1d3fa9409f252706acd0c4478282c7374c4ab3d52

Observation f1477ac4-7b3e-43c7-9e51-78000f8b228f · outbound

This paper cites an unresolved cited work.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-14T11:46:31.874087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-14T11:46:29.744066Z digest=sha256:a2c22941720b45fe0e93e4a9f2123f26e0e934741f7bdfe66000d54f68ddab89

Observation f048094e-e66a-4e83-9a46-670416b3cf26 · outbound

This paper cites Dudik, D.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Dudik, D

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:46:31.843720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-14T11:46:29.751133Z digest=sha256:1ccb095c518d00e73fcb52572552ef6252881aed4b725b8d54254a812d976d84

Observation d4341ded-05a6-40df-9260-59191525b859 · outbound

This paper cites Ertefaie and R.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Ertefaie and R

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:46:31.818799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-14T11:46:29.756664Z digest=sha256:52cef9948465ffc6bf880d42b712b2f445390002d33650c138f1bc798bb84d81

Observation e8741784-62bd-4a8e-ab96-26a3abe57435 · outbound

This paper cites Farajtabar, Y.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Farajtabar, Y

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:46:31.790363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-14T11:46:29.763061Z digest=sha256:95c5a2a0979f935ab576a2d2c5a82c003df84f9f6cb21e9caf2fbef4aa68af5f

Observation 187b1056-ac8d-418a-8ffe-06d61e441410 · outbound

This paper cites Gottesman, F.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Gottesman, F

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:46:31.771074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-14T11:46:29.768468Z digest=sha256:dfb7fe09a1e51579ee4ada6532de933c874e0640765c72e858c1959e139e4198

Observation 7188bc26-2d3a-4d1c-aee4-7ba21dd13e0c · outbound

This paper cites Gy \"o rfi, M.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Gy \"o rfi, M

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:46:31.747914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-14T11:46:29.779224Z digest=sha256:1a38aa2bde3c3e27620e12ab92b56bc345537775adb3f5381888ff551d008399

Observation 27e3deb6-450b-44d1-be84-81f952661e85 · outbound

This paper cites an unresolved cited work.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-14T11:46:31.724609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-14T11:46:29.784699Z digest=sha256:17cf04694371d2fee4fb31481f58530c790ecd1736e58899c2b829f18275c45f

Observation 96fac50b-3a5c-45d8-83ac-e8a784e66047 · outbound

This paper cites an unresolved cited work.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-14T11:46:31.701751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-14T11:46:29.791151Z digest=sha256:755ed7019cad3da7390f2e8b8e09a594a9c4ddafa14d223fd50fca59eaf08f50

Observation bc4b7f22-3f27-4757-b432-6e2ae0f6a4f5 · outbound

This paper cites an unresolved cited work.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-14T11:46:31.681744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-14T11:46:29.797140Z digest=sha256:b72347efdca3cde7f0a875b810ade795bfd4844eef805beb83d4fecd8e1afcca

Observation 40afc250-0db6-4f50-8d13-411bf036bc69 · outbound

This paper cites an unresolved cited work.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-14T11:46:31.658624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-14T11:46:29.805065Z digest=sha256:75a8204eca79f72fbc3eaba000d843031334c8182118a4c3ee378e672bb6c2a1

Observation daa26224-7df5-44e2-a5a5-a44414a2197e · outbound

This paper cites Hernan and J.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Hernan and J

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:46:31.637722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-14T11:46:29.812436Z digest=sha256:b004c89f5bc8b42728ba4f6531453ab443b3991549295902200b1c6ddd97e559

Observation 582a21ef-731e-4fcd-855a-f837ff193c4e · outbound

This paper cites Hirano, G.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Hirano, G

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:46:31.588084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-14T11:46:29.818987Z digest=sha256:c6ade43ba70df1612ae8fbf08ec62acb6991364a59ca5fb3a291a8b3aff73121

Observation e5cfc05c-9586-4051-ba4f-7268017bb553 · outbound

This paper cites Deep Neural Networks Learn Non-Smooth Functions Effectively.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Deep Neural Networks Learn Non-Smooth Functions Effectively

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-08-14T11:46:30.378997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-14T11:46:29.826640Z digest=sha256:212d6a07de26ddfd96141cf7e4d89a1124ed055dcee31464ea3b3caf56356c0a

Observation e833b080-1600-4f94-a90c-6b76babb06d3 · outbound

This paper cites Jiang and L.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Jiang and L

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:46:31.560180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-14T11:46:29.832031Z digest=sha256:556e0aaee2826c7d392d528e34859e080dbe96920878987ef815b589cecdd776

Observation 4130b4d5-b87b-4fce-a4c4-8a8211510a98 · outbound

This paper cites Kallus and M.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Kallus and M

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:46:31.534839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-14T11:46:29.839456Z digest=sha256:115a22ba59dfe65bda2d7f3948ee1bf325ba608c41c5ef99c14e6371408397ac

Observation 2784b18d-a65b-4730-9624-81580d3ef43e · outbound

This paper cites Non-Parametric Inference Adaptive to Intrinsic Dimension.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Non-Parametric Inference Adaptive to Intrinsic Dimension

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-08-14T11:46:30.339063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-14T11:46:29.845940Z digest=sha256:3c852e3004a8b6dd955d218519c4573007fca1a9a55e89d35f4f996b6ea37f21

Observation d974909d-c01f-419c-a4c3-b1691fdb1f99 · outbound

This paper cites an unresolved cited work.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-14T11:46:31.501943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-14T11:46:29.852119Z digest=sha256:eb619fc7a1ea402df447d6c3cf3f4d2d3d69634b226f5216047a32cf363d587d

Observation 7a009ad2-ec69-4248-8704-ab35dee85572 · outbound

This paper cites an unresolved cited work.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-14T11:46:31.465329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-14T11:46:29.857083Z digest=sha256:748ba3747966673954f9087f1dfef8f15dd9007f89d0cb4bf58463ef6e81a29f

Observation d225f321-f385-411e-a58b-87a7943f1e5c · outbound

This paper cites Lagoudakis and R.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Lagoudakis and R

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:46:31.430843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-14T11:46:29.862257Z digest=sha256:50dcceb914020aa83c2396eba9bb04c996a8897606f5904ba5cf85b874e310a0

Observation dd1a6b8b-b1df-4d79-9ef3-40a075870779 · outbound

This paper cites an unresolved cited work.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-14T11:46:31.397232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-14T11:46:29.873067Z digest=sha256:19059b2665e26a8ab51d8f40e186f77e1231e6c56239a8ade441c4be1e1f664d

Observation c809529c-5f0f-4cc5-9162-a5d15a9a7ab7 · outbound

This paper cites an unresolved cited work.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-14T11:46:31.375862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-14T11:46:29.878243Z digest=sha256:5623d94a26e75d6ff73b1be665785c9eec25b4565b56ba0ed3a0e9b1c015e22b

Observation 2b630a9a-167b-4935-acd9-17134ff9891c · outbound

This paper cites Li and J.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Li and J

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:46:31.354132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-14T11:46:29.887722Z digest=sha256:9b19dd41864d82df2a84a196cb027e7b8578862767d49cd8d224c8eddb903f3e

Observation bfb89cd7-e8a2-4360-bac8-2c7669c1f636 · outbound

This paper cites an unresolved cited work.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-14T11:46:31.326902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-14T11:46:29.894740Z digest=sha256:a80cf48e9f67e47c7c3fc04b8145a5945d3cb5306e230b458813f6dc5ab2f005

Observation 6d863ec6-1291-4917-a0a0-a21422839f2b · outbound

This paper cites an unresolved cited work.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-08-14T11:46:31.291614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-14T11:46:29.903268Z digest=sha256:071aef5fabd6ad49f3839a47663fe048da7869ee70263be843a67acd51efe134

Observation f89ec68f-efb6-4d90-a9ee-d020a1f21e7e · outbound

This paper cites an unresolved cited work.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-14T11:46:31.265746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-14T11:46:29.911261Z digest=sha256:ff55926009335fd5d4036640971a226cece69269ee9396987b37ad8aae1e1f6f

Observation a33cf5bf-3eee-4da9-adc9-09fc9f8b6d5a · outbound

This paper cites Mandel, Y.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Mandel, Y

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:46:31.240629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-14T11:46:29.916643Z digest=sha256:52eda461b6c53f638520062a4c780ffe201b0d2ebf19041d887806d3c3624e9f

Observation e2f2b559-f3fc-4543-b051-9461d4f8e323 · outbound

This paper cites Mannor, D.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Mannor, D

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:46:31.214132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-14T11:46:29.927349Z digest=sha256:008b92c2bcb6876d59d022ede90045def97cb2e0b28b7255d278150eb9588583

Observation d8b0a424-78a6-47d9-91d7-319aa22a519c · outbound

This paper cites Munos, T.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Munos, T

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:46:31.192089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-14T11:46:29.933559Z digest=sha256:ec4b6cbe0da813e5aea11cb57dca76f9fe50078b8db058f5eda703a41f119789

Observation 822b46da-14c6-4ba8-bfef-6d0673854268 · outbound

This paper cites an unresolved cited work.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-08-14T11:46:31.166768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-14T11:46:29.940144Z digest=sha256:92c8378d7981116141a8e24e5cf521bca259b332c9f544c7cc904d52f48be706

Observation 3c527b1f-0f8d-4bc6-9832-a66756345532 · outbound

This paper cites an unresolved cited work.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-14T11:46:31.129731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-14T11:46:29.950509Z digest=sha256:dd55506396761dbddeef2170685f6e9b6ab438b9833ab3aba32463f6528ff4c3

Observation 20279ca8-88b9-4e63-b109-6c0971aecc98 · outbound

This paper cites an unresolved cited work.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Unresolved cited work

Reference 47

Resolution
unresolved
raw_fallback, observed 2026-08-14T11:46:31.095052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-14T11:46:29.958145Z digest=sha256:f147c5e78d36d391561a8fa37fce92684e11e54ef8dafc80abdfc364cacc39d5

Observation 51abad15-06c2-498f-9d52-25c699b96e12 · outbound

This paper cites Precup, R.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Precup, R

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:46:31.069128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-14T11:46:29.965527Z digest=sha256:f2ead567d700f92d2154121095b29fce14eb8fc90f3c1cdf0e4d1c0cb7513b2e

Observation 4d4a39b5-8936-4af7-8593-97bd6094a46c · outbound

This paper cites Rahimi and B.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Rahimi and B

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:46:31.040737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-14T11:46:29.973681Z digest=sha256:c33b1f061b204205e955c1eedfa09b11e68d7a9d00a760e514656697f222ea0b

Observation a29adcc7-471b-4a9a-bd40-02b089af0ab7 · outbound

This paper cites an unresolved cited work.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Unresolved cited work

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-14T11:46:29.978792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T11:46:29.978792Z digest=sha256:3d16fcad1de6fddb853129ca0e58066aa0445fa1dceeb735b8ef9272e518c68f

Observation c5474ff6-ed48-4ff7-9f19-a2047e5088a0 · outbound

This paper cites an unresolved cited work.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-08-14T11:46:30.988725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-14T11:46:29.984207Z digest=sha256:a92dd9b28e6abdb123a0999085715e2a48e21eba8912d911c3954f10972bfb0e

Observation 3891080d-8386-47e1-9ee7-9440383b4ffa · outbound

This paper cites an unresolved cited work.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-08-14T11:46:30.961020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-14T11:46:29.991655Z digest=sha256:4f388040c7c70c533649ababc52f9172891080515e2990fd907a5070fc06ca4e

Observation f35b1f7a-ee20-4320-a4bf-ebfc7c5a3e12 · outbound

This paper cites Rotnitzky and S.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Rotnitzky and S

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:46:30.932614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-14T11:46:30.007990Z digest=sha256:ac4d9a85509bc67e3e6460c1963aa47202b7fe0efa1fcf37eef914472cb9d481

Observation 93237f4b-a33a-4e24-a832-e67a5bd7ced2 · outbound

This paper cites Identification, Doubly Robust Estimation, and Semiparametric Efficiency Theory of Nonignorable Missing Data With a Shadow Variable.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Identification, Doubly Robust Estimation, and Semiparametric Efficiency Theory of Nonignorable Missing Data With a Shadow Variable

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-14T11:46:30.017260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T11:46:30.017260Z digest=sha256:acb9bb68a876b9430fb964d395fa45a24b4917d2a55ee1317f7018382d0275c5

Observation 57dc18f1-35d7-4f56-992e-64068604e7f5 · outbound

This paper cites Scharfstein, A.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Scharfstein, A

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:46:30.908750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-14T11:46:30.027133Z digest=sha256:f8adc150caffd7e712de709906c79d5d1ce5a55d13e4242ff0547c07610c75f8

Observation c5646464-6853-4f32-929c-2a9e204ae2e7 · outbound

This paper cites an unresolved cited work.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Unresolved cited work

Reference 56

Resolution
unresolved
raw_fallback, observed 2026-08-14T11:46:30.886535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-14T11:46:30.036300Z digest=sha256:0ba817baa1dd3c959ba753b897728c158097dc3c12ed3f3e2640a15c207e34d8

Observation fdf71f61-1d43-4320-ba9b-e839a7b4f1d0 · outbound

This paper cites an unresolved cited work.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Unresolved cited work

Reference 57

Resolution
unresolved
raw_fallback, observed 2026-08-14T11:46:30.863362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-14T11:46:30.043767Z digest=sha256:43abac0c4156453661171ce8fe31481307a67899ad88cac20b4d1f804a1b1d72

Observation f07166fe-6c58-4096-a7d9-e03af9c18acb · outbound

This paper cites an unresolved cited work.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Unresolved cited work

Reference 58

Resolution
unresolved
raw_fallback, observed 2026-08-14T11:46:30.840593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-14T11:46:30.052525Z digest=sha256:d2ca1e88bce0dece84933dceddc6b135c04149c64296b5896b62e7dbdc30fb54

Observation a851e44f-fed0-4bbb-8988-eee8049152d0 · outbound

This paper cites Swaminathan and T.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Swaminathan and T

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:46:30.822877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-14T11:46:30.059703Z digest=sha256:d0906e7c9368abf144a2a9232edd1c3d4d5a30019b2a5b392162981e645389db

Observation 1abeb396-058d-45fe-abe3-bc2e6bc0f24f · outbound

This paper cites Thomas and E.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Thomas and E

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:46:30.800713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-14T11:46:30.067720Z digest=sha256:0a820bac32c21ee014d8434ebc390e0a9bbc02589091d1c6be453c2904076d46

Observation ea5317c0-baf0-426e-9596-c2fd244d828a · outbound

This paper cites Tripathi.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Tripathi

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:46:30.775764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-14T11:46:30.074811Z digest=sha256:171d57017f56ea43f252e6342a7ffa29e0ab9a0775e45aa268e2481d28eb0c7c

Observation b58a750b-1315-4ed8-b028-2bef40a88917 · outbound

This paper cites an unresolved cited work.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Unresolved cited work

Reference 62

Resolution
unresolved
raw_fallback, observed 2026-08-14T11:46:30.747367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-14T11:46:30.080935Z digest=sha256:25b7a89e4e246b9dced3a3a935757a7c32e372cd245265a6c1c001f6e57cc48c

Observation 50296aed-8de8-4c85-976e-022a1cdb4864 · outbound

This paper cites an unresolved cited work.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Unresolved cited work

Reference 63

Resolution
unresolved
raw_fallback, observed 2026-08-14T11:46:30.711610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-14T11:46:30.086598Z digest=sha256:11ce137f050a9d3b9610174bf0f2cf3a27d22bcc4a0bd02595f47f1a7b822f0b

Observation c7eb9fa1-801b-49f8-9920-8861b86df38e · outbound

This paper cites an unresolved cited work.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Unresolved cited work

Reference 64

Resolution
unresolved
raw_fallback, observed 2026-08-14T11:46:30.683598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-14T11:46:30.097210Z digest=sha256:14d6046a25d9f3661271f6fe29fc9ee940d7894cd4372f7fc921ad607c8ba666

Observation f816fdf0-9662-48fd-9c08-fb500a359bec · outbound

This paper cites an unresolved cited work.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Unresolved cited work

Reference 65

Resolution
unresolved
raw_fallback, observed 2026-08-14T11:46:30.664588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-14T11:46:30.104286Z digest=sha256:5c7bc2685876edd51f66dddbeecab1999049137fa4d7c8fad8fc79ef23ec5fd2

Observation 8a76e8bd-8928-4624-add7-95d5f7079e89 · outbound

This paper cites an unresolved cited work.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Unresolved cited work

Reference 66

Resolution
unresolved
raw_fallback, observed 2026-08-14T11:46:30.646811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-14T11:46:30.112904Z digest=sha256:475015cad618634792a385bc781d8391637773a56198241c4f67d1f262a686e9

Observation 686ae564-f31c-4a78-85a4-c6b9f51b6a91 · outbound

This paper cites Vermeulen.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Vermeulen

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:46:30.622835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-14T11:46:30.119558Z digest=sha256:17284a24548a7e1eb101705329958d9cc466af788b24d1faf7093bd9baa6fffc

Observation 1cbdc28f-9c5a-49e4-8aa7-5769a7bb7978 · outbound

This paper cites Adaptive Concentration of Regression Trees, with Application to Random Forests.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Adaptive Concentration of Regression Trees, with Application to Random Forests

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-14T11:46:30.130845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T11:46:30.130845Z digest=sha256:1bcceef047bce6f652f8ba21c07fa8a6e86656d1467bc361a4ae7dc6b5332941

Observation 9341a610-d4e4-4f3a-8465-1a1099db1c87 · outbound

This paper cites an unresolved cited work.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Unresolved cited work

Reference 69

Resolution
unresolved
raw_fallback, observed 2026-08-14T11:46:30.601862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-14T11:46:30.136827Z digest=sha256:5950b11f8e55cc3f57f2d0ce0df38f5d768dbc35757138a3d080389d6275a352

Observation ff344948-ef2b-4eac-a8dd-2c89ebf7e840 · outbound

This paper cites Yin and Y.-X.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Yin and Y.-X

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:46:30.569546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-14T11:46:30.144544Z digest=sha256:914969b73e50907ce23840da3e411ace57ce071e6b7411faaa364c36882a41dc

Observation a34c62ce-fd75-4651-bb54-fd8f7ef4f919 · outbound

This paper cites Zhang, A.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Zhang, A

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:46:30.542262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-14T11:46:30.161025Z digest=sha256:b3f8689dcd8c7ebcc04bd86b6e030dbe6933457d40b0b04c1f520584a4875329

Observation 36e41956-8ef9-4a85-84a7-a6e2bdd2c41c · outbound

This paper cites Zheng and M.

Double Reinforcement Learning for Efficient Off-Policy Evaluation in Markov Decision Processes Zheng and M

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T11:46:30.492426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-14T11:46:30.168252Z digest=sha256:bf6f623fd12b2f3e2994e1af7604446a652ce4d982775c71082763071e247510

Pith citing papers

No inbound Pith citation observations are available.