Pith. sign in

Paper Citation Record · LEDGER

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding

As of 19 August 2026, this Paper Citation Record lists 79 of 79 outbound references and 0 inbound Pith citation observations for arXiv:2505.00304.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.00304 v1

Coverage vector

measured 79 of 79 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T04:54:14.609941Z

measured 79 of 79 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

79 of 79 outbound references displayed

  • verified exact2
  • verified fuzzy47
  • unresolved30
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cd4073d0-0038-4952-adf9-63822c6ee9da · outbound

This paper cites , " * write output.state after.block = add.period write newline.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding , " * write output.state after.block = add.period write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T04:54:14.256109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:54:14.256109Z digest=sha256:f1ae6ae2886c6f11346353ed6e802ad623cfe66dae1c61251bb476c888737e81

Observation 919d6f0d-5353-48de-8699-761051d85bab · outbound

This paper cites write newline.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T04:54:14.261809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:54:14.261809Z digest=sha256:3cc3e79f35ee1f1c2d236e3687e0e7aa250bbebffad0802442b2a22257cfd8dd

Observation 410f8da7-edd5-4ad3-9709-f6359a423471 · outbound

This paper cites an unresolved cited work.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:54:15.823487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.266879Z digest=sha256:e3d68906177f8f804787b777a062a36392d1e6fc26e48cb44ac88f5d2d62b531

Observation 6fb5d4e1-7c06-417a-ad6f-51111b68b2b2 · outbound

This paper cites (2008), Learning near-optimal policies with Bellman-residual minimization based fitted policy iteration and a single sample path, Machine Learning, 71, 89--129.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding (2008), Learning near-optimal policies with Bellman-residual minimization based fitted policy iteration and a single sample path, Machine Learning, 71, 89--129

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:15.806643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.271354Z digest=sha256:c8135c2e37bfa52e437fce20d3ea897ee313aae39c90b9d5e7deed088e7c7b82

Observation 4e3b1f35-75ae-4d39-bac0-44115f865f80 · outbound

This paper cites (2017), Breaking the curse of dimensionality with convex neural networks, The Journal of Machine Learning Research, 18, 629--681.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding (2017), Breaking the curse of dimensionality with convex neural networks, The Journal of Machine Learning Research, 18, 629--681

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:15.790796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.275782Z digest=sha256:1007ddb092ef47be89776b720b42b3ff2156017410f7cc0bb7ea2c0bdd87f577

Observation 7b1ecb43-2fc0-4ad4-afc2-de323f09bd6d · outbound

This paper cites an unresolved cited work.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:54:15.775627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.280602Z digest=sha256:93676fc512b6207cb12734a3794b12e36ae8f1d8ddda0944b7b00c6d25667675

Observation f3845a6a-eff6-4af7-b773-d37f0917cade · outbound

This paper cites and Kallus, N.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding and Kallus, N

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:15.760428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.285139Z digest=sha256:260ba01769fe1515e5655bced7dc0eee87de37848bf2b463cfdd88f422b01cc4

Observation 2bfdfeb8-ce6d-418d-9481-3545b554ba51 · outbound

This paper cites (2021), Off-policy evaluation in infinite-horizon reinforcement learning with latent confounders, in International Conference on Artificial Intelligence and Statistics, PMLR, pp.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding (2021), Off-policy evaluation in infinite-horizon reinforcement learning with latent confounders, in International Conference on Artificial Intelligence and Statistics, PMLR, pp

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:15.745039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.289645Z digest=sha256:8ae3f3ce22ec5c091e9e017898313575f9335e09a06850f858dca7b448ef84a1

Observation b13a54ca-1684-4dbe-9b79-434d67a8905e · outbound

This paper cites and Kennedy, E.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding and Kennedy, E

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T04:54:14.294412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:54:14.294412Z digest=sha256:8218e09038e63031736c6a7381711afc13e0b3071ed41a73a63c39ce993fe629

Observation 466524b1-2ecf-4ea2-b1dd-f7190268ce50 · outbound

This paper cites u derl, J., Schmiedeberg, C., Castiglioni, L., Arr \'a nz Becker, O., Buhr, P., Fu , D., Ludwig, V., Schr \.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding u derl, J., Schmiedeberg, C., Castiglioni, L., Arr \'a nz Becker, O., Buhr, P., Fu , D., Ludwig, V., Schr \

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:15.717702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.298932Z digest=sha256:d3ee596d9c9ac7feb64eaee65d87f61e2c80443b33eab3310b32937d0bc9c874

Observation 11852540-b5e5-49c5-b25f-bf5ea33bb747 · outbound

This paper cites W., Yuan, Z., Zhou, S., Panerati, J., and Schoellig, A.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding W., Yuan, Z., Zhou, S., Panerati, J., and Schoellig, A

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:15.702282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.303552Z digest=sha256:be29c486990185a3082e401803be73e7f980204aa379051957feb4bfdbc14f2e

Observation ba310e6a-5471-49cb-849a-b1277b746cba · outbound

This paper cites an unresolved cited work.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:54:15.686011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.308223Z digest=sha256:db50c58128033579390698429412ef30efd025bcab6521e3175cb1bd6ed3bfa2

Observation 5ad80ee7-1129-4815-a00b-59980fb968fc · outbound

This paper cites Jump Interval-Learning for Individualized Decision Making.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding Jump Interval-Learning for Individualized Decision Making

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-08-16T04:54:14.824601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.312790Z digest=sha256:a23b6745891956ee8e0ff0e86842ff419e94266c4f073cc1c2c71f32d47c5f81

Observation 818cf5dd-c8e1-4f46-a3c2-c31add3c0577 · outbound

This paper cites (2022), Reinforcement learning from partial observation: Linear function approximation with provable sample efficiency, in International Conference on Machine Learning, PMLR, pp.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding (2022), Reinforcement learning from partial observation: Linear function approximation with provable sample efficiency, in International Conference on Machine Learning, PMLR, pp

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:15.670739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.317903Z digest=sha256:78bfec69826e35a057888111601e4136c0b04aa7b83797499ee466475ae48272

Observation d2bf2f97-5c20-4453-ad01-9ca77c8cee3c · outbound

This paper cites and Qi, Z.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding and Qi, Z

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:15.654763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.322320Z digest=sha256:ca6ac4b77932414c75189c75f7d9e509d14231637c8db284887f105cc081591d

Observation b0d9fe11-a3c4-40b0-8ad2-6ba8367df1f9 · outbound

This paper cites an unresolved cited work.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:54:15.638815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.326721Z digest=sha256:ff211e177be1e0afadde06e42ecdb277ae6d6bdb084baecac077d8cb85f9400a

Observation 4e3b8511-e91b-4714-bcc7-1e34e89d429f · outbound

This paper cites (2023), Semiparametric proximal causal inference, Journal of the American Statistical Association, 1--12.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding (2023), Semiparametric proximal causal inference, Journal of the American Statistical Association, 1--12

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:15.623843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.331138Z digest=sha256:b162a115e46c6da8ae5a630e5b0cf816c52772395355d107faaf2332a15f8091

Observation e9046a6a-e76e-46f2-93c4-d7e0f10f1179 · outbound

This paper cites and Tchetgen Tchetgen, E.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding and Tchetgen Tchetgen, E

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T04:54:14.335429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:54:14.335429Z digest=sha256:89a0b28050275a26b82d5974ecc46c89be0667cbb92b93966d7e91d9d5803937

Observation 29ef32ba-ce66-481e-a0ba-3e3a1705fd6d · outbound

This paper cites (2020), Minimax estimation of conditional moment models, Advances in Neural Information Processing Systems, 33, 12248--12262.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding (2020), Minimax estimation of conditional moment models, Advances in Neural Information Processing Systems, 33, 12248--12262

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:15.597816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.339985Z digest=sha256:22fb422300cb171e32598080244c80466f916924407571c697c04775776ec620

Observation 9867b87c-580b-44ac-83f4-0b8b48f2b823 · outbound

This paper cites (2011), On the completeness condition in nonparametric instrumental problems, Econometric Theory, 27, 460--471.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding (2011), On the completeness condition in nonparametric instrumental problems, Econometric Theory, 27, 460--471

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:15.583087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.344496Z digest=sha256:a27593ee574880a4f48427214fac3784221ea98b696e04182306639c3e3e3ff5

Observation ec68a9cb-3a54-4b1e-87c4-b4a7ba7c94a1 · outbound

This paper cites (2016), Regularized policy iteration with nonparametric function spaces, Journal of Machine Learning Research, 17, 1--66.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding (2016), Regularized policy iteration with nonparametric function spaces, Journal of Machine Learning Research, 17, 1--66

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:15.567752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.348905Z digest=sha256:dfb2d8f1c2ac5be81ec0e5de49eec1635ac36f6abd6960114e1e0d08619a99eb

Observation d2ed938a-2311-4043-9762-36a090b59f61 · outbound

This paper cites H., Moreau, Y., Murphy, S.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding H., Moreau, Y., Murphy, S

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:15.551560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.353247Z digest=sha256:f7bf7a46945b863cb15e64c81570ca7dbc43e976e0674b400a101e713e67373e

Observation a4e09640-c797-485e-8822-c14d95c61cff · outbound

This paper cites Offline Reinforcement Learning with Instrumental Variables in Confounded Markov Decision Processes.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding Offline Reinforcement Learning with Instrumental Variables in Confounded Markov Decision Processes

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T04:54:14.357740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:54:14.357740Z digest=sha256:24bc824573bc0c82f310411d4baf905110f6009f8645ee223509b53b0f19d9c6

Observation 97ab3828-eec7-472b-bb43-0f08ce7f9dc4 · outbound

This paper cites and Gu, S.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding and Gu, S

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:15.536508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.362574Z digest=sha256:75e4d331d8175fae205b2f8828e7f238eb762112d153aa02368ff1753e753af9

Observation 22989646-d206-43d8-bf02-e108ba7e524b · outbound

This paper cites an unresolved cited work.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:54:15.521397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.367001Z digest=sha256:d2a42aac20d9ac8e5cca22a3cb2ab8bac17a7fbf89a80f3cc5157608d852251d

Observation ef9c2fa8-2725-4236-b75d-1b1b5b19a955 · outbound

This paper cites an unresolved cited work.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:54:15.506348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.371400Z digest=sha256:023ca400a42b4326be99e127904869187b72e048c556a23da2c3ffd2930db47d

Observation 7a388575-444c-4319-a6c5-0497d81ae671 · outbound

This paper cites (2022), Provably efficient offline reinforcement learning for partially observable markov decision processes, in International Conference on Machine Learning, PMLR, pp.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding (2022), Provably efficient offline reinforcement learning for partially observable markov decision processes, in International Conference on Machine Learning, PMLR, pp

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:15.491679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.375883Z digest=sha256:f02ed86c69c5b91c014230922ccbe240433f7ee845eb0f22cc5670f13d50f349

Observation e918f253-246e-4e1c-9c86-be1d7007cfde · outbound

This paper cites Soft Actor-Critic Algorithms and Applications.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding Soft Actor-Critic Algorithms and Applications

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-16T04:54:14.380110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:54:14.380110Z digest=sha256:6bd33e1d68141711a5ba387bc4b65474cd557a0eb20bcdd93e8f877a07ea98ce

Observation a0c36dbc-7f04-4b38-b8ec-feabeff990fe · outbound

This paper cites (2021), Bootstrapping fitted q-evaluation for off-policy inference, in International Conference on Machine Learning, PMLR, pp.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding (2021), Bootstrapping fitted q-evaluation for off-policy inference, in International Conference on Machine Learning, PMLR, pp

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:15.476082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.384888Z digest=sha256:2e33ad4c140148fa457988aeb998af963e1392729ec46e50a732d2eefb73d8b0

Observation f542dcb4-c1a5-4151-b49c-b8e10f1fc051 · outbound

This paper cites W., Lazaric, A., Ghavamzadeh, M., and Munos, R.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding W., Lazaric, A., Ghavamzadeh, M., and Munos, R

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:15.461072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.390238Z digest=sha256:e9d13182ace334da57c2f7adef5c75f5638186f6e8b6fa9fb2965f2df6c6d558

Observation 2c93aa95-6b44-413d-b947-93e017b98bb9 · outbound

This paper cites A Policy Gradient Method for Confounded POMDPs.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding A Policy Gradient Method for Confounded POMDPs

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-16T04:54:14.394823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:54:14.394823Z digest=sha256:7a38569c6f9f2a11705675faaf2c4d67848e87a60c4f80d5f2e80a04aa636208

Observation 4c83a170-793b-4e33-9c36-92dee9e03d32 · outbound

This paper cites (2020), Sample-efficient reinforcement learning of undercomplete pomdps, Advances in Neural Information Processing Systems, 33, 18530--18539.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding (2020), Sample-efficient reinforcement learning of undercomplete pomdps, Advances in Neural Information Processing Systems, 33, 18530--18539

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:15.445717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.399509Z digest=sha256:4a80b4043bc212184e7cb1a93ea439a73fb359fc496b84eee5da036fe1bf0f17

Observation 679dbfa6-731c-40af-aaa2-0efcd270d2e8 · outbound

This paper cites (2022), Doubly robust distributionally robust off-policy evaluation and learning, in International Conference on Machine Learning, PMLR, pp.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding (2022), Doubly robust distributionally robust off-policy evaluation and learning, in International Conference on Machine Learning, PMLR, pp

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:15.431262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.403922Z digest=sha256:db57a94d5e5eaffd3ea158a000ab8c6eb6a61a324813be5885161ca7a8fab06a

Observation 41912fe6-7a55-40dc-bba5-46b1dde2d57c · outbound

This paper cites and Uehara, M.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding and Uehara, M

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:15.416571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.408243Z digest=sha256:995bdbe47c67759a6b2b4e6312c8004d0424c9af4a86f87fde573ebfdd0ba610

Observation ad1d724f-0a8b-40d4-a0f1-338d96381689 · outbound

This paper cites and Zhou, A.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding and Zhou, A

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:15.401782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.412788Z digest=sha256:23fb2aea8da8cfb41822f654f6113e78d34786d32f1a11c1528b505cc74e1e1f

Observation 770b2abc-3d1e-4920-bab1-b7e6034d2b1b · outbound

This paper cites an unresolved cited work.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:54:15.386488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.417271Z digest=sha256:31bacb693f670d047d3e4c67f1ecf84aa710874863d6b6e801430d7a2c81f5df

Observation 86f2fbbe-2e73-422d-82bb-c9ee1b033d39 · outbound

This paper cites B., Volkmann, J.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding B., Volkmann, J

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:15.371758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.421631Z digest=sha256:c131b353791d513576ffb80dd8edb7dc6bef48a861df9304ec9916a51e29e411

Observation 841944cf-839a-4b90-9e6d-a176aa198162 · outbound

This paper cites Offline Reinforcement Learning with Implicit Q-Learning.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding Offline Reinforcement Learning with Implicit Q-Learning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-16T04:54:14.426021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:54:14.426021Z digest=sha256:7f30e87f99433ba1597ea0c327acfbc1c1b85503d8843f2e191a493e72bf36c3

Observation 8864b3c9-f08c-4ef0-a21a-8980cb3ed206 · outbound

This paper cites (1989), Linear integral equations, vol.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding (1989), Linear integral equations, vol

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:15.357186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.430679Z digest=sha256:52b905f0e6af64778133ed7836afb188586e276ac98dff9f87b188fea3b75552

Observation 03655345-0d94-4581-9c58-68eb06c787cf · outbound

This paper cites (2020), Conservative q-learning for offline reinforcement learning, Advances in Neural Information Processing Systems, 33, 1179--1191.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding (2020), Conservative q-learning for offline reinforcement learning, Advances in Neural Information Processing Systems, 33, 1179--1191

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:15.342241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.435075Z digest=sha256:c894307d2468eb8a2dfd7ec244a42950894479995c95a3175cd9e0899c2290c7

Observation 8849dbc8-95c4-47ae-9032-6fe596954bd0 · outbound

This paper cites an unresolved cited work.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:54:15.327486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.439297Z digest=sha256:f9f0e8e050932582c30ba9e98e620ac91d82d8d45d6b8274e17922a9b84f8f28

Observation f88e3568-60d5-4d53-8989-91dbb3c9016c · outbound

This paper cites (2019), Batch policy learning under constraints, in International Conference on Machine Learning, PMLR, pp.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding (2019), Batch policy learning under constraints, in International Conference on Machine Learning, PMLR, pp

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:15.312552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.443453Z digest=sha256:a5f89575366e6fcd260d570b34164831e2a453d786c264fba73edc7baee17a09

Observation b7f6d36c-6ed5-467d-a10f-e67b1502c283 · outbound

This paper cites (2018), Deep reinforcement learning in continuous action spaces: a case study in the game of simulated curling, in International conference on machine learning, PMLR, pp.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding (2018), Deep reinforcement learning in continuous action spaces: a case study in the game of simulated curling, in International conference on machine learning, PMLR, pp

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:15.298204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.448080Z digest=sha256:daaeb22834acf91ed8a6470c650125983300f468601e37dc012ab1abcca9b781

Observation 8048a3b7-f55b-4017-9e54-1f6617564eb2 · outbound

This paper cites Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-16T04:54:14.452593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:54:14.452593Z digest=sha256:4c77288149e0de8bc1aed71fda550d7dbf4d9e1e42a8c41bc00ea819464a8281

Observation 14bd91e6-4c14-4b02-9245-b2dc72627b3f · outbound

This paper cites Asymptotic Theory for IV-Based Reinforcement Learning with Potential Endogeneity.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding Asymptotic Theory for IV-Based Reinforcement Learning with Potential Endogeneity

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-16T04:54:14.457368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:54:14.457368Z digest=sha256:270165132cfc221e9cb726211355978cbcc13f60dcdb5eeb51a8880d3005e7b6

Observation 70bb7568-6d3e-4162-86a0-2cb47449696a · outbound

This paper cites (2023), Quasi-optimal Reinforcement Learning with Continuous Actions, in The Eleventh International Conference on Learning Representations.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding (2023), Quasi-optimal Reinforcement Learning with Continuous Actions, in The Eleventh International Conference on Learning Representations

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:15.282669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.462291Z digest=sha256:a801e359cd1c0614f129cab13187c69e5f91cf15c2e7156df539e4c15428f2a2

Observation d0436792-f256-4aa7-bdde-4aeb45b17928 · outbound

This paper cites (2021), Off-policy estimation of long-term average outcomes with applications to mobile health, Journal of the American Statistical Association, 116, 382--391.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding (2021), Off-policy estimation of long-term average outcomes with applications to mobile health, Journal of the American Statistical Association, 116, 382--391

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:15.265633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.466677Z digest=sha256:1797f647607a8df00d2d30e461dced6c5d95537a47a5e9da159b443ec5ab111a

Observation af9514e0-1a5c-4152-88ea-09e1ae96af7c · outbound

This paper cites an unresolved cited work.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:54:15.250381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.470965Z digest=sha256:8ba69cc665bcdb68717867f3ce5d4c94c8488bea7eaf116a33d8b2f18d026ff7

Observation 73b4f8ed-f353-4222-9a7e-006ea9ebf242 · outbound

This paper cites Continuous control with deep reinforcement learning.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding Continuous control with deep reinforcement learning

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-16T04:54:14.475241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:54:14.475241Z digest=sha256:b7961ad90baca1b8d73f734fb37831440beebb8d89774bbc3c24bab4185ca75f

Observation 0e94eefd-784f-4d04-9bfa-9acd74d2424f · outbound

This paper cites (2018), Breaking the curse of horizon: Infinite-horizon off-policy estimation, Advances in neural information processing systems, 31.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding (2018), Breaking the curse of horizon: Infinite-horizon off-policy estimation, Advances in neural information processing systems, 31

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:15.234853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.479814Z digest=sha256:85f334e6b35f8fdf3d75112717b98dc27451f779c1d8afb46889f505edcbf870

Observation 2a5f2215-1b5a-4ce8-9779-ade6115e83c1 · outbound

This paper cites Pessimism in the Face of Confounders: Provably Efficient Offline Reinforcement Learning in Partially Observable Markov Decision Processes.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding Pessimism in the Face of Confounders: Provably Efficient Offline Reinforcement Learning in Partially Observable Markov Decision Processes

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-16T04:54:14.484531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:54:14.484531Z digest=sha256:c46f492f7bb3e6da256e0aadc273628371a8b70fb02af794d64ecaf7cc925895

Observation 0f8924a3-0a2d-4df5-92b6-271946ed8cb1 · outbound

This paper cites J., Laber, E.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding J., Laber, E

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:15.219656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.489344Z digest=sha256:2e4090494322a2d5bea5b987d8ae3f7d55ba8e0739c8cd131b37ad50a553b724

Observation 2d7fbf26-4a78-4e5b-9105-730ec6aad78d · outbound

This paper cites an unresolved cited work.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:54:15.203969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.493813Z digest=sha256:d910dadf40553f3fb2bd4cac1cebdb536e95e977ddf7ea828b9584e4d0c54f04

Observation 356b302c-a8b7-4be8-aa11-789b56f312a8 · outbound

This paper cites an unresolved cited work.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding Unresolved cited work

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-16T04:54:14.498384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:54:14.498384Z digest=sha256:019a558086f96ebdf027195f4099087f6724f647e5ea57408c7c5707058c38e9

Observation 93f237ba-a115-49e9-98e9-9a4756e8fc95 · outbound

This paper cites A., Veness, J., Bellemare, M.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding A., Veness, J., Bellemare, M

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:15.179126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.503080Z digest=sha256:f73f8a023354778c07e017fea9320de6f5d3e360e9bdb84965e8f3763812c3f0

Observation c198ad41-3aaa-4cbb-9fb4-b77187b50c60 · outbound

This paper cites an unresolved cited work.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding Unresolved cited work

Reference 56

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:54:15.164405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.507453Z digest=sha256:89362ba50f5f4690281d42d2ae63e8cfee2bdc3fbe7547e723bf725dbd3383af

Observation ea625b0f-788b-4480-ac5c-fa78767a645b · outbound

This paper cites an unresolved cited work.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding Unresolved cited work

Reference 57

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:54:15.149351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.511795Z digest=sha256:c1bd6f186c0db6c070cc1e955394d2028a795bc325671a9f31caecdcfa3124aa

Observation e4a26aa4-66a6-4696-8ef2-d0abb2353646 · outbound

This paper cites (2000), Eligibility traces for off-policy policy evaluation, Computer Science Department Faculty Publication Series, 80.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding (2000), Eligibility traces for off-policy policy evaluation, Computer Science Department Faculty Publication Series, 80

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:15.134256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.516271Z digest=sha256:f9fe32620b3d5e22effd41faab9619dfbb3dd1282ff98526a0eddb7fd3c0b3d7

Observation eb3bbccb-0449-4776-98fa-674978f528d4 · outbound

This paper cites (2023), Proximal learning for individualized treatment regimes under unmeasured confounding, Journal of the American Statistical Association, 1--14.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding (2023), Proximal learning for individualized treatment regimes under unmeasured confounding, Journal of the American Statistical Association, 1--14

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:15.118195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.520939Z digest=sha256:4b2e6bab19a4815db957d931d64bed0e61fec4c89b09e3bbb24ff8daeb04db72

Observation d3d81530-ba02-40af-b8b8-e5fa33af6952 · outbound

This paper cites an unresolved cited work.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding Unresolved cited work

Reference 60

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:54:15.102803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.525416Z digest=sha256:fcaa0c775a1f32aa7ef55e9f7a7a5a6cebf8709b9d889d330f4f2d505e99e456

Observation 97084137-1bcb-48e2-900e-7e0333532a06 · outbound

This paper cites an unresolved cited work.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding Unresolved cited work

Reference 61

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:54:15.087809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.529849Z digest=sha256:c3fc5067d5da0225600e80df78211c27ae7a40f7b420f2482a03d35f87a279bf

Observation 07b18144-1ed5-4b4a-9859-895647d7f1fa · outbound

This paper cites (2022 c ), Off-policy confidence interval estimation with confounded markov decision process, Journal of the American Statistical Association, 1--12.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding (2022 c ), Off-policy confidence interval estimation with confounded markov decision process, Journal of the American Statistical Association, 1--12

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:15.071842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.534202Z digest=sha256:64a55d5b5ba75587f2db63e026fd84b2cdecd7fa2d767a29b63769b226d9870f

Observation 8eb2c479-d3e8-497c-8163-4f098ac330b4 · outbound

This paper cites an unresolved cited work.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding Unresolved cited work

Reference 63

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:54:15.056559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.538514Z digest=sha256:ff2b438e1ae3c0ede0d71106259d06cdc941a79ed72e5943365d1690ce8469ac

Observation a206f1f6-67d3-45de-b359-cffd17e5a38e · outbound

This paper cites An Introduction to Proximal Causal Learning.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding An Introduction to Proximal Causal Learning

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-16T04:54:14.542761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:54:14.542761Z digest=sha256:4c3e98094e2fbbeac8510ffe86765563f4f9ea78218364f41208beeadcceced8

Observation 0e6742b5-5241-4b4f-b984-ea5e1d6c44f3 · outbound

This paper cites and Brunskill, E.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding and Brunskill, E

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:15.037610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.547332Z digest=sha256:f2433c718c3dbe47b854c2de31ad75de314d246acbc997f923d6e998f224b016

Observation b22ce68a-5aad-4f34-a0d5-285f3cf0ffdd · outbound

This paper cites (2020), Minimax weight and q-function learning for off-policy evaluation, in International Conference on Machine Learning, PMLR, pp.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding (2020), Minimax weight and q-function learning for off-policy evaluation, in International Conference on Machine Learning, PMLR, pp

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:15.020578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.551671Z digest=sha256:90d72923a33301e01db4fec0f7fc3f2a59737686ffdbdce62678035ad3157346

Observation 0b3f6020-ec07-4882-a829-8ed66983dec8 · outbound

This paper cites (2024), Future-dependent value-based off-policy evaluation in pomdps, Advances in Neural Information Processing Systems, 36.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding (2024), Future-dependent value-based off-policy evaluation in pomdps, Advances in Neural Information Processing Systems, 36

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:15.000899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.556461Z digest=sha256:433cbc44b57d81881cc26517b24fdda9b695ec2ba0389f487ecd0ccda423b26a

Observation ed0aa84a-7f48-49c7-993a-2d307ef8202c · outbound

This paper cites and Groothuis-Oudshoorn, K.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding and Groothuis-Oudshoorn, K

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:14.985762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.560704Z digest=sha256:7379b7386566ca0714ab0b0a0889e725177198c24fcaa9352569c562292aa271

Observation aff82c72-59d1-47d5-907b-b6d4222f706c · outbound

This paper cites and Zou, S.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding and Zou, S

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:14.971136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.565105Z digest=sha256:6c4bfc4ffbd95ce12b7a79e8f70f0dd9621ad5cc55ba3672caf507c1fa5a5499

Observation 3e1da517-73bf-45f0-a826-4367b3fdb266 · outbound

This paper cites (2019), Towards optimal off-policy evaluation for reinforcement learning with marginalized importance sampling, Advances in neural information processing systems, 32.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding (2019), Towards optimal off-policy evaluation for reinforcement learning with marginalized importance sampling, Advances in neural information processing systems, 32

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:14.955283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.569342Z digest=sha256:d51578fe86c6cf91c7c68ae3e9dc59a801a7f9d258605ab430cc007364bd5b44

Observation 0a1d9ce9-2d9b-4746-8d99-b9c128bb581c · outbound

This paper cites (2023), An instrumental variable approach to confounded off-policy evaluation, in International Conference on Machine Learning, PMLR, pp.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding (2023), An instrumental variable approach to confounded off-policy evaluation, in International Conference on Machine Learning, PMLR, pp

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:14.939449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.573726Z digest=sha256:609d81d5c35ce5c497098004b28b6b5bf664193df748239257a30d3a40b4dff7

Observation b62d42a6-c9b9-4314-a023-203d356b22a9 · outbound

This paper cites and Bareinboim, E.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding and Bareinboim, E

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:14.924013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.578035Z digest=sha256:ca8b68a7419ada5ca110348d5a02d4b783cfa4a803b085b0b8c7930cc5b00b65

Observation 686d4d10-0ba5-4f96-a399-6911a249dcc7 · outbound

This paper cites (2020), Causal imitation learning with unobserved confounders, Advances in neural information processing systems, 33, 12263--12274.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding (2020), Causal imitation learning with unobserved confounders, Advances in neural information processing systems, 33, 12263--12274

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:14.907724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.582444Z digest=sha256:fc8d426f97fa455e666e92f7122e47dd17c1a3c19448cdd8dc235b22de68fa46

Observation 98108e71-ba89-4819-8277-83b9d4093515 · outbound

This paper cites On the Curses of Future and History in Future-dependent Value Functions for Off-policy Evaluation.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding On the Curses of Future and History in Future-dependent Value Functions for Off-policy Evaluation

Reference 74

Resolution
verified exact
local_arxiv, observed 2026-08-16T04:54:14.671602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.586780Z digest=sha256:73d5646ac98fa6ccd3cacc753508c6579e8ecc9352875cd839607bfd56e3ee52

Observation a9b94b24-e0ac-4371-8991-def5ab75dc3b · outbound

This paper cites (2024), Bi-Level Offline Policy Optimization with Limited Exploration, Advances in Neural Information Processing Systems, 36.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding (2024), Bi-Level Offline Policy Optimization with Limited Exploration, Advances in Neural Information Processing Systems, 36

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:14.891413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.592125Z digest=sha256:445618e7293cecbcddc80fd6b57f82f962e9c4639854afb45cadb5b3d62bfd6c

Observation 8d09bd4e-b1cd-40e3-9026-7d253601b194 · outbound

This paper cites (2024 a ), Policy learning for individualized treatment regimes on infinite time horizon, in Statistics in Precision Health: Theory, Methods and Applications, Springer, pp.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding (2024 a ), Policy learning for individualized treatment regimes on infinite time horizon, in Statistics in Precision Health: Theory, Methods and Applications, Springer, pp

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:14.874788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.596573Z digest=sha256:f8c2202521d17d9422b2dda9081a57b9bb94df68c7aff371136f5180fcbd6ccd

Observation 5c19e899-4c9a-4945-8ca5-103002d52080 · outbound

This paper cites Distributional Shift-Aware Off-Policy Interval Estimation: A Unified Error Quantification Framework.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding Distributional Shift-Aware Off-Policy Interval Estimation: A Unified Error Quantification Framework

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-16T04:54:14.601000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:54:14.601000Z digest=sha256:cb7ef4ceb2919689157b2cf4ff44695fd720555e0059e1d033e9afaa6be053e5

Observation 6fe62085-e851-42d9-827b-f2b8ace0bd17 · outbound

This paper cites (2024 b ), Estimating optimal infinite horizon dynamic treatment regimes via pt-learning, Journal of the American Statistical Association, 119, 625--638.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding (2024 b ), Estimating optimal infinite horizon dynamic treatment regimes via pt-learning, Journal of the American Statistical Association, 119, 625--638

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:14.857615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.605462Z digest=sha256:750bf9d4579d1c5c183b50021e5bef6f458522feadaa403b3e1b296c25ea5625

Observation 7798237d-f28a-4154-8fe1-f5271244b8da · outbound

This paper cites (2020), Safe, efficient, and comfortable velocity control based on reinforcement learning for autonomous driving, Transportation Research Part C: Emerging Technologies, 117, 102662.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding (2020), Safe, efficient, and comfortable velocity control based on reinforcement learning for autonomous driving, Transportation Research Part C: Emerging Technologies, 117, 102662

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:14.841403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.609941Z digest=sha256:e8abe78d56af0521c42f5ddcd3280267afc4f0c8cb1e6692be5515a359135a50

Pith citing papers

No inbound Pith citation observations are available.