Pith. sign in

Paper Citation Record · LEDGER

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding

As of 23 August 2026, this Paper Citation Record lists 79 of 79 outbound references and 0 inbound Pith citation observations for arXiv:2505.00304.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.00304 v1

Coverage vector

measured 79 of 79 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T04:54:14.609941Z

measured 79 of 79 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

79 of 79 outbound references displayed

  • verified exact2
  • verified fuzzy47
  • unresolved30
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cd4073d0-0038-4952-adf9-63822c6ee9da · outbound

This paper cites , " * write output.state after.block = add.period write newline.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding , " * write output.state after.block = add.period write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T04:54:14.256109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:54:14.256109Z digest=sha256:f1ae6ae2886c6f11346353ed6e802ad623cfe66dae1c61251bb476c888737e81

Observation 919d6f0d-5353-48de-8699-761051d85bab · outbound

This paper cites write newline.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T04:54:14.261809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:54:14.261809Z digest=sha256:3cc3e79f35ee1f1c2d236e3687e0e7aa250bbebffad0802442b2a22257cfd8dd

Observation 410f8da7-edd5-4ad3-9709-f6359a423471 · outbound

This paper cites an unresolved cited work.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:54:15.823487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.266879Z digest=sha256:d7b0e1221d85382833613dd1e883e79b40de38bd013ead60e494c7cc7173b4d6

Observation 6fb5d4e1-7c06-417a-ad6f-51111b68b2b2 · outbound

This paper cites (2008), Learning near-optimal policies with Bellman-residual minimization based fitted policy iteration and a single sample path, Machine Learning, 71, 89--129.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding (2008), Learning near-optimal policies with Bellman-residual minimization based fitted policy iteration and a single sample path, Machine Learning, 71, 89--129

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:15.806643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.271354Z digest=sha256:4677bdaf3d5b37afb159e5f2f06bd8f00d23908270fe782eb751739c78882f78

Observation 4e3b1f35-75ae-4d39-bac0-44115f865f80 · outbound

This paper cites (2017), Breaking the curse of dimensionality with convex neural networks, The Journal of Machine Learning Research, 18, 629--681.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding (2017), Breaking the curse of dimensionality with convex neural networks, The Journal of Machine Learning Research, 18, 629--681

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:15.790796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.275782Z digest=sha256:f261e9878a3c9f1bd5cae963957643316b06dd948c8051f4969beb6e76d070e5

Observation 7b1ecb43-2fc0-4ad4-afc2-de323f09bd6d · outbound

This paper cites an unresolved cited work.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:54:15.775627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.280602Z digest=sha256:c39cf0d827bd5a694d697c0f0454c2b93e3b47a49d4cce7ef5210808f225f405

Observation f3845a6a-eff6-4af7-b773-d37f0917cade · outbound

This paper cites and Kallus, N.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding and Kallus, N

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:15.760428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.285139Z digest=sha256:9eb010bbb1920a006c82f0edcd4e264a0052daa9cfcdeefd2d8d63717f670722

Observation 2bfdfeb8-ce6d-418d-9481-3545b554ba51 · outbound

This paper cites (2021), Off-policy evaluation in infinite-horizon reinforcement learning with latent confounders, in International Conference on Artificial Intelligence and Statistics, PMLR, pp.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding (2021), Off-policy evaluation in infinite-horizon reinforcement learning with latent confounders, in International Conference on Artificial Intelligence and Statistics, PMLR, pp

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:15.745039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.289645Z digest=sha256:7f61d81d9d8d802b15a9fe4177468faecc134cc9c60cfcd05b9ae4601fed8a73

Observation b13a54ca-1684-4dbe-9b79-434d67a8905e · outbound

This paper cites and Kennedy, E.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding and Kennedy, E

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T04:54:14.294412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:54:14.294412Z digest=sha256:8218e09038e63031736c6a7381711afc13e0b3071ed41a73a63c39ce993fe629

Observation 466524b1-2ecf-4ea2-b1dd-f7190268ce50 · outbound

This paper cites u derl, J., Schmiedeberg, C., Castiglioni, L., Arr \'a nz Becker, O., Buhr, P., Fu , D., Ludwig, V., Schr \.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding u derl, J., Schmiedeberg, C., Castiglioni, L., Arr \'a nz Becker, O., Buhr, P., Fu , D., Ludwig, V., Schr \

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:15.717702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.298932Z digest=sha256:8ff6f2853de138922d47bcc9f5d3cfb8b543d805b30659154f6f39cdae241746

Observation 11852540-b5e5-49c5-b25f-bf5ea33bb747 · outbound

This paper cites W., Yuan, Z., Zhou, S., Panerati, J., and Schoellig, A.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding W., Yuan, Z., Zhou, S., Panerati, J., and Schoellig, A

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:15.702282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.303552Z digest=sha256:d8502d661c6ce4bec380bae7a256720c1e9d9d60811886a2804a5692c027c303

Observation ba310e6a-5471-49cb-849a-b1277b746cba · outbound

This paper cites an unresolved cited work.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:54:15.686011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.308223Z digest=sha256:92acd72b01fb0553eab7e930356a75322e4f4daa758a6ee2f36d3f9d8143c48a

Observation 5ad80ee7-1129-4815-a00b-59980fb968fc · outbound

This paper cites Jump Interval-Learning for Individualized Decision Making.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding Jump Interval-Learning for Individualized Decision Making

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-08-16T04:54:14.824601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.312790Z digest=sha256:76fa3e630483efc0b71f0692cb7d4b7055a4a33711f09017e985bac295d5d628

Observation 818cf5dd-c8e1-4f46-a3c2-c31add3c0577 · outbound

This paper cites (2022), Reinforcement learning from partial observation: Linear function approximation with provable sample efficiency, in International Conference on Machine Learning, PMLR, pp.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding (2022), Reinforcement learning from partial observation: Linear function approximation with provable sample efficiency, in International Conference on Machine Learning, PMLR, pp

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:15.670739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.317903Z digest=sha256:94eb59fbaffc0138af309078ff5651179ea698503678e4804380c7bda26a1546

Observation d2bf2f97-5c20-4453-ad01-9ca77c8cee3c · outbound

This paper cites and Qi, Z.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding and Qi, Z

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:15.654763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.322320Z digest=sha256:44328dbf11a40d936b5c3ecfa321d59dd99c286f7e8cdd5350651b79d031d630

Observation b0d9fe11-a3c4-40b0-8ad2-6ba8367df1f9 · outbound

This paper cites an unresolved cited work.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:54:15.638815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.326721Z digest=sha256:d2b9418cd3a7844133bfe770209a38f0cc0dda08f5d441d4c0dcf86c0c444415

Observation 4e3b8511-e91b-4714-bcc7-1e34e89d429f · outbound

This paper cites (2023), Semiparametric proximal causal inference, Journal of the American Statistical Association, 1--12.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding (2023), Semiparametric proximal causal inference, Journal of the American Statistical Association, 1--12

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:15.623843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.331138Z digest=sha256:323bd2e76580f6d555a1e17a79ee2a2fe7fefbc6d295b2cab09e4e8a732fef29

Observation e9046a6a-e76e-46f2-93c4-d7e0f10f1179 · outbound

This paper cites and Tchetgen Tchetgen, E.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding and Tchetgen Tchetgen, E

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T04:54:14.335429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:54:14.335429Z digest=sha256:89a0b28050275a26b82d5974ecc46c89be0667cbb92b93966d7e91d9d5803937

Observation 29ef32ba-ce66-481e-a0ba-3e3a1705fd6d · outbound

This paper cites (2020), Minimax estimation of conditional moment models, Advances in Neural Information Processing Systems, 33, 12248--12262.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding (2020), Minimax estimation of conditional moment models, Advances in Neural Information Processing Systems, 33, 12248--12262

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:15.597816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.339985Z digest=sha256:9b91eb786692c833c1a8eb55b6f4e8d92779eb8dc658c8f9ecc08ff25bca1983

Observation 9867b87c-580b-44ac-83f4-0b8b48f2b823 · outbound

This paper cites (2011), On the completeness condition in nonparametric instrumental problems, Econometric Theory, 27, 460--471.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding (2011), On the completeness condition in nonparametric instrumental problems, Econometric Theory, 27, 460--471

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:15.583087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.344496Z digest=sha256:573f3511ef89350b5c190e258c7cfae4feeba122f8260a6becd96fe5e0d1d7bb

Observation ec68a9cb-3a54-4b1e-87c4-b4a7ba7c94a1 · outbound

This paper cites (2016), Regularized policy iteration with nonparametric function spaces, Journal of Machine Learning Research, 17, 1--66.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding (2016), Regularized policy iteration with nonparametric function spaces, Journal of Machine Learning Research, 17, 1--66

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:15.567752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.348905Z digest=sha256:d63650049bad5476d6776eb4511f68a0e050ebed0c15b9bbfc9580f1f51d2580

Observation d2ed938a-2311-4043-9762-36a090b59f61 · outbound

This paper cites H., Moreau, Y., Murphy, S.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding H., Moreau, Y., Murphy, S

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:15.551560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.353247Z digest=sha256:ed0822e83aa7e10bca70e7430b0167b2b81ee62553a35f07431bc34c4b14640b

Observation a4e09640-c797-485e-8822-c14d95c61cff · outbound

This paper cites Offline Reinforcement Learning with Instrumental Variables in Confounded Markov Decision Processes.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding Offline Reinforcement Learning with Instrumental Variables in Confounded Markov Decision Processes

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T04:54:14.357740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:54:14.357740Z digest=sha256:ad68b1817ba11874bfd0da306fea65bba314a312b25b4fe00226a0245ce01bb6

Observation 97ab3828-eec7-472b-bb43-0f08ce7f9dc4 · outbound

This paper cites and Gu, S.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding and Gu, S

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:15.536508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.362574Z digest=sha256:3d54b697c756f8219309f5cf938f03d05b2b81798055685459572dcdc4c0628a

Observation 22989646-d206-43d8-bf02-e108ba7e524b · outbound

This paper cites an unresolved cited work.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:54:15.521397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.367001Z digest=sha256:13e13178091c478e1e29aed4b73d06b1d7aa58965d960f2b499f9e6daa670f23

Observation ef9c2fa8-2725-4236-b75d-1b1b5b19a955 · outbound

This paper cites an unresolved cited work.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:54:15.506348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.371400Z digest=sha256:5a2ce626cbed049b94ef29e5b47d24574ba8c23b76e6bb7bae5b691215b358b3

Observation 7a388575-444c-4319-a6c5-0497d81ae671 · outbound

This paper cites (2022), Provably efficient offline reinforcement learning for partially observable markov decision processes, in International Conference on Machine Learning, PMLR, pp.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding (2022), Provably efficient offline reinforcement learning for partially observable markov decision processes, in International Conference on Machine Learning, PMLR, pp

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:15.491679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.375883Z digest=sha256:2723df9d0c159a8489b3b56d9d73a834b4ad1c812709382003626dce3b8a0269

Observation e918f253-246e-4e1c-9c86-be1d7007cfde · outbound

This paper cites Soft Actor-Critic Algorithms and Applications.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding Soft Actor-Critic Algorithms and Applications

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-16T04:54:14.380110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:54:14.380110Z digest=sha256:6bd33e1d68141711a5ba387bc4b65474cd557a0eb20bcdd93e8f877a07ea98ce

Observation a0c36dbc-7f04-4b38-b8ec-feabeff990fe · outbound

This paper cites (2021), Bootstrapping fitted q-evaluation for off-policy inference, in International Conference on Machine Learning, PMLR, pp.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding (2021), Bootstrapping fitted q-evaluation for off-policy inference, in International Conference on Machine Learning, PMLR, pp

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:15.476082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.384888Z digest=sha256:bb3b1e19ddc503b31c56d9a246cd035a920f6e4e8464cb02d17eb06ebdbba820

Observation f542dcb4-c1a5-4151-b49c-b8e10f1fc051 · outbound

This paper cites W., Lazaric, A., Ghavamzadeh, M., and Munos, R.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding W., Lazaric, A., Ghavamzadeh, M., and Munos, R

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:15.461072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.390238Z digest=sha256:609639027e5376e1f43dd61837347c0fac4574b24c1d7d2854ca68ab31cfbc48

Observation 2c93aa95-6b44-413d-b947-93e017b98bb9 · outbound

This paper cites A Policy Gradient Method for Confounded POMDPs.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding A Policy Gradient Method for Confounded POMDPs

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-16T04:54:14.394823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:54:14.394823Z digest=sha256:f79f9229b589d3f2e66bc48f39aba9c4702605c69cb72dbd370e28053c567044

Observation 4c83a170-793b-4e33-9c36-92dee9e03d32 · outbound

This paper cites (2020), Sample-efficient reinforcement learning of undercomplete pomdps, Advances in Neural Information Processing Systems, 33, 18530--18539.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding (2020), Sample-efficient reinforcement learning of undercomplete pomdps, Advances in Neural Information Processing Systems, 33, 18530--18539

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:15.445717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.399509Z digest=sha256:b39d238f6f966e72c6b1cae1114dce6a4e8e7d6c05ebdaa1d44b59623988ab39

Observation 679dbfa6-731c-40af-aaa2-0efcd270d2e8 · outbound

This paper cites (2022), Doubly robust distributionally robust off-policy evaluation and learning, in International Conference on Machine Learning, PMLR, pp.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding (2022), Doubly robust distributionally robust off-policy evaluation and learning, in International Conference on Machine Learning, PMLR, pp

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:15.431262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.403922Z digest=sha256:36ddb1dee3d1d692bd8a3ce3a5aa3fe97782b269c9f35712c3f01dfe790de004

Observation 41912fe6-7a55-40dc-bba5-46b1dde2d57c · outbound

This paper cites and Uehara, M.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding and Uehara, M

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:15.416571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.408243Z digest=sha256:f5e4f30d1de4cd1197eed73eff6c5841b6de7b28441963e9cf1e60f141ff5d4d

Observation ad1d724f-0a8b-40d4-a0f1-338d96381689 · outbound

This paper cites and Zhou, A.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding and Zhou, A

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:15.401782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.412788Z digest=sha256:6e75850331f112cb074b18e8425b9c4d8e2ad3e77a12544f5300c16234bc5ef5

Observation 770b2abc-3d1e-4920-bab1-b7e6034d2b1b · outbound

This paper cites an unresolved cited work.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:54:15.386488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.417271Z digest=sha256:6a786961ecffd35438dd93412681cddcab6944a81dcbb4720a4fcb2542d87448

Observation 86f2fbbe-2e73-422d-82bb-c9ee1b033d39 · outbound

This paper cites B., Volkmann, J.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding B., Volkmann, J

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:15.371758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.421631Z digest=sha256:25dfa353486640597391f0398505c382b755636b4d18ca6ac2944ef3b5524be9

Observation 841944cf-839a-4b90-9e6d-a176aa198162 · outbound

This paper cites Offline Reinforcement Learning with Implicit Q-Learning.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding Offline Reinforcement Learning with Implicit Q-Learning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-16T04:54:14.426021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:54:14.426021Z digest=sha256:7f30e87f99433ba1597ea0c327acfbc1c1b85503d8843f2e191a493e72bf36c3

Observation 8864b3c9-f08c-4ef0-a21a-8980cb3ed206 · outbound

This paper cites (1989), Linear integral equations, vol.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding (1989), Linear integral equations, vol

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:15.357186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.430679Z digest=sha256:bb37903144438b75ed31c1423f7e342908a2af73b64d24f3e1f4393a3beff7d0

Observation 03655345-0d94-4581-9c58-68eb06c787cf · outbound

This paper cites (2020), Conservative q-learning for offline reinforcement learning, Advances in Neural Information Processing Systems, 33, 1179--1191.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding (2020), Conservative q-learning for offline reinforcement learning, Advances in Neural Information Processing Systems, 33, 1179--1191

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:15.342241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.435075Z digest=sha256:de09292bb7412c14c7b515c848671c10e62af328b39953253ca469bc173fe1ae

Observation 8849dbc8-95c4-47ae-9032-6fe596954bd0 · outbound

This paper cites an unresolved cited work.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:54:15.327486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.439297Z digest=sha256:96b1cc6821570980d4f9bb5f362b2602538531a0c2d8a47647ac1bcee6c6c2f3

Observation f88e3568-60d5-4d53-8989-91dbb3c9016c · outbound

This paper cites (2019), Batch policy learning under constraints, in International Conference on Machine Learning, PMLR, pp.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding (2019), Batch policy learning under constraints, in International Conference on Machine Learning, PMLR, pp

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:15.312552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.443453Z digest=sha256:3e29c1364286ed1fdd31ad4d51c010824b7f5a2debe0ebde07e338c6b62e9711

Observation b7f6d36c-6ed5-467d-a10f-e67b1502c283 · outbound

This paper cites (2018), Deep reinforcement learning in continuous action spaces: a case study in the game of simulated curling, in International conference on machine learning, PMLR, pp.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding (2018), Deep reinforcement learning in continuous action spaces: a case study in the game of simulated curling, in International conference on machine learning, PMLR, pp

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:15.298204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.448080Z digest=sha256:0be59f260d08e85d393cf2adc8cf8552e273b2f7c1cc444121366d7eeb19e77c

Observation 8048a3b7-f55b-4017-9e54-1f6617564eb2 · outbound

This paper cites Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-16T04:54:14.452593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:54:14.452593Z digest=sha256:5674623d3f92e09f8df46e19f1c26826364722bab33ab1da2977a520d2eba082

Observation 14bd91e6-4c14-4b02-9245-b2dc72627b3f · outbound

This paper cites Asymptotic Theory for IV-Based Reinforcement Learning with Potential Endogeneity.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding Asymptotic Theory for IV-Based Reinforcement Learning with Potential Endogeneity

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-16T04:54:14.457368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:54:14.457368Z digest=sha256:e69a129e9379021417d076644e3319a43dccf074cbd6a70ec23d1e7112740424

Observation 70bb7568-6d3e-4162-86a0-2cb47449696a · outbound

This paper cites (2023), Quasi-optimal Reinforcement Learning with Continuous Actions, in The Eleventh International Conference on Learning Representations.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding (2023), Quasi-optimal Reinforcement Learning with Continuous Actions, in The Eleventh International Conference on Learning Representations

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:15.282669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.462291Z digest=sha256:ddced3f15ab50861516f976728f0cad19fe521f8ef2c00ccdf6d10b43aceb8f8

Observation d0436792-f256-4aa7-bdde-4aeb45b17928 · outbound

This paper cites (2021), Off-policy estimation of long-term average outcomes with applications to mobile health, Journal of the American Statistical Association, 116, 382--391.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding (2021), Off-policy estimation of long-term average outcomes with applications to mobile health, Journal of the American Statistical Association, 116, 382--391

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:15.265633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.466677Z digest=sha256:1a18dcbddcae59a3d7615d61b957c3c24904e10069597b28b5a3ae5129bf1941

Observation af9514e0-1a5c-4152-88ea-09e1ae96af7c · outbound

This paper cites an unresolved cited work.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:54:15.250381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.470965Z digest=sha256:a180c87be1bc2777c9fa04e79a5687732999002ebc9d318b77f6222abcc41b38

Observation 73b4f8ed-f353-4222-9a7e-006ea9ebf242 · outbound

This paper cites Continuous control with deep reinforcement learning.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding Continuous control with deep reinforcement learning

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-16T04:54:14.475241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:54:14.475241Z digest=sha256:b7961ad90baca1b8d73f734fb37831440beebb8d89774bbc3c24bab4185ca75f

Observation 0e94eefd-784f-4d04-9bfa-9acd74d2424f · outbound

This paper cites (2018), Breaking the curse of horizon: Infinite-horizon off-policy estimation, Advances in neural information processing systems, 31.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding (2018), Breaking the curse of horizon: Infinite-horizon off-policy estimation, Advances in neural information processing systems, 31

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:15.234853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.479814Z digest=sha256:72c4e46052489d51d8e174b0c50880b69e941eb56d31f31683579cb35890c36c

Observation 2a5f2215-1b5a-4ce8-9779-ade6115e83c1 · outbound

This paper cites Pessimism in the Face of Confounders: Provably Efficient Offline Reinforcement Learning in Partially Observable Markov Decision Processes.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding Pessimism in the Face of Confounders: Provably Efficient Offline Reinforcement Learning in Partially Observable Markov Decision Processes

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-16T04:54:14.484531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:54:14.484531Z digest=sha256:35e8e51b49beb9c1f290a7e304ba0b2f242ff759de2a7a774c28c17222a478e8

Observation 0f8924a3-0a2d-4df5-92b6-271946ed8cb1 · outbound

This paper cites J., Laber, E.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding J., Laber, E

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:15.219656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.489344Z digest=sha256:49e55b92d72ae1a2b1ccbfd5d517079596e04c7d8889c8e6b62d5c1095d31a0d

Observation 2d7fbf26-4a78-4e5b-9105-730ec6aad78d · outbound

This paper cites an unresolved cited work.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:54:15.203969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.493813Z digest=sha256:a714a0a81d0ceec5fd4e14242f5d6af7fa25044aab9b0df03984bc5225ad9d95

Observation 356b302c-a8b7-4be8-aa11-789b56f312a8 · outbound

This paper cites an unresolved cited work.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding Unresolved cited work

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-16T04:54:14.498384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:54:14.498384Z digest=sha256:019a558086f96ebdf027195f4099087f6724f647e5ea57408c7c5707058c38e9

Observation 93f237ba-a115-49e9-98e9-9a4756e8fc95 · outbound

This paper cites A., Veness, J., Bellemare, M.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding A., Veness, J., Bellemare, M

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:15.179126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.503080Z digest=sha256:abb93ca737388a7ae44b40d681666f35b1989ac75c9fd1f9d1c853d761c785cf

Observation c198ad41-3aaa-4cbb-9fb4-b77187b50c60 · outbound

This paper cites an unresolved cited work.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding Unresolved cited work

Reference 56

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:54:15.164405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.507453Z digest=sha256:feaddb585976327e5cba83c28231ca3b1a1c02a6568be089571adcc7fe8da964

Observation ea625b0f-788b-4480-ac5c-fa78767a645b · outbound

This paper cites an unresolved cited work.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding Unresolved cited work

Reference 57

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:54:15.149351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.511795Z digest=sha256:cb7d3983d25fdae57f123b713b91e1d3d1b353eff850292c6ea86d163d07328c

Observation e4a26aa4-66a6-4696-8ef2-d0abb2353646 · outbound

This paper cites (2000), Eligibility traces for off-policy policy evaluation, Computer Science Department Faculty Publication Series, 80.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding (2000), Eligibility traces for off-policy policy evaluation, Computer Science Department Faculty Publication Series, 80

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:15.134256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.516271Z digest=sha256:987ac53843456593f4daf96a2c827d96d4647e3f64a9c9d6a9d3f6fc4d8ba7fe

Observation eb3bbccb-0449-4776-98fa-674978f528d4 · outbound

This paper cites (2023), Proximal learning for individualized treatment regimes under unmeasured confounding, Journal of the American Statistical Association, 1--14.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding (2023), Proximal learning for individualized treatment regimes under unmeasured confounding, Journal of the American Statistical Association, 1--14

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:15.118195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.520939Z digest=sha256:295c7d476fc989af0df6d14d2cf799d6fbb2bd8972be3a554a38b43d2fd81b9d

Observation d3d81530-ba02-40af-b8b8-e5fa33af6952 · outbound

This paper cites an unresolved cited work.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding Unresolved cited work

Reference 60

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:54:15.102803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.525416Z digest=sha256:b98edaed634d66abe488c465b511e808e7b0060f1647032553103408ef120df0

Observation 97084137-1bcb-48e2-900e-7e0333532a06 · outbound

This paper cites an unresolved cited work.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding Unresolved cited work

Reference 61

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:54:15.087809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.529849Z digest=sha256:2d49254ac825b7260023f29e16a1e9f25ecb040db61f99618b996cbb8c68a3b7

Observation 07b18144-1ed5-4b4a-9859-895647d7f1fa · outbound

This paper cites (2022 c ), Off-policy confidence interval estimation with confounded markov decision process, Journal of the American Statistical Association, 1--12.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding (2022 c ), Off-policy confidence interval estimation with confounded markov decision process, Journal of the American Statistical Association, 1--12

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:15.071842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.534202Z digest=sha256:4bad1c6a481a4946bfbe67f1760d75d45e189ee06ab038277336b1482eab406b

Observation 8eb2c479-d3e8-497c-8163-4f098ac330b4 · outbound

This paper cites an unresolved cited work.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding Unresolved cited work

Reference 63

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:54:15.056559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.538514Z digest=sha256:e89e147b5d0af3e801de3b1fe18200327651997007f76c447cf41e036c3bc396

Observation a206f1f6-67d3-45de-b359-cffd17e5a38e · outbound

This paper cites An Introduction to Proximal Causal Learning.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding An Introduction to Proximal Causal Learning

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-16T04:54:14.542761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:54:14.542761Z digest=sha256:4c3e98094e2fbbeac8510ffe86765563f4f9ea78218364f41208beeadcceced8

Observation 0e6742b5-5241-4b4f-b984-ea5e1d6c44f3 · outbound

This paper cites and Brunskill, E.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding and Brunskill, E

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:15.037610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.547332Z digest=sha256:2551a4214c20c59803c1c05ce3f110cc7ce8c7cfe44e1df16dc09e58cc1e9133

Observation b22ce68a-5aad-4f34-a0d5-285f3cf0ffdd · outbound

This paper cites (2020), Minimax weight and q-function learning for off-policy evaluation, in International Conference on Machine Learning, PMLR, pp.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding (2020), Minimax weight and q-function learning for off-policy evaluation, in International Conference on Machine Learning, PMLR, pp

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:15.020578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.551671Z digest=sha256:00c1f725c6b8d29bab444ef56452372416aeed5ea4b3f569e181f79b0bce4c99

Observation 0b3f6020-ec07-4882-a829-8ed66983dec8 · outbound

This paper cites (2024), Future-dependent value-based off-policy evaluation in pomdps, Advances in Neural Information Processing Systems, 36.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding (2024), Future-dependent value-based off-policy evaluation in pomdps, Advances in Neural Information Processing Systems, 36

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:15.000899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.556461Z digest=sha256:c3cc819d425082513266c53587b46dd9cc9c7864395a8bb82abadefd04f1b194

Observation ed0aa84a-7f48-49c7-993a-2d307ef8202c · outbound

This paper cites and Groothuis-Oudshoorn, K.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding and Groothuis-Oudshoorn, K

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:14.985762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.560704Z digest=sha256:6721d48511aa98b63101e2be485fc3c4d6c66742bba99320587d46b52e7f85d7

Observation aff82c72-59d1-47d5-907b-b6d4222f706c · outbound

This paper cites and Zou, S.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding and Zou, S

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:14.971136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.565105Z digest=sha256:50c5c238debd6340e873f4cd1a402d05e523da0361d5b8ab6bbf5032bed3aa83

Observation 3e1da517-73bf-45f0-a826-4367b3fdb266 · outbound

This paper cites (2019), Towards optimal off-policy evaluation for reinforcement learning with marginalized importance sampling, Advances in neural information processing systems, 32.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding (2019), Towards optimal off-policy evaluation for reinforcement learning with marginalized importance sampling, Advances in neural information processing systems, 32

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:14.955283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.569342Z digest=sha256:aa05dae283c3dd88d81c0dfa08113fed32563954e2bcaee5b7718e304c135c48

Observation 0a1d9ce9-2d9b-4746-8d99-b9c128bb581c · outbound

This paper cites (2023), An instrumental variable approach to confounded off-policy evaluation, in International Conference on Machine Learning, PMLR, pp.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding (2023), An instrumental variable approach to confounded off-policy evaluation, in International Conference on Machine Learning, PMLR, pp

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:14.939449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.573726Z digest=sha256:793ab33731baeb17178512ed9ab96ac0a35bd770003bae5228d2dedca29daea7

Observation b62d42a6-c9b9-4314-a023-203d356b22a9 · outbound

This paper cites and Bareinboim, E.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding and Bareinboim, E

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:14.924013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.578035Z digest=sha256:0110ba609e9a9ed84a385707f97fad23df351474b3ec78838e9cb2676fccfbdf

Observation 686d4d10-0ba5-4f96-a399-6911a249dcc7 · outbound

This paper cites (2020), Causal imitation learning with unobserved confounders, Advances in neural information processing systems, 33, 12263--12274.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding (2020), Causal imitation learning with unobserved confounders, Advances in neural information processing systems, 33, 12263--12274

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:14.907724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.582444Z digest=sha256:a8ceb2bef02f24847be3aa1c834e9425ce180f157ed26a4e0a96f4885bb5354a

Observation 98108e71-ba89-4819-8277-83b9d4093515 · outbound

This paper cites On the Curses of Future and History in Future-dependent Value Functions for Off-policy Evaluation.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding On the Curses of Future and History in Future-dependent Value Functions for Off-policy Evaluation

Reference 74

Resolution
verified exact
local_arxiv, observed 2026-08-16T04:54:14.671602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.586780Z digest=sha256:17f1272e05f87405050aa4f26b69261c989a964fa87391ce7c9485393996cc32

Observation a9b94b24-e0ac-4371-8991-def5ab75dc3b · outbound

This paper cites (2024), Bi-Level Offline Policy Optimization with Limited Exploration, Advances in Neural Information Processing Systems, 36.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding (2024), Bi-Level Offline Policy Optimization with Limited Exploration, Advances in Neural Information Processing Systems, 36

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:14.891413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.592125Z digest=sha256:bdd514eb17fba547c3154ddefbf12aa4096e1e0f8eb1986c66717677eada0af8

Observation 8d09bd4e-b1cd-40e3-9026-7d253601b194 · outbound

This paper cites (2024 a ), Policy learning for individualized treatment regimes on infinite time horizon, in Statistics in Precision Health: Theory, Methods and Applications, Springer, pp.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding (2024 a ), Policy learning for individualized treatment regimes on infinite time horizon, in Statistics in Precision Health: Theory, Methods and Applications, Springer, pp

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:14.874788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.596573Z digest=sha256:5ed4628101df372cc113e7b4d10499b801e69a49c891a172d334bb0841488c14

Observation 5c19e899-4c9a-4945-8ca5-103002d52080 · outbound

This paper cites Distributional Shift-Aware Off-Policy Interval Estimation: A Unified Error Quantification Framework.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding Distributional Shift-Aware Off-Policy Interval Estimation: A Unified Error Quantification Framework

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-16T04:54:14.601000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:54:14.601000Z digest=sha256:34a4cee9feea566bcb548ca0653c391382fa520fcc87e623dbbcdc24dce3fd5e

Observation 6fe62085-e851-42d9-827b-f2b8ace0bd17 · outbound

This paper cites (2024 b ), Estimating optimal infinite horizon dynamic treatment regimes via pt-learning, Journal of the American Statistical Association, 119, 625--638.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding (2024 b ), Estimating optimal infinite horizon dynamic treatment regimes via pt-learning, Journal of the American Statistical Association, 119, 625--638

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:14.857615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.605462Z digest=sha256:d948ffce9d66bd7947e2490d090e08739749759ac1295c55521f79191176f2bc

Observation 7798237d-f28a-4154-8fe1-f5271244b8da · outbound

This paper cites (2020), Safe, efficient, and comfortable velocity control based on reinforcement learning for autonomous driving, Transportation Research Part C: Emerging Technologies, 117, 102662.

Reinforcement Learning with Continuous Actions Under Unmeasured Confounding (2020), Safe, efficient, and comfortable velocity control based on reinforcement learning for autonomous driving, Transportation Research Part C: Emerging Technologies, 117, 102662

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:54:14.841403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-16T04:54:14.609941Z digest=sha256:d6d6ccfa65e4585d0db9936f490cbb7a99d33ae46cd39a040c32eb12f52c3c4c

Pith citing papers

No inbound Pith citation observations are available.