Pith. sign in

Paper Citation Record · LEDGER

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning

As of 13 August 2026, this Paper Citation Record lists 100 of 106 outbound references and 1 inbound Pith citation observation for arXiv:2412.05783.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.05783 v1

Coverage vector

measured 100 of 106 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T20:26:49.464366Z

measured 101 of 101 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T08:36:27.422590Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

100 of 106 outbound references displayed

  • verified exact8
  • verified fuzzy32
  • unresolved60
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 5cf9ca6b-303e-49b7-860a-b97f038ec39c · outbound

This paper cites Formulation and estimation of dynamic models using panel data.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Formulation and estimation of dynamic models using panel data

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T20:26:49.013868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:26:49.013868Z digest=sha256:f198ffa9a1caed50346e2fc15080865d6db7ef539c3fd82f2bf5350d1ef6c1f5

Observation 81eab347-b38c-4953-a397-f6fa7bbcb21d · outbound

This paper cites Doubly robust identification for causal panel data models.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Doubly robust identification for causal panel data models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T20:26:49.019604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:26:49.019604Z digest=sha256:bfd974003750d483f7088a6d6e807079a8c405ee45c5a83baddf4775cfe33a82

Observation d392ba94-52de-418b-9353-b0106c23bf1f · outbound

This paper cites Design-based analysis in difference-in-differences settings with staggered adoption.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Design-based analysis in difference-in-differences settings with staggered adoption

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T20:26:49.024472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:26:49.024472Z digest=sha256:fee9d54fa0b7285389860d354b34ef13d0dc246726c1098fbf15fce814fd2425

Observation a13dec1d-9588-44aa-bf28-610180a1d069 · outbound

This paper cites Econometric analysis of panel data, volume 4.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Econometric analysis of panel data, volume 4

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T20:26:49.029217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:26:49.029217Z digest=sha256:3b89d0b649e1a06869c9dfcef70e43103b39f072041a690723bf8745e889acf0

Observation 856b2df3-9142-4198-a91c-0af226994e1d · outbound

This paper cites Genetic risk profiles for cancer susceptibility and therapy response.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Genetic risk profiles for cancer susceptibility and therapy response

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T20:26:49.033801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:26:49.033801Z digest=sha256:3eada87d50bdb76b3d446239f1bebf86654ba6cd83c9db6452cb54819e397fa1

Observation b4fc580b-0d56-486a-9196-350a637475a3 · outbound

This paper cites Proximal reinforcement learning: Efficient off-policy evaluation in partially observed markov decision processes.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Proximal reinforcement learning: Efficient off-policy evaluation in partially observed markov decision processes

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T20:26:49.038799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:26:49.038799Z digest=sha256:f5ed6913a14b33dc77dc95786994a555815144ad5a8c495e1fa1c669754fab4e

Observation 2d2851fd-f9da-422f-a698-a9b11f3f2593 · outbound

This paper cites Off-policy Evaluation in Doubly Inhomogeneous Environments.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Off-policy Evaluation in Doubly Inhomogeneous Environments

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-08-11T20:26:50.168208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T20:26:49.043818Z digest=sha256:217429a93d9ec57bd7ca2894e187e67aebf5dd29ee0d0d2f676ce0a4c1a805bc

Observation e6aa3b11-ad83-48f3-aa5e-d6c55cdf5ca9 · outbound

This paper cites Time series deconfounder: Estimating treatment effects over time in the presence of hidden confounders.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Time series deconfounder: Estimating treatment effects over time in the presence of hidden confounders

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T20:26:49.048680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:26:49.048680Z digest=sha256:ebd23915efa8423ff2014eded535b449703a5ec1d3cd0146eb6c872ff8f93143

Observation f3703407-ced8-4883-8a64-918f469106e4 · outbound

This paper cites Robust fitted-q-evaluation and iteration under sequentially exogenous unobserved confounders.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Robust fitted-q-evaluation and iteration under sequentially exogenous unobserved confounders

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T20:26:49.052749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:26:49.052749Z digest=sha256:f05ffad8153bf8583c9e8300dc7c61830a7d26c84d2769ea0cc4997f9068dc94

Observation 2f731838-1fa3-45b9-bd8d-fc1dcab8f1d9 · outbound

This paper cites Treatment effects in interactive fixed effects models with a small number of time periods.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Treatment effects in interactive fixed effects models with a small number of time periods

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T20:26:49.057216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:26:49.057216Z digest=sha256:2418003075b35d73fc1b04dfa656195cc0df653fcec1b8e54a21103989071be8

Observation 64d47dda-039d-4779-9afd-7c65d9605d74 · outbound

This paper cites Statistical inference.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Statistical inference

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T20:26:49.061405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:26:49.061405Z digest=sha256:bfb05b5b5465a6e296beae6dbb3e01e51cc39ba98a592fa0a136f78344c6f518

Observation 674bde1d-3223-409c-babe-49b6fe6a5f80 · outbound

This paper cites Universal off-policy evaluation.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Universal off-policy evaluation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T20:26:49.065314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:26:49.065314Z digest=sha256:ef40d00f753649be5ca827cbb9ca5e2456210362437cce041b7540e32c1ba995

Observation ea0ba1f1-cf76-4fa8-ad75-515556cd94b5 · outbound

This paper cites Testing for the markov property in time series.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Testing for the markov property in time series

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T20:26:49.069222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:26:49.069222Z digest=sha256:cb20c29fc1cbe06a139766ad6028c9b156e0bfad8495767bab9d27ca06173ed3

Observation 61ddd041-b55b-4d1e-932d-44105215db94 · outbound

This paper cites Information-Theoretic Considerations in Batch Reinforcement Learning.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Information-Theoretic Considerations in Batch Reinforcement Learning

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-08-11T20:26:50.069475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T20:26:49.073582Z digest=sha256:7ec91c882cd4aa824c65284872c1fa776cdca00f97a42c9db59c277280cb928d

Observation 61dbe7c9-4d78-4c87-b6a3-b8ea0c5428b2 · outbound

This paper cites On well-posedness and minimax optimal rates of nonparametric q-function estimation in off-policy evaluation.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning On well-posedness and minimax optimal rates of nonparametric q-function estimation in off-policy evaluation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T20:26:49.078519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:26:49.078519Z digest=sha256:9ecf57e0a3943cda5c13b572a2c0c3cfc872f77600506938dd9adce1af959b87

Observation 79a87a3e-8b07-4c7e-9a83-f545bcf34291 · outbound

This paper cites On instrumental variable regression for deep offline policy evaluation.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning On instrumental variable regression for deep offline policy evaluation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T20:26:49.083127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:26:49.083127Z digest=sha256:0abc87019097031591470f4dbea00c13494d26916cd7bb63bb68a642de2c280f

Observation cca7315a-4c7e-47dc-9384-2d46fd4f03d5 · outbound

This paper cites Double/debiased machine learning for treatment and structural parameters.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Double/debiased machine learning for treatment and structural parameters

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T20:26:49.087629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:26:49.087629Z digest=sha256:3fb5ecdd003c95b1a68a30538e92b864b3652d97d46c47b56bbba76ae8381188

Observation 54065e70-f661-480c-b050-cb993b19469e · outbound

This paper cites A crash course in good and bad controls.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning A crash course in good and bad controls

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T20:26:49.092271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:26:49.092271Z digest=sha256:470d447d11e63b804a0070624f0f8479f2b0bf8036ae32931d212484b28f1b1f

Observation b3647527-68eb-42b6-ad3c-2c247dc2fb07 · outbound

This paper cites Coindice: Off-policy confidence interval estimation.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Coindice: Off-policy confidence interval estimation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T20:26:49.096926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:26:49.096926Z digest=sha256:a73518a160c69e88d45da7d240a59cc228fae2b75fd2657f5409c6c3f00a4dee

Observation ad05cb02-d180-4a4a-9508-f57b7f20b71e · outbound

This paper cites Comment: Reflections on the deconfounder, 2019.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Comment: Reflections on the deconfounder, 2019

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T20:26:49.101485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:26:49.101485Z digest=sha256:d0dbeef6b950eae1fafdd470a38d8f293859b55e5cc47c20b54acc44375954d3

Observation 68dde355-1d55-49a3-b452-15c5bef507d5 · outbound

This paper cites Two-way fixed effects estimators with heterogeneous treatment effects.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Two-way fixed effects estimators with heterogeneous treatment effects

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T20:26:49.105859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:26:49.105859Z digest=sha256:d526f9d2617d8d2dd969207a9288f1dc08ce742cddc1a1d689b65ae3a869ec9e

Observation 9ea2261c-b5ee-4ba1-aa62-e36811700efe · outbound

This paper cites Counterfactual inference in sequential experiments.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Counterfactual inference in sequential experiments

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-08-11T20:26:50.049057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T20:26:49.110445Z digest=sha256:d9ececa82d2a1d56cc0503bb5119364630eb3e7eeb6962e0b5e5e6f81a7272c1

Observation 722ea65d-52ee-42d8-9ce1-5a3902d93c9f · outbound

This paper cites A theoretical analysis of deep q-learning.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning A theoretical analysis of deep q-learning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T20:26:49.115428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:26:49.115428Z digest=sha256:2d18c3c2c9ed9677d86803420276da20a78cd63b86a9f58fba515766f582d1b9

Observation 6dc54f2f-3ac3-4e0a-ba59-30af9c519f98 · outbound

This paper cites More robust doubly robust off-policy evaluation.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning More robust doubly robust off-policy evaluation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T20:26:49.119846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:26:49.119846Z digest=sha256:d35043962a0cde0daabc70b2be3f075fed2fef474b9bcaf6728bd37ff7902732

Observation 1aa730a7-d898-46c9-9e90-bd224f8486c2 · outbound

This paper cites Deep neural networks for estimation and inference.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Deep neural networks for estimation and inference

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T20:26:49.124324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:26:49.124324Z digest=sha256:c626d3994d3b941ffbbee53bdc096683dbcf2d2df3321e9e1aea3c27de9b8b36

Observation d241ef39-edf2-4287-b0e8-deb304ea1fc6 · outbound

This paper cites Non-parametric panel data models with interactive fixed effects.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Non-parametric panel data models with interactive fixed effects

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T20:26:49.128856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:26:49.128856Z digest=sha256:a52665cd430e8c075de991fd94225d1975d427b8a7e4651d1c12ee254cb6d87a

Observation 4e1aa98c-71ab-44ca-8711-72bbcd9c208f · outbound

This paper cites Offline Reinforcement Learning with Instrumental Variables in Confounded Markov Decision Processes.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Offline Reinforcement Learning with Instrumental Variables in Confounded Markov Decision Processes

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-08-11T20:26:50.028564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T20:26:49.133163Z digest=sha256:0508340a4fa848cf9e7296ba63ea25e2c5924816ce710ba4baf028c255029443

Observation 71f7fafc-e76d-4c41-856e-74588a2a464e · outbound

This paper cites Prediction of treatment response for combined chemo-and radiation therapy for non-small cell lung cancer patients using a bio-mathematical model.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Prediction of treatment response for combined chemo-and radiation therapy for non-small cell lung cancer patients using a bio-mathematical model

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T20:26:49.138132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:26:49.138132Z digest=sha256:aeac0384615da2b0cd5d1e23f6da46ff9403d96d89fe908e1b8bf788c66ee06a

Observation af85bcdc-9ff7-4a0b-bd0c-dc89ad0653d8 · outbound

This paper cites Issues in assessing the contribution of research and development to productivity growth.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Issues in assessing the contribution of research and development to productivity growth

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T20:26:49.142440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:26:49.142440Z digest=sha256:66c695ef7c88ab6599b6126ae985ff325bc0385f798a8cb45f552135fc26bef5

Observation 4ca90b1d-0ee0-404d-a32e-8736af8529a9 · outbound

This paper cites Richard Guo, Anton Rask Lundborg, and Qingyuan Zhao.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Richard Guo, Anton Rask Lundborg, and Qingyuan Zhao

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T20:26:49.147024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:26:49.147024Z digest=sha256:b4776f387eda896474673ba2d62c2e5cc7db9383f496fefd6fc237de046574b2

Observation 52c89603-207b-4f0a-b66d-ed48e4f6765f · outbound

This paper cites Dream to Control: Learning Behaviors by Latent Imagination.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Dream to Control: Learning Behaviors by Latent Imagination

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T20:26:49.151455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:26:49.151455Z digest=sha256:560cf2733043f385e125303ea37f0033f0ce2b935d081b77884dda61a03d0cc1

Observation 3bfd73e2-7b4c-4d1b-a9ba-66f424d966f0 · outbound

This paper cites Learning latent dynamics for planning from pixels.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Learning latent dynamics for planning from pixels

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T20:26:49.156301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:26:49.156301Z digest=sha256:8ada31d19d4853fca36e03e7aecf0160693aa34bb68e4d4d38c09c704fa0f3f2

Observation fa09841b-afa9-4af0-8f09-c107a472fb7d · outbound

This paper cites Bootstrapping fitted q-evaluation for off-policy inference.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Bootstrapping fitted q-evaluation for off-policy inference

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T20:26:49.160798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:26:49.160798Z digest=sha256:ae88d8de75b2994dca535071486363f7ecc326ca657c06ed30af0fc6965e8055

Observation 52f1e374-03d7-4990-9a8d-02939d46e023 · outbound

This paper cites Sequential Deconfounding for Causal Inference with Unobserved Confounders.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Sequential Deconfounding for Causal Inference with Unobserved Confounders

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-08-11T20:26:49.994127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T20:26:49.165490Z digest=sha256:7379e8cf31426891fbd3b37538770d021c19e2799a7afd20056af8dc9a4e00bc

Observation 7fe94afa-707d-4d74-8f53-c5f8b8746612 · outbound

This paper cites Neural collaborative filtering.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Neural collaborative filtering

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T20:26:49.170520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:26:49.170520Z digest=sha256:4d22dc93afda3d6ce0500c608a7fa13ffa77bb623111c6ac664c8e6492df81aa

Observation 1b616d7c-1356-4f6b-8bf4-a49d8e0f5007 · outbound

This paper cites A Policy Gradient Method for Confounded POMDPs.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning A Policy Gradient Method for Confounded POMDPs

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T20:26:49.174755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:26:49.174755Z digest=sha256:128251f090c43f43c923883d3a8c039ad862915aec586d39d31235ea5f992f2e

Observation e4cadad5-e646-4614-a495-325ffc3a13c8 · outbound

This paper cites Collaborative filtering for implicit feedback datasets.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Collaborative filtering for implicit feedback datasets

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T20:26:49.178967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:26:49.178967Z digest=sha256:83fa69b887b38b4b43e0a63bc8e93889ba3f585cbe869c664f94276654ed8f63

Observation f18883af-ed44-4dd2-9508-340088c4bad5 · outbound

This paper cites On the use of two-way fixed effects regression models for causal inference with panel data.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning On the use of two-way fixed effects regression models for causal inference with panel data

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T20:26:49.182754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:26:49.182754Z digest=sha256:2c18721278f16f0527b03cb7df7c4b256627c8d037ad94f20e55dfeb9764fb17

Observation 32386744-8488-4fe7-9487-63338c368b07 · outbound

This paper cites Off-policy evaluation via off-policy classification.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Off-policy evaluation via off-policy classification

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:26:50.799584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T20:26:49.186790Z digest=sha256:0ebc8245ea674020099f1e51999df741cb887cf5f4480c76828a49f2e5600067

Observation a127d31e-b7f1-4795-8641-129d4520de1b · outbound

This paper cites When to trust your model: Model-based policy optimization.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning When to trust your model: Model-based policy optimization

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T20:26:49.190846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:26:49.190846Z digest=sha256:43473a050b516856598230f158837a3064dde245136870c7021b29a5fb8cff0a

Observation 71aaaba8-9046-4a8e-94b2-8f8cf63173aa · outbound

This paper cites A survey on knowledge graphs: Representation, acquisition, and applications.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning A survey on knowledge graphs: Representation, acquisition, and applications

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T20:26:49.194737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:26:49.194737Z digest=sha256:19abf269a941176e3b76d00cd132677521afd97246a35d1463468385a44ad76f

Observation 0b532d18-1e37-4f82-a1f3-29eeea0c962c · outbound

This paper cites A Note on Loss Functions and Error Compounding in Model-based Reinforcement Learning.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning A Note on Loss Functions and Error Compounding in Model-based Reinforcement Learning

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T20:26:49.199064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:26:49.199064Z digest=sha256:d136aadc820fbb85bddf0ee85b783b89750cf8e4a345fc06aa59f7299c00668c

Observation a49847dc-e071-44ab-8e47-3da895f4dd04 · outbound

This paper cites Doubly robust off-policy value evaluation for reinforcement learning.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Doubly robust off-policy value evaluation for reinforcement learning

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:26:50.766712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T20:26:49.203564Z digest=sha256:72affb1009697f725444a226911be91d3551813ca253e31360f92319f3b75a87

Observation af55c0cd-0b26-43a5-a67a-2265db6863e3 · outbound

This paper cites Mimic-iii, a freely accessible critical care database.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Mimic-iii, a freely accessible critical care database

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T20:26:49.208020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:26:49.208020Z digest=sha256:c3362cdda836db9d48fcf996ab6b988a7eb62c90ec4793f5fa5da0e2989b7049

Observation 64ee11ac-76dd-476a-87f0-6ae58685e7f9 · outbound

This paper cites Efficiently breaking the curse of horizon in off-policy evaluation with double reinforcement learning.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Efficiently breaking the curse of horizon in off-policy evaluation with double reinforcement learning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T20:26:49.212413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:26:49.212413Z digest=sha256:96ee17a2ab040b6b8bcca144cc128267a9fc45157e8ff555ecedafe179bb7952

Observation 0e964bf1-c5cb-41ef-a054-15e682554714 · outbound

This paper cites Confounding-robust policy evaluation in infinite-horizon reinforcement learning.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Confounding-robust policy evaluation in infinite-horizon reinforcement learning

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:26:50.734020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T20:26:49.216658Z digest=sha256:ff9ca06740da3e8d4ac34e61e7167fb7f2556c1ba1e475b93c2781d7a20c50eb

Observation 94a0e4f1-6eb6-483d-a136-302449c5190e · outbound

This paper cites Offline Policy Evaluation and Optimization under Confounding.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Offline Policy Evaluation and Optimization under Confounding

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-08-11T20:26:49.943952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T20:26:49.221467Z digest=sha256:48acdbb0412816d4081e7b40372578bb8e75591d6859fdb2935b1811eafc35ee

Observation 3b6da3e7-58ef-4f93-be19-d93779cdf21d · outbound

This paper cites Learning mixtures of markov chains and mdps.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Learning mixtures of markov chains and mdps

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:26:50.719425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T20:26:49.227412Z digest=sha256:aeac7a5c2832aa8f1b37e2cdcedc4aac9d37c1d2c072d493a70bceee404cf089

Observation 0e270572-5815-4174-bd97-59858a67f107 · outbound

This paper cites Asymptotic Theory for IV-Based Reinforcement Learning with Potential Endogeneity.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Asymptotic Theory for IV-Based Reinforcement Learning with Potential Endogeneity

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T20:26:49.232155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:26:49.232155Z digest=sha256:59e081e864e753e22b931ddf1aa254afaf3d790222e730be97ef33ee94aeab54

Observation ff919683-2e4f-47ab-94b7-05a99ee2375c · outbound

This paper cites Instrumental Variable Value Iteration for Causal Offline Reinforcement Learning.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Instrumental Variable Value Iteration for Causal Offline Reinforcement Learning

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T20:26:49.237033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:26:49.237033Z digest=sha256:1aba441fd5a62d57b9115f61214f4162708b6f882c98bc45b92e7e5afc621cdb

Observation d548c79c-5569-492f-bacf-4f08e31279a2 · outbound

This paper cites Off-policy estimation of long-term average outcomes with applications to mobile health.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Off-policy estimation of long-term average outcomes with applications to mobile health

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:26:50.705579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T20:26:49.241606Z digest=sha256:9c07cb588ed655508412dd12451f7836577ac63a566e9cd48b5a18eef0c8556b

Observation b53a4e2a-c74e-4d51-9b97-02e1eeb4febb · outbound

This paper cites Batch policy learning in average reward markov decision processes.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Batch policy learning in average reward markov decision processes

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:26:50.690923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T20:26:49.245977Z digest=sha256:064c6f6246f49acdbdc7d2f69ca511b63107d0c6bdb76db5045c9cac5f88efff

Observation 4e5d4ad9-54f4-44a1-ac39-e5d5a85ef690 · outbound

This paper cites Forecasting treatment responses over time using recurrent marginal structural networks.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Forecasting treatment responses over time using recurrent marginal structural networks

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T20:26:49.250321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:26:49.250321Z digest=sha256:d91f76b7f2d303dadf4839969fbc200051a0713d92a73a3ab59f4aca8227eef4

Observation 65385c22-c2a2-4eff-b63e-d9162e098533 · outbound

This paper cites Breaking the curse of horizon: Infinite-horizon off-policy estimation.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Breaking the curse of horizon: Infinite-horizon off-policy estimation

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T20:26:49.254963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:26:49.254963Z digest=sha256:15d1ba9fb8e18fc2f4d6b086d4f0997afcd66986c9103cc90e482f57588979d2

Observation 0cc27e4d-b8ac-4b9f-bd0f-8969dd803f30 · outbound

This paper cites Provably good batch off-policy reinforcement learning without great exploration.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Provably good batch off-policy reinforcement learning without great exploration

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:26:50.658294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T20:26:49.259473Z digest=sha256:de41340f07589e55c2610da6779db669c51b09e50b2a3fa9aa83b490ee1681f3

Observation 6020fd7c-11f3-408c-8ef3-d621c217b820 · outbound

This paper cites Causal effect inference with deep latent-variable models.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Causal effect inference with deep latent-variable models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T20:26:49.264028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:26:49.264028Z digest=sha256:2a212d337dca14c146fac6c9d2a4c61c1264ee6710860c21ade07415480563be

Observation 011c4092-9749-427a-9ef8-26338579198c · outbound

This paper cites Deconfounding Reinforcement Learning in Observational Settings.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Deconfounding Reinforcement Learning in Observational Settings

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-11T20:26:49.268555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:26:49.268555Z digest=sha256:54c4b802d391e6b617e4304f8592d390b1c35d3c4448c34fd171ea5f716614f5

Observation fd481239-bf69-4c47-9a24-db357de512c5 · outbound

This paper cites Pessimism in the face of confounders: Provably efficient offline reinforcement learning in partially observable markov decision processes.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Pessimism in the face of confounders: Provably efficient offline reinforcement learning in partially observable markov decision processes

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:26:50.636096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T20:26:49.273622Z digest=sha256:62415edbc38e985c481eb30ed8b257a82a3e405ed8aebb5c3211998230ceb145

Observation 4a51b86d-dc62-426d-afdc-9fa3765d7c6b · outbound

This paper cites Estimating causal peer influence in homophilous social networks by inferring latent locations.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Estimating causal peer influence in homophilous social networks by inferring latent locations

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:26:50.623463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T20:26:49.278278Z digest=sha256:ca50c295994ed1aed4b88a5d805c9a7ec4e5051f42cb02a91086e6491759f719

Observation 32153b46-476a-4467-aeae-7b83e851470e · outbound

This paper cites Off-policy evaluation for episodic partially observable markov decision processes under non-parametric models.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Off-policy evaluation for episodic partially observable markov decision processes under non-parametric models

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:26:50.609868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T20:26:49.282744Z digest=sha256:8ce4232d17d28b56dcfac2ef83fb210977a1a1b9a440466bf0203a4f17833ac7

Observation fc41d0e5-9d07-4f21-8904-5012e55c1384 · outbound

This paper cites Empirical production function free of management bias.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Empirical production function free of management bias

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:26:50.595983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T20:26:49.286681Z digest=sha256:e0b52cf22827f1a56e5e7c6200085aceee6fb2a636d8f44e3871fa9cd595416f

Observation 56e8f4d9-725a-4e58-8577-1a4e69e10266 · outbound

This paper cites A Spectral Approach to Off-Policy Evaluation for POMDPs.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning A Spectral Approach to Off-Policy Evaluation for POMDPs

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-11T20:26:49.290906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:26:49.290906Z digest=sha256:2e277e95c6f031fa85a929663c80906f8a7293717974759477be31ada48dff66

Observation 37f99efe-7629-4a6a-9394-86ee51e38913 · outbound

This paper cites Off-policy policy evaluation for sequential decisions under unobserved confounding.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Off-policy policy evaluation for sequential decisions under unobserved confounding

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:26:50.582136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T20:26:49.295005Z digest=sha256:06b62da2c22a8253c1b478820fc771da2c9309fb99c0b0ce0bc7534728e458d8

Observation 54907508-141e-4877-a4e4-f4268c6c760a · outbound

This paper cites A Novel Embedding Model for Knowledge Base Completion Based on Convolutional Neural Network.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning A Novel Embedding Model for Knowledge Base Completion Based on Convolutional Neural Network

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-11T20:26:49.299285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:26:49.299285Z digest=sha256:2070f46a649f8748efd3d3488947ef9431be263e97009e18d3de5ff3a927c429

Observation d57eeb98-a197-4cfd-92f9-518be18c6ab6 · outbound

This paper cites A review of relational machine learning for knowledge graphs.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning A review of relational machine learning for knowledge graphs

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:26:50.567741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T20:26:49.304112Z digest=sha256:b94568eadd50a37842b9101ec3922a152318d2f4a5ff6184bf0302405b18a65d

Observation eca033d3-c3ed-4dac-b4a0-4093929bb8f9 · outbound

This paper cites Automated cars meet human drivers: responsible human-robot coordination and the ethics of mixed traffic.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Automated cars meet human drivers: responsible human-robot coordination and the ethics of mixed traffic

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:26:50.553282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T20:26:49.308697Z digest=sha256:e1cc609341c1d79d4ab9845021573a31e7b1aec1abf585463435c4bc25524e1c

Observation 0502d09e-fd3a-4c68-99bd-16bd2ce1a339 · outbound

This paper cites blessings of multiple causes.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning blessings of multiple causes

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-11T20:26:49.313162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:26:49.313162Z digest=sha256:7d11320f94fcaa4187feaea549675c5f561808b6c325e439fb17a385f4570f99

Observation cd511ddc-54a8-4ae9-8bf4-841a0c66b884 · outbound

This paper cites Counterexamples to "The Blessings of Multiple Causes" by Wang and Blei.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Counterexamples to "The Blessings of Multiple Causes" by Wang and Blei

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-11T20:26:49.317636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:26:49.317636Z digest=sha256:a0062f8b76ebf4fad2728291850f003cc70b4c856d11f27e1a6f083df8e576e5

Observation f03a237d-00ca-423b-81c7-15fef16315ff · outbound

This paper cites A critical look at the consistency of causal estimation with deep latent variable models.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning A critical look at the consistency of causal estimation with deep latent variable models

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:26:50.539284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T20:26:49.322630Z digest=sha256:a4ecd49f727161488164f47058e54a877c63a1f60dc6884cd9824dd3f29a61f0

Observation 0df06961-fb28-404c-9c24-a728d438e613 · outbound

This paper cites Doubly robust difference-in-differences estimators.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Doubly robust difference-in-differences estimators

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:26:50.525737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T20:26:49.327210Z digest=sha256:4e6cc9ecf0a018ed29fef4148269cb2601d7df6b4db7cc380678d97ef0abe51c

Observation a4e00b95-682b-4952-895e-f771f7975623 · outbound

This paper cites Importance resampling for off-policy prediction.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Importance resampling for off-policy prediction

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:26:50.511247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T20:26:49.331634Z digest=sha256:e1223e4205e47d0ff8536ef152057738adcee7cfb0c2179ddb5e6d04b0c54e80

Observation 76918f35-ec26-4eb2-a005-7527cb92a1ce · outbound

This paper cites Nonparametric regression using deep neural networks with relu activation function.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Nonparametric regression using deep neural networks with relu activation function

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-11T20:26:49.336054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:26:49.336054Z digest=sha256:dc2a691a01529f11e984b86449144dab8fb7bbaf37903de18afef62e3c5f1369

Observation c94b25b5-0eb0-49ce-b030-b6d9692cd316 · outbound

This paper cites On counterfactual inference with unobserved confounding.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning On counterfactual inference with unobserved confounding

Reference 73

Resolution
verified exact
local_arxiv, observed 2026-08-11T20:26:49.773072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T20:26:49.340724Z digest=sha256:d6ec165b82a602283457e6403f2b102fb3e59d280be0059255a8ced8a93388bd

Observation 067ac510-944f-4d18-b6cb-84dc2de34cb7 · outbound

This paper cites Does the markov decision process fit the data: Testing for the markov property in sequential decision making.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Does the markov decision process fit the data: Testing for the markov property in sequential decision making

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:26:50.495984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T20:26:49.345607Z digest=sha256:5e55432a1f0bfd1e81bb216793716330deedc3c91b0c57260a714d93331e9571

Observation dba602c9-a56e-4e02-af6c-fd7fbd4fcb16 · outbound

This paper cites Deeply-debiased off-policy interval estimation.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Deeply-debiased off-policy interval estimation

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:26:50.481440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T20:26:49.350052Z digest=sha256:1b070856ad7c894ee6bdf2bb1b332068d0d1784c5838075920f0967184ac649f

Observation 01a0abe7-28ec-40c9-b6fa-76afeb142d23 · outbound

This paper cites A minimax learning approach to off-policy evaluation in confounded partially observable markov decision processes.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning A minimax learning approach to off-policy evaluation in confounded partially observable markov decision processes

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:26:50.467036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T20:26:49.354184Z digest=sha256:9c3ee50cfe481c92e0ee65bc0bbb37beffe07e90bf0b1ba5b0636102fac841f1

Observation 9397cc0a-09a2-462f-8e3e-0ea8651c291e · outbound

This paper cites Statistical inference of the value function for reinforcement learning in infinite-horizon settings.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Statistical inference of the value function for reinforcement learning in infinite-horizon settings

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:26:50.451972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T20:26:49.358732Z digest=sha256:a63270d3690291d825aa50526d83fab4719bbbee78add4eec3f0518c81247d68

Observation f75aea1e-95db-497e-83dd-6c2ecff17798 · outbound

This paper cites Off-policy confidence interval estimation with confounded markov decision process.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Off-policy confidence interval estimation with confounded markov decision process

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:26:50.437357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T20:26:49.363172Z digest=sha256:890b3b1687c5d5bc44a9056730a046607275bef50cd69d582d0cb1040d364c48

Observation 611f730e-97af-4690-94ad-937ecc2f2c3f · outbound

This paper cites Mediation pathway selection with unmeasured mediator-outcome confounding.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Mediation pathway selection with unmeasured mediator-outcome confounding

Reference 79

Resolution
verified exact
raw_fallback, observed 2026-08-11T20:26:49.752062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T20:26:49.367414Z digest=sha256:fa81b03f74b32efbcfb48ebd351df0fc19a0d1e1732fc5f598a0f094066e328b

Observation 5a207d18-b0e6-48f3-9c1d-8d6dfc330435 · outbound

This paper cites Reasoning with neural tensor networks for knowledge base completion.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Reasoning with neural tensor networks for knowledge base completion

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-11T20:26:49.371847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:26:49.371847Z digest=sha256:b0868ac7714438514b50a373a6ba5f2dcf5ed3b6689f43a0ab59d905c91d43e9

Observation 8bf13f72-5146-4824-b96b-f6c32cf73b4b · outbound

This paper cites Reinforcement learning: An introduction.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Reinforcement learning: An introduction

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-11T20:26:49.376356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:26:49.376356Z digest=sha256:6906a928aa537d0d2d3ebd9f4f167c64abf2f40dad85f412d28a3da5eba5e1f7

Observation e254b7c3-ddd9-4d3c-a61e-477b512f9f93 · outbound

This paper cites Doubly Robust Bias Reduction in Infinite Horizon Off-Policy Estimation.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Doubly Robust Bias Reduction in Infinite Horizon Off-Policy Estimation

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-11T20:26:49.381181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:26:49.381181Z digest=sha256:98a21ae365b002961a9aa980bb0543b7b724c59ded736c3343ca16149b514e2e

Observation ff5acfdd-eea0-4bea-915a-509ac1b12643 · outbound

This paper cites An Introduction to Proximal Causal Learning.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning An Introduction to Proximal Causal Learning

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-11T20:26:49.385666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:26:49.385666Z digest=sha256:2490eb6c31102eb6e75125164cbc14fffaf55025e1f2cb87bcd30f69fc810238

Observation f53839b7-c0a4-4ed0-a2b4-dde6c539d20f · outbound

This paper cites Off-policy evaluation in partially observable environments.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Off-policy evaluation in partially observable environments

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:26:50.405119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T20:26:49.390375Z digest=sha256:3511af8b3e0176638d3382b48de31ba4b70b75fe34bc1f874dab61c84ec2be4e

Observation 2b35f4a5-8d54-4921-9d90-dacbeb46223a · outbound

This paper cites Data-efficient off-policy policy evaluation for reinforcement learning.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Data-efficient off-policy policy evaluation for reinforcement learning

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:26:50.391180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T20:26:49.394915Z digest=sha256:951c6cb868d22356bd72b07369ae75d4af6e7019b02690d73732f42073a98a7b

Observation 3538c71c-a3a8-4ff6-975c-c9a846b1b822 · outbound

This paper cites High-confidence off-policy evaluation.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning High-confidence off-policy evaluation

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:26:50.376590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T20:26:49.399443Z digest=sha256:1edf8a55cd17b0b21bfbd5755f26b0713ee7ceb499c339488491d5facf858871

Observation aa3f4e79-c057-494b-94b1-6dd011a90948 · outbound

This paper cites Implicit Causal Models for Genome-wide Association Studies.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Implicit Causal Models for Genome-wide Association Studies

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-11T20:26:49.403988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:26:49.403988Z digest=sha256:fb3e0fc0d92345d728793e1cb7f7b72ddf119818ce0192c3f55cdcc3069f28f5

Observation 8e91b53c-552f-45e5-ad6e-80d394d79687 · outbound

This paper cites Pessimistic Model-based Offline Reinforcement Learning under Partial Coverage.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Pessimistic Model-based Offline Reinforcement Learning under Partial Coverage

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-11T20:26:49.408990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:26:49.408990Z digest=sha256:b485155c26d3d5f56d02a008f68ae84076b4cf45ca5acab40c2be3d3e0f51d82

Observation b6b9adc8-d1cd-4b59-b49e-c42fafb9b975 · outbound

This paper cites Minimax weight and q-function learning for off-policy evaluation.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Minimax weight and q-function learning for off-policy evaluation

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-11T20:26:49.413790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:26:49.413790Z digest=sha256:4b9ca0a842dee8a50a386013d4261804a74d45833a475d731b06f752bf07936f

Observation f501c8cd-c13f-4e91-ab43-6980c84cd4be · outbound

This paper cites Using embeddings to correct for unobserved confounding in networks.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Using embeddings to correct for unobserved confounding in networks

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:26:50.352098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T20:26:49.418982Z digest=sha256:7f39e26f329aa5497d50239a489740221c28b347da2f41b397597d24d13c4d30

Observation da6374d8-9829-4a5f-b237-bcc3689aef74 · outbound

This paper cites Adapting text embeddings for causal inference.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Adapting text embeddings for causal inference

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-11T20:26:49.423736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:26:49.423736Z digest=sha256:0cd520c1d524716238ba74643a095ce1aa8a312431629f7f90917db38feb9db8

Observation 3b16c69a-bd9d-4fc2-a840-e839d600de9d · outbound

This paper cites Relational deep learning: A deep latent variable model for link prediction.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Relational deep learning: A deep latent variable model for link prediction

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:26:50.329080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T20:26:49.428544Z digest=sha256:fae249d5f23030fbc3010fd110c56b768cca2330689cd426cb04dc87b1182ab7

Observation 1f3480fa-a383-465c-8444-211ed4b6ea20 · outbound

This paper cites Blessing from Human-AI Interaction: Super Reinforcement Learning in Confounded Environments.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Blessing from Human-AI Interaction: Super Reinforcement Learning in Confounded Environments

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-11T20:26:49.433207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:26:49.433207Z digest=sha256:3f889b5aa37cc08921a6f049ddb6c505394eabf4b3f69b105c55b6fee15ca73e

Observation 4b0b5083-cc16-4f7f-bcf0-a517aa9b41f3 · outbound

This paper cites Provably efficient causal reinforcement learning with confounded observational data.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Provably efficient causal reinforcement learning with confounded observational data

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:26:50.315863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T20:26:49.438145Z digest=sha256:d2b643a463468b1bccf33a35f2b01985e1676002b0ad8c2be41af6f85f17eace

Observation bdf0cca4-42c4-44f5-a381-04d74a0cc6f8 · outbound

This paper cites The blessings of multiple causes.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning The blessings of multiple causes

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:26:50.302272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T20:26:49.442778Z digest=sha256:6519c9b9c0ecbd32ebf207464586b7a9f57cd1b4e762b27d586d254bb4c4e8b9

Observation 62992a5b-55ea-4fb3-a59e-7beabeab7809 · outbound

This paper cites The Deconfounded Recommender: A Causal Inference Approach to Recommendation.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning The Deconfounded Recommender: A Causal Inference Approach to Recommendation

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-11T20:26:49.447182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:26:49.447182Z digest=sha256:34bead60bc98e460d04061e121747713aa79c61ad067e5559a1f793496189b17

Observation 358dc732-754a-4111-847a-086c856d366f · outbound

This paper cites Semiparametrically efficient off-policy evaluation in linear markov decision processes.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Semiparametrically efficient off-policy evaluation in linear markov decision processes

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:26:50.288702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T20:26:49.451301Z digest=sha256:0d39fc03e22502e17b035240dc5ed9623b61e0bd3386e98a5b927bdc08896508

Observation 3c3c5a1d-a062-4258-96ee-c74991263fb6 · outbound

This paper cites Towards optimal off-policy evaluation for reinforcement learning with marginalized importance sampling.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Towards optimal off-policy evaluation for reinforcement learning with marginalized importance sampling

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:26:50.274909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T20:26:49.455826Z digest=sha256:c5f1e6f1f9768e337ad46cf404c2ace4310ac157716bb35557db62935baf29fd

Observation eb66cdb4-473c-43a7-a7f0-558b3da967f7 · outbound

This paper cites An instrumental variable approach to confounded off-policy evaluation.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning An instrumental variable approach to confounded off-policy evaluation

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T20:26:50.260239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T20:26:49.460191Z digest=sha256:46bae08594b3379d428f57a02fe5ba74a009f1a8675cdb098fa69435cd8e8a8b

Observation 7f363e52-78a9-4d07-b6c1-8d1bd082a64d · outbound

This paper cites Strategic Decision-Making in the Presence of Information Asymmetry: Provably Efficient RL with Algorithmic Instruments.

Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning Strategic Decision-Making in the Presence of Information Asymmetry: Provably Efficient RL with Algorithmic Instruments

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-11T20:26:49.464366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:26:49.464366Z digest=sha256:e348d517f3a45483ea4c13f3de5de0336260e6f7756e13e4db69f7f8a943d432

Pith citing papers

Observation 76ef2f32-1274-4e25-b5f0-bf677d993731 · inbound

Training Large Language Models for Self-Explanation Faithfulness cites this paper.

Training Large Language Models for Self-Explanation Faithfulness Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning

Reference 104

Resolution
verified exact
local_arxiv, observed 2026-08-01T08:38:36.968420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-01T08:36:27.422590Z digest=sha256:35b5f44150d3535bb814955911b35772b57f9cff06793b9455288657611bf2be