SCOREBED isolates EIG double intractability in a policy-independent score-matching stage, then trains design policies with a singly intractable gradient estimator, enabling cheap multi-policy selection.
The sample size required in importance sampling , volume =
4 Pith papers cite this work, alongside 12 external citations. Polarity classification is still indexing.
representative citing papers
Derives non-asymptotic error bounds for standard, defensive, and self-normalized importance sampling with random KDE proposals from geometrically ergodic Markov chains, separating n^{-1/2} Monte Carlo error from MIAE/MISE proposal error.
Derives Õ(d β² A² / ε⁴) oracle complexity for AIS estimating normalizing constant Z to relative error ε and introduces reverse diffusion sampler for geometric paths with large action.
Proposes causal reinforcement learning (CRL) as a framework that decomposes RL environments into structural causal models to unify online, off-policy, and causal learning while defining new tasks including generalized policy learning and counterfactual learning.
citing papers explorer
-
Bayesian Experimental Design via Score Matching
SCOREBED isolates EIG double intractability in a policy-independent score-matching stage, then trains design policies with a singly intractable gradient estimator, enabling cheap multi-policy selection.
-
Error Bounds for Importance Sampling with Estimated Proposal Distributions
Derives non-asymptotic error bounds for standard, defensive, and self-normalized importance sampling with random KDE proposals from geometrically ergodic Markov chains, separating n^{-1/2} Monte Carlo error from MIAE/MISE proposal error.
-
Complexity Analysis of Normalizing Constant Estimation: from Jarzynski Equality to Annealed Importance Sampling and beyond
Derives Õ(d β² A² / ε⁴) oracle complexity for AIS estimating normalizing constant Z to relative error ε and introduces reverse diffusion sampler for geometric paths with large action.
-
An Introduction to Causal Reinforcement Learning
Proposes causal reinforcement learning (CRL) as a framework that decomposes RL environments into structural causal models to unify online, off-policy, and causal learning while defining new tasks including generalized policy learning and counterfactual learning.