Pith. sign in

REVIEW 1 cited by

Bayesian Sequential Optimal Experimental Design for Nonlinear Models Using Policy Gradient Reinforcement Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2110.15335 v2 pith:FCBU67S6 submitted 2021-10-28 cs.LG stat.COstat.MEstat.ML

classification cs.LGstat.COstat.MEstat.ML
keywords policydesignsoeddesignsgradientoptimalsequentialbatch
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present a mathematical framework and computational methods to optimally design a finite number of sequential experiments. We formulate this sequential optimal experimental design (sOED) problem as a finite-horizon partially observable Markov decision process (POMDP) in a Bayesian setting and with information-theoretic utilities. It is built to accommodate continuous random variables, general non-Gaussian posteriors, and expensive nonlinear forward models. sOED then seeks an optimal design policy that incorporates elements of both feedback and lookahead, generalizing the suboptimal batch and greedy designs. We solve for the sOED policy numerically via policy gradient (PG) methods from reinforcement learning, and derive and prove the PG expression for sOED. Adopting an actor-critic approach, we parameterize the policy and value functions using deep neural networks and improve them using gradient estimates produced from simulated episodes of designs and observations. The overall PG-sOED method is validated on a linear-Gaussian benchmark, and its advantages over batch and greedy designs are demonstrated through a contaminant source inversion problem in a convection-diffusion field.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Active Learning of Model Discrepancy with Bayesian Experimental Design

    cs.LG 2025-02 conditional novelty 6.0 of 10

    A hybrid framework alternates Bayesian experimental design for physics parameters with gradient-based calibration of a neural network model-discrepancy term, gated by an ensemble Kalman information-gain indicator.

Pith tools