Pith. sign in

REVIEW 1 cited by

An Efficient On-Policy Deep Learning Framework for Stochastic Optimal Control

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.05163 v3 pith:ZN6TUICT submitted 2024-10-07 cs.LG math.OC

classification cs.LGmath.OC
keywords controlon-policystochasticmethodoptimalproblemsacceleratesadjoint
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present a novel on-policy algorithm for solving stochastic optimal control (SOC) problems. By leveraging the Girsanov theorem, our method directly computes on-policy gradients of the SOC objective without expensive backpropagation through stochastic differential equations or adjoint problem solutions. This approach significantly accelerates the optimization of neural network control policies while scaling efficiently to high-dimensional problems and long time horizons. We evaluate our method on classical SOC benchmarks as well as applications to sampling from unnormalized distributions via Schr\"odinger-F\"ollmer processes and fine-tuning pre-trained diffusion models. Experimental results demonstrate substantial improvements in both computational speed and memory efficiency compared to existing approaches.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Neural feedback approximation for stochastic control with degenerate diffusions: error estimates and numerical analysis

    math.OC 2026-07 conditional novelty 6.0 of 10

    Direct neural feedback learning for time-discrete stochastic control admits an averaged value-error bound without transition-density assumptions, covering degenerate and deterministic dynamics.

Pith tools