REVIEW 3 major objections 3 minor 161 references
Towards Adaptive External Communication in Autonomous Vehicles: A Conceptual Design Framework
T0 review · 3 major / 3 minor · reviewed 2026-08-05 · deepseek-v4-flash
Pith's one-line read Adaptive car-to-pedestrian displays need a three-layer design framework so the vehicle's message can change with who is around and what is happening.
desk verdict A plausible three-layer framework for adaptive eHMIs, but the only part we can actually check is the abstract; the provided full text is a different paper, so the central utility claim is unverified. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central organizing object is the three-layer Input–Processing–Output decomposition, using the cyber-physical system as a structuring lens. It separates perception (what the vehicle detects about road actors and context), decision (how the system chooses communication), and expression (how the interface outputs the message), so that each layer can vary independently and be analysed for its contribution to communication success.
What would settle it
Take two groups of designers: one uses the framework, one does not; ask both to design an eHMI for the same mixed-traffic scenario, then measure pedestrian comprehension and crossing decisions in a simulator across child, adult, older adult, and cyclist actors in day and night contexts. If framework-designed interfaces are not reliably better, the claim that the framework helps design, analyse, and assess adaptive eHMIs fails.
Extended reading notes
Core claim
The central discovery is a conceptual ordering rather than a measured result: adaptive eHMI design can be decomposed into what the system detects (Input), how it decides to communicate (Processing), and how it presents the communication (Output). The paper argues that most current eHMIs are reactive—fixed messages that do not adjust—and that treating the interface as part of a cyber-physical loop makes variability across road actors and contexts a design parameter rather than an afterthought. Developed through theory-led abstraction and expert discussion, the framework is offered as a structured tool for researchers and designers to design, analyse, and assess adaptive communication strategi
Load-bearing premise
The load-bearing premise is that expert-derived categories of road actors and contexts are complete enough that a designer can map any real situation onto Input–Processing–Output choices, and that this mapping improves real eHMIs; the paper offers no user data to establish that.
Editorial extensions
If this is right
- eHMI design becomes a configurable pipeline rather than a fixed display, so a single interface can vary message content, timing, and modality by road actor and context.
- Researchers can attribute communication outcomes to specific layers, making studies of adaptive eHMIs easier to compare and reproduce.
- Scalability is treated as a design requirement: the system does not need a new fixed message for every situation if the Processing layer can generate appropriate outputs from detected inputs.
- Inclusivity becomes part of the Processing layer's job, since adaptation must account for differences in perception, literacy, and mobility among road actors.
- Wider adoption of adaptive eHMIs would confront design teams with ethical choices about sensing people and deciding who receives extra cautionary information.
Reading between the lines
- If the framework is right, a concrete prediction follows that the paper does not test: a system designed through Input–Processing–Output will communicate better than a static interface in mixed-traffic settings with varied actors and contexts.
- The cyber-physical framing implies evaluation should cover the entire perception–decision–output loop; a display that fails because the vehicle misdetected an actor is still an adaptive-eHMI failure under this view.
- One testable extension is to turn the framework into an audit matrix (road actor × context × message type) and use it to identify gaps in existing eHMI proposals.
- Note on the supplied record: the full text under this title is a different manuscript on stochastic optimal control; it neither elaborates nor tests the eHMI framework, so the framework's support rests on the abstract's description of theory-led abstraction and expert discussion.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript arXiv:2508.12518 presents, in its abstract, a conceptual design framework for adaptive external human-machine interfaces (eHMIs) for autonomous vehicles. The framework is described as having three layers—Input, Processing, Output—developed through theory-led abstraction and expert discussion, and is claimed to help researchers and designers systematically design, analyse, and assess adaptive communication strategies, resolving longstanding limitations in eHMI research. The full text supplied for review, however, is arXiv:2508.12511v2, a paper on trust-region stochastic optimal control and measure transport. It contains no description of eHMIs, no Input/Processing/Output framework, and no expert-discussion methodology. Consequently, the development of the framework, its definitions, its application, and its claimed benefits cannot be checked from the submitted manuscript.
Significance. If properly developed and validated, a three-layer framework for adaptive eHMIs could provide a useful structuring vocabulary for a fragmented research area, especially with attention to scalability and inclusivity. The conceptual decomposition into Input, Processing, and Output is a reasonable starting point, and the abstract's emphasis on dynamic adjustment to road actors and context is timely. However, the current manuscript provides no evidence that the framework is descriptively adequate or practically useful: there is no worked example, no user or design study, no systematic classification of existing eHMI implementations, and no comparison with alternative taxonomies. The significance therefore remains entirely hypothetical at this stage.
major comments (3)
- [Full text / Abstract] The supplied full text is arXiv:2508.12511v2, a paper on trust-region stochastic optimal control, not the eHMI manuscript announced in the abstract. This mismatch removes the body of the paper under review: the Input/Processing/Output framework, the theory-led abstraction, the expert discussion, and the claims about resolving longstanding limitations have no supporting text. This is load-bearing because no reviewer can verify that the framework is derived as claimed or that the layers are coherently defined. The paper must be resubmitted with the correct full text.
- [Abstract, central utility claim] The claim that the framework 'helps researchers and designers think systematically' and provides a 'structured tool to design, analyse, and assess' adaptive eHMIs is unsupported. Theory-led abstraction and expert discussion can generate a taxonomy, but they do not establish that the taxonomy maps cleanly onto real road-actor variability or that it is practically useful. No user study, design exercise, or systematic classification of existing eHMI implementations is reported. A concrete test would be to apply the framework to existing eHMI designs spanning multi-party interactions, occluded pedestrians, group behaviour, and time-varying contexts, and to show that it represents them without omission or conflation.
- [Abstract, 'resolving longstanding limitations'] The abstract does not identify which longstanding limitations in eHMI research are resolved, nor the mechanism by which the framework resolves them. This makes the claim unfalsifiable. The authors should enumerate the limitations and map each to a specific feature of the Input/Processing/Output layers, ideally demonstrating the mapping on known hard cases such as group behaviour, occluded road users, and context-dependent communication norms.
minor comments (3)
- [Abstract / Definitions] The term 'adaptive' is used without a precise definition. It could mean reactive parameter changes, learned adaptation, or policy-level adjustment. The introduction of the correct full text should define this explicitly.
- [References] Once the correct full text is provided, the paper should engage with existing eHMI surveys and taxonomies to position the contribution within the field; the current abstract provides no related-work grounding.
- [Notation] The labels Input, Processing, and Output need explicit definitions and examples. Without the correct body, terms like 'what the system detects' are ambiguous and cannot be operationalized.
Circularity Check
No significant circularity: the abstract presents an expert-derived taxonomy whose utility is unverified, but no derivation or prediction reduces to its inputs.
full rationale
The abstract makes no quantitative prediction, fits no parameter, and contains no equation in which an output is equal to an input by construction. The proposed Input/Processing/Output framework is a structuring lens (a standard detect-decide-act decomposition), and the claim that it 'helps researchers and designers think systematically' is a stated design hypothesis grounded in 'theory-led abstraction and expert discussion.' That claim is unsupported by user studies or deployed-system evidence, but lack of empirical validation is a correctness risk, not circularity. There are no load-bearing self-citations, no imported uniqueness theorems, and no fitted values renamed as predictions. The provided full text is a different arXiv paper (2508.12511 on trust-region stochastic optimal control), so it cannot corroborate or contradict the eHMI abstract; this is a document-integrity / input-mismatch issue rather than circular reasoning. Honest finding: no circularity identified in the abstract's derivation chain.
Assumptions & free parameters
assumptions (3)
- domain assumption Adaptive eHMIs are a desirable or necessary direction for autonomous vehicle external communication.
- domain assumption The cyber-physical system lens provides a valid structuring of eHMI design.
- domain assumption Theory-led abstraction and expert discussion yield a framework that is useful for design, analysis, and assessment.
invented entities (1)
-
Three-layer adaptive eHMI framework (Input, Processing, Output)
Cite this review
Pith. "Pith review of Towards Adaptive External Communication in Autonomous Vehicles: A Conceptual Design Framework." pith.science (2026). https://pith.science/paper/NUI62EGP
@misc{pith2026250812518,
author = {Pith},
title = {Pith review of: Towards Adaptive External Communication in Autonomous Vehicles: A Conceptual Design Framework},
year = {2026},
howpublished = {\url{https://pith.science/paper/NUI62EGP}},
note = {Machine review of arXiv:2508.12518}
}
read the original abstract
External Human-Machine Interfaces (eHMIs) are key to facilitating interaction between autonomous vehicles and external road actors, yet most remain reactive and do not account for scalability and inclusivity. This paper introduces a conceptual design framework for adaptive eHMIs-interfaces that dynamically adjust communication as road actors vary and context shifts. Using the cyber-physical system as a structuring lens, the framework comprises three layers: Input (what the system detects), Processing (how the system decides), and Output (how the system communicates). Developed through theory-led abstraction and expert discussion, the framework helps researchers and designers think systematically about adaptive eHMIs and provides a structured tool to design, analyse, and assess adaptive communication strategies. We show how such systems may resolve longstanding limitations in eHMI research while raising new ethical and technical considerations.
Reference graph
Works this paper leans on
-
[1]
Abdolmaleki, R
A. Abdolmaleki, R. Lioutikov, J. R. Peters, N. Lau, L. Pualo Reis, and G. Neumann. Model- based relative entropy stochastic search.Advances in Neural Information Processing Systems, 28, 2015
2015
-
[2]
A. Abdolmaleki, J. T. Springenberg, J. Degrave, S. Bohez, Y . Tassa, D. Belov, N. Heess, and M. Riedmiller. Relative entropy regularized policy iteration.arXiv preprint arXiv:1812.02256, 2018
arXiv 2018
-
[3]
A. Abdolmaleki, J. T. Springenberg, Y . Tassa, R. Munos, N. Heess, and M. Riedmiller. Maximum a posteriori policy optimisation.arXiv preprint arXiv:1806.06920, 2018
arXiv 2018
-
[4]
Achiam, D
J. Achiam, D. Held, A. Tamar, and P. Abbeel. Constrained policy optimization. InInternational conference on machine learning, pages 22–31. PMLR, 2017
2017
-
[5]
T. Akhound-Sadegh, J. Lee, A. J. Bose, V . De Bortoli, A. Doucet, M. M. Bronstein, D. Beaini, S. Ravanbakhsh, K. Neklyudov, and A. Tong. Progressive inference-time annealing of diffusion models for sampling from Boltzmann densities.arXiv preprint arXiv:2506.16471, 2025
arXiv 2025
-
[6]
T. Akhound-Sadegh, J. Rector-Brooks, A. J. Bose, S. Mittal, P. Lemos, C.-H. Liu, M. Sendera, S. Ravanbakhsh, G. Gidel, Y . Bengio, et al. Iterated denoising energy matching for sampling from Boltzmann densities.arXiv preprint arXiv:2402.06121, 2024
arXiv 2024
-
[7]
Akrour, J
R. Akrour, J. Pajarinen, J. Peters, and G. Neumann. Projections for approximate policy iteration algorithms. InInternational Conference on Machine Learning, pages 181–190. PMLR, 2019
2019
-
[8]
M. S. Albergo and E. Vanden-Eijnden. NETS: A non-equilibrium transport sampler.arXiv preprint arXiv:2410.02711, 2024
arXiv 2024
Show all 161 references
-
[9]
Arenz, P
O. Arenz, P. Dahlinger, Z. Ye, M. V olpp, and G. Neumann. A unified perspective on natural gradient variational inference with Gaussian mixture models.arXiv preprint arXiv:2209.11533, 2022
2022 arXiv
-
[10]
Arenz, M
O. Arenz, M. Zhong, and G. Neumann. Trust-region variational inference with Gaussian mixture models.Journal of Machine Learning Research, 21(163):1–60, 2020
2020
-
[11]
C. Beck, S. Becker, P. Grohs, N. Jaafari, and A. Jentzen. Solving the Kolmogorov PDE by means of deep learning.Journal of Scientific Computing, 88:1–28, 2021
2021
-
[12]
Becker, N
P. Becker, N. Freymuth, S. Thilges, F. Otto, and G. Neumann. Troll: Trust regions improve reinforcement learning for large language models.arXiv preprint arXiv:2510.03817, 2025
2025
-
[13]
Bellman.Dynamic programming
R. Bellman.Dynamic programming. Princeton University Press, 1957
1957
-
[14]
Berner, M
J. Berner, M. Dablander, and P. Grohs. Numerically solving parametric families of high- dimensional Kolmogorov partial differential equations via deep learning.Advances in Neural Information Processing Systems, 33:16615–16627, 2020
2020
-
[15]
Berner, L
J. Berner, L. Richter, M. Sendera, J. Rector-Brooks, and N. Malkin. From discrete-time policies to continuous-time diffusion samplers: Asymptotic equivalences and faster training. arXiv preprint arXiv:2501.06148, 2025
2025
-
[16]
Berner, L
J. Berner, L. Richter, and K. Ullrich. An optimal control perspective on diffusion-based generative modeling.Transactions on Machine Learning Research, 2024
2024
-
[17]
Black, M
K. Black, M. Janner, Y . Du, I. Kostrikov, and S. Levine. Training diffusion models with reinforcement learning. InThe Twelfth International Conference on Learning Representations, 2024
2024
-
[18]
Blessing, J
D. Blessing, J. Berner, L. Richter, and G. Neumann. Underdamped diffusion bridges with appli- cations to sampling. InThe Thirteenth International Conference on Learning Representations, 2025
2025
-
[19]
Blessing, X
D. Blessing, X. Jia, J. Esslinger, F. Vargas, and G. Neumann. Beyond ELBOs: A large-scale evaluation of variational methods for sampling.arXiv preprint arXiv:2406.07423, 2024. 11
2024 arXiv
-
[20]
Blessing, X
D. Blessing, X. Jia, and G. Neumann. End-to-end learning of Gaussian mixture priors for diffusion sampler.arXiv preprint arXiv:2503.00524, 2025
2025 arXiv
-
[21]
P. G. Bolhuis, D. Chandler, C. Dellago, and P. L. Geissler. Transition path sampling: Throwing ropes over rough mountain passes, in the dark.Annual review of physical chemistry, 53(1):291– 318, 2002
2002
-
[22]
Bradbury, R
J. Bradbury, R. Frostig, P. Hawkins, M. J. Johnson, C. Leary, D. Maclaurin, G. Necula, A. Paszke, J. VanderPlas, S. Wanderman-Milne, et al. Jax: Autograd and xla.Astrophysics Source Code Library, pages ascl–2111, 2021
2021
-
[23]
Brekelmans, V
R. Brekelmans, V . Masrani, F. Wood, G. V . Steeg, and A. Galstyan. All in the exponential fam- ily: Bregman duality in thermodynamic variational inference.arXiv preprint arXiv:2007.00642, 2020
2007 arXiv
-
[24]
R. P. Brent. An algorithm with guaranteed convergence for finding a zero of a function.The computer journal, 14(4):422–425, 1971
1971
-
[25]
Celik, Z
O. Celik, Z. Li, D. Blessing, G. Li, D. Palanicek, J. Peters, G. Chalvatzaki, and G. Neu- mann. DIME: Diffusion-based maximum entropy reinforcement learning.arXiv preprint arXiv:2502.02316, 2025
2025 arXiv
-
[26]
J. Chen, L. Richter, J. Berner, D. Blessing, G. Neumann, and A. Anandkumar. Sequential controlled Langevin diffusions. InThe Thirteenth International Conference on Learning Representations, 2025
2025
-
[27]
Chetrite and H
R. Chetrite and H. Touchette. Variational and optimal control representations of condi- tioned and driven processes.Journal of Statistical Mechanics: Theory and Experiment, 2015(12):P12001, 2015
2015
-
[28]
J. Choi, Y . Chen, M. Tao, and G.-H. Liu. Non-equilibrium annealed adjoint sampler.arXiv preprint arXiv:2506.18165, 2025
2025
-
[29]
Clark, P
K. Clark, P. Vicol, K. Swersky, and D. J. Fleet. Directly fine-tuning diffusion models on differentiable rewards. InThe Twelfth International Conference on Learning Representations, 2024
2024
-
[30]
A. R. Conn, N. I. Gould, and P. L. Toint.Trust region methods. SIAM, 2000
2000
-
[31]
G. E. Crooks. Measuring thermodynamic length.Physical Review Letters, 99(10):100602, 2007
2007
-
[32]
M. Cuturi. Sinkhorn distances: Lightspeed computation of optimal transport.Advances in neural information processing systems, 26, 2013
2013
-
[33]
Cuturi, L
M. Cuturi, L. Meng-Papaxanthos, Y . Tian, C. Bunne, G. Davis, and O. Teboul. Optimal trans- port tools (OTT): A jax toolbox for all things Wasserstein.arXiv preprint arXiv:2201.12324, 2022
2022 arXiv
-
[34]
P. Dai Pra. A stochastic control approach to reciprocal diffusion processes.Applied mathemat- ics and Optimization, 23(1):313–329, 1991
1991
-
[35]
Dai Pra, L
P. Dai Pra, L. Meneghini, and W. J. Runggaldier. Connections between stochastic control and dynamic games.Mathematics of Control, Signals and Systems, 9:303–326, 1996
1996
-
[36]
A. Das, D. C. Rose, J. P. Garrahan, and D. T. Limmer. Reinforcement learning of rare diffusive dynamics.The Journal of Chemical Physics, 155(13), 2021
2021
-
[37]
De Bortoli, J
V . De Bortoli, J. Thornton, J. Heng, and A. Doucet. Diffusion Schrödinger bridge with applications to score-based generative modeling.Advances in Neural Information Processing Systems, 34:17695–17709, 2021
2021
-
[38]
Dellago, P
C. Dellago, P. G. Bolhuis, and D. Chandler. Efficient transition path sampling: Application to lennard-jones cluster rearrangements.The Journal of chemical physics, 108(22):9236–9245, 1998. 12
1998
-
[39]
K. Didi, F. Vargas, S. V . Mathis, V . Dutordoir, E. Mathieu, U. J. Komorowska, and P. Lio. A framework for conditional diffusion modelling with applications in motif scaffolding for protein design.arXiv preprint arXiv:2312.09236, 2023
2023 arXiv
-
[40]
Z. Ding, Y . Jiao, X. Lu, Z. Yang, and C. Yuan. Sampling via Föllmer flow.arXiv preprint arXiv:2311.03660, 2023
2023 arXiv
-
[41]
Domingo-Enrich
C. Domingo-Enrich. A taxonomy of loss functions for stochastic optimal control.arXiv preprint arXiv:2410.00345, 2024
2024 arXiv
-
[42]
Domingo-Enrich, M
C. Domingo-Enrich, M. Drozdzal, B. Karrer, and R. T. Q. Chen. Adjoint matching: Fine-tuning flow and diffusion generative models with memoryless stochastic optimal control. InThe Thirteenth International Conference on Learning Representations, 2025
2025
-
[43]
Domingo-Enrich, J
C. Domingo-Enrich, J. Han, B. Amos, J. Bruna, and R. T. Q. Chen. Stochastic optimal control matching. InThe Thirty-eighth Annual Conference on Neural Information Processing Systems, 2024
2024
-
[44]
Doucet, W
A. Doucet, W. Grathwohl, A. G. d. G. Matthews, and H. Strathmann. Score-based diffusion meets annealed importance sampling. InAdvances in Neural Information Processing Systems, 2022
2022
-
[45]
Y . Du, M. Plainer, R. Brekelmans, C. Duan, F. Noe, C. P. Gomes, A. Aspuru-Guzik, and K. Neklyudov. Doob’s lagrangian: A sample-efficient variational approach to transition path sampling. InThe Thirty-eighth Annual Conference on Neural Information Processing Systems
-
[46]
Y . Du, M. Plainer, R. Brekelmans, C. Duan, F. Noe, C. P. Gomes, A. Aspuru-Guzik, and K. Neklyudov. Doob’s lagrangian: A sample-efficient variational approach to transition path sampling. InThe Thirty-eighth Annual Conference on Neural Information Processing Systems, 2024
2024
-
[47]
Erives, B
E. Erives, B. Jing, P. Holderrieth, and T. Jaakkola. Continuously tempered diffusion samplers. InFrontiers in Probabilistic Inference: Learning meets Sampling, 2025
2025
-
[48]
Y . Fan, O. Watkins, Y . Du, H. Liu, M. Ryu, C. Boutilier, P. Abbeel, M. Ghavamzadeh, K. Lee, and K. Lee. Dpok: Reinforcement learning for fine-tuning text-to-image diffusion models. arXiv preprint arXiv:2305.16381, 2023
2023 arXiv
-
[49]
M. F. Faulkner and S. Livingstone. Sampling algorithms in statistical physics: a guide for statistics and machine learning.Statistical Science, 39(1):137–164, 2024
2024
-
[50]
Fleming and R
W. Fleming and R. Rishel.Deterministic and Stochastic Optimal Control. Applications of mathematics. Springer, 1975
1975
-
[51]
W. H. Fleming and H. M. Soner.Controlled Markov processes and viscosity solutions, volume 25. Springer Science & Business Media, 2006
2006
-
[52]
S. Fu, N. Tamir, S. Sundaram, L. Chai, R. Zhang, T. Dekel, and P. Isola. Dreamsim: Learning new dimensions of human visual similarity using synthetic data.arXiv preprint arXiv:2306.09344, 2023
2023 arXiv
-
[53]
Geffner and J
T. Geffner and J. Domke. MCMC variational inference via uncorrected hamiltonian annealing. Advances in Neural Information Processing Systems, 34:639–651, 2021
2021
-
[54]
Geffner and J
T. Geffner and J. Domke. Langevin diffusion variational inference. InInternational Conference on Artificial Intelligence and Statistics, pages 576–593. PMLR, 2023
2023
-
[55]
Gelman, J
A. Gelman, J. Carlin, H. Stern, D. Dunson, A. Vehtari, and D. Rubin.Bayesian Data Analysis, Third Edition. Chapman & Hall/CRC Texts in Statistical Science. Taylor & Francis, 2013
2013
-
[56]
Grenioux, M
L. Grenioux, M. Noble, and M. Gabrié. Improving the evaluation of samplers on multi-modal targets.arXiv preprint arXiv:2504.08916, 2025
2025 arXiv
-
[57]
Gritsaev, N
T. Gritsaev, N. Morozov, K. Tamogashev, D. Tiapkin, S. Samsonov, A. Naumov, D. Vetrov, and N. Malkin. Adaptive destruction processes for diffusion samplers.arXiv preprint arXiv:2506.01541, 2025. 13
2025 arXiv
-
[58]
W. Guo, M. Tao, and Y . Chen. Complexity analysis of normalizing constant estimation: from Jarzynski equality to annealed importance sampling and beyond.arXiv preprint arXiv:2502.04575, 2025
2025 arXiv
-
[59]
J. Han, A. Jentzen, and W. E. Solving high-dimensional partial differential equations using deep learning.Proceedings of the National Academy of Sciences, 115(34):8505–8510, 2018
2018
-
[60]
Hartmann, O
C. Hartmann, O. Kebiri, L. Neureither, and L. Richter. Variational approach to rare event simulation using least-squares regression.Chaos: An Interdisciplinary Journal of Nonlinear Science, 29(6), 2019
2019
-
[61]
Hartmann and L
C. Hartmann and L. Richter. Nonasymptotic bounds for suboptimal importance sampling. SIAM/ASA Journal on Uncertainty Quantification, 12(2):309–346, 2024
2024
-
[62]
Hartmann, L
C. Hartmann, L. Richter, C. Schütte, and W. Zhang. Variational characterization of free energy: Theory and algorithms.Entropy, 19(11), 2017
2017
-
[63]
Hartmann and C
C. Hartmann and C. Schütte. Efficient rare event simulation by optimal nonequilibrium forcing. Journal of Statistical Mechanics: Theory and Experiment, 2012(11):P11004, 2012
2012
-
[64]
Havens, B
A. Havens, B. K. Miller, B. Yan, C. Domingo-Enrich, A. Sriram, B. Wood, D. Levine, B. Hu, B. Amos, B. Karrer, et al. Adjoint sampling: Highly scalable diffusion samplers via adjoint matching.arXiv preprint arXiv:2504.11713, 2025
2025 arXiv
-
[65]
J. He, W. Chen, M. Zhang, D. Barber, and J. M. Hernández-Lobato. Training neural samplers with reverse diffusive kl divergence.arXiv preprint arXiv:2410.12456, 2024
2024 arXiv
-
[66]
J. He, Y . Du, F. Vargas, Y . Wang, C. P. Gomes, J. M. Hernández-Lobato, and E. Vanden-Eijnden. FEAT: Free energy estimators with adaptive transport.arXiv preprint arXiv:2504.11516, 2025
2025
-
[67]
J. He, Y . Du, F. Vargas, D. Zhang, S. Padhy, R. OuYang, C. Gomes, and J. M. Hernández- Lobato. No trick, no treat: Pursuits and challenges towards simulation-free training of neural samplers.arXiv preprint arXiv:2502.06685, 2025
2025 arXiv
-
[68]
Hendrycks and K
D. Hendrycks and K. Gimpel. Gaussian error linear units (gelus).arXiv preprint arXiv:1606.08415, 2016
2016 arXiv
-
[69]
Hénin, T
J. Hénin, T. Lelièvre, M. R. Shirts, O. Valsson, and L. Delemotte. Enhanced sampling methods for molecular dynamics simulations.arXiv preprint arXiv:2202.04164, 2022
2022 arXiv
-
[70]
Hertrich and R
J. Hertrich and R. Gruhlke. Importance corrected neural jko sampling.arXiv preprint arXiv:2407.20444, 2024
2024 arXiv
-
[71]
Hessel, A
J. Hessel, A. Holtzman, M. Forbes, R. L. Bras, and Y . Choi. Clipscore: A reference-free evaluation metric for image captioning.arXiv preprint arXiv:2104.08718, 2021
2021 arXiv
-
[72]
Holderrieth, M
P. Holderrieth, M. S. Albergo, and T. Jaakkola. Leaps: A discrete neural sampler via locally equivariant networks.arXiv preprint arXiv:2502.10843, 2025
2025 arXiv
-
[73]
Holdijk, Y
L. Holdijk, Y . Du, F. Hooft, P. Jaini, B. Ensing, and M. Welling. Stochastic optimal control for collective variable free sampling of molecular transition paths. InThirty-seventh Conference on Neural Information Processing Systems, 2023
2023
-
[74]
Holdijk, Y
L. Holdijk, Y . Du, P. Jaini, F. Hooft, B. Ensing, and M. Welling. Path integral stochastic optimal control for sampling transition paths. InICML 2022 2nd AI for Science Workshop, 2022
2022
-
[75]
Huang, Y
J. Huang, Y . Jiao, L. Kang, X. Liao, J. Liu, and Y . Liu. Schrödinger-Föllmer sampler: sampling without ergodicity.arXiv preprint arXiv:2106.10880, 2021
2021 arXiv
-
[76]
Huang, H
X. Huang, H. Dong, Y . Hao, Y . Ma, and T. Zhang. Monte Carlo sampling without isoperimetry: A reverse diffusion approach.arXiv preprint arXiv:2307.02037, 2023
2023 arXiv
-
[77]
Izrailev, S
S. Izrailev, S. Stepaniants, B. Isralewitz, D. Kosztin, H. Lu, F. Molnar, W. Wriggers, and K. Schulten. Steered molecular dynamics. InComputational Molecular Dynamics: Chal- lenges, Methods, Ideas: Proceedings of the 2nd International Symposium on Algorithms for Macromolecular...
1997
-
[78]
H. J. Kappen and H. C. Ruiz. Adaptive importance sampling for control and inference.Journal of Statistical Physics, 162(5):1244–1266, 2016
2016
-
[79]
M. Kim, S. Choi, T. Yun, E. Bengio, L. Feng, J. Rector-Brooks, S. Ahn, J. Park, N. Malkin, and Y . Bengio. Adaptive teachers for amortized samplers.arXiv preprint arXiv:2410.01432, 2024
2024 arXiv
-
[80]
M. Kim, K. Seong, D. Woo, S. Ahn, and M. Kim. On scalable and efficient training of diffusion samplers.arXiv preprint arXiv:2505.19552, 2025
2025
-
[81]
D. P. Kingma. Adam: A method for stochastic optimization.arXiv preprint arXiv:1412.6980, 2014
2014 arXiv
-
[82]
G.-H. Liu, J. Choi, Y . Chen, B. K. Miller, and R. T. Chen. Adjoint schrödinger bridge sampler. arXiv preprint arXiv:2506.22565, 2025
2025
-
[83]
J. Liu, G. Liu, J. Liang, Y . Li, J. Liu, X. Wang, P. Wan, D. Zhang, and W. Ouyang. Flow-grpo: Training flow matching models via online rl, 2025
2025
-
[84]
Z. Liu, T. Z. Xiao, W. Liu, Y . Bengio, and D. Zhang. Efficient diversity-preserving diffusion alignment via gradient-informed GFlowNets, 2025
2025
-
[85]
W. Meng, Q. Zheng, Y . Shi, and G. Pan. An off-policy trust region policy optimization method with monotonic improvement guarantee for deep reinforcement learning.IEEE Transactions on Neural Networks and Learning Systems, 33(5):2223–2235, 2021
2021
-
[86]
L. I. Midgley, V . Stimper, G. N. Simm, B. Schölkopf, and J. M. Hernández-Lobato. Flow annealed importance sampling bootstrap.arXiv preprint arXiv:2208.01893, 2022
2022 arXiv
-
[87]
R. M. Neal. Probabilistic inference using Markov chain Monte Carlo methods. 1993
1993
-
[88]
Noble, L
M. Noble, L. Grenioux, M. Gabrié, and A. O. Durmus. Learned reference-based diffusion sampling for multi-modal distributions.arXiv preprint arXiv:2410.19449, 2024
2024 arXiv
-
[89]
Nüsken and L
N. Nüsken and L. Richter. Solving high-dimensional Hamilton–Jacobi–Bellman PDEs using neural networks: perspectives from the theory of controlled diffusions and measures on path space.Partial differential equations and applications, 2:1–48, 2021
2021
-
[90]
F. Otto, P. Becker, N. A. Vien, H. C. Ziesche, and G. Neumann. Differentiable trust region layers for deep reinforcement learning.arXiv preprint arXiv:2101.09207, 2021
2021 arXiv
-
[91]
OuYang, B
R. OuYang, B. Qiang, and J. M. Hernández-Lobato. BNEM: A Boltzmann sampler based on bootstrapped noised energy matching.arXiv preprint arXiv:2409.09787, 2024
2024
-
[92]
Pajarinen, H
J. Pajarinen, H. L. Thai, R. Akrour, J. Peters, and G. Neumann. Compatible natural gradient policy search.Machine Learning, 108:1443–1466, 2019
2019
-
[93]
M. Pavon. Stochastic control and nonequilibrium thermodynamical systems.Applied Mathe- matics and Optimization, 19(1):187–202, 1989
1989
-
[94]
M. Pavon. On local entropy, stochastic control and deep neural networks.arXiv preprint arXiv:2204.13049, 2022
2022 arXiv
-
[95]
Peters, K
J. Peters, K. Mulling, and Y . Altun. Relative entropy policy search. InProceedings of the AAAI Conference on Artificial Intelligence, volume 24, pages 1607–1612, 2010
2010
-
[96]
Pham.Continuous-time Stochastic Control and Optimization with Financial Applications
H. Pham.Continuous-time Stochastic Control and Optimization with Financial Applications. Stochastic Modelling and Applied Probability. Springer Berlin Heidelberg, 2009
2009
-
[97]
Richter.Solving high-dimensional PDEs, approximation of path space measures and importance sampling of diffusions
L. Richter.Solving high-dimensional PDEs, approximation of path space measures and importance sampling of diffusions. PhD thesis, BTU Cottbus-Senftenberg, 2021
2021
-
[98]
Richter and J
L. Richter and J. Berner. Robust SDE-based variational formulations for solving linear PDEs via deep learning. InInternational Conference on Machine Learning, pages 18649–18666. PMLR, 2022. 15
2022
-
[99]
Richter and J
L. Richter and J. Berner. Improved sampling via learned diffusions. InInternational Conference on Learning Representations, 2024
2024
-
[100]
Richter, A
L. Richter, A. Boustati, N. Nüsken, F. Ruiz, and O. D. Akyildiz. VarGrad: A low-variance gradient estimator for variational inference.Advances in Neural Information Processing Systems, 33:13481–13492, 2020
2020
-
[101]
Richter, L
L. Richter, L. Sallandt, and N. Nüsken. Solving high-dimensional parabolic PDEs using the tensor train format. InInternational Conference on Machine Learning, pages 8998–9009. PMLR, 2021
2021
-
[102]
Richter, L
L. Richter, L. Sallandt, and N. Nüsken. From continuous-time formulations to discretization schemes: tensor trains and robust regression for bsdes and parabolic pdes.Journal of Machine Learning Research, 25(248):1–40, 2024
2024
-
[103]
Rissanen, R
S. Rissanen, R. OuYang, J. He, W. Chen, M. Heinonen, A. Solin, and J. M. Hernández-Lobato. Progressive tempering sampler with diffusion.arXiv preprint arXiv:2506.05231, 2025
2025
-
[104]
Rombach, A
R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer. High-resolution image synthesis with latent diffusion models. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 10684–10695, 2022
2022
-
[105]
D. C. Rose, J. F. Mair, and J. P. Garrahan. A reinforcement learning approach to rare trajectory sampling.New Journal of Physics, 23(1):013013, 2021
2021
-
[106]
R. Y . Rubinstein and D. P. Kroese.The cross-entropy method: a unified approach to combi- natorial optimization, Monte-Carlo simulation and machine learning. Springer Science & Business Media, 2013
2013
-
[107]
Sabate Vidales, D
M. Sabate Vidales, D. Šiška, and L. Szpruch. Unbiased deep solvers for linear parametric PDEs.Applied Mathematical Finance, 28(4):299–329, 2021
2021
-
[108]
Salamon and R
P. Salamon and R. S. Berry. Thermodynamic length and dissipated availability.Physical Review Letters, 51(13):1127, 1983
1983
-
[109]
Sanokowski, W
S. Sanokowski, W. Berghammer, M. Ennemoser, H. P. Wang, S. Hochreiter, and S. Lehner. Scalable discrete diffusion samplers: Combinatorial optimization and statistical physics.arXiv preprint arXiv:2502.08696, 2025
2025 arXiv
-
[110]
Sanokowski, L
S. Sanokowski, L. Gruber, C. Bartmann, S. Hochreiter, and S. Lehner. Rethinking losses for diffusion bridge samplers.arXiv preprint arXiv:2506.10982, 2025
2025
-
[111]
Schopmans and P
H. Schopmans and P. Friederich. Temperature-annealed boltzmann generators.arXiv preprint arXiv:2501.19077, 2025
2025 arXiv
-
[112]
Schulman, S
J. Schulman, S. Levine, P. Abbeel, M. Jordan, and P. Moritz. Trust region policy optimization. InInternational conference on machine learning, pages 1889–1897. PMLR, 2015
2015
-
[113]
Schulman, F
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov. Proximal policy optimization algorithms.arXiv preprint arXiv:1707.06347, 2017
2017 arXiv
-
[114]
Sendera, M
M. Sendera, M. Kim, S. Mittal, P. Lemos, L. Scimeca, J. Rector-Brooks, A. Adam, Y . Bengio, and N. Malkin. Improved off-policy training of diffusion samplers.Advances in Neural Information Processing Systems, 37:81016–81045, 2024
2024
-
[115]
Seong, S
K. Seong, S. Park, S. Kim, W. Y . Kim, and S. Ahn. Transition path sampling with improved off-policy training of diffusion path samplers.arXiv preprint arXiv:2405.19961, 2024
2024 arXiv
-
[116]
Z. Shi, L. Yu, T. Xie, and C. Zhang. Diffusion-PINN sampler.arXiv preprint arXiv:2410.15336, 2024
2024 arXiv
-
[117]
A. N. Singh, A. Das, and D. T. Limmer. Variational path sampling of rare dynamical events. Annual Review of Physical Chemistry, 76, 2025
2025
-
[118]
A. N. Singh and D. T. Limmer. Variational deep learning of equilibrium transition path ensembles.The Journal of Chemical Physics, 159(2), 2023. 16
2023
-
[119]
J. Sun, J. Berner, L. Richter, M. Zeinhofer, J. Müller, K. Azizzadenesheli, and A. Anand- kumar. Dynamical measure transport and neural PDE solvers for sampling.arXiv preprint arXiv:2407.07873, 2024
2024 arXiv
-
[120]
Y . Sun, D. Wierstra, T. Schaul, and J. Schmidhuber. Efficient natural evolution strategies. In Proceedings of the 11th Annual conference on Genetic and evolutionary computation, pages 539–546, 2009
2009
-
[121]
S. Syed, A. Bouchard-Côté, K. Chern, and A. Doucet. Optimised annealed sequential Monte Carlo samplers.arXiv preprint arXiv:2408.12057, 2024
2024
-
[122]
C. B. Tan, A. J. Bose, C. Lin, L. Klein, M. M. Bronstein, and A. Tong. Scalable equilibrium sampling with sequential Boltzmann generators.arXiv preprint arXiv:2502.18462, 2025
2025
-
[123]
H. Y . Tan, S. Osher, and W. Li. Noise-free sampling algorithms via regularized Wasserstein proximals.arXiv preprint arXiv:2308.14945, 2023
2023 arXiv
-
[124]
Tancik, P
M. Tancik, P. Srinivasan, B. Mildenhall, S. Fridovich-Keil, N. Raghavan, U. Singhal, R. Ra- mamoorthi, J. Barron, and R. Ng. Fourier features let networks learn high frequency functions in low dimensional domains.Advances in neural information processing systems, 33:7537– 7547, 2020
2020
-
[125]
Thalmeier, H
D. Thalmeier, H. J. Kappen, S. Totaro, and V . Gómez. Adaptive smoothing for path integral control.Journal of Machine Learning Research, 21(191):1–37, 2020
2020
-
[126]
A. Thin, N. Kotelevskii, A. Durmus, E. Moulines, M. Panov, and A. Doucet. Monte Carlo variational auto-encoders. InInternational Conference on Machine Learning, 2021
2021
-
[127]
Tzen and M
B. Tzen and M. Raginsky. Theoretical guarantees for sampling and inference in generative models with latent diffusions. InConference on Learning Theory, pages 3084–3114. PMLR, 2019
2019
-
[128]
Uehara, Y
M. Uehara, Y . Zhao, K. Black, E. Hajiramezanali, G. Scalia, N. L. Diamant, A. M. Tseng, T. Biancalani, and S. Levine. Fine-tuning of continuous-time diffusion models as entropy- regularized control.arXiv preprint arXiv:2402.15194, 2024
2024 arXiv
-
[129]
Van Handel
R. Van Handel. Stochastic calculus, filtering, and stochastic control.Course notes., URL http://www. princeton. edu/rvan/acm217/ACM217. pdf, 14, 2007
2007
-
[130]
Vanden-Eijnden et al
E. Vanden-Eijnden et al. Transition-path theory and path-finding algorithms for the study of rare events.Annual review of physical chemistry, 61:391–420, 2010
2010
-
[131]
Vargas, W
F. Vargas, W. Grathwohl, and A. Doucet. Denoising diffusion samplers.arXiv preprint arXiv:2302.13834, 2023
2023 arXiv
-
[132]
Vargas, A
F. Vargas, A. Ovsianas, D. Fernandes, M. Girolami, N. D. Lawrence, and N. Nüsken. Bayesian learning via neural Schrödinger–Föllmer flows.Statistics and Computing, 33(1):3, 2023
2023
-
[133]
Vargas, S
F. Vargas, S. Padhy, D. Blessing, and N. Nüsken. Transport meets variational inference: Controlled Monte Carlo diffusions. InThe Twelfth International Conference on Learning Representations, 2024
2024
-
[134]
Venkatraman, M
S. Venkatraman, M. Jain, L. Scimeca, M. Kim, M. Sendera, M. Hasan, L. Rowe, S. Mittal, P. Lemos, E. Bengio, et al. Amortizing intractable inference in diffusion models for vision, language, and control.arXiv preprint arXiv:2405.20971, 2024
2024 arXiv
-
[135]
von Klitzing, D
C. von Klitzing, D. Blessing, H. Schopmans, P. Friederich, and G. Neumann. Learning boltzmann generators via constrained mass transport.arXiv preprint arXiv:2510.18460, 2025
2025
-
[136]
C. Wang, X. Zhang, K. Cui, W. Zhao, Y . Guan, and T. Yu. Importance weighted score matching for diffusion samplers with enhanced mode coverage.arXiv preprint arXiv:2505.19431, 2025
2025 arXiv
-
[137]
Wierstra, T
D. Wierstra, T. Schaul, T. Glasmachers, Y . Sun, J. Peters, and J. Schmidhuber. Natural evolution strategies.The Journal of Machine Learning Research, 15(1):949–980, 2014. 17
2014
-
[138]
H. Wu, J. Köhler, and F. Noé. Stochastic normalizing flows.Advances in neural information processing systems, 33:5933–5944, 2020
2020
-
[139]
L. Wu, Y . Han, C. A. Naesseth, and J. P. Cunningham. Reverse diffusion Sequential Monte Carlo samplers.arXiv preprint arXiv:2508.05926, 2025
2025
-
[140]
X. Wu, Y . Hao, K. Sun, Y . Chen, F. Zhu, R. Zhao, and H. Li. Human preference score v2: A solid benchmark for evaluating human preferences of text-to-image synthesis.arXiv preprint arXiv:2306.09341, 2023
2023 arXiv
-
[141]
Y . Wu, E. Mansimov, R. B. Grosse, S. Liao, and J. Ba. Scalable trust-region method for deep reinforcement learning using Kronecker-factored approximation.Advances in neural information processing systems, 30, 2017
2017
-
[142]
H. Xu, J. Xuan, G. Zhang, and J. Lu. Trust region policy optimization via entropy regularization for kullback–leibler divergence constraint.Neurocomputing, 589:127716, 2024
2024
-
[143]
J. Xu, X. Liu, Y . Wu, Y . Tong, Q. Li, M. Ding, J. Tang, and Y . Dong. Imagereward: Learning and evaluating human preferences for text-to-image generation. InThirty-seventh Conference on Neural Information Processing Systems, 2023
2023
-
[144]
J. Yan, H. Touchette, and G. M. Rotskoff. Learning nonequilibrium control forces to character- ize dynamical phase transitions.Physical Review E, 105(2):024115, 2022
2022
-
[145]
T.-Y . Yang, J. Rosca, K. Narasimhan, and P. J. Ramadge. Projection-based constrained policy optimization.arXiv preprint arXiv:2010.03152, 2020
2010 arXiv
-
[146]
S. Yoon, H. Hwang, H. Jeong, D. K. Shin, C.-S. Park, S. Kweon, and F. C. Park. Value gradient sampler: Sampling as sequential decision making.arXiv preprint arXiv:2502.13280, 2025
2025
-
[147]
Zhang, R
D. Zhang, R. T. Chen, C.-H. Liu, A. Courville, and Y . Bengio. Diffusion generative flow samplers: Improving learning signals through partial trajectory optimization. InThe Twelfth International Conference on Learning Representations, 2024
2024
-
[148]
Zhang, Y
D. Zhang, Y . Zhang, J. Gu, R. Zhang, J. Susskind, N. Jaitly, and S. Zhai. Improving GFlowNets for text-to-image diffusion alignment.arXiv preprint arXiv:2406.00633, 2024
2024 arXiv
-
[149]
Zhang, K
G. Zhang, K. Hsu, J. Li, C. Finn, and R. Grosse. Differentiable annealed importance sampling and the perils of gradient noise. InAdvances in Neural Information Processing Systems, 2021
2021
-
[150]
Zhang, P
L. Zhang, P. Potaptchik, G. Deligiannidis, A. Doucet, H.-D. Dau, and S. Syed. Generalised parallel tempering: Flexible replica exchange via flows and diffusions. InFrontiers in Proba- bilistic Inference: Learning meets Sampling, 2025
2025
-
[151]
Zhang, P
L. Zhang, P. Potaptchik, J. He, Y . Du, A. Doucet, F. Vargas, H.-D. Dau, and S. Syed. Acceler- ated parallel tempering via neural transports.arXiv preprint arXiv:2502.10328, 2025
2025
-
[152]
Zhang and Y
Q. Zhang and Y . Chen. Path Integral Sampler: a stochastic control approach for sampling. In International Conference on Learning Representations, 2022
2022
-
[153]
Zhang, H
W. Zhang, H. Wang, C. Hartmann, M. Weber, and C. Schütte. Applications of the cross-entropy method to importance sampling and optimal control of diffusions.SIAM Journal on Scientific Computing, 36(6):A2654–A2672, 2014
2014
-
[154]
Zhang, L
X. Zhang, L. Wang, J. Helwig, Y . Luo, C. Fu, Y . Xie, M. Liu, Y . Lin, Z. Xu, K. Yan, et al. Artificial intelligence for science in quantum, atomistic, and continuum systems.arXiv preprint arXiv:2307.08423, 2023
2023 arXiv
-
[155]
M. Zhou, J. Han, and J. Lu. Actor-critic method for high dimensional static Hamilton–Jacobi– Bellman partial differential equations based on neural networks.SIAM Journal on Scientific Computing, 43(6):A4043–A4066, 2021
2021
-
[156]
reward hacking
Y . Zhu, W. Guo, J. Choi, G.-H. Liu, Y . Chen, and M. Tao. Mdns: Masked diffusion neural sampler via stochastic optimal control, 2025. 18 Appendix A Assumptions and auxiliary results 21 A.1 Additional notation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . ....
2025
-
[157]
[84, 148] consider alternative algorithms that learn the value functions
considers GRPO for flow matching fine-tuning. [84, 148] consider alternative algorithms that learn the value functions. Diffusion-based sampling from unnormalized densities.Early work on sampling from unnormal- ized densities based on a Schrödinger-Föllmer diffusions dates bac...
-
[158]
(Connection between solution and value function) The solution can be written as u∗ =�σ ⊤�V
-
[159]
the uncontrolled path measurePsatisfies dQ dP (X) = e−W(X,0) Z(X 0) withZ(X 0) =E e−W(X,0) |X0 .(38)
(Optimal change of measure) The Radon-Nikodym derivative of the optimal path measure Q w.r.t. the uncontrolled path measurePsatisfies dQ dP (X) = e−W(X,0) Z(X 0) withZ(X 0) =E e−W(X,0) |X0 .(38)
-
[160]
(PDE for value function) The value function V is the solution to the Hamilton-Jacobi-Bellman (HJB) equation (∂t +L)V(x, t)− 1 2 ∥(σ⊤∇V)(x, t)∥2 +f(x, t) = 0, V(x, T) =g(x),(39) 24 where L:= 1 2 � d i,j=1(σσ ⊤)ij∂x� ∂x� + � d i=1 bi∂x� denotes the infinitesimal generator of the...
-
[161]
log dP u��� dP u (X u��� ) 2# − E log dP u��� dP u (X u��� ) 2 (94b) =E
(Estimator for value function) For every (x, t)�� d �[0, T] the value function can be written as V(x, t) =�log� � e−W(X,t)��Xt =x � ,whereXis the solution of the uncontrolled SDE in(34). Combining the expressions for u∗ and V in Thm. D.1, we directly obtain the path integral r...
2000
Reviewed August 5, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.