Pith. sign in

REVIEW 4 major objections 5 minor 26 references

A physics-informed reinforcement learning approach for the interfacial area transport in two-phase flow

T0 review · 4 major / 5 minor · reviewed 2026-08-14 · deepseek-v4-flash

Pith's one-line read A physics-informed reinforcement-learning framework predicts axial bubble surface area in two-phase flow, reporting a 6.556% relative root-mean-square error with a hybrid physics-plus-data reward.

desk verdict The headline accuracy figure is an in-sample fit because the test case gets measured void-fraction and pressure profiles, but the reward-design comparison is a legitimate contribution. read the letter →

arxiv 1908.02750 v2 pith:4JKGMUDC submitted 2019-08-06 physics.comp-ph cs.GTcs.LGphysics.flu-dyn

classification physics.comp-phcs.GTcs.LGphysics.flu-dyn
keywords reinforcementlearningDeepDeterministicPolicyGradienttwo-phaseflowinterfacialareatransportconcentrationMarkovdecisionprocessbubblyphysics-informedreward
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper proposes a physics-informed reinforcement-learning framework (PIRLF) that predicts how bubble surface area per unit volume—the interfacial area concentration—changes along a vertical pipe in bubbly two-phase flow. The flow is modeled as a Markov decision process, and the agent's only lever is a continuous rescaling of the bubble Sauter mean diameter; the interfacial area is then computed from the measured void fraction through $a_i = 6\alpha/D_{\mathrm{sm}}$. What makes the approach physics-informed is a reward function that compares the agent's proposed change with the one-group Interfacial Area Transport Equation, optionally combined with a data-aided reward that compares with experimental measurements at the same axial positions. In leave-one-condition-out tests on seven vertical upward air-water bubbly flows, the hybrid reward gives a relative root-mean-square error of 6.556%, with the best performance at an inform frequency around 0.6. If the claim holds, reinforcement learning steered by a mechanistic model plus sparse measurements can produce accurate, steady interfacial-area profiles without solving the full two-phase flow equations directly.

What carries the argument

The load-bearing object is the Markov decision process built around the identity $a_i = 6\alpha/D_{\mathrm{sm}}$, which converts the agent's action on bubble size into the quantity being predicted. The one-group Interfacial Area Transport Equation supplies the physics-informed part of the reward by predicting the axial change of $a_i$ from bubble coalescence and breakup source terms; the data-aided part supplies the same comparison against experimental measurements at matching axial positions. DDPG is the continuous-control algorithm that learns the policy mapping states to diameter-rescaling actions, and the inform frequency $\epsilon$ decides how often the physics-informed feedback is delivered.

What would settle it

Run PIRLF on a held-out vertical bubbly-flow condition with only the inlet condition supplied—initial bubble size, void fraction, pressure, and gas and liquid flow rates—and compare the full axial interfacial-area profile with measurements. If the relative root-mean-square error rises far above 6.556% or the predicted curve becomes unsteady, the reported performance does not extend to forward prediction and instead describes within-data tuning.

Watch

Extended reading notes

Core claim

The central claim is that interfacial area transport in vertical upward bubbly flow can be treated as a stochastic process and solved as a continuous-control reinforcement-learning problem. The state encodes gas density $\rho_g$, void fraction $\alpha$, interfacial area concentration $a_i$, bubble velocity $v_g$, and Sauter mean diameter $D_{\mathrm{sm}}$; the action is a multiplicative change to $D_{\mathrm{sm}}$; and the reward is the negative absolute deviation between the predicted and reference IAC change. Using the Deep Deterministic Policy Gradient algorithm and the one-group IATE with bubble interaction source terms, the framework achieves an rRMSE of 6.556% when the reward blends the physics-informed and data-aided signals, and the sensitivity study shows an optimal inform frequency $\epsilon$ near 0.6. The paper also shows that reward design changes behavior: physics-only rewards give steady but less accurate curves, data-only rewards hit measured points but can oscillate between them, and the hybrid reward gives both accuracy and steadiness.

Load-bearing premise

The framework assumes the measured void fraction, gas density, and local pressure at every axial position are known inputs, so the reported 6.556% error applies only when those data profiles are supplied, not when predicting from inlet conditions alone.

Editorial extensions

If this is right

  • With the hybrid reward and an inform frequency around 0.6, the framework is claimed to give predictions that are both close to the experimental values and steady between measurement positions, combining the strengths of the physics-only and data-only rewards.
  • Reward type is the main performance driver: the physics-informed reward inherits the steadiness—and the errors—of the mechanistic model, while the data-aided reward tracks measured points but can oscillate in unmeasured regions.
  • The current environment covers only adiabatic air-water bubbly flow in a 50.8 mm vertical tube; extending to other regimes and geometries requires a richer state set and more general IATE closure models, as the paper itself states.
  • Because the action only rescales the Sauter mean diameter and the interfacial area follows from $a_i = 6\alpha/D_{\mathrm{sm}}$, the framework's practical output is a tuned bubble-size profile consistent with the measured void-fraction trend along the tube.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The reported 6.556% relative root-mean-square error is an in-envelope measure: with measured void fraction, gas density, and pressure supplied at every axial position, the agent is fitting the interfacial area curve inside the data envelope rather than predicting it from inlet conditions; a forward-prediction benchmark that supplies only the inlet state has not been run, so its error is unknown.
  • The hybrid reward design is portable: any setting where a mechanistic model gives the right trend but wrong magnitude, and sparse measurements pin the magnitude, could use the same physics-plus-data reward with a re-tuned inform frequency.
  • A stress test on a bubbly-to-slug transition condition, using the same state set and reward, would separate the limits of the one-group IATE closure from the limits of the RL machinery; the paper identifies the IATE model as a ceiling but does not run such a case.
  • The data-aided reward treats experimental measurements as true values; weighting the reward by measurement uncertainty would be a natural modification for noisier industrial two-phase-flow data.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper proposes a physics-informed reinforcement learning framework (PIRLF) for predicting the axial evolution of interfacial area concentration (IAC) in vertical upward bubbly air-water flow. The two-phase flow is cast as a Markov decision process, with state variables including void fraction, gas density, IAC, bubble velocity, and Sauter mean diameter; the action is a multiplicative change to the Sauter mean diameter; and rewards are defined from the interfacial area transport equation (IATE), from experimental data, or as a hybrid of the two. The framework is solved with Deep Deterministic Policy Gradients (DDPG) and evaluated on seven flow conditions from Wu et al., with one condition held out as a test case, reporting an rRMSE of 6.556% for the hybrid reward and an optimal inform frequency of about 0.6. The paper also discusses possible extensions of the framework.

Significance. If the claimed predictive accuracy were established under a genuinely forward-looking protocol, this work would be a useful step toward data-informed closures for the two-fluid model and would demonstrate a concrete way to combine mechanistic models with reinforcement learning. The MDP formulation, the reward design space, and the sensitivity analysis of the inform frequency are conceptually interesting and could stimulate further work in the multiphase-flow machine-learning community. However, the current evaluation does not support the headline predictive claim: the test case receives measured void fraction and pressure at every axial position, so the reported accuracy is essentially a within-data fit rather than a forward prediction. The paper would also be strengthened by providing code or data to make the training procedure reproducible.

major comments (4)
  1. [Sec. 2.1.2 and Sec. 3] The headline rRMSE of 6.556% is computed under an evaluation protocol that does not support the stated forward-prediction claim. In Sec. 2.1.2, the void fraction and gas density are "calculated based on the local pressure provided based on the experimental data" at each axial position, and Eq. (1) gives ai = 6α/Dsm. Because α is taken from the same experimental database that provides the target ai, the reported error primarily measures whether the DDPG policy can choose a Dsm sequence that reproduces the experimental Dsm values within the data envelope; it does not measure the ability to predict the IAC profile for an unseen flow condition where the local α and pressure are unknown. The abstract and conclusion therefore overstate what is demonstrated. The evaluation should be rerun in a genuinely predictive mode, for example by supplying only inlet conditions and solving for α and pressure with a hydrodynamic model, or the claims should be explicitly limited to within-data interpolation.
  2. [Sec. 3, Eq. (12)] The single reported rRMSE comes from one randomly selected test flow condition out of seven, with no reporting of the other six possible folds or the variance across random seeds. Since Sec. 3 states that "one of the flow conditions is randomly selected as the test case," a single draw cannot establish the robust performance implied by the phrase "good performance." Please report all seven leave-one-out folds, the mean and standard deviation of rRMSE, and the sensitivity of the results to random seeds in the DDPG training.
  3. [Sec. 4.1, Figs. 5-6] The claim that PIRLF gives "satisfying estimations" is not supported by any baseline comparison. The manuscript does not report the rRMSE of the IATE-only prediction obtained by integrating Eq. (11) with the same measured conditions, nor a simple interpolator, nor a policy with constant Dsm tied to the initial condition. Without such baselines, the reader cannot judge whether the RL agent adds predictive value over the existing mechanistic model or over the curve that simply follows the measured α profile. Please include these comparisons in the revised analysis.
  4. [Sec. 2.1.4, Eq. (4) and Sec. 2.2] The "physics-informed" reward is not an independent physical constraint in the current setting. The IATE source terms in Eqs. (8)-(10) involve coefficients CRC, CWE, and CTI that were estimated from experimental data, and the transport equation is evaluated with the same measured α and pressure profiles used to define the target. Consequently, comparing Δai,pred with Δai,IATE in Eq. (4) measures consistency with a data-fitted model rather than physical correctness. The authors should at least quantify the sensitivity of the final predictions to the fitted source-term coefficients and should report how much the physics-informed component changes the result relative to a policy trained only on the data-aided reward.
minor comments (5)
  1. [Sec. 4.2, Fig. 7] The caption of Fig. 7 labels the curve as the "hybrid reward function," but the text in Sec. 4.2 states that the rRMSEs were generated with the physics-informed reward function (Type-1); please reconcile the caption and the text.
  2. [Eq. (3)] The Bayesian form of the transition probability is written in a compressed way; the terms p(a|s',s) and p(s'|s) should be defined, and it should be clarified whether p(a|s) is intended as the marginal over s'.
  3. [Sec. 3, Eq. (12)] The phrase "rooted-mean-square error" should be "root-mean-square error" to match standard terminology.
  4. [Sec. 2.1.1 and 2.1.2] The state set is described inconsistently as including "air density ρg" in Sec. 2.1.1 and "gas density" in Sec. 2.1.2; please standardize the terminology and symbols.
  5. [Sec. 2.1.3] The text introduces two possible transition probability functions, but the experiments do not state which one was actually implemented; please specify which transition model was used in the reported results.

Circularity Check

2 steps flagged · score 7.0 of 10

The headline 6.556% rRMSE is achieved by feeding the test case the measured axial void fraction and pressure and rewarding the agent against the measured IAC; the prediction therefore interpolates the experimental IAC envelope rather than demonstrating forward predictive ability.

  1. self definitional [Section 2.1.2, Eq. (1); Section 3 evaluation protocol; Eq. (12)]
    "ai = 6α Dsm (1) ... the void fraction and gas density are calculated based on the local pressure provided based on the experimental data."

    At each evaluation position the void fraction α is taken from the same experimental run that supplies the target ai,exp. Since Eq. (1) forces ai,pred = 6α_exp / Dsm, and the action only changes Dsm by a multiplicative ratio, the agent's output is the experimental α profile divided by a fitted Dsm sequence. The rRMSE of Eq. (12) therefore measures how closely the RL policy reproduces the experimental Dsm values inside the measured axial envelope; it is not a test of predicting IAC for a case where α and P are unknown.

  2. fitted input called prediction [Section 2.1.4, Eq. (5); Section 4.2, Eq. (12)]
    "The measurement values can be compared with the predicting values at the same axial positions, r = {−| ai,pred−ai,exp| zpred =zexp 0 otherwise (5)"

    The data-aided (and therefore hybrid) reward is literally the absolute difference between the predicted and experimental IAC at the same axial locations, and the performance metric is the rRMSE of exactly those same quantities. Training the policy to maximize this reward is equivalent to fitting Dsm so that 6α_exp/Dsm approximates ai,exp. The reported 6.556% is thus an in-sample curve-fit accuracy against the test condition's measured data, not an out-of-sample forward prediction; the held-out condition's own α and P profile is supplied to the state during evaluation.

full rationale

The central numerical claim of the paper is that PIRLF 'shows a good performance with rRMSE of 6.556%.' The paper's own environment definition makes this number a within-data fit. Section 2.1.2 states that void fraction and gas density are computed from local pressures provided by the experimental data, and the predicted IAC is defined by Eq. (1) as ai = 6α/Dsm. Thus the measured void fraction profile of the test case enters directly into the predicted quantity. The only degree of freedom controlled by the RL agent is Dsm, and Eq. (5) rewards the agent for making 6α_exp/Dsm coincide with ai,exp at every measurement station. Consequently, the rRMSE of Eq. (12) compares the agent's fitted Dsm trajectory with the experimental Dsm implied by ai,exp and α_exp; it does not validate prediction of IAC from inlet conditions alone. The physics-informed reward is not an independent guard: it compares the agent's Δai with the IATE of Wu et al. [3], whose source-term coefficients were estimated from experimental data and whose test conditions are drawn from the same database. These observations do not make the framework useless, but they do mean the abstract's headline accuracy is a data-envelope interpolation and the paper overstates its predictive reach.

Assumptions & free parameters 3 free parameters · 5 assumptions · 0 invented entities

The central claim rests on two pillars: the MDP formulation (with the Markov property and a deterministic transition) and the IATE as physics reference (with coefficients fitted to the same type of data the framework is tested on). No new physical entities are introduced.

free parameters (3)
  • Inform frequency epsilon = approximately 0.6
    Eq. 4; the sensitivity study in Sec. 4.2 selects the epsilon that minimizes rRMSE on the same dataset, so it is a fitted tuning parameter rather than a derived constant.
  • IATE source-term coefficients CRC, CWE, CTI = not given in this paper; estimated from experimental data in Wu et al. [3]
    Eqs. 8-10; these coefficients are fit to bubbly-flow data and are adopted as the physics reference, so the 'physics-informed' reward is itself data-calibrated.
  • DDPG hyperparameters (learning rates 1e-4/1e-3, gamma=0.99, OU noise sigma=0.5, 3 hidden layers of 50 units) = as listed in Appendix A
    Chosen by hand following Lillicrap et al. [2]; no sensitivity analysis is reported in this paper, and performance depends on them.
assumptions (5)
  • domain assumption Two-phase flow development satisfies the Markov property
    Stated in Sec. 1 and 2.1 to justify the MDP formulation; never tested. The implemented transition (Eq. 2) is deterministic, so the stochastic premise is unused.
  • domain assumption Wu et al. [3] IATE source-term closures apply to the present database
    Sec. 2.2 adopts Eqs. 8-10 with coefficients fitted to a similar bubbly-flow database; the paper gives no independent verification of this transfer.
  • domain assumption Cross-sectionally averaged covariance terms are negligible
    Sec. 2.2, following Hibiki and Ishii [15]; adopted without re-derivation.
  • domain assumption Experimental local pressure and void fraction are available at every axial position during prediction
    Sec. 2.1.2 states void fraction and gas density are calculated from local pressure provided from experimental data; this is the key data-conditioning assumption that makes the reported rRMSE possible.
  • domain assumption The environment update rules for void fraction and bubble velocity from local pressure are sufficient
    Sec. 2.1.2; the explicit update equations are not given, so the reader must assume the environment is consistent with drift-flux or continuity relations.

how reviews work

0 comments
Cite this review

Pith. "Pith review of A physics-informed reinforcement learning approach for the interfacial area transport in two-phase flow." pith.science (2026). https://pith.science/paper/4JKGMUDC

@misc{pith2026190802750,
  author       = {Pith},
  title        = {Pith review of: A physics-informed reinforcement learning approach for the interfacial area transport in two-phase flow},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/4JKGMUDC}},
  note         = {Machine review of arXiv:1908.02750}
}
read the original abstract

The prediction of interfacial structure in two-phase flow systems is difficult and challenging. In this paper, a novel physics-informed reinforcement learning-aided framework (PIRLF) for the interfacial area transport is proposed. A Markov Decision Process that describes the bubble transport is established by assuming that the development of two-phase flow is a stochastic process with Markov property. The framework aims to capture the complexity of two-phase flow using the advantage of reinforcement learning (RL) in discovering complex patterns with the help of the physical model (Interfacial Area Transport Equation) as reference. The details of the framework design are described including the design of the environment and the algorithm used in solving the RL problem. The performance of the PIRLF is tested through experiments using the experimental database for vertical upward bubbly air-water flows. The result shows a good performance of PIRLF with rRMSE of 6.556%. The case studies on the PIRLF performance also show that the type of reward function that is related to the physical model can affect the framework performance. Based on the study, the optimal reward function is established. The approaches to extending the capability of PIRLF are discussed, which can be a reference for the further development of this methodology.

Figures

Figures reproduced from arXiv: 1908.02750 by the authors.

Figure 1
Figure 1. Formulation of two-phase flow problem into a Markov Decision Process. [PITH_FULL_IMAGE:figures/full_fig_p006_1.png] view at source ↗
Figure 2
Figure 2. The Physics-informed Reinforcement Learning. [PITH_FULL_IMAGE:figures/full_fig_p007_2.png] view at source ↗
Figure 3
Figure 3. The MDP environment that describes the interfacial area change for vertical [PITH_FULL_IMAGE:figures/full_fig_p012_3.png] view at source ↗
Figures from the paper (4 more)
Figure 4
Figure 4. Figure 4: The implementation flow of the framework. [PITH_FULL_IMAGE:figures/full_fig_p014_4.png]
Figure 5
Figure 5. Figure 5: Prediction results against experimental data by PIRLF with left): physics [PITH_FULL_IMAGE:figures/full_fig_p016_5.png]
Figure 6
Figure 6. Figure 6: Prediction results against experimental data by PIRLF hybrid reward function. [PITH_FULL_IMAGE:figures/full_fig_p016_6.png]
Figure 7
Figure 7. Figure 7: Prediction results against experimental data by PIRLF hybrid reward function. [PITH_FULL_IMAGE:figures/full_fig_p017_7.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

26 extracted references · 22 canonical work pages

  1. [1]

    Kocamustafaogullari, M

    G. Kocamustafaogullari, M. Ishii, Foundation of the interfacial area transport equation and its closure relations, International Journal of Heat and Mass Transfer 38 (3) (1995) 481–493

  2. [2]

    T. P. Lillicrap, J. J. Hunt, A. Pritzel, N. Heess, T. Erez, Y. Tassa, D. Sil- ver, D. Wierstra, Continuous control with deep reinforcement learning, arXiv preprint arXiv:1509.02971 (2015)

  3. [3]

    Q. Wu, S. Kim, M. Ishii, S. Beus, One-group interfacial area transport in vertical bubbly flow, International Journal of Heat and Mass Transfer 41 (8-9) (1998) 1103–1112

  4. [4]

    TRACE, Theory manual

    V. TRACE, Theory manual. field equations, solution methods and phys- ical models. u. s (2007)

  5. [5]

    Ishii, K

    M. Ishii, K. Mishima, Two-fluid model and hydrodynamic constitutive relations, Nuclear Engineering and design 82 (2-3) (1984) 107–126

  6. [6]

    M. S. Bernard, Implementation of the interfacial area transport equation in trace for boiling two-phase flows, Ph.D. thesis, Pennsylvania State University (2014)

  7. [7]

    Durbin, A stochastic model of two-particle dispersion and concentra- tion fluctuations in homogeneous turbulence, Journal of Fluid Mechanics 100 (2) (1980) 279–302

    P. Durbin, A stochastic model of two-particle dispersion and concentra- tion fluctuations in homogeneous turbulence, Journal of Fluid Mechanics 100 (2) (1980) 279–302. 20

  8. [8]

    Novikov, Two-particle description of turbulence, markov property, and intermittency, Physics of Fluids A: Fluid Dynamics 1 (2) (1989) 326–330

    E. Novikov, Two-particle description of turbulence, markov property, and intermittency, Physics of Fluids A: Fluid Dynamics 1 (2) (1989) 326–330

Show all 26 references
  1. [9]

    Pedrizzetti, E

    G. Pedrizzetti, E. A. Novikov, On markov modelling of turbulence, Jour- nal of Fluid Mechanics 280 (1) (1994) 69–93

  2. [10]

    R. S. Sutton, A. G. Barto, et al., Introduction to reinforcement learning, Vol. 2, MIT press Cambridge, 1998

  3. [11]

    Bellman, A markovian decision process, Journal of mathematics and mechanics (1957) 679–684

    R. Bellman, A markovian decision process, Journal of mathematics and mechanics (1957) 679–684

  4. [12]

    Z. Dang, G. Wang, P. Ju, X. Yang, R. Bean, M. Ishii, S. Bajorek, M. Bernard, Experimental study of interfacial characteristics of vertical upward air-water two-phase flow in 25.4 mm id round pipe, International Journal of Heat and Mass Transfer 108 (2017) 1825–1838

  5. [13]

    Z. Dang, M. Ishii, Two-phase interfacial structure of bubbly-to- slug transition flows in a 12.7 mm id vertical tube, arXiv preprint arXiv:2009.01779 (2020)

  6. [14]

    Hibiki, M

    T. Hibiki, M. Ishii, Development of one-group interfacial area transport equation in bubbly flow systems, International Journal of Heat and Mass Transfer 45 (11) (2002) 2351–2372

  7. [15]

    Hibiki, M

    T. Hibiki, M. Ishii, Interfacial area transport equations for gas-liquid flow, The Journal of Computational Multiphase Flows 1 (1) (2009) 1– 22

  8. [16]

    Silver, G

    D. Silver, G. Lever, N. Heess, T. Degris, D. Wierstra, M. Riedmiller, Deterministic policy gradient algorithms, 2014

  9. [17]

    Bhatnagar, M

    S. Bhatnagar, M. Ghavamzadeh, M. Lee, R. S. Sutton, Incremental natural actor-critic algorithms, in: Advances in neural information pro- cessing systems, 2008, pp. 105–112

  10. [18]

    V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, et al., Human-level control through deep reinforcement learning, Nature 518 (7540) (2015) 529–533. 21

  11. [19]

    Hibiki, M

    T. Hibiki, M. Ishii, Z. Xiao, Axial interfacial area transport of vertical bubbly flows, International Journal of Heat and Mass Transfer 44 (10) (2001) 1869–1888

  12. [20]

    B. Ozar, C. Brooks, T. Hibiki, M. Ishii, Interfacial area transport of ver- tical upward steam–water two-phase flow in an annular channel at ele- vated pressures, International Journal of Heat and Mass Transfer 57 (2) (2013) 504–518

  13. [21]

    J. Du, Z. Dang, Y. Zhao, C. Zhao, H. Bo, M. Ishii, Experimental study of vibration effects on local interfacial parameters in boiling flow, Inter- national Journal of Heat and Mass Transfer 151 (2020) 119369

  14. [22]

    X. Fu, M. Ishii, Two-group interfacial area transport in vertical air– water flow: I. mechanistic model, Nuclear Engineering and Design 219 (2) (2003) 143–168

  15. [23]

    Sun, Two-group interfacial area transport equation for a confined test section (2001)

    X. Sun, Two-group interfacial area transport equation for a confined test section (2001)

  16. [24]

    T. S. Worosz, Interfacial area transport equation for bubbly to cap- bubbly transition flows (2015)

  17. [25]

    Haarnoja, A

    T. Haarnoja, A. Zhou, P. Abbeel, S. Levine, Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor, arXiv preprint arXiv:1801.01290 (2018)

  18. [26]

    D. P. Kingma, J. Ba, Adam: A method for stochastic optimization, arXiv preprint arXiv:1412.6980 (2014). 22

Pith tools

Reviewed August 14, 2026 · model on record in the stance chip above.