REVIEW 2 major objections 6 minor 43 references
A General Model Validation and Testing Tool
T0 review · 2 major / 6 minor · reviewed 2026-08-14 · deepseek-v4-flash
Pith's one-line read A single Bayesian probability unifies model validation
desk verdict A useful unifying framework for validation metrics, with a fixable gap in the Bayes-factor equivalence and somewhat overbroad universality claims. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The agreement kernel $\Theta(B(\hat{z},z))$ — the indicator that the user's Boolean agreement function is true — and the marginalization integral that defines the BVM carry the argument. The four BVM inputs (comparison values $\hat{z},z$; model and data probability densities $\rho(\hat{z}|M,D)$ and $\rho(z|D)$; comparison value function $f(\hat{z},z)$; and agreement function $B(f)$) specify any validation scenario, and the BVM Ratio $R(B)$ extends the Bayes factor to arbitrary agreement rules. This machinery allows the paper to represent each standard metric by identifying its implied comparison values and agreement function, and to construct new compound agreement functions for multidimensional or multi-criteria validation.
What would settle it
Find any well-defined validation metric from the literature that compares two probability distributions with a value not expressible as the expectation of a Boolean function of a comparison value, and show it cannot be written as $\int \rho(\hat{z}|M,D) \, \Theta(B(\hat{z},z)) \, \rho(z|D) \, d\hat{z} dz$ for any choice of the four BVM inputs; alternatively, compute the BVM and a classical metric on the same model-data pair with a deliberately misspecified model output distribution and show the two metrics rank models differently even under identical definitions of agreement.
Extended reading notes
Core claim
The central discovery is that validation can be reduced to a marginal probability: agreement between a model and data is the probability that a user-defined Boolean function $B(f(\hat{z},z))$ is true, averaged over the joint uncertainty in the model output $\hat{z}$ and the data $z$. When the data distribution is independent of the model, this takes the form $p(A|M,D) = \int \rho(\hat{z}|M,D) \, \Theta(B(\hat{z},z)) \, \rho(z|D) \, d\hat{z} dz$, and equivalently as $\int \rho(f|M,D) \, \Theta(B(f)) \, df$ after propagating uncertainty through the comparison function. The paper shows that reliability, probability of agreement, the frequentist metric, the area metric, pdf comparison metrics, statistical hypothesis testing, and Bayesian model testing all emerge as special cases for particular choices of comparison values, probability distributions, comparison functions, and agreement functions. It further constructs the BVM Ratio, $R(B) = p(A|M,D,B)/p(A|M',D,B)$ times a prior ratio, which generalizes the Bayes factor to model selection under arbitrary definitions of agreement.
Load-bearing premise
The load-bearing premise is that the user can supply an accurate probability distribution for the model output, obtained by forward propagation of all parameter and input uncertainties; if that distribution is unavailable or misspecified, the BVM's probability of agreement is only as trustworthy as that input, and none of the claims of generality replace it.
Editorial extensions
If this is right
- All standard validation metrics can be reported as probabilities between 0 and 1, so their uncertainties become comparable in the same quantitative units.
- Metrics that look different coincide under explicit conditions: the frequentist metric equals the reliability metric under a natural tolerance, and Bayesian model testing equals the improved reliability metric under exact agreement.
- Model selection can be performed under any user-defined agreement rule via the BVM Ratio, not only under exact data likelihoods.
- Compound Boolean agreement functions allow multidimensional or multi-criteria validation to be expressed as a single probability.
- A statistical-power BVM can avoid type I and type II errors when both model and data probability densities are available, and its resolving power improves with confidence sets rather than confidence intervals.
Reading between the lines
- This unification implies that choosing a validation metric is effectively choosing an agreement function; disagreements between analysts over validity can be understood as disagreements over $B$, and reporting $B$ alongside the probability would make them explicit.
- The approach's practical reach is bounded by the quality of the model output probability distribution; if uncertainty propagation is miscalibrated, the BVM returns a precise-looking probability that inherits that error, so the framework directs attention back to uncertainty quantification and calibration.
- The $(\gamma,\epsilon)$ Boolean example suggests testable applications to high-dimensional or non-visualizable model-data comparisons, where design of agreement rules substitutes for visual inspection.
- A natural extension the paper leaves open is learning the four inputs — especially the agreement function or its tolerance parameters — from data, effectively calibrating the validator itself.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a 'Bayesian Validation Metric' (BVM), defined as the marginal probability that a user-specified Boolean agreement function B(f(ẑ,z)) holds, where ẑ and z are model and data comparison values with joint distribution ρ(ẑ,z|M,D). The authors claim that the BVM reproduces all standard validation metrics (reliability, probability of agreement, frequentist, area, pdf comparison, statistical hypothesis testing, and Bayesian model testing) as special cases, that it satisfies the six desirable validation criteria of Liu et al., and that the BVM Ratio generalizes Bayesian model testing to arbitrary definitions of agreement. Three examples are given: a statistical power BVM, a compound (⟨ε⟩,β_D) BVM, and a (γ,ε) BVM used for model selection.
Significance. The BVM is a genuinely useful unifying perspective: the core marginalization in Eq. (2) is simple, transparent, and correctly captures the idea that validation is a probabilistic statement about a user-defined agreement concept. The paper organizes a large literature into a common framework (Tables 1 and 2), explicitly tests the framework on nonstandard compound Booleans, and emphasizes the statistical responsibility of stating the agreement definition. These are real contributions. However, the strength of the paper's central claim is undermined by two issues: the derivation equating the BVM to Bayesian model testing silently drops the data distribution (Eq. (55)), and the exact-agreement BVM is a density rather than a probability, contradicting the foundational definition. The equivalence to Bayesian model testing is therefore not established as stated, and the 'generalizes Bayesian model testing' claim requires the data to be treated as a point mass or another explicit redefinition.
major comments (2)
- [Appendix A.6, Eq. (55)] The derivation of the equivalence to Bayesian model testing is not valid as written. Starting from p(A|M,D) = ∫ p(Ŷ=Y|x,α,M,D) ρ(x,α) ρ(Y|D) dx dα dY, the second line removes the ρ(Y|D) factor and the Y integral, yielding ∫ p(Ŷ≡Y|x,α,M,D) ρ(x,α) dx dα. This step is valid only when ρ(Y|D) = δ(Y−Y_obs), i.e., the validation data are known with complete certainty. Without that point-mass condition, the first line is an expectation of the model likelihood over the data distribution, not the likelihood evaluated at the observed data. The manuscript does not state that a delta function is being imposed in A.6; it treats the elimination of ρ(Y|D) as a routine integration. This is load-bearing because the claim that the BVM Ratio generalizes Bayesian model testing depends on this equality. The text should either explicitly impose the certain-data case or qualify the generalization claim to the uncertain-data setting.
- [Section 3.2, Eqs. (10)–(12); Section 4.1, Table 1] Under exact agreement with continuous comparison values, the BVM is not a probability in [0,1] but a density proportional to dẑ: Eq. (12) gives p(A|M,D) = ρ(ẑ≡z|M,D)dẑ. The authors note this and argue that the Bayes factor avoids the issue because the measures drop out, but Table 1 and Section 4.1 still call p(A|M,D) = p(Ŷ≡Y|M,D) 'the Bayesian evidence' and treat it as a probability. This contradicts the definition in Section 2 that the BVM is a probability in [0,1]. Since the BVM ratio (13)–(14) is a ratio of these objects, the interpretation of R(B) as a ratio of probabilities is also affected when B demands exact equality. The manuscript should either restrict the 'probability' claim and explicitly treat exact agreement as a limiting density, or define the BVM to include an implicit discretization/measure convention from the start.
minor comments (6)
- [Section 3.3] The claim that the BVM meets the fifth desirable criterion of Liu et al. (artificially widening distributions should not increase validation rates) is supported only by assuming the user is not engaging in misconduct, and Section 4.1 later shows that the reliability-metric representation admits widening. The conclusion 'the BVM was shown to obey all of the desired validation metric criteria' should be softened or the discussion clarified.
- [Section 4.1, Table 1 rows 'Area' and 'Pdf Comp.'] For the area metric and pdf comparison metrics, the BVM representation uses point-mass distributions and reduces to a deterministic threshold around the existing metric value. This is a valid 'special case', but it is a degenerate one that adds little beyond wrapping the metric in an indicator function; the language 'represents' should be tempered or the generalization to uncertain cdfs/pdfs should be foregrounded.
- [Section 5.1, Eq. (46)] The statement that the statistical power BVM 'removes the possibility of both type I and type II errors' is imprecise: the test avoids the null-hypothesis testing framework rather than eliminating error probabilities within it. The exposition should distinguish between avoiding the framework and removing errors in the classical sense.
- [Section 5.2, Eq. (19)] The Monte Carlo estimate uses K = 3000 samples for an indicator integrand, but no Monte Carlo standard error is reported. A simple binomial confidence interval would make the reported values like P(A|⟨ε⟩) = 0.99 more interpretable.
- [Section 5.3, Eq. (25)] The 'averaged Boolean BVM ratio' marginalizes over a uniform p(γ,ϵ) on the tested volume; the result depends on the arbitrarily selected (γ,ϵ) range. This dependence should be acknowledged in the text as part of the definition of agreement, since different volumes can give different R(B).
- [Section 3.2, Eq. (8)] The notation δ_{ẑ,z} for a Kronecker delta with continuous labels is nonstandard and could confuse readers; a brief explanation or alternative notation (e.g., an indicator of equality) would help.
Circularity Check
No significant circularity: the BVM is a definitional framework with explicit input specifications rather than fitted predictions, and the main caveat (the Appendix A.6 Bayes-factor equivalence) is a derivation gap, not a circular reduction.
full rationale
The BVM is introduced as a definition in Eqs. (2)-(5): p(A|M,D) is the user-specified integral of an agreement kernel over stated model and data distributions. Table 1 and Table 2 then represent standard validation metrics by substituting explicit comparison values, probability densities, comparison functions, and agreement functions; these substitutions are identities by the paper's own definitions, and none of the examples fits a parameter to a subset of data and then reports a closely related quantity as a prediction. The model/data distributions and tolerance parameters in Section 5 are hand-set for toy illustrations, not inferred from the quantities being validated. There are no load-bearing self-citations or imported uniqueness theorems that force the BVM's form. The one substantive concern in the paper is not circularity: in Appendix A.6, Eq. (55) moves from an integral containing ρ(Y|D) to ρ(Ŷ≡Y|M,D)dŶ without explicitly imposing ρ(Y|D)=δ(Y−Y_obs); the Bayes-factor identity is therefore valid only in the complete-certainty special case, and the 'generalizes Bayesian model testing' claim needs that point-mass assumption stated. This is a mathematical/expository gap that should be corrected, but it does not make the BVM's core derivation equivalent to its own inputs, so the circularity score remains 0.
Assumptions & free parameters
free parameters (5)
- epsilon (agreement threshold) =
0.46 in Section 5.2; varied 0 to 1 with step 0.01 in Section 5.3
- beta_D tolerance band 95% +/- 4% =
0.91 to 0.99
- gamma and m in (gamma, epsilon) Boolean =
gamma 75%-100%; m=5
- K = 3000 MC samples =
3000
- Parameter standard deviations in (gamma, epsilon) example =
(0.1, 0.05, 0.005, 0.0005)
assumptions (4)
- domain assumption Model output pdf and data pdf are available and can be quantified through uncertainty propagation.
- domain assumption The model and data uncertainties are conditionally independent, rho(z|M,z_hat,D) = rho(z|D).
- standard math All integrals are finite and well-defined, and the agreement kernel is a Boolean indicator.
- ad hoc to paper The user can specify a meaningful comparison function and agreement function for the context.
invented entities (3)
-
Bayesian Validation Metric (BVM) as a named framework
-
BVM Ratio
-
Statistical power BVM
Cite this review
Pith. "Pith review of A General Model Validation and Testing Tool." pith.science (2026). https://pith.science/paper/HDGRF6I6
@misc{pith2026190811251,
author = {Pith},
title = {Pith review of: A General Model Validation and Testing Tool},
year = {2026},
howpublished = {\url{https://pith.science/paper/HDGRF6I6}},
note = {Machine review of arXiv:1908.11251}
}
read the original abstract
We construct and propose the "Bayesian Validation Metric" (BVM) as a general model validation and testing tool. We find the BVM to be capable of representing all of the standard validation metrics (square error, reliability, probability of agreement, frequentist, area, probability density comparison, statistical hypothesis testing, and Bayesian model testing) as special cases and find that it can be used to improve, generalize, or further quantify their uncertainties. Thus, the BVM allows us to assess the similarities and differences between existing validation metrics in a new light. The BVM has the capacity to allow users to invent and select models according to novel validation requirements. We formulate and test a few novel compound validation metrics that improve upon other validation metrics in the literature. Further, we construct the BVM Ratio for the purpose of quantifying model selection under user defined definitions of agreement in the presence or absence of uncertainty. This construction generalizes the Bayesian model testing framework.
Figures
Figures from the paper (3 more)
Reference graph
Works this paper leans on
-
[1]
W. L. Oberkampf, T. G. Trucano, and C. Hirsch, 2004, Verification, Validation, and Predictive Capa- bility in Computational Engineering and Physics, Appl. Mech. Rev., 57(3), pp. 345384
work page 2004
-
[2]
D. Sornette, A. B. Davis, K. Ide, K. R. Vixie, V. Pisarenko, and J. R. Kamm, 2007, Algorithm for Model Validation: Theory and Applications, Proc. Natl. Acad. Sci. U.S.A., 104(16), pp. 65626567
work page 2007
-
[3]
Calibration, validation, and sensitivity analysis: Whats what
T.G. Trucano, L.P. Swiler, T. Igusa, W. L. Oberkampf, and M. Pilch, 2006, “Calibration, validation, and sensitivity analysis: Whats what”, Reliab. Eng. Syst. Saf. 91, pp. 1331-1357
work page 2006
-
[4]
M. C. Kennedy and A. O’Hagan, Bayesian calibration of computer models, J. R. Stat. Soc.: Ser. B (Stat Methodol.), vol. 63, no. 3, pp. 425-464, 2001
work page 2001
-
[5]
O. P. Maˆ ıtre, O. M. Knio, Spectral Methods for Uncertainty Quantification. New York: Springer Sci- ence+Business Media, 2010
work page 2010
-
[6]
Updated November 2015 (Version 6.3)
Adams, B.M., Bauman, L.E., Bohnhoff, W.J., Dalbey, K.R., Ebeida, M.S., Eddy, J.P., Eldred, M.S., Hough, P.D., Hu, K.T., Jakeman, J.D., Stephens, J.A., Swiler, L.P., Vigil, D.M., and Wildey, T.M., ”Dakota, A Multilevel Parallel Object-Oriented Framework for Design Optimization, Parameter Estima- tion, Uncertainty Quantification, and Sensitivity Analysis: Ver...
work page 2014
-
[7]
The Uncertainty Quantification Toolkit (UQTk)
R. Ghanem, D. Higdon, and H.Owhadi, “The Uncertainty Quantification Toolkit (UQTk)” Handbook of Uncertainty Quantification, Springer, 2016
work page 2016
-
[8]
B. Debusschere, N. Habib, N. Najm, P. P´ ebay, O. M. Knio, R. Ghanem, and O. P. Maˆ ıtre, Numerical Challenges in the Use of Polynomial Chaos Representations for Stochastic Processes, SIAM J. Sci. Comput. vol. 26, no. 2, pp. 698-719, 2005
work page 2005
Show all 43 references
-
[9]
Gilks, S
W. Gilks, S. Richardson, and D. J. Spiegelhalter, Markov chain Monte Carlo in practice. London, UK: Chapman and Hall, 1996
1996
-
[10]
Metropolis, A
N. Metropolis, A. W. Rosenbluth, M. N. Rosenbluth, A. H. Teller, and E. Teller, Equation of state calculations by fast computing machines, J. Chem. Phys., vol. 21, no. 6, pp. 1087, 1953
1953
-
[11]
Sankararaman and S
S. Sankararaman and S. Mahadevan, Integration of model verification, validation, and calibration for uncertainty quantification in engineering systems, Reliab. Eng. Syst. Saf., vol. 138, pp. 194-209, 2015
2015
-
[12]
Sankararaman and S
S. Sankararaman and S. Mahadevan, Model validation under epistemic uncertainty, Reliab. Eng. Syst. Saf. vol. 96, no. 9, pp. 1232-1241, 2011
2011
-
[13]
Mahadevan and R
S. Mahadevan and R. Rebba, Validation of reliability computational models using Bayes networks, Reliab. Eng. Syst. Saf. vol. 87, no. 2, pp. 223-232, 2005
2005
-
[14]
Li and S
C. Li and S. Mahadevan, Role of calibration, validation, and relevance in multi-level uncertainty inte- gration, Reliab. Eng. Syst. Saf., vol. 148, pp. 32-43, 2016
2016
-
[15]
Ferson, W
S. Ferson, W. L. Oberkampf, and L. Ginzburg, Model validation and predictive capability for the thermal challenge problem, Comput. Methods Appl. Mech. Eng., vol. 197, no. 29, pp. 2408-2430, 2008
2008
-
[16]
Roy and W
C. Roy and W. Oberkampf, A comprehensive framework for verification, validation, and uncertainty quantification in scientific computing, Comput. Methods Appl. Mech. Engrg., vol. 200, pp. 2131-2144, 2011
2011
-
[17]
Ling and S
Y. Ling and S. Mahadevan, Quantitative model validation techniques: New insights, Reliab. Eng. Syst. Saf., vol. 111, pp. 217, 2013
2013
-
[18]
W. Li, W. Chen, Z. Jiang, Z. Lu, and Yu Liu, New validation metrics for models with multiple correlated responses, Reliab. Eng. Syst. Saf., vol. 127, pp. 1-11, 2014
2014
-
[19]
D. Wu, Z. Lun, Y. Wang, and L. Cheng, Model validation and calibration based on component functions of model output, Reliab. Eng. Syst. Saf., vol. 140, pp. 59-70, 2015
2015
-
[20]
L. Zhao, Z. Lu, W. Yun, and W. Wang, Validation metric based on Mahalanobis distance for models with multiple correlated responses, Reliab. Eng. Syst. Saf., vol. 159, pp. 80-89, 2017
2017
-
[21]
Stefano and B
M. Stefano and B. Sudret, UQLab user manual - Polynomial chaos expansions. Report UQLab-V0.9-104, Chair of Risk, Safety and Uncertainty Quantification, ETH Zurich, 2015
2015
-
[22]
MUQ: MIT Uncertainty Quantification Library
M. Parno and A. Davis, “MUQ: MIT Uncertainty Quantification Library”, http://muq.mit.edu/home, 2018
2018
-
[23]
Wang and H
C. Wang and H. G. Matthies, Novel model calibration method via non-probabilistic interval character- ization and Bayesian theory, Reliab. Eng. Syst. Saf., vol 183, pp 84-92, 2019
2019
-
[24]
Rebba and S
R. Rebba and S. Mahadevan, Computational methods for model reliability assessment, Reliab. Eng. Syst. Saf., vol 93, no. 8, pp 1197-1207, 2008
2008
-
[25]
N. T. Stevens, Assessment and comparison of Continuous Measurement Systems, Ph. D. Thesis, Dept. Statistics, Univ. of Waterloo, Waterloo, Ontario, Canada, 2014. 20
2014
-
[26]
Sankararaman and S
S. Sankararaman and S. Mahadevan, Assessing the reliability of computational models under uncer- tainty. In: The 54th AIAA/ASME/ASCE/AHS/ASC structures, structural dynamics, and materials conference; 2013
2013
-
[27]
W. L. Oberkampf, M. F. Barone, Measures of agreement between computation and experiment: vali- dation metrics, J. Comput. Phys. vol. 217, no. 1, pp. 5-36, 2006
2006
-
[28]
Oberkampf, M.F
W.L. Oberkampf, M.F. Barone, Measures of agreement between computation and experiment: validation metrics, AIAA Paper 2004 2626
2004
-
[29]
Data Analysis A Bayesian Tutorial second edition
D. Sivia and J. Skilling, “Data Analysis A Bayesian Tutorial second edition”, Oxford University Press, Oxford, UK, 2006
2006
-
[30]
Placek, Bayesian Detection and Characterization of Extra-Solar Planets Via Photometric Variations, Ph
B. Placek, Bayesian Detection and Characterization of Extra-Solar Planets Via Photometric Variations, Ph. D. Thesis, Dept. Physics, Univ. at Alb. (S.U.N.Y.), Albany, NY, USA, 2014
2014
-
[31]
A. E. Gelfand and D. K. Dey, Bayesian model choice: asymptotics and exact calculations, J. R. Stat. Soc. Ser. B (Methodol.), pp. 501-514, 1994
1994
-
[32]
Geweke, Bayesian model comparison and validation, Am
J. Geweke, Bayesian model comparison and validation, Am. Econ. Rev. 2007;97(2):604
2007
-
[33]
Zhang, S
R. Zhang, S. Mahadevan S, Bayesian methodology for reliability model acceptance, Reliab. Eng. Syst. Saf., vol. 80, no. 1, pp. 95-103, 2003
2003
-
[34]
Y. Liu, W. Chen, P. Arendt, and H. Z. Huang, Toward a better understanding of model validation metrics, Trans ASME J. Mech. Des. vol. 133, no. 7, pp. 071005, 2011
2011
-
[35]
K. A. Maupin, L. P. Swiler, and N. W. Porter, Validation metrics for deterministic and probabilistic data. Journal of Verification, Validation and Uncertainty Quantification, vol. 3, no. 3, pp. 031002, 2018
2018
-
[36]
G. Lee, W. Kim, H. Oh, B. D. Youn, and N. H. Kim, Review of statistical model calibration and validation from the perspective of uncertainty structures, Structural and Multidisciplinary Optimization 91, pp. 1-26, 2019
2019
-
[37]
Mullins, Y
J. Mullins, Y. Ling, S. Mahadevan, L. Sun, A. Strachan, Separation of aleatory and epistemic uncertainty in probabilistic model validation, Reliab. Eng. Syst. Saf., vol. 147, pp. 49-59, 2016
2016
-
[38]
Pedregosa, G
F. Pedregosa, G. Varoquaux, A. Gramfort, V. Michel, B. Thirion, O. Grisel, M. Blondel, P. Prettenhofer, R. Weiss, V. Dubourg, J. Vanderplas, A. Passos, D. Cournapeau, M. Brucher, M. Perrot, E. Duchesnay, Scikit-learn: Machine Learning in Python, Journal of Machine Learning Res...
2011
-
[39]
Foundations of Probability Theory, Statistical Inference, and Statistical Theories of Science Vol. II
E. T. Jaynes. In “Foundations of Probability Theory, Statistical Inference, and Statistical Theories of Science Vol. II”, D. Reidel Publishing Company, Dordrecht-Holland, pp. 175-257, 1976
1976
-
[40]
E. T. Jaynes, Probability Theory: The Logic of Science, UK Cambridge: Cambridge University Press, 2003
2003
-
[41]
Feroz and M
F. Feroz and M. P. Hobson, Multimodal nested sampling: an efficient and robust alternative to Markov chain Monte Carlo methods for astronomical data analyses, Monthly Notices of the Royal Astronomical Society, vol. 384, no. 2, pp. 449-463, 2008
2008
-
[42]
Caticha, Entropic Inference and the Foundations of Physics (monograph commis- sioned by the 11th Brazilian Meeting on Bayesian Statistics - EBEB-2012)
A. Caticha, Entropic Inference and the Foundations of Physics (monograph commis- sioned by the 11th Brazilian Meeting on Bayesian Statistics - EBEB-2012). 2012. URL http://www.albany.edu/physics/ACaticha-EIFP-book.pdf
2012
-
[43]
probability of agreement
K. Knuth, Optimal Data-Based Binning for Histograms, arxiv:physics/0605197v2, https://arxiv.org/abs/physics/0605197v2, 2013. 21 A Deriving the other validation metrics from the BVM In the following subsections we will show some of the special cases of the Bayesian validation m...
2013 arXiv
Reviewed August 14, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.