REVIEW 4 minor 62 references
A Censored Transformed Model for Proportional Outcomes with Boundary Mass and an Application to Loss Given Default Modeling
T0 review · 0 major / 4 minor · reviewed 2026-06-26 · grok-4.3
Pith's one-line read The zero-one censored transformed normal model combines a censored Gaussian with an affine-logit transform to handle proportional outcomes that pile up at 0 and 1.
desk verdict The paper introduces the ZOC-TN model for proportional data with boundary mass and shows solid empirical results on mortgage LGD. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The zero-one censored transformed normal (ZOC-TN) specification, formed by applying a two-parameter affine-logit transformation to a censored Gaussian variable.
What would settle it
A dataset of proportional outcomes whose conditional densities on (0,1) systematically deviate from the shapes generated by any affine-logit transform of a normal, or where the boosted ZOC-TN model with frailty fails to improve out-of-sample log-score or calibration relative to the listed benchmark models on a held-out mortgage portfolio.
Extended reading notes
Core claim
The ZOC-TN model represents the interior distribution on (0,1) as an affine-logit transformation of a Gaussian random variable that is censored at the boundaries, thereby generating a flexible family of densities with atoms at 0 and 1; the transformation parameters are explicitly characterized, asymptotic properties are established, and the resulting estimator remains numerically stable. When embedded in a tree-boosted framework with an added spatio-temporal frailty Gaussian process, the model delivers the strongest predictive performance on loss-given-default data.
Load-bearing premise
The interior distribution on (0,1) is adequately described by the two-parameter affine-logit transformation of a Gaussian random variable.
Editorial extensions
If this is right
- The model can represent a wider range of unimodal and bimodal interior densities than standard beta or logit-normal alternatives while retaining only a small number of parameters.
- Large-sample theory supplies consistent estimators and asymptotic normality for inference on the transformation parameters and regression coefficients.
- Tree boosting incorporates covariate interactions and nonlinearities without destroying the boundary-mass structure.
- The added spatio-temporal frailty term absorbs residual geographic and temporal dependence that would otherwise bias predictions.
- In loss-given-default applications the combined specification produces lower out-of-sample error than several common benchmark models.
Reading between the lines
- The same transformation could be tested on other bounded fractional outcomes such as vote shares or recovery rates in different asset classes.
- If the affine-logit-Gaussian assumption holds only approximately, the model may still serve as a convenient starting point for semiparametric refinements.
- The performance gain from the frailty term suggests that ignoring space-time clustering in mortgage portfolios systematically understates tail loss risk.
- Numerical stability of the likelihood may allow routine use on datasets orders of magnitude larger than the mortgage sample examined.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces the zero-one censored transformed normal (ZOC-TN) model for proportional responses with probability mass at the boundaries 0 and 1. The model combines a censored Gaussian random variable with a two-parameter affine-logit transformation applied to the interior (0,1) interval. The authors characterize the transformation parameters, establish large-sample properties, and relate the specification to broader classes of interior distributions. Theoretical and experimental results are presented showing that the ZOC-TN model captures a wider range of qualitative density shapes than several benchmark models while remaining parsimonious, computationally efficient, and numerically stable. Extensions are proposed to incorporate nonlinearities via tree-boosting and residual spatio-temporal variability via a frailty Gaussian process. The model is applied to loss given default (LGD) modeling on a large dataset of U.S. residential mortgages, where a tree-boosted ZOC-TN model with spatio-temporal frailty delivers the strongest out-of-sample performance.
Significance. If the claimed theoretical properties and empirical superiority hold, the ZOC-TN model offers a parsimonious and stable alternative for bounded proportional outcomes with boundary masses, with direct relevance to credit-risk applications such as LGD. The tree-boosting and frailty extensions address practical features like covariate nonlinearity and unmodeled space-time dependence. Explicit credit is due for the parameter characterization, large-sample results, and the emphasis on numerical stability and computational efficiency, which support reproducibility and implementation.
minor comments (4)
- The abstract and introduction assert that the model captures a wider range of qualitative density shapes; the corresponding section or figure comparing the attainable shapes (e.g., via parameter sweeps) should explicitly list the benchmark models and the metrics or visual criteria used for the comparison.
- In the application section, basic descriptive statistics of the U.S. residential mortgage LGD dataset (sample size, time span, geographic coverage, and proportion of boundary observations) should be reported to allow readers to assess the relevance of the out-of-sample results.
- Notation for the two transformation parameters and the censoring mechanism should be introduced with a single, self-contained equation block early in the model section to improve readability for readers unfamiliar with transformed-normal constructions.
- The large-sample properties are stated to be established; the statement of the regularity conditions (e.g., on the covariate design or the interior density) should be collected in one location rather than dispersed across the theoretical development.
Simulated Author's Rebuttal
We thank the referee for the positive assessment of our manuscript and the recommendation for minor revision. The referee's summary accurately captures the key elements of the ZOC-TN model, its theoretical properties, extensions, and empirical application to LGD data.
Circularity Check
No significant circularity
full rationale
The paper introduces the ZOC-TN model as a distinct construction that combines a censored Gaussian random variable with a two-parameter affine-logit transformation on (0,1), then characterizes the transformation parameters and derives large-sample properties using standard statistical arguments. These steps are independent of the target claims about density shape flexibility and out-of-sample performance; the latter are established via explicit comparisons to benchmark models and empirical application rather than by re-expressing fitted quantities as predictions. No self-citation chains, self-definitional reductions, or fitted-input-as-prediction patterns appear in the derivation chain.
Assumptions & free parameters
free parameters (1)
- two transformation parameters
assumptions (1)
- domain assumption Large-sample properties of the estimators hold under standard regularity conditions for censored and transformed models.
Cite this review
Pith. "Pith review of A Censored Transformed Model for Proportional Outcomes with Boundary Mass and an Application to Loss Given Default Modeling." pith.science (2026). https://pith.science/paper/BSZ76NFS
@misc{pith2026260621515,
author = {Pith},
title = {Pith review of: A Censored Transformed Model for Proportional Outcomes with Boundary Mass and an Application to Loss Given Default Modeling},
year = {2026},
howpublished = {\url{https://pith.science/paper/BSZ76NFS}},
note = {Machine review of arXiv:2606.21515}
}
read the original abstract
We introduce the zero-one censored transformed normal (ZOC-TN) model for proportional responses with potential probability mass at the boundaries 0 and 1. The model combines a censored Gaussian variable with a two-parameter affine-logit transformation on the interior (0,1). We characterize the transformation parameters, establish large-sample properties, and relate the affine-logit specification to broader classes of interior distributions. Theoretical and experimental results demonstrate that the proposed model can capture a wider range of qualitative density shapes than several benchmark models while remaining parsimonious, computationally efficient, and numerically stable. Furthermore, the ZOC-TN model can be extended (i) to account for nonlinearities and interactions in a tree-boosting machine learning framework and (ii) to explicitly model residual spatio-temporal variability. We apply the ZOC-TN model to loss given default (LGD) modeling for a large dataset of U.S. residential mortgages and compare it to multiple benchmark models. We find that a tree-boosted ZOC-TN model with a spatio-temporal frailty Gaussian process delivers the strongest out-of-sample performance, indicating that mortgage losses are shaped by nonlinear covariate effects and by unaccounted-for space-time variation.
Figures
Figures from the paper (3 more)
Reference graph
Works this paper leans on
-
[1]
Williams, Christopher KI and Rasmussen, Carl Edward , year=
-
[2]
Leng, Kun and Li, Emmy and Eser, Rana and Piergies, Antonia and Sit, Rene and Tan, Michelle and Neff, Norma and Li, Song Hua and Rodriguez, Roberta Diehl and Suemoto, Claudia Kimie and Leite, Renata Elaine Paraizo and Ehrenberg, Alexander J and Pasqualucci, Carlos A and Seeley, William W and Spina, Salvatore and Heinsen, Helmut and Grinberg, Lea T and Kam...
-
[3]
and Khoo, Celyn L
Geissinger, Emilie A. and Khoo, Celyn L. L. and Richmond, Isabella C. and Faulkner, Sally J. M. and Schneider, David C. , title =. Ecosphere , volume =
-
[4]
and Weedon, James T
Douma, Jacob C. and Weedon, James T. , title =
-
[5]
Mindy Leow and Christophe Mues , journal =
-
[6]
Tony Bellotti and Jonathan Crook , journal =
-
[7]
2026 , journal=
Scalable and robust regression models for continuous proportional data , author=. 2026 , journal=
2026
-
[8]
2017 , month = dec, url =
2017
Show all 62 references
-
[9]
2016 , pages =
Tianqi Chen and Carlos Guestrin , title =. 2016 , pages =
2016
-
[10]
Journal of Machine Learning Research , year =
Fabio Sigrist , title =. Journal of Machine Learning Research , year =
-
[11]
Scholes , journal=
Fischer Black and Myron S. Scholes , journal=. 1973 , volume=
1973
-
[12]
1993 , month=
Real Estate Economics , author=. 1993 , month=
1993
-
[13]
Kau and Donald C
James B. Kau and Donald C. Keenan , journal =. Patterns of rational default , volume =
-
[14]
and Capone, Charles A
Ambrose, Brent W. and Capone, Charles A. and Deng, Yongheng , journal =
-
[15]
Calem and Michael LaCour-Little , journal =
Paul S. Calem and Michael LaCour-Little , journal =. Risk-based capital requirements for mortgage loans , volume =
-
[16]
2009 , journal =
Loss given default of high loan-to-value residential mortgages , author =. 2009 , journal =
2009
-
[17]
2021 , month=
Real Estate Economics , author=. 2021 , month=
2021
-
[18]
1975 , month=
Econometrica , author=. 1975 , month=
1975
-
[19]
, journal =
Gupton, Greg M. , journal =
-
[20]
Sigrist, Fabio and Stahel, Werner A. , year=
-
[21]
Forecasting bank loans loss-given-default , volume =
Jo. Forecasting bank loans loss-given-default , volume =. Journal of Banking & Finance , number =
-
[22]
and Wooldridge, Jeffrey M
Papke, Leslie E. and Wooldridge, Jeffrey M. , journal =
-
[23]
Tong and Christophe Mues and Lyn Thomas , journal =
Edward N.C. Tong and Christophe Mues and Lyn Thomas , journal =. A zero-adjusted gamma model for mortgage loan loss given default , volume =
-
[24]
Predicting bank loan recovery rates with a mixed continuous-discrete model , volume =
Calabrese, Raffaella , journal =. Predicting bank loan recovery rates with a mixed continuous-discrete model , volume =
-
[25]
Enhancing two-stage modelling methodology for loss given default with support vector machines , volume =
Xiao Yao and Jonathan Crook and Galina Andreeva , journal =. Enhancing two-stage modelling methodology for loss given default with support vector machines , volume =
-
[26]
Suykens, Johan and Vandewalle, Joos , year =
-
[27]
Tobback, Ellen and Martens, David and Van Gestel, Tony and Baesens, Bart , year =
-
[28]
Benchmarking regression algorithms for loss given default modeling , volume =
Gert Loterman and Iain Brown and David Martens and Christophe Mues and Bart Baesens , journal =. Benchmarking regression algorithms for loss given default modeling , volume =
-
[29]
Opening the black box -- Quantile neural networks for loss given default prediction , volume =
Ralf Kellner and Maximilian Nagl and Daniel R. Opening the black box -- Quantile neural networks for loss given default prediction , volume =. Journal of Banking & Finance , pages =
-
[30]
and Chomsisengphet, Souphala and Sanders, Anthony B
Agarwal, Sumit and Ambrose, Brent W. and Chomsisengphet, Souphala and Sanders, Anthony B. , journal =
-
[31]
2014 , month=
The Journal of Real Estate Finance and Economics , author=. 2014 , month=
2014
-
[32]
Kelley Pace and Luca Zanin , journal =
Raffaella Calabrese and Timothy Dombrowski and Antoine Mandel and R. Kelley Pace and Luca Zanin , journal =
-
[33]
European Journal of Operational Research , title =
Pascal K. European Journal of Operational Research , title =
-
[34]
James Tobin , journal =
-
[35]
2026 , howpublished =
2026
-
[36]
Extended-support beta regression for [0, 1] responses , volume=
Kosmidis, Ioannis and Zeileis, Achim , year=. Extended-support beta regression for [0, 1] responses , volume=
-
[37]
Trevor Hastie and Robert Tibshirani and Jerome Friedman , title =
-
[38]
Jerome Friedman and Trevor Hastie and Robert Tibshirani , journal =
-
[39]
Silvia Ferrari and Francisco Cribari-Neto , journal =
-
[40]
Ospina, Raydonal and Ferrari, Silvia L. P. , journal =. Inflated beta distributions , volume =
-
[41]
and Ramalho, Joaquim J.S
Ramalho, Esmeralda A. and Ramalho, Joaquim J.S. and Murteira, José M.R. , title =. Journal of Economic Surveys , volume =
-
[42]
Smithson, Michael and Verkuilen, Jay , year =
-
[43]
2025 , pages=
Journal of the American Statistical Association , author=. 2025 , pages=
2025
-
[44]
A. V. Vecchia , journal =. Estimation and Model Identification for Continuous Spatial Processes , volume =
-
[45]
Statistical Science , publisher=
Katzfuss, Matthias and Guinness, Joseph , year=. Statistical Science , publisher=
-
[46]
Sigrist, Fabio , year =
-
[47]
The American Statistician , publisher=
Kleiber, Christian and Zeileis, Achim , year=. The American Statistician , publisher=
-
[48]
James Bradbury and Roy Frostig and Peter Hawkins and Matthew James Johnson and Yash Katariya and Chris Leary and Dougal Maclaurin and George Necula and Adam Paszke and Jake Vander
-
[49]
and Haberland, Matt and Reddy, Tyler and Cournapeau, David and Burovski, Evgeni and Peterson, Pearu and Weckesser, Warren and Bright, Jonathan and
Virtanen, Pauli and Gommers, Ralf and Oliphant, Travis E. and Haberland, Matt and Reddy, Tyler and Cournapeau, David and Burovski, Evgeni and Peterson, Pearu and Weckesser, Warren and Bright, Jonathan and. Nature Methods , volume =
-
[50]
and Nocedal, Jorge , title =
Zhu, Ciyou and Byrd, Richard H. and Nocedal, Jorge , title =
-
[51]
Advances in Neural Information Processing Systems , volume =
Lundberg, Scott and Lee, Su-In , title =. Advances in Neural Information Processing Systems , volume =
-
[52]
Austin Harrison and Dan Immergluck , journal =
-
[53]
Econometrica: journal of the Econometric Society , pages=
Estimation of relationships for limited dependent variables , author=. Econometrica: journal of the Econometric Society , pages=. 1958 , publisher=
1958
-
[54]
Roger Koenker and José A. F. Machado , journal =
-
[55]
Lorentz, G. G. , year=
-
[56]
, journal=
Farouki, Rida T. , journal=
-
[57]
, title =
Scheuerer, Michael and Hamill, Thomas M. , title =. Monthly Weather Review , volume =
-
[58]
Expert Systems with Applications , volume=
Gradient and Newton boosting for classification and regression , author=. Expert Systems with Applications , volume=. 2021 , publisher=
2021
-
[59]
Journal of the American Statistical Association , volume=
Hierarchical nearest-neighbor Gaussian process models for large geostatistical datasets , author=. Journal of the American Statistical Association , volume=. 2016 , publisher=
2016
-
[60]
Annals of statistics , pages=
Greedy function approximation: a gradient boosting machine , author=. Annals of statistics , pages=. 2001 , publisher=
2001
-
[61]
International Journal of Forecasting , volume=
Forecasting with trees , author=. International Journal of Forecasting , volume=. 2022 , publisher=
2022
-
[62]
Neural Information Processing Systems Datasets and Benchmarks Track , year=
Why do tree-based models still outperform deep learning on tabular data? , author=. Neural Information Processing Systems Datasets and Benchmarks Track , year=
Reviewed June 26, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.