Pith. sign in

REVIEW 4 minor 62 references

A Censored Transformed Model for Proportional Outcomes with Boundary Mass and an Application to Loss Given Default Modeling

T0 review · 0 major / 4 minor · reviewed 2026-06-26 · grok-4.3

Pith's one-line read The zero-one censored transformed normal model combines a censored Gaussian with an affine-logit transform to handle proportional outcomes that pile up at 0 and 1.

desk verdict The paper introduces the ZOC-TN model for proportional data with boundary mass and shows solid empirical results on mortgage LGD. read the letter →

arxiv 2606.21515 v1 pith:BSZ76NFS submitted 2026-06-19 stat.ME q-fin.RMq-fin.STstat.APstat.ML

classification stat.MEq-fin.RMq-fin.STstat.APstat.ML
keywords zero-onecensoredtransformednormalproportionaloutcomesboundarymasslossgivendefaulttreeboostingspatio-temporalfrailtyaffine-logittransformation
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper introduces the ZOC-TN model for responses in [0,1] that can have positive probability at the endpoints. A censored Gaussian is mapped through a two-parameter affine-logit function on the open interval (0,1), which lets the density take many different shapes while the whole specification stays simple. The authors derive the parameter mapping, prove large-sample consistency, and show how the model extends to tree boosting for nonlinear effects and to a spatio-temporal Gaussian process frailty term. In a large U.S. mortgage loss-given-default dataset the boosted version with the frailty term produces the best out-of-sample predictions among the models compared.

What carries the argument

The zero-one censored transformed normal (ZOC-TN) specification, formed by applying a two-parameter affine-logit transformation to a censored Gaussian variable.

What would settle it

A dataset of proportional outcomes whose conditional densities on (0,1) systematically deviate from the shapes generated by any affine-logit transform of a normal, or where the boosted ZOC-TN model with frailty fails to improve out-of-sample log-score or calibration relative to the listed benchmark models on a held-out mortgage portfolio.

Watch

Extended reading notes

Core claim

The ZOC-TN model represents the interior distribution on (0,1) as an affine-logit transformation of a Gaussian random variable that is censored at the boundaries, thereby generating a flexible family of densities with atoms at 0 and 1; the transformation parameters are explicitly characterized, asymptotic properties are established, and the resulting estimator remains numerically stable. When embedded in a tree-boosted framework with an added spatio-temporal frailty Gaussian process, the model delivers the strongest predictive performance on loss-given-default data.

Load-bearing premise

The interior distribution on (0,1) is adequately described by the two-parameter affine-logit transformation of a Gaussian random variable.

Editorial extensions

If this is right

  • The model can represent a wider range of unimodal and bimodal interior densities than standard beta or logit-normal alternatives while retaining only a small number of parameters.
  • Large-sample theory supplies consistent estimators and asymptotic normality for inference on the transformation parameters and regression coefficients.
  • Tree boosting incorporates covariate interactions and nonlinearities without destroying the boundary-mass structure.
  • The added spatio-temporal frailty term absorbs residual geographic and temporal dependence that would otherwise bias predictions.
  • In loss-given-default applications the combined specification produces lower out-of-sample error than several common benchmark models.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The same transformation could be tested on other bounded fractional outcomes such as vote shares or recovery rates in different asset classes.
  • If the affine-logit-Gaussian assumption holds only approximately, the model may still serve as a convenient starting point for semiparametric refinements.
  • The performance gain from the frailty term suggests that ignoring space-time clustering in mortgage portfolios systematically understates tail loss risk.
  • Numerical stability of the likelihood may allow routine use on datasets orders of magnitude larger than the mortgage sample examined.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

0 major / 4 minor

Summary. The paper introduces the zero-one censored transformed normal (ZOC-TN) model for proportional responses with probability mass at the boundaries 0 and 1. The model combines a censored Gaussian random variable with a two-parameter affine-logit transformation applied to the interior (0,1) interval. The authors characterize the transformation parameters, establish large-sample properties, and relate the specification to broader classes of interior distributions. Theoretical and experimental results are presented showing that the ZOC-TN model captures a wider range of qualitative density shapes than several benchmark models while remaining parsimonious, computationally efficient, and numerically stable. Extensions are proposed to incorporate nonlinearities via tree-boosting and residual spatio-temporal variability via a frailty Gaussian process. The model is applied to loss given default (LGD) modeling on a large dataset of U.S. residential mortgages, where a tree-boosted ZOC-TN model with spatio-temporal frailty delivers the strongest out-of-sample performance.

Significance. If the claimed theoretical properties and empirical superiority hold, the ZOC-TN model offers a parsimonious and stable alternative for bounded proportional outcomes with boundary masses, with direct relevance to credit-risk applications such as LGD. The tree-boosting and frailty extensions address practical features like covariate nonlinearity and unmodeled space-time dependence. Explicit credit is due for the parameter characterization, large-sample results, and the emphasis on numerical stability and computational efficiency, which support reproducibility and implementation.

minor comments (4)
  1. The abstract and introduction assert that the model captures a wider range of qualitative density shapes; the corresponding section or figure comparing the attainable shapes (e.g., via parameter sweeps) should explicitly list the benchmark models and the metrics or visual criteria used for the comparison.
  2. In the application section, basic descriptive statistics of the U.S. residential mortgage LGD dataset (sample size, time span, geographic coverage, and proportion of boundary observations) should be reported to allow readers to assess the relevance of the out-of-sample results.
  3. Notation for the two transformation parameters and the censoring mechanism should be introduced with a single, self-contained equation block early in the model section to improve readability for readers unfamiliar with transformed-normal constructions.
  4. The large-sample properties are stated to be established; the statement of the regularity conditions (e.g., on the covariate design or the interior density) should be collected in one location rather than dispersed across the theoretical development.

Simulated Author's Rebuttal

0 responses · 0 unresolved

We thank the referee for the positive assessment of our manuscript and the recommendation for minor revision. The referee's summary accurately captures the key elements of the ZOC-TN model, its theoretical properties, extensions, and empirical application to LGD data.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity

full rationale

The paper introduces the ZOC-TN model as a distinct construction that combines a censored Gaussian random variable with a two-parameter affine-logit transformation on (0,1), then characterizes the transformation parameters and derives large-sample properties using standard statistical arguments. These steps are independent of the target claims about density shape flexibility and out-of-sample performance; the latter are established via explicit comparisons to benchmark models and empirical application rather than by re-expressing fitted quantities as predictions. No self-citation chains, self-definitional reductions, or fitted-input-as-prediction patterns appear in the derivation chain.

Assumptions & free parameters 1 free parameters · 1 assumptions · 0 invented entities

The model rests on standard censored-regression assumptions plus the specific choice of affine-logit link; no new physical entities are postulated and the two transformation parameters are estimated from data.

free parameters (1)
  • two transformation parameters
    Parameters of the affine-logit transformation on (0,1) that are estimated from data to control the interior density shape.
assumptions (1)
  • domain assumption Large-sample properties of the estimators hold under standard regularity conditions for censored and transformed models.
    The abstract states that large-sample properties are established.

how reviews work

0 comments
Cite this review

Pith. "Pith review of A Censored Transformed Model for Proportional Outcomes with Boundary Mass and an Application to Loss Given Default Modeling." pith.science (2026). https://pith.science/paper/BSZ76NFS

@misc{pith2026260621515,
  author       = {Pith},
  title        = {Pith review of: A Censored Transformed Model for Proportional Outcomes with Boundary Mass and an Application to Loss Given Default Modeling},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/BSZ76NFS}},
  note         = {Machine review of arXiv:2606.21515}
}
read the original abstract

We introduce the zero-one censored transformed normal (ZOC-TN) model for proportional responses with potential probability mass at the boundaries 0 and 1. The model combines a censored Gaussian variable with a two-parameter affine-logit transformation on the interior (0,1). We characterize the transformation parameters, establish large-sample properties, and relate the affine-logit specification to broader classes of interior distributions. Theoretical and experimental results demonstrate that the proposed model can capture a wider range of qualitative density shapes than several benchmark models while remaining parsimonious, computationally efficient, and numerically stable. Furthermore, the ZOC-TN model can be extended (i) to account for nonlinearities and interactions in a tree-boosting machine learning framework and (ii) to explicitly model residual spatio-temporal variability. We apply the ZOC-TN model to loss given default (LGD) modeling for a large dataset of U.S. residential mortgages and compare it to multiple benchmark models. We find that a tree-boosted ZOC-TN model with a spatio-temporal frailty Gaussian process delivers the strongest out-of-sample performance, indicating that mortgage losses are shaped by nonlinear covariate effects and by unaccounted-for space-time variation.

Figures

Figures reproduced from arXiv: 2606.21515 by the authors.

Figure 1
Figure 1. Empirical LGD distribution. Yellow bars represent boundary observations. [PITH_FULL_IMAGE:figures/full_fig_p010_1.png] view at source ↗
Figure 2
Figure 2. Left plot: Heatmap of average LGDs by ZIP3 locality (no data for gray regions). Right plot: [PITH_FULL_IMAGE:figures/full_fig_p010_2.png] view at source ↗
Figure 3
Figure 3. Suspended rootograms for independent linear models compared in Table [PITH_FULL_IMAGE:figures/full_fig_p011_3.png] view at source ↗
Figures from the paper (3 more)
Figure 4
Figure 4. Figure 4: One-year-ahead prediction accuracy of ZOC-TN models vs. time. [PITH_FULL_IMAGE:figures/full_fig_p014_4.png]
Figure 5
Figure 5. Figure 5: Posterior mean of latent Gaussian process in spatial tree-boosting model. [PITH_FULL_IMAGE:figures/full_fig_p015_5.png]
Figure 6
Figure 6. Figure 6: SHAP dependence plots for the four most important numerical covariates in the spatio [PITH_FULL_IMAGE:figures/full_fig_p016_6.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

62 extracted references

  1. [1]

    Williams, Christopher KI and Rasmussen, Carl Edward , year=

  2. [2]

    Leng, Kun and Li, Emmy and Eser, Rana and Piergies, Antonia and Sit, Rene and Tan, Michelle and Neff, Norma and Li, Song Hua and Rodriguez, Roberta Diehl and Suemoto, Claudia Kimie and Leite, Renata Elaine Paraizo and Ehrenberg, Alexander J and Pasqualucci, Carlos A and Seeley, William W and Spina, Salvatore and Heinsen, Helmut and Grinberg, Lea T and Kam...

  3. [3]

    and Khoo, Celyn L

    Geissinger, Emilie A. and Khoo, Celyn L. L. and Richmond, Isabella C. and Faulkner, Sally J. M. and Schneider, David C. , title =. Ecosphere , volume =

  4. [4]

    and Weedon, James T

    Douma, Jacob C. and Weedon, James T. , title =

  5. [5]

    Mindy Leow and Christophe Mues , journal =

  6. [6]

    Tony Bellotti and Jonathan Crook , journal =

  7. [7]

    2026 , journal=

    Scalable and robust regression models for continuous proportional data , author=. 2026 , journal=

  8. [8]

    2017 , month = dec, url =

Show all 62 references
  1. [9]

    2016 , pages =

    Tianqi Chen and Carlos Guestrin , title =. 2016 , pages =

  2. [10]

    Journal of Machine Learning Research , year =

    Fabio Sigrist , title =. Journal of Machine Learning Research , year =

  3. [11]

    Scholes , journal=

    Fischer Black and Myron S. Scholes , journal=. 1973 , volume=

  4. [12]

    1993 , month=

    Real Estate Economics , author=. 1993 , month=

  5. [13]

    Kau and Donald C

    James B. Kau and Donald C. Keenan , journal =. Patterns of rational default , volume =

  6. [14]

    and Capone, Charles A

    Ambrose, Brent W. and Capone, Charles A. and Deng, Yongheng , journal =

  7. [15]

    Calem and Michael LaCour-Little , journal =

    Paul S. Calem and Michael LaCour-Little , journal =. Risk-based capital requirements for mortgage loans , volume =

  8. [16]

    2009 , journal =

    Loss given default of high loan-to-value residential mortgages , author =. 2009 , journal =

  9. [17]

    2021 , month=

    Real Estate Economics , author=. 2021 , month=

  10. [18]

    1975 , month=

    Econometrica , author=. 1975 , month=

  11. [19]

    , journal =

    Gupton, Greg M. , journal =

  12. [20]

    Sigrist, Fabio and Stahel, Werner A. , year=

  13. [21]

    Forecasting bank loans loss-given-default , volume =

    Jo. Forecasting bank loans loss-given-default , volume =. Journal of Banking & Finance , number =

  14. [22]

    and Wooldridge, Jeffrey M

    Papke, Leslie E. and Wooldridge, Jeffrey M. , journal =

  15. [23]

    Tong and Christophe Mues and Lyn Thomas , journal =

    Edward N.C. Tong and Christophe Mues and Lyn Thomas , journal =. A zero-adjusted gamma model for mortgage loan loss given default , volume =

  16. [24]

    Predicting bank loan recovery rates with a mixed continuous-discrete model , volume =

    Calabrese, Raffaella , journal =. Predicting bank loan recovery rates with a mixed continuous-discrete model , volume =

  17. [25]

    Enhancing two-stage modelling methodology for loss given default with support vector machines , volume =

    Xiao Yao and Jonathan Crook and Galina Andreeva , journal =. Enhancing two-stage modelling methodology for loss given default with support vector machines , volume =

  18. [26]

    Suykens, Johan and Vandewalle, Joos , year =

  19. [27]

    Tobback, Ellen and Martens, David and Van Gestel, Tony and Baesens, Bart , year =

  20. [28]

    Benchmarking regression algorithms for loss given default modeling , volume =

    Gert Loterman and Iain Brown and David Martens and Christophe Mues and Bart Baesens , journal =. Benchmarking regression algorithms for loss given default modeling , volume =

  21. [29]

    Opening the black box -- Quantile neural networks for loss given default prediction , volume =

    Ralf Kellner and Maximilian Nagl and Daniel R. Opening the black box -- Quantile neural networks for loss given default prediction , volume =. Journal of Banking & Finance , pages =

  22. [30]

    and Chomsisengphet, Souphala and Sanders, Anthony B

    Agarwal, Sumit and Ambrose, Brent W. and Chomsisengphet, Souphala and Sanders, Anthony B. , journal =

  23. [31]

    2014 , month=

    The Journal of Real Estate Finance and Economics , author=. 2014 , month=

  24. [32]

    Kelley Pace and Luca Zanin , journal =

    Raffaella Calabrese and Timothy Dombrowski and Antoine Mandel and R. Kelley Pace and Luca Zanin , journal =

  25. [33]

    European Journal of Operational Research , title =

    Pascal K. European Journal of Operational Research , title =

  26. [34]

    James Tobin , journal =

  27. [35]

    2026 , howpublished =

  28. [36]

    Extended-support beta regression for [0, 1] responses , volume=

    Kosmidis, Ioannis and Zeileis, Achim , year=. Extended-support beta regression for [0, 1] responses , volume=

  29. [37]

    Trevor Hastie and Robert Tibshirani and Jerome Friedman , title =

  30. [38]

    Jerome Friedman and Trevor Hastie and Robert Tibshirani , journal =

  31. [39]

    Silvia Ferrari and Francisco Cribari-Neto , journal =

  32. [40]

    Ospina, Raydonal and Ferrari, Silvia L. P. , journal =. Inflated beta distributions , volume =

  33. [41]

    and Ramalho, Joaquim J.S

    Ramalho, Esmeralda A. and Ramalho, Joaquim J.S. and Murteira, José M.R. , title =. Journal of Economic Surveys , volume =

  34. [42]

    Smithson, Michael and Verkuilen, Jay , year =

  35. [43]

    2025 , pages=

    Journal of the American Statistical Association , author=. 2025 , pages=

  36. [44]

    A. V. Vecchia , journal =. Estimation and Model Identification for Continuous Spatial Processes , volume =

  37. [45]

    Statistical Science , publisher=

    Katzfuss, Matthias and Guinness, Joseph , year=. Statistical Science , publisher=

  38. [46]

    Sigrist, Fabio , year =

  39. [47]

    The American Statistician , publisher=

    Kleiber, Christian and Zeileis, Achim , year=. The American Statistician , publisher=

  40. [48]

    James Bradbury and Roy Frostig and Peter Hawkins and Matthew James Johnson and Yash Katariya and Chris Leary and Dougal Maclaurin and George Necula and Adam Paszke and Jake Vander

  41. [49]

    and Haberland, Matt and Reddy, Tyler and Cournapeau, David and Burovski, Evgeni and Peterson, Pearu and Weckesser, Warren and Bright, Jonathan and

    Virtanen, Pauli and Gommers, Ralf and Oliphant, Travis E. and Haberland, Matt and Reddy, Tyler and Cournapeau, David and Burovski, Evgeni and Peterson, Pearu and Weckesser, Warren and Bright, Jonathan and. Nature Methods , volume =

  42. [50]

    and Nocedal, Jorge , title =

    Zhu, Ciyou and Byrd, Richard H. and Nocedal, Jorge , title =

  43. [51]

    Advances in Neural Information Processing Systems , volume =

    Lundberg, Scott and Lee, Su-In , title =. Advances in Neural Information Processing Systems , volume =

  44. [52]

    Austin Harrison and Dan Immergluck , journal =

  45. [53]

    Econometrica: journal of the Econometric Society , pages=

    Estimation of relationships for limited dependent variables , author=. Econometrica: journal of the Econometric Society , pages=. 1958 , publisher=

  46. [54]

    Roger Koenker and José A. F. Machado , journal =

  47. [55]

    Lorentz, G. G. , year=

  48. [56]

    , journal=

    Farouki, Rida T. , journal=

  49. [57]

    , title =

    Scheuerer, Michael and Hamill, Thomas M. , title =. Monthly Weather Review , volume =

  50. [58]

    Expert Systems with Applications , volume=

    Gradient and Newton boosting for classification and regression , author=. Expert Systems with Applications , volume=. 2021 , publisher=

  51. [59]

    Journal of the American Statistical Association , volume=

    Hierarchical nearest-neighbor Gaussian process models for large geostatistical datasets , author=. Journal of the American Statistical Association , volume=. 2016 , publisher=

  52. [60]

    Annals of statistics , pages=

    Greedy function approximation: a gradient boosting machine , author=. Annals of statistics , pages=. 2001 , publisher=

  53. [61]

    International Journal of Forecasting , volume=

    Forecasting with trees , author=. International Journal of Forecasting , volume=. 2022 , publisher=

  54. [62]

    Neural Information Processing Systems Datasets and Benchmarks Track , year=

    Why do tree-based models still outperform deep learning on tabular data? , author=. Neural Information Processing Systems Datasets and Benchmarks Track , year=

Pith tools

Reviewed June 26, 2026 · model on record in the stance chip above.