REVIEW 4 major objections 4 minor 11 references
Probabilistic Assessment of Engineered Timber Reusability after Moisture Exposure
T0 review · 4 major / 4 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read This paper argues that moisture-exposed engineered timber can be graded into three reuse levels from just wet-cycle count and grain orientation, with a Bayesian model supplying the probability for each level.
desk verdict Useful raw moisture-cycle flexural data, but the central 'two-input prediction' claim is circular: the reuse labels are derived from the same variables used as predictors, so the reported accuracy is an in-sample fit. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the residual-performance metric R_i = 0.8(E_i/E_0) + 0.2(sigma_max,i/sigma_0), computed against unexposed controls of the same grain orientation, combined with a hierarchical Bayesian multinomial logistic model. The model uses horseshoe priors to force irrelevant predictors' coefficients toward zero and treats the lower reuse-level boundary tau_2 as an unknown parameter with a Uniform(0.70, 0.85) prior, so the decision rule is inferred jointly with the regression coefficients through Markov-chain Monte Carlo using the No-U-Turn sampler. This setup lets the output be a probability per reuse level, not a single deterministic grade, and it is what allows the two-variable minimal predictor set to be discovered rather than assumed.
What would settle it
A direct check would be to take a larger batch of engineered timber with known service histories, measure the two binary predictors, apply the paper's decision rule, and compare the assigned reuse levels against full destructive testing and long-term field performance; the claim would be weakened if specimens with identical wet-cycle count and grain orientation but different density, moisture content, or prior load history show materially different residual performance, because that would mean the two-variable model is omitting a predictive dimension.
Extended reading notes
Core claim
The paper's central discovery is that a probabilistic classifier trained on destructive flexural tests can reproduce the observed distribution of moisture-damaged engineered timber into reuse levels using only two binary inputs: how many soaking-and-drying cycles a member has endured and whether the face grain runs parallel or perpendicular to the load. All other field-measurable candidates, density, moisture content, specimen size, and surface hardness, shrink toward zero under the horseshoe prior. The paper reports that a single wet-dry cycle leaves about 70% of specimens above the Level-1 threshold of R = 0.90, while two cycles pull the mean residual down to roughly 0.78 and move many specimens into lower levels; the inferred lower decision boundary is about 0.76 with a 95% HDI of 0.73-0.79. On this basis the paper claims to offer the first probabilistic framework for classifying moisture-exposed engineered timber within the MMC reuse hierarchy and proposes a three-zone decision rule based on the probability of Level-1 status.
Load-bearing premise
The whole classifier stands on the assumption that the weighted residual-performance score R, with 80% weight on stiffness and 20% on strength and thresholds at 0.90 and about 0.75, is the correct operational definition of reusability; if those weights or thresholds are wrong, or if laboratory soaking and oven-drying do not represent real service wetting, the probability labels have no fixed meaning.
Editorial extensions
If this is right
- If the two-input classifier generalises, on-site reuse triage reduces to recording cycle history and grain orientation and reading off P(Level 1), with no need for hardness or density meters.
- Quantified decision boundaries make it possible to write a standard: redeploy at P >= 0.70, non-destructive check at 0.40-0.70, downgrade below 0.40.
- The paper's estimate that replacing moisture-limit rules with this classifier would roughly double the volume of salvaged timber reused without structural downgrading follows directly from the fitted level proportions.
- Every component gets an explicit uncertainty interval, so a specifier can state the probability that a given member is below the direct-reuse threshold rather than asserting a binary pass/fail.
Reading between the lines
- A natural next step the paper does not run is a cost-benefit comparison: the model trades an off-by-one-level error rate of about 33% against the cost of destructive testing, and a decision-theoretic extension could optimise the P(Level 1) thresholds for a given tolerance for over- or under-grading.
- Because the horseshoe prior selected only two categorical features, the framework is directly portable to other engineered wood products, but the coefficient values and thresholds would need refitting on glulam or LVL data before field use.
- The paper's own text is internally inconsistent about the test material and sample counts, with the methods describing 50 mm plywood coupons and repeated 'three replicates' groups while Table 1 lists n=10 per group and the conclusion speaks of 72 spruce CLT specimens, so the demonstrated result should be read as small-coupon plywood evidence pending confirmation on structural-scale CLT.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper develops a Bayesian multinomial logistic model to classify the reusability of moisture-exposed engineered timber into three levels defined by a residual-performance metric R. R is computed as a weighted combination of retained flexural modulus and maximum load relative to orientation-matched controls (Eq. 1). The authors fix an upper threshold tau1 = 0.90, infer a lower threshold tau2 from the posterior, assign each specimen a reuse level by thresholding R, and then fit a two-predictor model using wet-cycle group and grain orientation. They report 67% classification accuracy, a multiclass Brier score of 0.432, and conclude that only two binary inputs are needed to reproduce the observed level distribution, positioning the framework as a streamlined on-site triage tool.
Significance. If the central claim were valid, the paper would offer a practically valuable contribution: a simple, two-variable probabilistic rule for triaging moisture-exposed engineered timber, with quantified uncertainty, in a domain where standards are lacking. The modelling machinery (horseshoe priors, NUTS, posterior predictive checks) is standard and appropriate for the intended inferential task. However, the significance is conditional on resolving a fundamental validation problem: the reuse labels are constructed from the same experimental variables that later serve as predictors, so the reported accuracy and probability estimates are in-sample re-descriptions of the label-construction rule rather than evidence of predictive capacity. The manuscript also contains multiple unresolved internal inconsistencies in the reported data and material identity, which further reduce confidence in the empirical basis.
major comments (4)
- [Secs. 2.3, 3.3, 3.5] The central claim that only wet-cycle count and grain orientation reproduce the observed level distribution is not supported because the reuse labels are derived from the same variables that are then used as predictors. In Eq. (1), R is computed relative to orientation-matched controls; in Sec. 3.3, tau2 is inferred from the posterior of R; and the labels L1/L2/L3 are assigned by thresholding R at tau1 = 0.90 and tau2 = 0.75. The multinomial model in Eq. (2) uses exactly cycle group and grain orientation as covariates. Consequently, the 67% accuracy and Brier score reported in Sec. 3.5 are computed on the training data and quantify in-sample fit, not predictive reuse classification. Without an independent ground-truth label (e.g., expert grading or a separate validation set), the two-predictor rule is confounded with the label construction and cannot support the abstract's and conclusions' claims.
- [Sec. 3.2 and Table 1] The descriptive statistics are internally inconsistent. The text in Sec. 3.2 reports n = 24 per group with standard deviations 0.03, 0.07, and 0.11, while Table 1 reports n = 10 with standard deviations 0.118, 0.097, and 0.268. Additionally, Sec. 3.3 gives class counts L1: 40, L2: 21, L3: 12, which sum to 73, while Sec. 5 states that flexural tests were performed on 72 specimens. These discrepancies affect the interpretation of the posterior proportions (0.55, 0.28, 0.17) and prevent a reader from reproducing the analysis.
- [Secs. 2.1 and 5] The material under study is identified inconsistently: Sec. 2.1 and Sec. 2.2 describe plywood specimens, while Sec. 5 states that the tests were conducted on '72 spruce CLT specimens.' CLT and plywood have different lamella orientations, thicknesses, and failure modes, so this is not a terminological triviality. The material identity is load-bearing for the external-validity claim that the framework applies to engineered timber in MMC reuse hierarchies.
- [Abstract and Sec. 3.4] The abstract claims the model predicts reuse levels from five field-measurable features: density, moisture content, specimen size, grain orientation, and surface hardness. However, the methods section never describes measurements of density, moisture content at the time of testing, or surface hardness, and Eq. (2) includes only two predictors. The claim that horseshoe shrinkage 'retains' two predictors and 'attenuates' three others is therefore unsubstantiated, because the other three features are not present in the reported experimental design or model specification.
minor comments (4)
- [Sec. 2.4] The convergence diagnostic is written as 'R ≤ 1.01'; the standard notation is R-hat (or 'R-hat'), and the value should be reported with its definition to avoid confusion with the residual-performance metric R introduced in Eq. (1).
- [Sec. 3.4 and Fig. 7] The text uses 'HID' in the phrase 'Posterior 95% HID'; this should be 'HDI' (highest density interval), consistent with Sec. 3.3.
- [Sec. 2.1] The specimen preparation states that three replicates were prepared for each group and orientation, but Sec. 3.2 and Table 1 report group sample sizes of 24 or 10. The relation between the number of replicates and the reported n is not explained.
- [Sec. 3.1] The flexural results are reported as means without confidence intervals or measures of variability around the mean, which is important because the residual-performance metric R is constructed from these means and the subsequent Bayesian inference propagates uncertainty from the posterior only, not from the flexural measurement error.
Circularity Check
Reuse-level predictions are in-sample fits: the lower threshold is inferred from the same residual-performance values used to define the labels, and the reported accuracy is computed on the training data.
-
fitted input called prediction
[Section 3.5, 'Classification accuracy and uncertainty' (labels defined in Sec. 3.3; model in Sec. 2.4)]
"Using posterior-mean logits, the categorical predictions yield the confusion matrix shown in Fig. 8. Overall accuracy ((TP+TN)/N) is 67%; misclassifications are almost exclusively off-by-one-level—only one Set 2 sample (4% of that group) is erroneously labelled Level 1."
These 'predictions' are posterior-mean logits evaluated on the same specimens used to fit the multinomial model of Eq. (2); the text describes no held-out set or cross-validation. The target labels are thresholds of the residual-performance metric R at 0.90 and 0.75, where 0.75 was itself inferred from the posterior of the same R values (Sec. 3.3). Reported accuracy and Brier score therefore measure in-sample agreement of a fitted model with labels derived from the fitted quantity, not out-of-sample predictive skill.
-
fitted input called prediction
[Section 3.3, 'Posterior estimates of τ₂ and level proportions']
"The posterior median for τ₂ converged at 0.76 with a 95 % highest-density interval (HDI) of 0.73–0.79, validating the provisional value of 0.75 adopted in the subsequent analyses. Implementing τ₁/τ₂ = 0.90/0.75 yields the class counts L1: 40, L2: 21, L3: 12"
τ₂ is estimated from the residual-performance values R, and the same R values are then thresholded at τ₂=0.75 to create the Level 1/2/3 labels. The counts 40/21/12 are thus outputs of a cut-point fitted to the outcome, not independent ground-truth grading. Any classifier trained on these labels and evaluated on them is being judged against a target that the analysis itself constructed, making the reported level distribution partly an artifact of the fitted threshold.
1 more flagged steps
-
other
[Section 5, Conclusions]
"Remarkably, only two binary inputs, wet-cycle count and grain orientation, were needed to reproduce the observed level distribution."
Eq. (2) is fit, via MCMC, with exactly these two binary predictors to the level distribution obtained by thresholding R. 'Reproducing the observed level distribution' is the fitting objective, so the conclusion restates the in-sample fit rather than demonstrating that the two inputs predict reuse levels in new or independently graded specimens. Without an external test set or an independent grading standard, the two-input rule is not shown to be predictive.
full rationale
The paper's derivation chain is not circular in the self-citation sense: none of its references are authored by the present authors, and Eq. (1)'s residual-performance metric is a reasonable operationalization of retained capacity. The circularity is statistical and constructional. The reuse levels (1/2/3) are defined by thresholding R at 0.90 and at τ₂, but τ₂ is inferred from the posterior of the very R values that are then thresholded (Sec. 3.3). The multinomial logistic model of Eq. (2) is fitted to those constructed labels using only the two experimental factors (cycle group and grain orientation), and the reported 67% accuracy and Brier score are computed on the same training data (Sec. 3.5). Consequently, the central claim that two binary inputs 'reproduce the observed level distribution' is a re-description of the in-sample fit, not an out-of-sample prediction. The honest reading is that the paper demonstrates a plausible in-sample association between moisture cycling, grain orientation, and its own residual-performance-based reuse labels; it does not yet validate a predictive two-input rule against independent ground truth.
Assumptions & free parameters
free parameters (4)
- Lower reuse threshold tau2 =
posterior median 0.76, provisional 0.75
- Residual-performance weights omega_E, omega_sigma =
0.8, 0.2
- Level-1 threshold tau1 =
0.90
- Orientation-specific baselines E0, sigma0 =
means from control groups (e.g., 10.59 GPa, 113.31 MPa longitudinal)
assumptions (4)
- standard math Bayesian probability calculus and NUTS convergence diagnostics justify posterior summaries.
- domain assumption Laboratory soak-dry cycles represent realistic in-service moisture exposure.
- domain assumption The residual-performance metric R with weights 0.8 and 0.2 is a valid measure of structural reusability.
- ad hoc to paper Reuse labels derived from R and thresholds tau1=0.90, tau2=0.75 are ground truth.
Cite this review
Pith. "Pith review of Probabilistic Assessment of Engineered Timber Reusability after Moisture Exposure." pith.science (2026). https://pith.science/paper/XJVBKZWK
@misc{pith2026250611061,
author = {Pith},
title = {Pith review of: Probabilistic Assessment of Engineered Timber Reusability after Moisture Exposure},
year = {2026},
howpublished = {\url{https://pith.science/paper/XJVBKZWK}},
note = {Machine review of arXiv:2506.11061}
}
read the original abstract
Engineered timber is pivotal to low-carbon construction, but moisture uptake during its service life can compromise structural reliability and impede reuse within a circular economy model. Despite growing interest, quantitative standards for classifying the reusability of moisture-exposed timber are still lacking. This study develops a probabilistic framework to determine the post-exposure reusability of engineered timber. Laminated specimens were soaked to full saturation, dried to 25% moisture content, and subjected to destructive three-point flexural testing. Structural integrity was quantified by a residual-performance metric that assigns 80% weight to the retained flexural modulus and 20% to the retained maximum load, benchmarked against unexposed controls. A hierarchical Bayesian multinomial logistic model with horseshoe priors, calibrated through Markov-Chain Monte-Carlo sampling, jointly infers the decision threshold separating three Modern Methods of Construction (MMC) reuse levels and predicts those levels from five field-measurable features: density, moisture content, specimen size, grain orientation, and surface hardness. Results indicate that a single wet-dry cycle preserves 70% of specimens above the 0.90 residual-performance threshold (Level 1), whereas repeated cycling lowers the mean residual to 0.78 and reallocates many specimens to Levels 2-3. The proposed framework yields quantified decision boundaries and a streamlined on-site testing protocol, providing a foundation for robust quality assurance standards.
Reference graph
Works this paper leans on
-
[1]
Ding, Y., Pang, Z., Lan, K., Yao, Y., Panzarasa, G., Xu, L., ... & Hu, L. (2022). Emerging engineered wood for building applications. Chemical Reviews, 123(5), 1843-1888
work page 2022
-
[2]
Kyjaková, L., Mandičák, T., & Mesároš, P. (2014). Modern methods of constructions and their components. Journal of Engineering and Architecture, 2(1), 27-35
work page 2014
-
[3]
Maqbool, R., Namaghi, J. R., Rashid, Y., & Altuwaim, A. (2023). How modern methods of construction would support to meet the sustainable construction 2025 targets, the answer is still unclear. Ain Shams Engineering Journal, 14(4), 101943
work page 2023
-
[4]
Rose, C., & Isaac, P. (2024). Reusing wood from demolition in mass timber products. The Structural Engineer: journal of the Institution of Structural Engineer, 102(6), 36-38
work page 2024
-
[5]
Godina, M., Gowler, P., Rose, C. M., Wiegand, E., Mills, H. F., Koronaki, A., ... & Shah, D. U. (2025). Strategies for salvaging and repurposing timber elements from existing build- ings in the UK. Journal of Cleaner Production, 489, 144629
work page 2025
-
[6]
Ghobadi, M., & Sepasgozar, S. M. (2023). Circular economy strategies in modern timber construction as a potential response to climate change. Journal of Building Engineering, 77, 107229
work page 2023
-
[7]
Tamke, M., Svilans, T., Huber, J. A., Wuyts, W., & Thomsen, M. R. (2024, June). Non - Destructive Assessment of Reclaimed Timber Elements Using CT Scanning: Methods and Computational Modelling Framework. In The International Conference on Net -Zero Civil Infrastructures: Innovations in Materials, Structures, and Management Practices (NTZR) (pp. 1275-1288)...
work page 2024
-
[8]
Shirmohammadi, M., Leggate, W., & Redman, A. (2021). Effects of moisture ingress and egress on the performance and service life of mass timber products in buildings: a review. Construction and Building Materials, 290, 123176
work page 2021
Show all 11 references
-
[9]
Niklewski, J., Isaksson, T., Frühwald Hansson, E., & Thelandersson, S. (2018). Moisture conditions of rain-exposed glue-laminated timber members: the effect of different detailing. Wood Material Science & Engineering, 13(3), 129-140
2018
-
[10]
Schmidt, E., & Riggio, M. (2019). Monitoring moisture performance of cross -laminated timber building elements during construction. Buildings, 9(6), 144
2019
-
[11]
S., Branco, J
Sousa, H. S., Branco, J. M., & Lourenço, P. B. (2016). A holistic methodology for probabil- istic safety assessment of timber elements combining onsite and laboratory data. Interna- tional Journal of Architectural Heritage, 10(5), 526-538
2016
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.