{"id":"9ae7a59e-e5a4-4006-b2cd-cf0965733d36","arxiv_id":"2506.11061","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"high","formal_verification":"none","parameter_count":4,"one_line_summary":"A Bayesian classifier with two binary inputs assigns moisture-exposed plywood to three reuse levels, but the supporting data are internally inconsistent and the evaluation is circular.","lead":"This paper builds a Bayesian model to classify whether engineered timber can be reused after getting wet, based on a small set of lab tests. The authors say a single wet-dry cycle keeps 70% of specimens above a 0.90 residual-performance threshold, but the data and methods have internal inconsistencies that weaken the result.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Two-input predictive claim is not tested: reuse labels are constructed from the same experimental variables used as predictors, leaving reported accuracy as an in-sample, partly tautological result.","rationale":"The reader's weakest assumption centres on the residual-performance metric R as the operational definition of reusability, including the weighting and thresholds. My review identifies a related but distinct load-bearing problem: the reuse labels are constructed from the same cycle and orientation variables that are then used as predictors, and tau2 is inferred from the same R values that define the labels. This makes the two-input predictive claim at least partly tautological and the reported accuracy an in-sample fit rather than a predictive validation. The reader also flags that tau2 is inferred from the same residual values and that accuracy is computed on training data, so there is substantial overlap; I would frame it as the label construction being endogenous to the predictors. A held-out or externally labelled test would directly settle whether the two-input rule generalizes. Because the reader already recommends REJECT and my concern reinforces that rejection rather than altering it, the verdict should remain unchanged. I do not raise objections beyond this central issue, such as the specimen-count inconsistencies, because they are secondary to the argument's logical structure, though they would matter for any revision.","tokens_in":6500,"tokens_out":2138,"duration_ms":21609,"concrete_test":"Hold out one full cycle-orientation condition (e.g., all Cycle-1 cross-grain specimens) before fitting; infer tau2 and coefficients on the remaining specimens, then compute the confusion matrix on the held-out specimens. Compare against a trivial benchmark that assigns each held-out specimen to the majority reuse level of its cycle-orientation cell computed from the training set. If held-out accuracy is not materially above that cell-majority benchmark, or if the re-estimated tau2 shifts beyond its reported HDI across folds, the two-input rule is an artifact of in-sample label construction. Independently, re-derive labels from an external assessment (e.g., visual grading or residual capacity on a separate set of beams) and rerun the classifier on those labels.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that wet-cycle count and grain orientation alone reproduce the reuse-level distribution—is not supported by the evidence, because the level labels are generated from the very variables that serve as predictors. In Sec. 2.3, residual performance R_i is computed relative to orientation-matched controls (Eq. 1), and in Sec. 3.3 the lower threshold tau2 is inferred from the posterior distribution of R, after which R is thresholded to assign Levels 1/2/3. Cycle group and grain orientation are the only experimental manipulations (Sec. 2.1), and the multinomial model in Eq. 2 uses exactly these two variables. Thus the classifier is asked to predict a deterministic relabeling of the data used to fit it: a specimen's cycle group and orientation largely determine its R, and tau2 is learned from the same R values. The reported 67% accuracy (Sec. 3.5) is computed on the training data—posterior-mean logits on the same specimens—so it quantifies in-sample fit, not predictive reuse classification. The abstract's five-feature claim and the conclusion's 'reproduce the observed level distribution' overstate what the experiment can establish. Without an independent ground-truth label—e.g., expert grading or a separate destructive-validation set—the two-predictor rule is confounded with the label construction. The material mismatch (plywood in Sec. 2.1 vs. 'spruce CLT' in Sec. 5) further weakens external validity, but the circular label-predictor relationship is the most load-bearing issue.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper develops a Bayesian multinomial logistic model to classify the reusability of moisture-exposed engineered timber into three levels defined by a residual-performance metric R. R is computed as a weighted combination of retained flexural modulus and maximum load relative to orientation-matched controls (Eq. 1). The authors fix an upper threshold tau1 = 0.90, infer a lower threshold tau2 from the posterior, assign each specimen a reuse level by thresholding R, and then fit a two-predictor model using wet-cycle group and grain orientation. They report 67% classification accuracy, a multiclass Brier score of 0.432, and conclude that only two binary inputs are needed to reproduce the observed level distribution, positioning the framework as a streamlined on-site triage tool.","tokens_in":6780,"tokens_out":2519,"duration_ms":25447,"significance":"If the central claim were valid, the paper would offer a practically valuable contribution: a simple, two-variable probabilistic rule for triaging moisture-exposed engineered timber, with quantified uncertainty, in a domain where standards are lacking. The modelling machinery (horseshoe priors, NUTS, posterior predictive checks) is standard and appropriate for the intended inferential task. However, the significance is conditional on resolving a fundamental validation problem: the reuse labels are constructed from the same experimental variables that later serve as predictors, so the reported accuracy and probability estimates are in-sample re-descriptions of the label-construction rule rather than evidence of predictive capacity. The manuscript also contains multiple unresolved internal inconsistencies in the reported data and material identity, which further reduce confidence in the empirical basis.","major_comments":[{"comment":"The central claim that only wet-cycle count and grain orientation reproduce the observed level distribution is not supported because the reuse labels are derived from the same variables that are then used as predictors. In Eq. (1), R is computed relative to orientation-matched controls; in Sec. 3.3, tau2 is inferred from the posterior of R; and the labels L1/L2/L3 are assigned by thresholding R at tau1 = 0.90 and tau2 = 0.75. The multinomial model in Eq. (2) uses exactly cycle group and grain orientation as covariates. Consequently, the 67% accuracy and Brier score reported in Sec. 3.5 are computed on the training data and quantify in-sample fit, not predictive reuse classification. Without an independent ground-truth label (e.g., expert grading or a separate validation set), the two-predictor rule is confounded with the label construction and cannot support the abstract's and conclusions' claims.","section":"Secs. 2.3, 3.3, 3.5"},{"comment":"The descriptive statistics are internally inconsistent. The text in Sec. 3.2 reports n = 24 per group with standard deviations 0.03, 0.07, and 0.11, while Table 1 reports n = 10 with standard deviations 0.118, 0.097, and 0.268. Additionally, Sec. 3.3 gives class counts L1: 40, L2: 21, L3: 12, which sum to 73, while Sec. 5 states that flexural tests were performed on 72 specimens. These discrepancies affect the interpretation of the posterior proportions (0.55, 0.28, 0.17) and prevent a reader from reproducing the analysis.","section":"Sec. 3.2 and Table 1"},{"comment":"The material under study is identified inconsistently: Sec. 2.1 and Sec. 2.2 describe plywood specimens, while Sec. 5 states that the tests were conducted on '72 spruce CLT specimens.' CLT and plywood have different lamella orientations, thicknesses, and failure modes, so this is not a terminological triviality. The material identity is load-bearing for the external-validity claim that the framework applies to engineered timber in MMC reuse hierarchies.","section":"Secs. 2.1 and 5"},{"comment":"The abstract claims the model predicts reuse levels from five field-measurable features: density, moisture content, specimen size, grain orientation, and surface hardness. However, the methods section never describes measurements of density, moisture content at the time of testing, or surface hardness, and Eq. (2) includes only two predictors. The claim that horseshoe shrinkage 'retains' two predictors and 'attenuates' three others is therefore unsubstantiated, because the other three features are not present in the reported experimental design or model specification.","section":"Abstract and Sec. 3.4"}],"minor_comments":[{"comment":"The convergence diagnostic is written as 'R ≤ 1.01'; the standard notation is R-hat (or 'R-hat'), and the value should be reported with its definition to avoid confusion with the residual-performance metric R introduced in Eq. (1).","section":"Sec. 2.4"},{"comment":"The text uses 'HID' in the phrase 'Posterior 95% HID'; this should be 'HDI' (highest density interval), consistent with Sec. 3.3.","section":"Sec. 3.4 and Fig. 7"},{"comment":"The specimen preparation states that three replicates were prepared for each group and orientation, but Sec. 3.2 and Table 1 report group sample sizes of 24 or 10. The relation between the number of replicates and the reported n is not explained.","section":"Sec. 2.1"},{"comment":"The flexural results are reported as means without confidence intervals or measures of variability around the mean, which is important because the residual-performance metric R is constructed from these means and the subsequent Bayesian inference propagates uncertainty from the posterior only, not from the flexural measurement error.","section":"Sec. 3.1"}],"recommendation":"reject","confidential_remarks":"The paper has a potentially useful practical goal, but the core validation argument is circular, and the reported data contain unresolved numerical and material inconsistencies. In my view, the circularity cannot be fixed by a revision within the current scope: it requires new data with independent ground-truth labels or a genuinely held-out validation protocol, and the material identity must be clarified before the empirical claims can be assessed. I therefore recommend rejection rather than major revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Honestly, the paper is worth a look for its raw data but not for its headline claims. The useful new thing is a small set of flexural measurements on plywood that went through one or two soak-dry cycles, with separate results for long-grain and cross-grain. That is a real empirical contribution, and the idea of turning residual performance into a probabilistic reuse triage is sensible. The Bayesian multinomial logistic with horseshoe priors is not new, but applying it to this decision problem is a reasonable move.\n\nWhere it unravels is that the 'predictions' are not predictions. The reuse labels are created by thresholding the residual performance R, and R is computed relative to controls that share the same grain orientation. Then the model uses cycle group and orientation as predictors. Tau2 is inferred from the same R values, so the labels are a deterministic relabeling of the data the model is fit to. The reported 67% accuracy and Brier score are in-sample numbers; they don't show that a two-input rule can triage real timber. The claim that two binary inputs 'reproduce the observed level distribution' is close to tautological.\n\nThere are also plain inconsistencies: the material is called plywood in Methods and 'spruce CLT' in Conclusions; n is 24 in text but 10 in Table 1; class counts sum to 73 while the text says 72. No data or code is provided, so none of this can be checked. The abstract's five features are reduced to two in the paper, which is fine, but the framing overstates the evidence.\n\nOn the positive side, the paper honestly lists its limitations at the end, and the raw flexural data could be reused by someone doing a careful study with proper independent labels.\n\nMy take: a serious editor would send this to review to get the methodological circularity and inconsistencies on the record, but it is not acceptable as is. It would need a major revision with released data, out-of-sample validation, and a benchmark against simple threshold rules.","headline":"Useful raw moisture-cycle flexural data, but the central 'two-input prediction' claim is circular: the reuse labels are derived from the same variables used as predictors, so the reported accuracy is an in-sample fit.","tokens_in":7383,"tokens_out":2448,"would_cite":false,"duration_ms":21954,"reading_group":"maybe","serious_thinker":"no","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62F15","62J12","62P30"],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper argues that moisture-exposed engineered timber can be graded into three reuse levels from just wet-cycle count and grain orientation, with a Bayesian model supplying the probability for each level.","keywords":["engineered timber","reusability","moisture exposure","Bayesian multinomial logistic","horseshoe prior","residual performance","circular economy","modern methods of construction"],"falsifier":"A direct check would be to take a larger batch of engineered timber with known service histories, measure the two binary predictors, apply the paper's decision rule, and compare the assigned reuse levels against full destructive testing and long-term field performance; the claim would be weakened if specimens with identical wet-cycle count and grain orientation but different density, moisture content, or prior load history show materially different residual performance, because that would mean the two-variable model is omitting a predictive dimension.","tokens_in":6200,"feed_emoji":"🪵","tokens_out":5941,"duration_ms":55632,"temperature":0.7,"pith_summary":"This paper argues that the reusability of moisture-exposed engineered timber can be classified probabilistically rather than by case-by-case inspection. It builds a residual-performance score R that gives 80% weight to retained flexural stiffness and 20% to retained strength, then fits a Bayesian multinomial-logistic model to specimens that underwent zero, one, or two wet-dry cycles. The fitted model assigns each specimen a probability of falling into one of three Modern Methods of Construction (MMC) reuse levels, with the lower decision boundary inferred from data rather than fixed. The paper's headline finding is that two binary features, wet-cycle count and grain orientation, carry nearly all the predictive information, making on-site triage a matter of looking up a probability. If true, this would give the construction industry a quantitative, auditable basis for reuse decisions instead of conservative disposal.","feed_headline":"Two binary inputs predict whether wet timber can be reused","feed_subtitle":"A Bayesian classifier turns cycle count and grain orientation into a reuse probability, replacing case-by-case inspection.","key_machinery":"The load-bearing object is the residual-performance metric R_i = 0.8(E_i/E_0) + 0.2(sigma_max,i/sigma_0), computed against unexposed controls of the same grain orientation, combined with a hierarchical Bayesian multinomial logistic model. The model uses horseshoe priors to force irrelevant predictors' coefficients toward zero and treats the lower reuse-level boundary tau_2 as an unknown parameter with a Uniform(0.70, 0.85) prior, so the decision rule is inferred jointly with the regression coefficients through Markov-chain Monte Carlo using the No-U-Turn sampler. This setup lets the output be a probability per reuse level, not a single deterministic grade, and it is what allows the two-variable minimal predictor set to be discovered rather than assumed.","core_discovery":"The paper's central discovery is that a probabilistic classifier trained on destructive flexural tests can reproduce the observed distribution of moisture-damaged engineered timber into reuse levels using only two binary inputs: how many soaking-and-drying cycles a member has endured and whether the face grain runs parallel or perpendicular to the load. All other field-measurable candidates, density, moisture content, specimen size, and surface hardness, shrink toward zero under the horseshoe prior. The paper reports that a single wet-dry cycle leaves about 70% of specimens above the Level-1 threshold of R = 0.90, while two cycles pull the mean residual down to roughly 0.78 and move many specimens into lower levels; the inferred lower decision boundary is about 0.76 with a 95% HDI of 0.73-0.79. On this basis the paper claims to offer the first probabilistic framework for classifying moisture-exposed engineered timber within the MMC reuse hierarchy and proposes a three-zone decision rule based on the probability of Level-1 status.","pith_inferences":["A natural next step the paper does not run is a cost-benefit comparison: the model trades an off-by-one-level error rate of about 33% against the cost of destructive testing, and a decision-theoretic extension could optimise the P(Level 1) thresholds for a given tolerance for over- or under-grading.","Because the horseshoe prior selected only two categorical features, the framework is directly portable to other engineered wood products, but the coefficient values and thresholds would need refitting on glulam or LVL data before field use.","The paper's own text is internally inconsistent about the test material and sample counts, with the methods describing 50 mm plywood coupons and repeated 'three replicates' groups while Table 1 lists n=10 per group and the conclusion speaks of 72 spruce CLT specimens, so the demonstrated result should be read as small-coupon plywood evidence pending confirmation on structural-scale CLT."],"forward_implications":["If the two-input classifier generalises, on-site reuse triage reduces to recording cycle history and grain orientation and reading off P(Level 1), with no need for hardness or density meters.","Quantified decision boundaries make it possible to write a standard: redeploy at P >= 0.70, non-destructive check at 0.40-0.70, downgrade below 0.40.","The paper's estimate that replacing moisture-limit rules with this classifier would roughly double the volume of salvaged timber reused without structural downgrading follows directly from the fitted level proportions.","Every component gets an explicit uncertainty interval, so a specifier can state the probability that a given member is below the direct-reuse threshold rather than asserting a binary pass/fail."],"supporting_citations":[{"why":"Establishes the carbon-storage and emission-avoidance case for engineered wood that motivates the reuse imperative.","marker":"[1]"},{"why":"Describes upcycling reclaimed wood into mass secondary timber, the reuse context the classifier targets.","marker":"[4]"},{"why":"Reviews moisture-ingress effects on mass timber performance and service life, the degradation mechanism being modelled.","marker":"[8]"},{"why":"Documents moisture conditions in exposed glulam and motivates the need for quantitative exposure standards.","marker":"[9]"},{"why":"Is the industry guidance saying there are no simple rules for assessing moisture impact on timber-framed construction, the gap the paper fills.","marker":"[10]"},{"why":"Supplies the rationale for probabilistic safety assessment of timber combining onsite and laboratory data.","marker":"[11]"}],"fun_headline_variants":["Two binary inputs classify timber reuse after moisture","Cycle count and grain orientation predict timber reusability","After one wet-dry cycle, 70% of timber is reusable","Bayesian model trims five field inputs to two for reuse test","Moisture-exposed timber reuse predicted by two simple factors"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole classifier stands on the assumption that the weighted residual-performance score R, with 80% weight on stiffness and 20% on strength and thresholds at 0.90 and about 0.75, is the correct operational definition of reusability; if those weights or thresholds are wrong, or if laboratory soaking and oven-drying do not represent real service wetting, the probability labels have no fixed meaning.","fun_headline_variants_meta":{"raw":{"variants":["Two binary inputs classify timber reuse after moisture","Cycle count and grain orientation predict timber reusability","After one wet-dry cycle, 70% of timber is reusable","Bayesian model trims five field inputs to two for reuse test","Moisture-exposed timber reuse predicted by two simple factors"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000208,"raw_usage":{"total_tokens":1435,"prompt_tokens":1005,"completion_tokens":430,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":621,"completion_tokens_details":{"reasoning_tokens":347}},"tokens_in":621,"tokens_out":430,"duration_ms":4740,"temperature":1.0,"reasoning_tokens":347,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T12:45:11.114065+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A direct check would be to take a larger batch of engineered timber with known service histories, measure the two binary predictors, apply the paper's decision rule, and compare the assigned reuse levels against full destructive testing and long-term field performance; the claim would be weakened if specimens with identical wet-cycle count and grain orientation but different density, moisture content, or prior load history show materially different residual performance, because that would mean the two-variable model is omitting a predictive dimension.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Establishes the carbon-storage and emission-avoidance case for engineered wood that motivates the reuse imperative."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Describes upcycling reclaimed wood into mass secondary timber, the reuse context the classifier targets."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Reviews moisture-ingress effects on mass timber performance and service life, the degradation mechanism being modelled."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Documents moisture conditions in exposed glulam and motivates the need for quantitative exposure standards."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Is the industry guidance saying there are no simple rules for assessing moisture impact on timber-framed construction, the gap the paper fills."},{"cited_title":"S., Branco, J","cited_arxiv_id":null,"evidence_quote":"Supplies the rationale for probabilistic safety assessment of timber combining onsite and laboratory data."}],"review_version":1}