{"id":"01a4d19c-f568-4867-8431-32e768a675a9","arxiv_id":"2506.10606","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Combining geometric pore descriptors with pore network model predictions via fitted power laws improves in-sample prediction of CFD-computed air flow through paper cutouts.","lead":"The paper fits power-law regressions that combine geometric descriptors of the pore space with pore network model flow predictions to approximate full CFD air flow simulations. The combined models improve in-sample agreement and identify which morphological features the network model misses, but the dataset is too small to confirm the improvements out of sample.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"In-sample R^2 with up to six parameters on 11–12 cutouts does not establish 'predictive power'; the central claim needs leave-one-out validation before descriptor-relevance conclusions can be trusted.","rationale":"The reader's CONDITIONAL verdict is appropriate, and the main condition should be out-of-sample validation. The reader's stated weakest assumption is that CFD fluxes are an adequate ground truth despite the factor 4–5 offset from Gurley measurements. That is a real concern for external validity, but it is secondary to the internal statistical problem: the paper's own evidence for 'improved predictive power' is based on the same data used for fitting. Even if CFD were accepted as the ground truth, the reported R^2 values are training fits with up to six parameters on 11–12 points, so they cannot support the predictive claim or the descriptor-relevance ranking. This is why I focus on leave-one-out cross-validation as the decisive check. The reader's rationale does mention the lack of hold-out validation and the small sample size, so there is partial agreement, but the weakest_assumption field emphasizes CFD rather than in-sample evaluation. I agree with the reader's CONDITIONAL verdict because the paper is an honest empirical study whose conclusions can be rescued by validation; if the LOOCV results fail, the verdict should move toward REJECT of the strong predictive claim.","tokens_in":23039,"tokens_out":5243,"duration_ms":63929,"concrete_test":"Perform leave-one-out cross-validation for each sample separately and for the joint dataset, for all models v(1), v(2), v(2,1), v(2,2), v(3), v(4), v(5), v(6). For each left-out cutout, fit the model on the remaining cutouts using the same log-linear least-squares procedure (Section 5.2) and predict the left-out log(v_CFD). Then compute leave-one-out R^2 and MAPE using Eq. (27). Report these values alongside the in-sample values and the resulting ranking of models and descriptors. If the leave-one-out R^2 for v(4) or v(5) in the compressed sample does not remain substantially above that for v(1) or v(3), or if the tortuosity-vs-surface-area ordering reverses, the claim of improved predictive power is not supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The abstract claims that combining PNM fluxes with geometric descriptors 'significantly improves the predictive power' of PNM, and Section 6 draws sample-specific conclusions about which descriptors matter (tortuosity for uncompressed, surface area for compressed). The evidence for these claims is entirely in-sample: Section 5.2 explicitly states that the same data are used to fit and to evaluate R^2 and MAPE, with 12 cutouts for uncompressed paper and 11 for compressed paper after one cutout is excluded (Section 2.1). Some models have nearly as many parameters as data points: v(2) has 5 coefficients, v(5) and v(6) have 4–6 coefficients. Under these conditions, the reported improvements (e.g., R^2 rising from 0.39 to 0.89 when porosity is added to PNM in the compressed sample) can be dominated by overfitting rather than by a genuine structure-property relationship. The comparison between v(2,1) and v(2,2) that identifies tortuosity versus surface area as sample-specific drivers relies on R^2 differences of 0.01–0.02 (0.70 vs. 0.71; 0.97 vs. 0.98), which are within the noise expected from fitting several competing models to 11 points. Because the central claim is explicitly about predictive power, the absence of any out-of-sample check is the most load-bearing weakness: if the ranking of descriptors changes under cross-validation, the paper's headline conclusions do not hold, regardless of whether CFD is a perfect ground truth.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper addresses prediction of laminar air flow through thin paper sheets from μ-CT images, using two samples (uncompressed and calendered paper). The authors compute volume fluxes per cutout with CFD, approximate them with a pore network model (PNM), and then fit six power-law regression models that combine PNM fluxes with geometric descriptors (porosity, specific surface area, geodesic tortuosity mean/standard deviation, and median pore radius). Models are fitted separately for each sample and jointly for both samples, and are evaluated by in-sample R² and MAPE on log-transformed fluxes. The main findings are that adding geometric descriptors improves agreement with CFD relative to PNM alone or porosity alone, that the most useful descriptor differs between samples (geodesic tortuosity for uncompressed paper, specific surface area for compressed paper), and that a high correlation between two descriptors does not imply that one is redundant. The manuscript explicitly acknowledges the empirical nature of the regressions and the use of the same data for fitting and evaluation.","tokens_in":23370,"tokens_out":5010,"duration_ms":56144,"significance":"If the reported structure-property relationships survive out-of-sample testing, the work would be a practically relevant step toward inexpensive prediction of paper air permeance from tomographic data. The paper is transparent about the empirical status of the fitted laws, publishes all fitted coefficients in the supplementary tables, and carefully checks the sensitivity of PNM predictions to the conduit shape. The claim that descriptor correlation does not imply redundancy is interesting and falsifiable. However, the central evidence currently consists of in-sample fits to 11–12 cutouts with up to six parameters, so the significance of the descriptor-ranking conclusions depends on the cross-validation requested below.","major_comments":[{"comment":"The central claim that combining PNM fluxes with geometric descriptors 'significantly improves the predictive power' rests entirely on in-sample R² and MAPE. The paper explicitly states in §5.2 that the same data are used to fit and evaluate, and with 12 cutouts for uncompressed paper and 11 for compressed paper, models such as v(5) and v(6) contain up to six coefficients, so in-sample improvement is expected even if additional predictors are pure noise. In particular, the sample-specific conclusions about tortuosity versus surface area in §6.1 are based on R² differences of 0.70 vs. 0.71 and 0.97 vs. 0.98, which are well within the noise of fitting several competing models to roughly 11 points. I request leave-one-out cross-validation (or an equivalent out-of-sample protocol) for all models, with cross-validated R² and MAPE, and standard errors or confidence intervals for the fitted exponents. Without this, 'predictive power' is not established and the descriptor-relevance ranking may not transfer to new cutouts.","section":"Section 5.2, Eq. (27); Section 6.1; Tables S1–S3"},{"comment":"The paper treats CFD-computed fluxes as ground truth for air flow even though v_CFD and experimental Gurley fluxes differ by a factor of four to five. The regression prefactors can absorb a constant multiplicative offset, but the manuscript does not demonstrate that the offset is constant across the morphological range; if the CFD/experiment discrepancy depends on porosity or connectivity, the fitted descriptor exponents target a biased quantity. I ask the authors to quantify the CFD-versus-experiment relationship beyond medians and quartiles, for example with a pointwise scatter plot or a ratio-versus-porosity plot, or to explicitly restrict all claims to 'CFD flux' rather than measured air permeance. The exclusion of one compressed cutout for an 'implausibly large' CFD porosity (§2.1) is a related robustness concern, because a mesh artifact in one cutout raises the possibility of systematic mesh-induced bias in the remaining cutouts.","section":"Section 3.1, Figure 2; Section 2.1"},{"comment":"The targeted cutout design (eight cutouts near the mean porosity, two dense, two open) means the reported R² values are not estimates for a random sample of the paper sheet. The authors acknowledge in §4.2 that this design boosts the importance of other descriptors in explaining flow variation, but the in-sample model comparisons are still affected: the apparent success of a descriptor can depend on this stratified design. Please report results under a design-aware analysis, for example weighted R² or leave-one-group-out cross-validation where the three porosity strata are the groups, and state more cautiously that the descriptor ranking is specific to this stratified set.","section":"Section 2.1; Section 4.2"},{"comment":"The generalization analysis in §6.3 fits a single model on the combined data and evaluates on the same two samples; this is still in-sample evaluation and cannot be read as evidence that the relationships generalize to other paper grades. The conclusion that 'the decrease in MAPE between v(1) and v(4) is consistent between both samples' should be rephrased as a statement about the fitted data, not about predictive generalization, unless an out-of-sample protocol is added.","section":"Section 6.3, Figure 10"}],"minor_comments":[{"comment":"There appears to be a parenthesis error in the MAPE expression; it should read |log(v_CFD,k) - log(v_k)| / |log(v_CFD,k)| (or the percentage version).","section":"Equation (27)"},{"comment":"The caption for Figure 5 assigns blue diamonds to the compressed sample and orange circles to the uncompressed sample, while the text in §3.2.3 assigns the colors the other way; please make these consistent.","section":"Figure 5 and Section 3.2.3"},{"comment":"The word 'particurlarly' should be 'particularly'.","section":"Section 2.1"},{"comment":"The phrase 'mean absolute percentage MAPE' should be 'mean absolute percentage error MAPE'.","section":"Section 6.1"},{"comment":"After Eq. (26), the notation for the predicted value log(v) is used without distinguishing it from the same symbol in Eq. (25); a hat or subscript would improve readability.","section":"General notation"}],"recommendation":"major_revision","confidential_remarks":"The paper is honest about its limitations, but the framing 'predictive power' overstates what in-sample R² values on 11–12 cutouts can support. The requested leave-one-out cross-validation is, in my view, the decisive test: if the descriptor ranking changes under cross-validation, the paper's headline conclusions should be substantially weakened. I would not require experimental validation of the CFD ground truth for a revised version, but the multiplicative-offset assumption should be stated and justified more explicitly. The manuscript is within scope for the journal and the methods are reproducible in principle, so I see major_revision as appropriate rather than rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper is a careful empirical study that fits power-law regressions combining PNM predictions with geometric descriptors from micro-CT to reproduce CFD fluxes for two paper samples. What is genuinely new is the sample-specific descriptor relevance—tortuosity matters for the uncompressed paper, specific surface area for the compressed one—and the point that highly correlated descriptors can still add information in a combined prediction. The methodology is clearly described, the conduit-shape sensitivity analysis is a useful robustness check, and the authors are admirably transparent about the empirical nature of the fitted exponents.\n\nThe soft spot is the central claim. The abstract says the approach \"significantly improves the predictive power\" of the PNM, but every R² and MAPE is computed on the same data used to fit, with 11–12 cutouts and up to six coefficients. That is in-sample fitting, not prediction. The descriptor-relevance comparison rests on R² differences of 0.01–0.02, which are noise at this sample size. There is no cross-validation, no confidence intervals, and the post-hoc exclusion of one cutout weakens the compressed-sample results. The factor-four-to-five offset between CFD and experimental Gurley fluxes also matters: the fitted prefactors absorb that scale, so the model could be learning an offset plus noise rather than pore-space physics. The authors acknowledge most of this in Section 5.2, but the abstract and conclusions do not carry the same caveat.\n\nWhat the paper does well is show a plausible route: PNM plus a few cheap geometric descriptors might be enough to approximate CFD-level fluxes, and the joint fit across both samples gives a weak but real indication that some improvements do not transfer. That is useful evidence, just not proof of predictive power.\n\nThis paper is for porous-media engineers and image-analysis researchers who work on structure-property links and want an inexpensive correction to pore-network simulations. It deserves a serious referee, but my recommendation is major revision: add leave-one-out or k-fold cross-validation, report uncertainty on coefficients and on the R² differences, and reframe the abstract to claim an in-sample improvement rather than predictive power. The CFD-versus-experiment offset should also be discussed as a limitation rather than waved away.","headline":"A careful, transparent in-sample demonstration that geometric descriptors can correct PNM flux predictions, but the 'predictive power' claim is overreaching without any out-of-sample validation.","tokens_in":23952,"tokens_out":1642,"would_cite":true,"duration_ms":19270,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":["47.56.+r"],"model":"deepseek-v4-flash","headline":"Combining cheap pore-network flow simulations with a few geometric pore-space descriptors yields substantially better predictions of air flux through paper, with different descriptors decisive for uncompressed and calendered sheets.","keywords":["air permeance","pore network model","computational fluid dynamics","geometric descriptors","power-law regression","calendered paper","μ-CT imaging","tortuosity"],"falsifier":"Fit the same six regressions to experimentally measured Gurley fluxes instead of CFD fluxes, on cutouts aligned with the measured sheet positions. Because the prefactor $c_0$ can absorb the four-to-fivefold scale offset, the central claim predicts that the descriptor rankings should survive; if geodesic tortuosity stops helping for uncompressed paper or specific surface area stops helping for compressed paper, the CFD-as-ground-truth premise is the weak link. Repeating the CFD on a finer mesh with compressible flow would test whether the four-to-fivefold offset and the excluded cutout's implausible porosity disappear.","tokens_in":22748,"feed_emoji":"💨","tokens_out":10653,"duration_ms":121542,"temperature":0.7,"pith_summary":"Air flow through paper is usually predicted either from a single easy-to-measure quantity like porosity or from expensive simulations, and this paper asks whether a cheap pore-network simulation corrected by a few geometry measurements can do much better. Using μ-CT images of uncompressed and calendered paper, it compares CFD-computed fluxes with six power-law regression models built from pore-network flux and descriptors such as geodesic tortuosity, specific surface area, and median pore radius. The central result is that adding the right descriptor markedly improves prediction accuracy, and the right descriptor is sample-specific: geodesic tortuosity for uncompressed paper ($R^2$ from 0.62 to 0.71), specific surface area for compressed paper ($R^2$ from 0.87 to 0.98). A second finding is that a descriptor strongly correlated with porosity can still add real predictive value, so high correlation does not imply redundancy.","feed_headline":"Pore shape data sharpen paper air-flow predictions","feed_subtitle":"A power-law mix of morphology and pore-network physics beats porosity alone and pinpoints what simulations miss.","key_machinery":"The central machinery is a family of six power-law regression models of the form $v = c_0 x_1^{c_1} \\cdots x_n^{c_n}$, made linear by taking logarithms and fitted by least squares to CFD fluxes. The fitted exponents on porosity $\\varepsilon$, specific surface area $S$, mean and standard deviation of geodesic tortuosity $\\mu(\\tau)$ and $\\sigma(\\tau)$, median pore radius $r_{\\max}$, and the PNM flux $v_{\\mathrm{PNM}}$ are the quantitative measure of how much each quantity corrects the pore-network prediction. The PNM flux itself comes from a graph representation of the pore space whose local conductances are computed with an analytical conduit formula, so the regressions inherit the physics of the network model and add morphology on top.","core_discovery":"On the paper's own terms, the claim is that pore-network (PNM) flux predictions can be corrected into close agreement with CFD reference fluxes by multiplying them by power-law factors in geometric descriptors of the pore space. For descriptor-only regressions, the uncompressed sample improves from $R^2=0.62$ with porosity alone to $R^2=0.71$ when tortuosity statistics are added; the compressed sample improves from $R^2=0.87$ with porosity alone to $R^2=0.98$ when specific surface area is added. When PNM flux is the base predictor, porosity as an extra factor raises the compressed-sample fit from $R^2=0.39$ to $R^2=0.89$. The paper further demonstrates that median pore radius remains useful even though it is strongly correlated with porosity, and that jointly fitted models lose the sample-specific gains, indicating that the corrections encode real microstructural differences rather than generic trends.","pith_inferences":["Because the fitted prefactor absorbs the four-to-fivefold CFD/experiment gap, the descriptor exponents are likely more transferable to real paper than the absolute predicted fluxes; refitting the same models to Gurley permeances would test this directly.","The same recipe—one fast network simulation plus a handful of image-based descriptors—could serve as a manufacturing proxy for local permeance, replacing case-by-case CFD on paper, filters, and gas-diffusion layers.","If compression shifts the decisive descriptor from tortuosity to surface area in paper, analogous processing-induced anisotropy in other fibrous sheets should be detectable by the same regression diagnostic before any tomographic segmentation is optimized."],"forward_implications":["Corrected PNM fluxes give a computationally cheap way to map local air-permeance variations across many cutouts of a paper sheet, which full CFD cannot cover at the same cost.","The optimal correction descriptor is grade-specific, so a single universal porosity-permeability curve should not be assumed across different paper grades.","The compressed sample's poor PNM-only fit ($R^2=0.39$) and its rescue by porosity identify pore-throat and surface-area features that the simplified conduit model misses.","Strong correlation with porosity does not make a descriptor redundant, so feature screening by correlation alone can discard useful morphology information."],"supporting_citations":[{"why":"Supplies the μ-CT image data and the uncompressed/compressed paper samples on which all simulations and descriptors are computed.","marker":"[16]"},{"why":"Provides the earlier comparison showing CFD and experimental fluxes differ by a factor of four to five, the basis for treating CFD as ground truth.","marker":"[20]"},{"why":"Documents the altered correlation structure of pore-space descriptors after compression, motivating the sample-specific descriptor analysis.","marker":"[11]"},{"why":"Supplies the pore-network extraction implementation that partitions the μ-CT pore space into graph vertices and edges.","marker":"[33]"},{"why":"Supplies the pore-network solver that computes PNM fluxes from the graph via mass-balance equations.","marker":"[37]"},{"why":"Provides the analytical conduit-conductance formula used to compute local flow through the network conduits.","marker":"[40]"},{"why":"Supports the claim that collinearity among regression descriptors does not by itself spoil fit quality.","marker":"[51]"}],"fun_headline_variants":["Physics plus pore geometry sharpens airflow forecasts","Morphology boosts pore-network airflow predictions","Pore shape data refine physics-based airflow model","Richer pore descriptors beat porosity alone for airflow"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that CFD-simulated fluxes are the correct target, even though they run four to five times below the measured Gurley fluxes and one cutout was excluded because its simulated porosity looked like a mesh artifact.","fun_headline_variants_meta":{"raw":{"variants":["Physics plus pore geometry sharpens airflow forecasts","Morphology boosts pore-network airflow predictions","Pore shape data refine physics-based airflow model","Richer pore descriptors beat porosity alone for airflow"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00018,"raw_usage":{"total_tokens":1275,"prompt_tokens":888,"completion_tokens":387,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":504,"completion_tokens_details":{"reasoning_tokens":330}},"tokens_in":504,"tokens_out":387,"duration_ms":5065,"temperature":1.0,"reasoning_tokens":330,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T04:22:04.215045+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Fit the same six regressions to experimentally measured Gurley fluxes instead of CFD fluxes, on cutouts aligned with the measured sheet positions. Because the prefactor $c_0$ can absorb the four-to-fivefold scale offset, the central claim predicts that the descriptor rankings should survive; if geodesic tortuosity stops helping for uncompressed paper or specific surface area stops helping for compressed paper, the CFD-as-ground-truth premise is the weak link. Repeating the CFD on a finer mesh with compressible flow would test whether the four-to-fivefold offset and the excluded cutout's implausible porosity disappear.","supporting_citations":[{"cited_title":"Neumann, E","cited_arxiv_id":null,"evidence_quote":"Supplies the μ-CT image data and the uncompressed/compressed paper samples on which all simulations and descriptors are computed."},{"cited_title":"Leitl, E","cited_arxiv_id":null,"evidence_quote":"Provides the earlier comparison showing CFD and experimental fluxes differ by a factor of four to five, the basis for treating CFD as ground truth."},{"cited_title":"Neumann, P","cited_arxiv_id":null,"evidence_quote":"Documents the altered correlation structure of pore-space descriptors after compression, motivating the sample-specific descriptor analysis."},{"cited_title":"Gostick, Z","cited_arxiv_id":null,"evidence_quote":"Supplies the pore-network extraction implementation that partitions the μ-CT pore space into graph vertices and edges."},{"cited_title":"Gostick, M","cited_arxiv_id":null,"evidence_quote":"Supplies the pore-network solver that computes PNM fluxes from the graph via mass-balance equations."},{"cited_title":"Akbari, D","cited_arxiv_id":null,"evidence_quote":"Provides the analytical conduit-conductance formula used to compute local flow through the network conduits."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supports the claim that collinearity among regression descriptors does not by itself spoil fit quality."}],"review_version":1}