{"id":"03c0f7a9-c26b-48f2-aacb-bc5e02ef0a15","arxiv_id":"2502.02624","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A GAN-based model upsamples low-statistics simulated muography images to look like 21 to 31 days of sampling, while a segmentation model identifies rebar and tendon ducts and, the authors claim, removes vertical smearing artifacts.","lead":"The authors train a machine learning model to turn noisy, short-exposure muography scans of concrete into clear images equivalent to weeks of data collection, and a second model to identify rebar and tendon ducts. The results are shown only on simulated scans, so the practical impact depends on how well the simulation matches real concrete.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Headline results rest entirely on Geant4/EcoMug simulation; with no real-data validation, the '21/31-day equivalence' and Dice improvements remain unverified simulation artifacts.","rationale":"I agree with the reader that the load-bearing assumption is simulation fidelity. The paper is a credible proof of concept within the simulation: the dataset is large and varied, the architecture is standard and appropriate, and the authors are transparent about the simulation-to-real gap and the class-imbalance limitations (Section 4.2). The smearing-removal result is a legitimate property of the segmentation model trained on geometric ground truth, not an artifact of the upsampling target; however, its transfer to real data is exactly what remains unvalidated. I considered whether the 'smearing removal' claim is circular because both models share the same simulated geometry, but that is better described as a domain-shift risk than a logical circularity. No internal inconsistency or mathematical error was found. The appropriate verdict remains CONDITIONAL, pending the real-data test described above and, ideally, release of code and data. The concern does not move the reader's verdict, so UNCHANGED is used.","tokens_in":18074,"tokens_out":6434,"duration_ms":60951,"concrete_test":"Apply the trained upsampling and segmentation models to experimental muography data from the reinforced concrete reference block of Niederleithinger et al. (2021), or from a new controlled scan, and compare the 1-day-equivalent outputs against the measured longer-exposure images and known/CT ground truth. If the real-data SSIM/PSNR/Dice improvements are substantially below the simulated values (e.g., the 1-day upsampled SSIM does not reach the 21-day unaltered SSIM, or the segmentation Dice for rebar or ducts drops by more than ~0.1), the simulation-fidelity assumption fails and the strong claims must be downgraded. A secondary check is to inspect the simulated and real image noise statistics (e.g., per-voxel variance vs. exposure) to test whether PoCA/2 mm voxelization faithfully reproduces the detector response.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim—that cWGAN-GP upsampling allows 1-day scans to match 21-day perceptual quality and 31-day noise quality, and that segmentation removes z-smearing—is established solely on simulated data. Section 2.1 describes a Geant4/EcoMug model of the Glasgow MIS detector, with PoCA reconstruction and 2 mm voxelization, and all training targets are the corresponding 100-day simulated images (or, for segmentation, the known geometry). The evaluation metrics (SSIM, PSNR, Dice) are computed against these same simulated targets. If the simulation's muon flux, angular acceptance, multiple-scattering treatment, detector efficiency, or PoCA approximation differ from the physical MIS detector and real concrete, the learned mapping may simply be inverting simulation-specific noise and artifact statistics. The authors themselves state in Section 4.3 that 'it is important to verify these models on real-world data' and that testing on real data is future work. No real-data test, no error bars, no code/data release, and no comparison to classical denoising or upsampling baselines are provided. Consequently, the reported equivalence to 21/31 days and the Dice improvements are, at present, simulation artifacts rather than demonstrated field capabilities. The conditional acceptance is appropriate only if this external validity gap is explicitly disclosed, which the paper does, but the headline claims should be framed as simulation-based predictions.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript applies a conditional Wasserstein GAN with gradient penalty (cWGAN-GP) to upsample undersampled cosmic-ray muography images of reinforced concrete, and a second cWGAN-GP for semantic segmentation of features such as rebar grids and tendon ducts. Training and evaluation use a large Geant4/EcoMug simulation dataset of 700 concrete block designs with equivalent sampling times from 1 to 100 days. The authors report that 1-day upsampled images reach SSIM and PSNR levels equivalent to 21-day and 31-day unaltered images, and that segmentation Dice coefficients improve for several structural classes, with the segmentation model also learning to ignore z-plane smearing artifacts. The paper explicitly acknowledges in Section 4.3 that testing on real-world data is future work, but the abstract and conclusions present the headline results as demonstrated capabilities.","tokens_in":18339,"tokens_out":7626,"duration_ms":75639,"significance":"If the simulation-based results transfer to real detectors, the work offers a practical route to reducing muography acquisition times and automating feature detection in reinforced concrete. Strengths of the paper include a large and varied paired simulation dataset, the use of semantic segmentation as an interpretable evaluation tool, and candid discussion of limitations, with the authors identifying the need for real-data validation. The main weakness is that all headline quantities—the 21/31-day equivalence, Dice improvements, and smearing removal—are computed entirely within the Geant4/EcoMug simulation, with no real scans, no non-learning baselines, and no error bars; the claims are therefore currently simulation-based predictions rather than established capabilities.","major_comments":[{"comment":"The paper overstates simulation-based results as demonstrated capabilities. The abstract states that the results 'demonstrate significant improvements in both acquisition speed and image quality,' and the conclusions similarly report the 21-day SSIM and 31-day PSNR equivalences as findings. However, Section 4.3 concedes that 'it is important to verify these models on real-world data' and that testing on real data is future work. All metrics are computed against simulated 100-day targets generated by the same Geant4/EcoMug model. The headline claims should be explicitly framed as simulation-based predictions, with the external-validity caveat stated in the abstract and conclusions, not only in a later future-work paragraph.","section":"Abstract and Section 4.3"},{"comment":"No non-learning baseline is provided, so the cWGAN-GP is not shown to be responsible for the reported improvements. The equivalence claims (1-day upsampled SSIM matching 21-day unaltered input, and PSNR matching 31-day input) are relative only to the unaltered simulated inputs and the 100-day ground truth. A simple classical denoiser or smoother (e.g., Gaussian or median filtering, non-local means) applied to the 1-day inputs would provide a fair comparison and would test whether the gains are specific to the learned upsampling. Without such a baseline, the contribution of the GAN architecture is not isolated.","section":"Section 3.1, Figure 1"},{"comment":"The segmentation model is described as an 'independent evaluation metric,' but it is not independent of the simulation pipeline. It is trained on 100-day simulated muography images with geometric ground truths derived from the same Monte Carlo dataset that generated the upsampling training data. Dice improvements on upsampled images may therefore reflect that upsampled outputs are closer to the 100-day segmentation training distribution, rather than that true feature information has been recovered. Relatedly, the claim that segmentation 'mitigate[s] or entirely remove[s] z-plane smearing artifacts' is too strong: the segmentation model learns to ignore certain artifacts because its training labels contain no smearing, but the underlying muographic images are not deconvolved. Please rephrase the smearing-removal claim as a learned artifact-classification effect rather than an image-reconstruction result.","section":"Section 3.2, Figures 3 and 4"},{"comment":"Equation (6) is technically incorrect as written: it places the gradient penalty term in the generator loss, but in WGAN-GP the gradient penalty regularizes the discriminator (critic) and is part of the critic's loss, not the generator's. The text preceding Eq. (6) states that 'the generator's loss function, LG, is updated to include the gradient penalty,' which would not correspond to the standard WGAN-GP training procedure. Please correct the equation and the associated text, or clarify if a non-standard objective was used and explain why.","section":"Section 2.2, Eq. (6)"},{"comment":"The paper reports averages over the 6900-image test set without error bars, confidence intervals, or statistical tests. Point estimates such as '1-day upsampled images exhibited SSIM and PSNR scores of 0.88 and 37 dB' and the Dice differences in Figure 3 cannot be assessed for significance. Reporting standard deviations or confidence intervals across the test images, and where relevant a paired significance test for Dice improvements, would substantially strengthen the claims and allow comparison with future work.","section":"Sections 3.1 and 3.2, Figures 1 and 3"}],"minor_comments":[{"comment":"The dataset size description is internally inconsistent: the text says '70,000 500 × 500 images, each with 100 different versions,' which would total 7,000,000 images, but 700 unique blocks times 100 days yields 70,000 images. Please clarify the number of unique geometries, the number of cumulative-day images per geometry, and the total dataset size.","section":"Section 2, first paragraph after the sample list"},{"comment":"The values of the hyperparameters λpixel and λGP are not reported, and the learning-rate schedule is described only qualitatively ('reduced ... by an order of magnitude every 25 epochs'). Providing these values would improve reproducibility.","section":"Section 2.2"},{"comment":"There is a typo, 'sampled from from that of,' and other minor grammatical issues throughout. A thorough proofreading pass is recommended.","section":"Section 2.1.2"},{"comment":"The discussion attributes the low air-void Dice (0.1265) to class imbalance, but no class-frequency statistics are given. A table of pixel percentages per class in the training/test sets would make the imbalance argument quantitative.","section":"Section 4.2, air void discussion"},{"comment":"The figure captions and axis labels are occasionally redundant (e.g., repeating 'Dice-Sørensen Coefficient' on the left of all four subplots in Figure 3) and the Dice-difference subplots use inconsistent y-axis scales. Harmonizing the axes would ease cross-class comparison.","section":"Section 3, Figures 1 and 3"}],"recommendation":"major_revision","confidential_remarks":"The manuscript header indicates it was already accepted by an MDPI journal, but I reviewed the arXiv version. The main risk is that the headline claims will be read as demonstrated field capabilities rather than simulation-based predictions; a major-revision request for reframing, additional baselines, and statistical reporting is appropriate for a journal submission at this stage."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Worth a look if you work on muography or on simulation-to-real transfer in imaging. The paper shows, within a Geant4/EcoMug simulation, that a cWGAN-GP can make 1-day muography images look like 21-31 day data on SSIM/PSNR, and that a U-Net segmentation model trained on clean geometry can pick out rebar and tendon ducts from the upsampled images. The catch: everything is simulated; there is no real detector data anywhere in the validation, a point the authors themselves acknowledge in Section 4.3.\n\nWhat is actually new: this is, as far as I know, the first application of GAN-based upsampling to muographic images of concrete infrastructure. The architecture is off-the-shelf pix2pix/WGAN-GP, but the dataset is large and thoughtfully varied: 700 simulated blocks with randomized rebar grids, tendon ducts, voids, and 'unknown' objects, each rendered at 100 sampling times. Using semantic segmentation as an evaluation tool is a good idea—it gives per-feature Dice scores and shows that the upsampler helps tendon ducts most, while the segmentation model ignores z-smearing shadows because it was trained on clean geometric labels.\n\nThe soft spots are the usual ones for simulation-only ML papers. No real-data validation (noted), no error bars on the averaged SSIM/PSNR/Dice curves, and no comparison to a non-learning baseline like total variation or non-local means. Without that baseline you cannot tell how much of the gain is the GAN as opposed to any reasonable denoiser. The circularity concern is real but not damning: the segmentation model learned to reproduce clean geometry, so its 'smearing removal' is expected rather than surprising; calling it 'unexpected' is an overstatement. Also, no code or data release, just 'available upon request.'\n\nThe central claim, though, holds up as a simulation result. The authors are explicit that real-world testing is future work and they list sensible next steps (3D context, class weighting for voids). If you need a citable example of ML applied to muography for concrete NDT, this is a reasonable one.\n\nMy recommendation: send it to peer review. It is not a desk reject. I would ask the authors for a baseline comparison, error bars, and a release of code/data so the simulation results can be independently checked.","headline":"A clean simulation-only proof of concept for GAN-based muography upsampling, with the natural caveat that the headline numbers remain unverified until someone runs it on real detector data.","tokens_in":18892,"tokens_out":2569,"would_cite":true,"duration_ms":28869,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A one-day muography scan of concrete can be upsampled to the image quality of a three-week scan, and a second network can erase the smearing artifacts that hide rebar and tendon ducts.","keywords":["muography","non-destructive testing","machine learning","image processing","upsampling","semantic segmentation","simulation","cosmic rays"],"falsifier":"Take a real reinforced-concrete specimen, run a one-day muography acquisition and then continue the same scan to 21 or 31 days, apply the trained upsampler to the one-day data, and compare the upsampled image to the actual long-exposure image using SSIM and PSNR; if the simulated equivalence does not reproduce on real detector data, the central claim collapses.","tokens_in":17849,"feed_emoji":"🏗️","tokens_out":10689,"duration_ms":95770,"temperature":0.7,"pith_summary":"The paper tries to establish that machine learning can make muography—a slow, noisy imaging method that uses cosmic-ray muons to see inside concrete—usable for inspecting bridges and other infrastructure. The authors train a conditional Wasserstein generative adversarial network with gradient penalty to turn severely undersampled scans into clean images, reporting that a simulated one-day scan can be upsampled to match the perceptual quality (SSIM) of a 21-day scan and the noise level (PSNR) of a 31-day scan. A second segmentation network trained on the same simulated data then shows that the cleaned images preserve and sometimes improve the visibility of rebar grids and tendon ducts, with Dice scores of 0.8174 and 0.8663 on those classes. The segmentation network also learns to ignore the z-direction smearing that normally obscures object boundaries in muography. If the results transfer from simulation to real detectors, inspection times could drop from weeks to days.","feed_headline":"One-day muon scans upsampled to match three weeks of quality","feed_subtitle":"The same network sharpens rebar and tendon ducts in concrete and erases the smearing that limits muography.","key_machinery":"The load-bearing mechanism is a paired-image conditional generative adversarial network: a generator with an encoder–decoder structure and skip connections turns a low-sampling-time muography slice into a 100-day image, while a discriminator plus a per-pixel L1 loss drives the output toward the simulated ground truth. A Wasserstein distance with gradient penalty stabilises training, and each input is re-sampled at a random equivalent sampling time each epoch so the model learns to upsample across the full range. A second cWGAN-GP is trained for five-class semantic segmentation on 100-day images only, using a combined cross-entropy and Dice loss, so it can evaluate the upsampler's feature-level effects without being biased by upsampled outputs. The training images come from Monte Carlo simulation of a scintillating-fibre muon tracker, with muon tracks reconstructed by the point-of-closest-approach algorithm and voxelised at 2 mm into 500×500 X–Y slices.","core_discovery":"On the paper's own terms, the central discovery is that a two-stage deep-learning pipeline can substantially undo the two main defects of muon scattering tomography of concrete: sparse statistics from short acquisition times and smearing along the detector-normal direction from the inverse imaging problem. Trained on a large Monte Carlo dataset of reinforced concrete blocks with known geometry, the upsampling model maps images corresponding to 1–99 days of muon exposure onto a 100-day ground truth. Averaged over the 6900-image test set, one-day inputs reach SSIM 0.88 and PSNR 37 dB after upsampling, matching a 21-day raw image perceptually and a 31-day raw image in noise; improvement shrinks as sampling time grows and converges near 50–85 days. The segmentation model, trained only on 100-day images, gives Dice coefficients of 0.8174 for rebar grids and 0.8663 for tendon ducts, and it learns to classify z-smear shadows as background, removing some of the worst artifacts. The authors frame the work as a step toward making muography practical for reinforced-concrete non-destructive evaluation.","pith_inferences":["If the simulation-to-real transfer holds, the same upsampler could shorten routine muography inspections of concrete from weeks to a single day, and the segmentation output could be used directly for locating rebar and tendon ducts.","The reported 'one-day equals 21/31 days' equivalence is measured against a simulated 100-day target; on real data the equivalence could shift, so the fairest check is to compare the upsampled one-day reconstruction with a genuinely long real acquisition of the same block.","Because the segmentation model already removes z-smear shadows, a single network trained directly on geometry ground truths might do both denoising and feature location, avoiding the small mismatch the authors observe between upsampled outputs and their training distribution.","A class-weighted loss would likely improve the very low air-void Dice score (0.1265), which currently limits the method as a defect-detection tool despite its success on rebar and ducts."],"forward_implications":["On the 6900-image test set, upsampled one-day inputs reach an average SSIM of 0.88 and PSNR of 37 dB, matching the perceptual quality of raw 21-day images and the noise level of raw 31-day images.","Upsampling improves segmentation most for tendon ducts and rebar at low sampling times, with Dice differences shrinking to near zero as equivalent sampling time reaches 50–70 days.","The segmentation network learns to ignore z-direction smearing shadows in X–Y slices, so it can report rebar at its true location even when the muography image shows a shadowed grid.","The benefit of upsampling is feature-dependent: thin 8–10 mm rebar can be partially washed out between 20 and 85 days, and air voids remain poorly detected, with Dice 0.1265 on 100-day inputs.","For PSNR, the upsampler converges with raw inputs only around 80–85 days, because it smooths away the exact pixel-level noise fluctuations present in the 100-day ground truth."],"supporting_citations":[{"why":"Provides the reference experimental muography scan of a reinforced concrete block that motivates the simulation scenario and the planned real-data validation.","marker":"[5]"},{"why":"Supplies the Monte Carlo simulation toolkit used to generate the concrete-block training images.","marker":"[12–14]"},{"why":"Generates the cosmic-ray muon events with realistic momentum, charge ratio, and angular distributions.","marker":"[15]"},{"why":"Supplies the point-of-closest-approach reconstruction that converts simulated muon tracks into scattering-angle volumes.","marker":"[24]"},{"why":"Provides the encoder–decoder generator design with skip connections that preserves detail in the upsampled images.","marker":"[26]"},{"why":"Defines the scintillating-fibre tracker geometry and readout that the simulations model.","marker":"[30]"},{"why":"Introduces the conditional GAN image-to-image translation formulation that the upsampling model builds on.","marker":"[32]"},{"why":"Supplies the Wasserstein distance used in the adversarial loss to stabilise training.","marker":"[33]"},{"why":"Adds the gradient penalty that enforces the Lipschitz constraint in the Wasserstein discriminator.","marker":"[34]"},{"why":"Provides the structural similarity index used to compare upsampled images with longer-exposure images.","marker":"[35]"}],"fun_headline_variants":["AI upscales 1-day muon scans to 21-day clarity","Deep learning upgrades muon scans: 1 day for 3 weeks of quality","Muography AI sharpens 1-day scans to match 3-week detail","Neural network removes muon smear and boosts scan resolution"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that Monte Carlo simulations of a muon tracker looking at concrete blocks are faithful enough to real cosmic-ray muography that a model trained on simulated 100-day targets will behave the same way when applied to real scans of reinforced concrete; the authors themselves note in Section 4.3 that validation on real data is future work.","fun_headline_variants_meta":{"raw":{"variants":["AI upscales 1-day muon scans to 21-day clarity","Deep learning upgrades muon scans: 1 day for 3 weeks of quality","Muography AI sharpens 1-day scans to match 3-week detail","Neural network removes muon smear and boosts scan resolution"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000674,"raw_usage":{"total_tokens":3144,"prompt_tokens":1097,"completion_tokens":2047,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":713,"completion_tokens_details":{"reasoning_tokens":1967}},"tokens_in":713,"tokens_out":2047,"duration_ms":15012,"temperature":1.0,"reasoning_tokens":1967,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-09T12:26:49.732359+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a real reinforced-concrete specimen, run a one-day muography acquisition and then continue the same scan to 21 or 31 days, apply the trained upsampler to the one-day data, and compare the upsampled image to the actual long-exposure image using SSIM and PSNR; if the simulated equivalence does not reproduce on real detector data, the central claim collapses.","supporting_citations":[{"cited_title":"Muon Tomography of the Interior of a Reinforced Concrete Block: First Experimental Proof of Concept","cited_arxiv_id":null,"evidence_quote":"Provides the reference experimental muography scan of a reinforced concrete block that motivates the simulation scenario and the planned real-data validation."},{"cited_title":"EcoMug: An Efficient COsmic MUon Generator for cosmic-ray muon applications","cited_arxiv_id":null,"evidence_quote":"Generates the cosmic-ray muon events with realistic momentum, charge ratio, and angular distributions."},{"cited_title":"Cosmic Ray Muon Radiography","cited_arxiv_id":null,"evidence_quote":"Supplies the point-of-closest-approach reconstruction that converts simulated muon tracks into scattering-angle volumes."}],"review_version":1}