{"id":"a23b9017-da7b-4a6c-a551-94370d100b83","arxiv_id":"2608.01746","paper_version":1,"verdict":"CONDITIONAL","confidence":"LOW","novelty_score":5.0,"correctness_risk":"high","formal_verification":"none","parameter_count":4,"one_line_summary":"Non-linear multi-exposure fusion, applied inside ptychography, is claimed to improve reconstruction robustness even though it breaks the standard linear intensity model.","lead":"The paper shows that merging multiple diffraction exposures with a non-linear, confidence-weighted scheme rather than a strictly linear one can improve ptychographic reconstruction on tabletop sources. This matters because it could let low-flux and broadband laboratory sources approach nanoscale imaging without a synchrotron.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Central claim unverified: no ground-truth reconstruction error or named phase-retrieval solver; reported SNR and visual sharpness cannot distinguish nonlinear regularizer from structured bias.","rationale":"I read the paper as proposing that non-linear exposure fusion can be inserted into the ptychographic pipeline without the usual radiometric-linearity guarantee, because the induced bias acts as a spatially adaptive spectral preconditioner. This is an interesting and testable claim. For it to hold, the fused intensity must be at least as good as a Poisson-linear fusion for recovering the true object, and that condition is the least secure part of the argument. The paper provides visual comparisons, convergence trajectories, SNR of a fused diffraction frame, and MTF of reconstructed phase, but with no named phase-retrieval algorithm and no ground-truth object, a bias-driven sharpening cannot be distinguished from genuine fidelity. The 13% radiometric deviation is not defined, the proof is deferred to absent supplementary notes, and the method's hyperparameters are unspecified. The printed pyramid reconstruction equation also raises a reproducibility question. I would not reject the underlying idea: Mertens-style exposure fusion is a credible source of implicit regularization, and the experimental reports may be consistent with it. But the strongest form of the central claim is currently unverified rather than demonstrated. Since the reader's conditional verdict already identifies this evidence gap, my stress-test pass does not move the verdict.","tokens_in":9525,"tokens_out":9614,"duration_ms":111952,"concrete_test":"Run a controlled numerical ptychography experiment on a known complex-valued phantom: generate multi-exposure low-flux diffraction frames with Poisson noise and detector saturation; fuse with (i) the paper's MNF using its exact equations and stated hyperparameters and (ii) a linear/Poisson-likelihood fusion baseline. Reconstruct both with the same named phase-retrieval solver. Compute relative RMSE and Fourier ring correlation against the phantom. If MNF's reconstruction error is not at least as good as the linear baseline across a range of fusion parameters, the claim that strict linearity is unnecessary fails. This test also locates the bias level at which the solver leaves the convergence basin.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim requires that the MNF-fused intensity, despite its ~13% radiometric deviation, drives the phase-retrieval solver to the true complex object with at least the fidelity of a linear/Poisson fusion. The paper never tests this. No phase-retrieval algorithm is named; no ground-truth object is shown; no RMSE, Fourier ring correlation, or other reconstruction error metric is reported. The quantitative results are SNR of a single fused diffraction pattern (Fig. 4e), noise-variance curves, line profiles, and MTF (Supp Fig. 6c), all of which can improve while the object estimate moves away from the truth: a low-noise but systematically biased diffraction pattern can be inverted into a sharper but wrong image. The '13% radiometric deviation' is not defined or connected to convergence. The theoretical proof of the spectral-preconditioner mechanism is deferred to Supplementary Notes 1 and 4, which are not included; the hyperparameters (σ, μ, Δ_l, L) are not given. Additionally, as printed, the pyramid reconstruction step (§3.4.4) appears to propagate unweighted Gaussian base layers rather than weight-modulated ones, so the implemented fusion may differ from the described one. Without code/data and a ground-truth benchmark, the observed sharpening cannot be attributed to the claimed mechanism.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes Multi-Scale Non-Linear Fusion (MNF) for ptychographic high-dynamic-range (HDR) imaging, replacing strict radiometric linearity with a weight-map-modulated Laplacian/Gaussian pyramid fusion of multi-exposure diffraction data. The authors claim that non-linear fusion acts as a spatially adaptive spectral preconditioner that filters stochastic gradient noise, allowing accurate reconstruction despite a ~13% radiometric deviation from linearity. They support the claim with two experimental demonstrations: a quasi-monochromatic HeNe laser setup and a broadband tabletop HHG XUV setup, reporting visual comparisons, line profiles, SNR of fused diffraction patterns, and MTF curves.","tokens_in":9844,"tokens_out":3712,"duration_ms":42484,"significance":"If substantiated, the central claim would be significant: it would relax a long-standing assumption in ptychographic HDR fusion—that the fused intensity must be linearly proportional to the squared wavefront modulus—and would offer a practical computational-photography-inspired route to improving resolution under photon-starved, broadband tabletop illumination. The paper addresses a real bottleneck (detector bit-depth limitations and dynamic range) and demonstrates two experimental implementations. The claimed O(N) complexity is a practical advantage. However, the validation as presented is incomplete: the main evidence is visual and SNR-based, no ground-truth reconstruction error is reported, the phase-retrieval algorithm is not named, key theoretical appendices are missing, and hyperparameters are not provided. These gaps must be filled before the central claim can be considered established.","major_comments":[{"comment":"The central claim is that MNF-fused diffraction intensities, despite a ~13% radiometric deviation, drive iterative phase retrieval to a more accurate object than linear/Poisson-based fusion. This is never directly tested. The manuscript reports no ground-truth object, no RMSE or Fourier ring correlation, and no named phase-retrieval algorithm. The quantitative result in Fig. 4e is the SNR of a single fused diffraction pattern, which can improve while the object estimate moves away from the truth: a low-noise but biased diffraction pattern can be inverted into a sharper but wrong image. Please provide a quantitative reconstruction-error benchmark against a known object (e.g., a lithographed test pattern), specify the iterative solver (ePIE, rPIE, ML-based, etc.) and its hyperparameters, and report reconstruction fidelity metrics for all compared fusion methods.","section":"§2.2, §2.3, Fig. 4e"},{"comment":"The claimed 13% radiometric deviation is not defined: is it per-pixel RMS, a spectral-band error, or a maximum deviation? More importantly, the manuscript does not connect this deviation to the convergence properties of the phase-retrieval solver. The theoretical derivation of the 'spectral preconditioner' mechanism is deferred to Supplementary Note 1, which is not included in the preprint, and the stability analysis to Supplementary Note 4. Without these appendices, the central theoretical claim is unsupported. The main text needs a self-contained statement of the assumptions, the convergence result, and the quantitative meaning of the 13% operating point.","section":"§2.2, Supplementary Fig. 3, Supplementary Note 1"},{"comment":"The MNF hyperparameters are not reported: the Gaussian weight midpoint μ and selectivity σ (Eq. 5), the quantization steps Δ_l per pyramid level, the number of pyramid levels L, and the Gaussian kernel width are all absent. The text states (end of §2.2) that reconstruction fidelity varies by less than 9% across hyperparameter settings, but no ranges or settings are given. This makes the method irreproducible and prevents assessment of the claimed robustness. Please provide the exact values used in both experiments, or at least a table of the full parameter set.","section":"§3.4.1, §3.4.3"},{"comment":"The description of the reconstruction step appears internally inconsistent. Eq. (9) reconstructs R_{l-1}^{(k)} from the up-sampled Gaussian image G_l^{(k)} plus the quantized, weight-modulated detail L̃_{l-1}^{(k)}. The text says 'the base structural information is propagated from the Gaussian pyramid, while the weight modulation is selectively applied to the detail layers.' If G_l^{(k)} is the unweighted Gaussian pyramid of the image, then the low-frequency base from saturated or noisy exposures is not suppressed by the weight maps, contradicting the purpose of the weight assessment in §3.4.1. If instead G_l^{(k)} denotes the weight Gaussian pyramid, the notation is confusing and the base is not an image. Please clarify the role of the base layer and explain how the low-frequency envelope is protected from saturation.","section":"§3.4.4, Eq. (8)-(9)"}],"minor_comments":[{"comment":"The term 'spectral preconditioner' is used prominently but is not formally defined in the main text; a short definition or pointer to the theoretical appendix would help.","section":"Abstract / §1"},{"comment":"Reference 'Jagatap and Hegde 2019' in the introduction appears to cite a title about photoinduced metamaterials, which does not match the ptychographic context; please verify and correct.","section":"References"},{"comment":"The forward model omits a detector noise variance term that is discussed later (background noise variance in Fig. 4b). Consider defining η(q) explicitly.","section":"§3.3, Eq. (3)"},{"comment":"The recursion for the Laplacian pyramid should specify the handling of the coarsest level (L_l for l=L) and the boundary conditions; otherwise the decomposition is not fully specified.","section":"§3.4.2"}],"recommendation":"major_revision","confidential_remarks":"The paper is potentially interesting but the validation falls short of supporting the central claim. The missing supplementary notes, the lack of a named phase-retrieval solver, and the absence of ground-truth reconstruction error are critical. I would also urge the editor to check that the claimed 'first demonstration' of multi-harmonic HDR ptychography is properly scoped against the existing literature (Kodgirwar et al., Takahashi et al.). Data and code availability statements are vague ('available upon request'); for a methods paper, the code should be deposited and linked."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: the transfer of Mertens-style exposure fusion into ptychographic HDR is a real idea, and the broadband HHG demonstration is the kind of result worth seeing. But the preprint currently argues on faith: the convergence and stability proofs live in Supplementary Notes 1 and 4 that aren't there, and the experiments never report a ground-truth reconstruction error.\n\nThe claim that strict radiometric linearity is not necessary—that nonlinear weights can act as a spatially adaptive spectral preconditioner—is genuinely new in this subfield. The implementation is a straightforward adaptation of exposure fusion (Gaussian confidence weights, Laplacian pyramid, plus a quantization step), but the application to coherent diffractive imaging is novel. The two experiments, HeNe quasi-monochromatic and multi-harmonic HHG XUV in grazing incidence, are real, and the HHG result is, as far as I know, a first.\n\nThe soft spots are exactly where the stress-test note lands. No phase-retrieval algorithm is named, no ground truth, no RMSE or FRC. The quantitative metrics are SNR of a single fused diffraction pattern and MTF, which can improve while the complex object drifts away from the truth. The '13% radiometric deviation' is a number, not an analysis; nothing ties it to the solver's convergence basin. The hyperparameters μ, σ, Δ_l, L are not given. Code and data are 'available upon request,' which in practice means not available. There is also a possible bug in §3.4.4: the recursion uses the unweighted Gaussian base G_l^(k) rather than a weight-modulated base, so the implemented fusion might not match the described one. That needs checking. And the Jagatap & Hegde citation looks mismatched—the title given is about metamaterials, which is not a phase-retrieval paper. I'd ask the authors to fix that.\n\nThese are fixable. The central idea is plausible; I'm not calling it wrong. But the current preprint doesn't give a reader enough to verify it.\n\nWho it's for: people working on lab-source ptychography and HDR for coherent imaging. It deserves a serious referee—the idea and the HHG demo are worth the time. I'd send it out, but with a clear request for the supplementary notes, a ground-truth benchmark, and the missing hyperparameters.","headline":"Non-linear exposure fusion for ptychographic HDR is a promising transfer, but the preprint's evidence is incomplete: no ground-truth reconstruction error, missing proofs, and a possible implementation mismatch.","tokens_in":10307,"tokens_out":1970,"would_cite":false,"duration_ms":22350,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper establishes that strict radiometric linearity is not a prerequisite for accurate ptychographic reconstruction: a multi-scale non-linear fusion (MNF) step acts as a spatially adaptive spectral preconditioner that filters stochasti","keywords":["ptychography","high-dynamic-range imaging","multi-exposure fusion","phase retrieval","non-linear spectral preconditioning","tabletop XUV source","broadband coherent diffractive imaging","lensless imaging"],"falsifier":"A synthetic ptychography experiment with a known ground-truth object can settle this: reconstruct the object from MNF-fused simulated diffraction data and from linear-fused data under identical Poisson noise. If the MNF reconstruction error is worse than linear fusion, or if the 13% non-linear bias drives the solver to a wrong local minimum, the central claim fails. The test requires specifying the phase-retrieval algorithm, which the paper currently omits.","tokens_in":9427,"feed_emoji":"🔬","tokens_out":7133,"duration_ms":73742,"temperature":0.7,"pith_summary":"Ptychography reconstructs a specimen from overlapping diffraction patterns, and its resolution in laboratory settings is limited by low photon flux, detector bit depth, and the huge dynamic range between the zero-order beam and weak high-frequency fringes. The standard remedy is multi-exposure high-dynamic-range (HDR) fusion, but current fusion rules insist the merged intensity stay linearly proportional to the squared wavefront to match Poisson likelihood models. This paper argues that linearity is not a prerequisite: it introduces a multi-scale non-linear fusion (MNF) step, borrowed from computational photography, whose pixel-wise confidence weights and Laplacian-pyramid quantization suppress shot noise while preserving the low-frequency envelope. The central claim is that the non-linear weights act as a spatially adaptive spectral preconditioner—an implicit regularizer that filters the gradient noise of iterative phase retrieval rather than distorting the physical signal. If true, tabletop ptychography can use broadband, multi-harmonic XUV sources without monochromatic filtering and still reach diffraction-limited resolution.","feed_headline":"Ptychography can drop strict linearity and gain resolution","feed_subtitle":"Multi-scale exposure fusion recovers high-frequency fringes under low photon flux and handles multi-harmonic XUV beams.","key_machinery":"Multi-Scale Non-Linear Fusion (MNF). The load-bearing object is the MNF pipeline: each raw exposure gets a Gaussian-likelihood weight map W_k(q) = exp(-(I - mu)^2/(2 sigma^2)); intensities and weights are decomposed into Laplacian and Gaussian pyramids; the detail coefficients are multiplied by the corresponding smoothed weights and quantized with a scale-adaptive step Delta_l; each exposure is then reconstructed from its preserved Gaussian base plus filtered detail layers and summed into I_MNF. The quantization is the denoising step, the pyramid keeps the global intensity envelope physically consistent, and the overall weight map is what the paper interprets as a spatially adaptive spectral","core_discovery":"The paper's central claim is that structural consistency—preserving the reliable parts of each exposure—matters more than strict statistical linearity in photon-starved ptychography. The authors show that an MNF-fused diffraction pattern, which deviates by about 13% from the ideal intensity proportional to the squared wavefront, reconstructs fine features that standard linear integration, structural patch decomposition, and variance-weighted Bayesian fusion either bury in noise or over-smooth. The mechanism they propose is gradient-descent dynamics: the non-linear fusion weights redistribute measurement confidence and quiet the rugged loss landscape caused by stochastic noise, so the solver","pith_inferences":["If the preconditioning interpretation is right, other non-linear image-fusion rules from computational photography could be transplanted into phase retrieval, and the reported 13% deviation tolerance becomes a testable budget for how non-linear a fusion rule may be before it breaks convergence.","The same mechanism might transfer to Fourier ptychography or other iterative inverse problems where detector dynamic range and photon starvation limit resolution, not just to scanning ptychography.","The convergence and stability proofs are deferred to supplementary notes; until those appear, the exact conditions under which MNF stays inside the solver's convergence basin remain the open load-bearing question.","A natural extension would be to learn the non-linear weight maps from data, potentially outperforming the hand-crafted Gaussian likelihood weights while preserving the spectral-preconditioning effect."],"forward_implications":["HDR ptychography can drop Poisson-linear fusion without sacrificing reconstruction accuracy; a radiometric deviation of roughly 13% from linearity is compensated by the noise suppression the fusion provides.","Multi-exposure fusion from computational photography becomes a valid component of the ptychographic pipeline, opening the door to perceptually motivated fusion rules in coherent diffractive imaging.","The approach enables HDR ptychography with tabletop multi-harmonic HHG XUV sources without pre-filtering to a single harmonic, broadening the effective spectral bandwidth for laboratory nanoscopy.","MNF has linear computational complexity O(N) and its reconstruction fidelity remains stable across hyperparameter settings (relative error variation below 9%), making it practical for high-throughput imaging.","MNF-reconstructed images preserve modulation transfer function values above the 50% threshold up to the Nyquist limit, recovering edges and phase boundaries that linear and parametric baselines over-smooth."],"supporting_citations":[{"why":"Introduces ptychographic phase retrieval with a moving aperture, the imaging framework the paper operates within.","marker":"Rodenburg and Faulkner 2004"},{"why":"Provides the iterative ptychography engine used as the reconstruction baseline for the fused diffraction data.","marker":"Maiden and Rodenburg 2009"},{"why":"Supplies the exposure-fusion method with Gaussian weights and Laplacian pyramids that MNF adapts to diffraction patterns.","marker":"Mertens et al. 2009"},{"why":"Establishes standard HDR ptychography via linear stitching, the linear-fusion convention MNF challenges.","marker":"Dierolf et al. 2010"},{"why":"Formulates maximum-likelihood/Poisson refinement in coherent diffractive imaging, the statistical framework from which MNF deliberately deviates.","marker":"Thibault and Guizar-Sicairos 2012"},{"why":"Treats non-linear pixel response as a detector error, the conventional position this paper directly opposes.","marker":"Driel et al. 2015"},{"why":"Supplies Structural Patch Decomposition, the non-linear but over-smoothing fusion baseline compared in the experiments.","marker":"Ma et al. 2017"},{"why":"Applies HDR ptychography in the photon-starved regime with mixed-mode detectors, the low-flux context of the central claim.","marker":"Giewekemeyer et al. 2014"},{"why":"Presents Bayesian multi-exposure HDR ptychography, the statistical-linear baseline that MNF outperforms in the experiments.","marker":"Kodgirwar et al. 2024"}],"fun_headline_variants":["Ptychography drops strict linearity, gains resolution","Non-linear fusion boosts ptychography resolution","Tabletop ptychography without linear fusion","Multi-scale fusion sharpens photon-starved ptychography"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The key assumption is that ptychographic reconstruction still converges to the true specimen when the input diffraction intensities are only approximately proportional to the squared wavefront—about 13 percent off—and that this convergence holds for the algorithm actually used; the proof is deferred to supplementary notes not included in the preprint.","fun_headline_variants_meta":{"raw":{"variants":["Ptychography drops strict linearity, gains resolution","Non-linear fusion boosts ptychography resolution","Tabletop ptychography without linear fusion","Multi-scale fusion sharpens photon-starved ptychography"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000174,"raw_usage":{"total_tokens":1082,"prompt_tokens":673,"completion_tokens":409,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":417,"completion_tokens_details":{"reasoning_tokens":358}},"tokens_in":417,"tokens_out":409,"duration_ms":4713,"temperature":1.0,"reasoning_tokens":358,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-04T21:40:19.401451+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A synthetic ptychography experiment with a known ground-truth object can settle this: reconstruct the object from MNF-fused simulated diffraction data and from linear-fused data under identical Poisson noise. If the MNF reconstruction error is worse than linear fusion, or if the 13% non-linear bias drives the solver to a wrong local minimum, the central claim fails. The test requires specifying the phase-retrieval algorithm, which the paper currently omits.","supporting_citations":[],"review_version":1}