{"id":"fff9a1ef-f8df-4f96-a2bd-f710aa86cb47","arxiv_id":"2501.14290","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"Phase retrieval from single Zernike phase contrast images is achieved by optimizing an untrained deep decoder against a compressive propagation forward model, without manually tuned regularization.","lead":"This paper uses an untrained neural network as a built-in image prior to recover quantitative phase images from Zernike phase contrast microscopy, removing manual tuning of regularization. The method is tested on simulated and real microscope images and compared with a regularized optimization baseline.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Forward-model fidelity is the load-bearing risk: only two experimental sample types test Eq. (5), so the 'various samples' claim rests largely on simulations that use the same H.","rationale":"We agree with the reader's weakest_assumption; we see it as the most load-bearing because it is a necessary condition for the other claims. The method's novelty (no manual regularization tuning) and the network prior are secondary: if H is wrong, even a perfect optimizer produces biased phase. The paper does provide genuine experimental support on beads and separated cells, so we do not argue that the model is false, only that the breadth of validation is insufficient. A focused experiment on the two simulated-only classes would settle the concern. The verdict remains CONDITIONAL.","tokens_in":6715,"tokens_out":7399,"duration_ms":69903,"concrete_test":"Acquire experimental PCM images of grouped COS-7 cells and phase resolution targets using the same system and compare UNN reconstructions against the DH ground-truth used in the numerical sections. If RMS errors remain close to the simulated values (grouped cells ~0.24 rad, resolution targets ~0.24 rad), the forward model is adequate; if errors are substantially higher, further model calibration or a pupil-function measurement (e.g., via a known phase target) is required before the 'various samples' claim holds.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—quantitative phase recovery without manual tuning—depends on the forward model H in Eq. (5) being an accurate description of the actual PCM. The authors state that numerical experiments use the same model for data generation and inversion, but experimental verification covers only 19 bead images and 6 separated-cell images. Grouped cells and resolution targets, which are central to the demonstration of breadth, are validated only in simulation through this identical H. If the pupil filters c and p, or the random-wavefront approximation to partial coherence, deviate from the real microscope for clustered or extended samples, the recovered phase will be biased no matter how well the untrained-network prior regularizes. Notably, the paper does not specify how c and p are calibrated or validated, so the experimental agreement on two simple sample types does not establish that H is faithful for the full claim.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript proposes a computational phase retrieval method for Zernike phase contrast microscopy (PCM) using an untrained deep decoder as an image prior. The estimated phase image is represented as the output of a deep decoder applied to a fixed random tensor, and the network weights are optimized to minimize the L2 distance between the observed PCM intensity image and the intensity predicted by a compressive-propagation forward model (Eq. 5). The method is tested on simulated PCM images of microbeads, separated cells, grouped cells, and resolution targets, and on experimental PCM images of microbeads and separated cells. The authors compare their results with a previously published regularization-based ADMM approach using three representative regularization strengths. They report lower root-mean-square phase errors for the untrained-network method in most cases and conclude that it eliminates manual regularization tuning while improving accuracy and robustness.","tokens_in":6914,"tokens_out":6405,"duration_ms":60221,"significance":"If the reported results hold, the method would provide a practical path to quantitative phase imaging from off-the-shelf PCM without per-sample regularization tuning, which would be useful for biological applications. The paper's strengths are its simple and reproducible formulation, the use of a fixed network architecture and optimizer across all experiments, and the inclusion of both simulated and experimental data. The authors also explicitly acknowledge that the numerical experiments use an exact forward model. However, the breadth of the central claim is not yet fully supported: the experimental validation covers only two simple sample types with a small number of images and no error bars, and the robustness of the method to forward-model mismatch is not assessed. The significance is therefore moderate and depends on whether the forward-model fidelity and the experimental evidence can be strengthened.","major_comments":[{"comment":"The simulation protocol uses the same forward model H for data generation and inversion (the authors state that 'the model used to solve the inverse problem is exact'), so the numerical comparisons in Table II cannot detect errors in the PCM model. The experimental validation is limited to 19 microbead images and 6 separated-cell images; grouped cells and resolution targets, which are central to the claimed breadth, are tested only in simulation. Since the forward model in Eq. (5) depends on the pupil filters c and p, the manuscript should explain how these filters are calibrated or validated for the experimental microscope, and should either add experimental data for a more complex sample type or clearly temper the 'various samples' claim in the conclusion. A perturbation or noise-sensitivity test of H would strengthen the robustness argument.","section":"III, Tables II and III; Eq. (5)"},{"comment":"The claim that the approach eliminates empirical hyperparameter tuning is broader than what is demonstrated. The method fixes the deep decoder channel count (k_i=128), the input canvas size, the upsampling schedule, the optimizer learning rate (3e-4), the number of epochs (4000), and the number of random wavefronts per step (M=200). No sensitivity study is reported for these choices, so it is not established that they transfer to other experimental setups without adjustment. The authors should clarify that the elimination refers specifically to the regularization parameters (rho, epsilon_TV, epsilon_l1) of the earlier ADMM method, and should provide a small sensitivity analysis for at least the learning rate and epoch count to support the robustness claim.","section":"II and V"},{"comment":"The experimental improvement over the best regularization baseline is modest for separated cells (UNN RMS 0.172 rad vs 0.199 rad for R-L) and is based on only six images, with no standard deviation or per-sample values reported. Tables II and III also average five restarts per sample without reporting the spread across restarts. Given the acknowledged hole artifacts (Section IV, Appendix B), the claim that the untrained network outperforms the regularization method in experiment should be supported by error bars or a statistical test; otherwise the reader cannot assess whether the difference is meaningful.","section":"III, Table III"}],"minor_comments":[{"comment":"The slope of the LeakyReLU activation is not specified; please state the value (e.g., 0.01) used in the experiments.","section":"II, Eq. (3)"},{"comment":"The symbol M is used both for the number of wavefronts in the forward model (4000 in simulation) and for the number used per optimization step (200); using M_sim and M_opt would avoid ambiguity.","section":"II, Eq. (5)"},{"comment":"For the 'High' regularization setting, the text says it was 'determined through empirically tuning to optimize performance,' but it does not say on which sample or metric. Please specify this so the comparison baseline is transparent.","section":"III, Table I"},{"comment":"The caption lists 'PCM images, ground truth phase images, and phase restorations' but the figure panels are not labeled with these categories; adding labels and a common phase colorbar would improve readability.","section":"Fig. 2 caption"},{"comment":"The suggestion for convergence analysis cites reference [8]; it would be helpful to state what specific result from that work transfers to the present setting, or to cite a reference that directly addresses convergence of deep-decoder optimization.","section":"IV, reference [8]"}],"recommendation":"major_revision","confidential_remarks":"The central idea is sound and within the journal's scope, but the experimental validation is thin relative to the strength of the conclusions. The editor may wish to request that the authors either temper the 'various samples' claim or add a third experimental sample type, and report variability in the tables. The comparison is only against the authors' own earlier regularization method; an independent baseline would increase confidence. The manuscript does not include a data or code availability statement, which would help reproducibility."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a genuine new combination, not a big idea. Untrained deep priors and compressive propagation both exist; putting them into Zernike PCM phase retrieval is new in the literature, and it works at least as well as the best of three fixed-regularization baselines across four sample classes in simulation and two in experiment. The authors are also unusually honest: they state plainly that the simulation uses the exact forward model, they show the hole artifacts, they discuss the low-frequency loss, and they give the computational cost.\n\nWhere I part company with the stress-test note: I don't think the absence of grouped-cell and resolution-target experiments is fatal. The simulation evidence is still informative, and the two experimental classes, while small, are real tests with DH ground truth. The small sample size does bother me: 6 separated-cell images, 19 beads, no standard deviations. Given that the central selling point is 'improved accuracy and robustness across various samples,' the robustness half of that claim rests mostly on exact-model simulations.\n\nThe bigger issue is the forward model H in Eq. (5). The paper cites [7] for the parameter settings but does not say how the pupil filters c and p were calibrated for the particular microscope used. If those filters are off, the recovered phase will be biased no matter what prior you use. For two simple sample types the experimental agreement suggests the model is roughly right, but 'roughly right' may not be enough for the quantitative claims. This is the part I would want a referee to push on.\n\nOne more thing: the no-tuning selling point is real but modest. The network architecture, learning rate, epoch count, and M are all fixed, and they presumably selected those once on some validation set. So they eliminated per-sample regularization tuning, not all tuning. That is an honest claim and worth having, but it should not be oversold.\n\nRecommendation: let it into peer review with the expectation of heavy-ish revision. Add more experimental samples, give error bars, and describe the pupil calibration. I would not want this desk-rejected; it is a legitimate method contribution and the limitations are openly discussed.","headline":"A useful, honest method paper that convincingly shows an untrained decoder prior can replace manual regularization in Zernike PCM phase retrieval—but the experimental validation is thinner than the breadth claims, and forward-model fidelity is the real risk.","tokens_in":7407,"tokens_out":2777,"would_cite":true,"duration_ms":27755,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":["42.30.Rx"],"model":"deepseek-v4-flash","headline":"An untrained neural network can turn ordinary Zernike phase-contrast microscope images into quantitative phase maps without manual tuning of regularization parameters.","keywords":["phase retrieval","Zernike phase contrast microscopy","untrained neural network","deep image prior","compressive propagation","quantitative phase imaging","computational microscopy"],"falsifier":"A direct calibration test: acquire PCM images of a well-characterized phase target (e.g., microfabricated pillars with known heights from profilometry) on the same microscope, run the proposed fixed-network pipeline, and compare recovered phase to the profilometry ground truth. If the recovered phase shows a systematic, sample-size-dependent bias that persists across random restarts, or if deliberately misestimating the pupil parameters c and p in H changes the reconstruction, then the forward model rather than the network prior is the limiting factor.","tokens_in":6564,"feed_emoji":"🔬","tokens_out":4884,"duration_ms":40052,"temperature":0.7,"pith_summary":"The paper claims that phase retrieval from a single Zernike phase-contrast microscopy (PCM) image can be made quantitative by replacing hand-tuned regularizers with an untrained deep decoder used as a structural prior. The method minimizes the error between the observed PCM intensity and the intensity predicted by a compressive-propagation forward model applied to the network's phase output. With fixed network architecture and optimizer settings, the authors report lower RMS phase error than a regularization-based baseline on simulated beads, cells, and resolution targets, and on experimental beads and separated cells. If correct, this means off-the-shelf PCM hardware can deliver quantitative phase imaging without per-sample parameter tuning, making QPI practical for routine biological use.","feed_headline":"Untrained network makes phase-contrast microscopy quantitative","feed_subtitle":"A fixed deep decoder replaces hand-tuned regularization, recovering phase maps from standard Zernike images.","key_machinery":"The load-bearing object is the pair (D, H): D is a deep decoder—a non-convolutional network with channel-wise linear layers, bilinear upsampling, ReLU and ChannelNorm, and a final LeakyReLU that encodes the prior that the phase image has near-zero background—and H is the compressive-propagation model of PCM, which approximates partially coherent illumination by averaging M random wavefronts propagated through a condenser annulus pupil c and phase-ring pupil p. The optimization runs Adam on the network weights W with 200 random wavefronts per step, 4000 epochs, and a learning rate of 3×$10^{-4}$. The deep decoder acts as a self-learned regularizer that restricts the solution space, while H links the phase estimate to the observed intensity.","core_discovery":"The central claim is that the estimated phase image θ = D(W; B0), produced by a fixed deep decoder with randomly initialized weights W and a fixed input tensor B0, can be optimized by minimizing ||g − H(D(W;B0))||$2^{2}$, where H is the compressive-propagation PCM forward model, to recover quantitative phase from a single PCM intensity image. The authors establish this by comparing RMS errors against the regularization-based method of Kurata et al. under three manually tuned regularization strengths, across four simulated sample classes and two experimental classes. They report the untrained-network method matches or beats the best-tuned regularization result in nearly every class, and is the only method that does not require sample-dependent hyperparameter selection. They further map the recoverable sample size-phase range, finding diameters up to about 40 µm at phases ≲1.5 rad and phases up to 2.88 rad for small samples.","pith_inferences":["If the forward model H is accurate only for thin, weakly scattering samples, the same fixed network will inherit that limitation; a testable extension is to calibrate H against a known standard and verify retrieval accuracy, quantifying model mismatch separately from network behavior.","The observed 'hole artifacts' suggest the deep decoder's initialization can trap optimization in a local basin; pretraining on a segmentation image helps partially, implying that a better initialization scheme or a convergence-aware optimizer could remove the residual artifacts entirely.","The authors' comparison with three fixed regularization strengths shows no single setting generalizes; this implies the advantage of the untrained network is not raw accuracy but adaptability—an inference testable by running the same benchmark with oracle-tuned per-sample regularization.","Since the network uses no convolutions, the role of the prior is purely statistical (smoothness and background sparsity); one could test whether an explicit hand-crafted prior with the same properties, such as a learned dictionary, achieves comparable results without a network."],"forward_implications":["If correct, any existing Zernike phase-contrast microscope can provide quantitative phase maps with no hardware modification and no per-image tuning.","The fixed-network recipe removes the user-dependent choice of regularization strength, making phase retrieval practical for non-specialists.","The recovered phase range and object-size limits are now characterized (about 40 µm diameter at phase ≲1.5 rad; up to 2.88 rad for small objects), giving users a concrete applicability envelope.","The method demonstrates that deep priors are compatible with incoherent illumination models, extending untrained-network phase retrieval beyond coherent setups.","Because the network is untrained, the approach requires no dataset and can be applied immediately to new sample types without retraining."],"supporting_citations":[{"why":"Supplies the compressive-propagation PCM forward model and the regularization-based phase retrieval baseline that the paper extends and compares against.","marker":"[7]"},{"why":"Introduces the deep image prior concept that motivates using an untrained network as a structural regularizer.","marker":"[13]"},{"why":"Provides the specific deep decoder architecture (non-convolutional, channel-wise linear layers) used as the network prior D.","marker":"[14]"},{"why":"Supplies the compressive-propagation approximation with random wavefronts that makes the forward model H computationally feasible.","marker":"[15]"},{"why":"The Adam optimizer used to minimize the objective over network weights.","marker":"[16]"},{"why":"Grounds the digital holography system used to obtain ground-truth phase distributions for the bead and cell samples.","marker":"[19]"}],"fun_headline_variants":["Untrained network makes Zernike phase contrast quantitative","Fixed deep decoder replaces manual tuning in phase retrieval","Untrained prior boosts phase retrieval from Zernike images","Phase retrieval without regularization tuning via deep prior","Zernike phase contrast goes quantitative with untrained network"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The forward model H in Eq. (5)—with known condenser and phase-ring pupil filters and randomly drawn wavefronts—accurately represents the actual microscope's optical behavior; the numerical experiments use the same H for both simulation and inversion, and only two experimental sample types test the model, so any mismatch would bias the recovered phase regardless of the network prior.","fun_headline_variants_meta":{"raw":{"variants":["Untrained network makes Zernike phase contrast quantitative","Fixed deep decoder replaces manual tuning in phase retrieval","Untrained prior boosts phase retrieval from Zernike images","Phase retrieval without regularization tuning via deep prior","Zernike phase contrast goes quantitative with untrained network"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000211,"raw_usage":{"total_tokens":1353,"prompt_tokens":824,"completion_tokens":529,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":440,"completion_tokens_details":{"reasoning_tokens":454}},"tokens_in":440,"tokens_out":529,"duration_ms":5208,"temperature":1.0,"reasoning_tokens":454,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T15:14:31.613128+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A direct calibration test: acquire PCM images of a well-characterized phase target (e.g., microfabricated pillars with known heights from profilometry) on the same microscope, run the proposed fixed-network pipeline, and compare recovered phase to the profilometry ground truth. If the recovered phase shows a systematic, sample-size-dependent bias that persists across random restarts, or if deliberately misestimating the pupil parameters c and p in H changes the reconstruction, then the forward model rather than the network prior is the limiting factor.","supporting_citations":[{"cited_title":"Kurata, K","cited_arxiv_id":null,"evidence_quote":"Supplies the compressive-propagation PCM forward model and the regularization-based phase retrieval baseline that the paper extends and compares against."},{"cited_title":"Ulyanov, A","cited_arxiv_id":null,"evidence_quote":"Introduces the deep image prior concept that motivates using an untrained network as a structural regularizer."},{"cited_title":"Horisaki, T","cited_arxiv_id":null,"evidence_quote":"Supplies the compressive-propagation approximation with random wavefronts that makes the forward model H computationally feasible."},{"cited_title":"Bhaduri, C","cited_arxiv_id":null,"evidence_quote":"Grounds the digital holography system used to obtain ground-truth phase distributions for the bead and cell samples."}],"review_version":1}