{"id":"7cf6d5ea-1d84-409e-8657-66069d39f065","arxiv_id":"2505.07839","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"An untrained neural network combined with physical wave-propagation modeling reconstructs terahertz single-pixel images down to about one seventh of the wavelength using a thick silicon modulator.","lead":"Terahertz imaging with a single-pixel detector and an untrained neural network resolves features about seven times smaller than the imaging wavelength. The method uses a standard thick silicon modulator, which could make sub-wavelength terahertz microscopy faster and more practical.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 118 µm feature is evanescently attenuated by many orders of magnitude across the 500 µm Si wafer, so the claimed λ/7 resolution must come from the network prior; no control experiment rules out prior-driven hallucination.","rationale":"The reader's weakest assumption already identifies the core physics: the forward-propagation distance and the evanescent attenuation of high spatial frequencies are what the claimed resolution depends on. My pass sharpens that into a concrete information-theoretic bound: the 118 µm component is attenuated by more than ten orders of magnitude over the 500 µm silicon path, so it cannot be recovered by any linear backpropagation operator from the recorded intensities. Therefore the reconstructed gap must be supplied by the untrained network's prior, and the paper's resolution claim is only as strong as the prior's ability to synthesize plausible binary structure. This is not an accusation of dishonesty; it is a structural gap in the evidence. The presented experiments use a symmetric, three-slit target whose structure is exactly the kind of pattern a deep-image prior will readily generate, and no control demonstrates that the 118 µm feature is actually constrained by the measurements. This concern is addressable: an isolated-feature target, a shuffled-measurement control, and a no-ASP ablation would separate measurement-driven recovery from prior-driven hallucination. Because the reader already assigned a conditional verdict based on missing quantitative grounding, my finding does not change the verdict; it identifies the specific experiment that should be required before the λ/7 claim is accepted.","tokens_in":17163,"tokens_out":7245,"duration_ms":104835,"concrete_test":"Acquire data for a binary target containing an isolated 118 µm feature (e.g., a single pinhole or a randomized arrangement of 118 µm bars) and run the identical untrained TSPIDL+ASP pipeline at SR = 3.125%, with a no-ASP control, a shuffled-measurement control, and at least five repeats to report peak-to-valley contrast and error bars. If the isolated feature is not recovered, or if the shuffled control still produces a plausible three-slit image, the λ/7 result is prior-driven rather than a verified system resolution.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that ASP backpropagation recovers 118 µm (λ/7) structure from single-pixel measurements through a 500 µm Si wafer. The spatial frequency corresponding to 118 µm is fx ≈ 8.47 mm⁻¹, giving a normalized spatial frequency u = λfx ≈ 7 in air and u ≈ 2 in silicon. In Eq. 3 the evanescent term decays as exp(−2πd/λ·√(u²−1)); over d = 0.5 mm inside silicon (d/λ_Si ≈ 2) this factor is below 10⁻¹⁰. The measured far-field intensities therefore contain essentially no information about that gap. Any 118 µm contrast in Fig. 4f3 must be contributed by the untrained network's implicit deep-image prior, not by the physical measurement. That is acceptable only if the paper claims a prior-based super-resolution demo for known binary targets; it does not establish a spatial resolution of the imaging system. The paper provides no control experiment — e.g., a non-periodic 118 µm feature, a shuffled-measurement control, or a no-ASP ablation — to separate data-driven recovery from hallucination. The dBP sweep in Fig. 4f1–f4 shows a refocusing trend but no quantitative contrast metric, error bars, or false-positive analysis. Additionally, Eqs. 2–3 are written for complex field propagation, while the network output is described as an intensity distribution; the missing phase/amplitude conversion is never specified, compounding the risk that the apparent refocusing is not a valid physical backpropagation.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes an untrained physics-constrained neural network, termed untrained TSPIDL, for terahertz single-pixel imaging (THz SPI). A 0.36 THz continuous-wave source illuminates an object, the transmitted field is encoded by a 500-μm-thick silicon photomodulator, and a far-field single-pixel detector records intensity. The network is optimized so that simulated single-pixel measurements from its output match actual detector readings, with an angular-spectrum propagation (ASP) model embedded in the output layer to backpropagate the diffracted field from the recording plane to the object plane. The authors claim a spatial resolution of 118 μm (about λ0/7) using this approach at a sampling ratio of 3.125%, and demonstrate refocusing of images at distances up to 6 mm. The central demonstration is a three-slit target with 118 μm separation recovered when the backpropagation distance is set to the known wafer thickness of 0.5 mm.","tokens_in":17443,"tokens_out":3530,"duration_ms":40877,"significance":"If the resolution claim were substantiated, this would be a valuable contribution: the method requires no training data, uses a thick silicon modulator rather than ultrathin photomodulators, and achieves very low sampling ratios. The integration of an untrained deep-image prior with a physical propagation model is a sensible and timely idea. However, the current evidence is not sufficient to establish that the recovered 118 μm features are information from the measurements rather than artifacts of the network's implicit prior. The paper also lacks quantitative rigor in the resolution assessment, with a single target, no error bars, and SSIM computed against HSPI reconstructions rather than ground truth.","major_comments":[{"comment":"The λ/7 resolution claim is not supported as a property of the imaging system because the 118 μm feature corresponds to a normalized spatial frequency u ≈ 7 in air (u ≈ 2 in silicon), and the evanescent transfer factor in Eq. (3), exp(−2πd/λ·√(u²−1)), is below 10⁻¹⁰ over d = 0.5 mm inside silicon. The measured far-field intensities therefore contain essentially no information about that gap, so the recovered contrast in Fig. 4f3 must come from the network's implicit deep-image prior. The authors need a control experiment—e.g., a non-periodic 118 μm feature, a shuffled-measurement control, or a no-ASP ablation—to distinguish measurement-driven recovery from prior-driven hallucination before claiming a spatial resolution of the imaging system.","section":"Methods, Eqs. (2)–(3); Results, Fig. 4f"},{"comment":"The resolution claim rests on visual separation of three slits in a single target, with no error bars, no repeated trials, and no quantitative contrast metric. The cross-sectional profiles in Fig. 4g show a trend, but the authors should report repeated measurements, a resolution metric (e.g., modulation depth or a Rayleigh-type criterion), and ideally compare against the known optical mask rather than only against HSPI reference images. The current evidence is anecdotal and cannot support a quantitative λ/7 resolution statement.","section":"Results, Fig. 4f1–f4 and Fig. 4g"},{"comment":"The ASP model is defined for complex field propagation, while the network output is described as an intensity distribution (Eq. (6): O0 = |E0|²). The manuscript never specifies how the missing phase is obtained or why applying the complex ASP propagator to an intensity-only estimate is physically valid. This missing step is load-bearing for the backpropagation claim, and the paper should clarify the field-to-intensity conversion and justify the intensity-only ASP approximation.","section":"Methods, Eqs. (2)–(3) and Eq. (6)"},{"comment":"The method requires prior knowledge of the forward propagation distance dFP, as the authors explicitly state, and the dBP sweep in Fig. 4f1–f4 shows that a wrong dBP produces artifacts (four apertures instead of three at dBP = 1.0 mm). This means the sub-diffraction result is conditional on feeding the correct object position into the algorithm. The paper should state this conditioning prominently in the abstract and conclusion, and it should discuss whether the method can be extended to unknown object depth or whether the claimed resolution only holds for known binary targets at a known distance.","section":"Results, paragraphs after Fig. 5 and Conclusion"}],"minor_comments":[{"comment":"The abstract advertises an ultralow sampling ratio of 1.5625%, but the λ/7 result in Fig. 4 uses SR = 3.125%. Please clarify which sampling ratio belongs to which claim and avoid implying that the sub-diffraction result was obtained at the lowest stated ratio.","section":"Abstract and Fig. 4 caption"},{"comment":"The caption refers to images 'at d = 0 mm' while the text states the object is at dFP = 0.5 mm from the back surface of the wafer; please make the distance definition consistent.","section":"Fig. 3 caption"},{"comment":"The sentence 'Even at a low SR = 1.5625%, the SSIM can reach 0.67 and the reconstructed images are acceptable (Figure 3a1,b1)' appears to cite the full-sampling HSPI reference images rather than the low-SR reconstructions; the figure references should be corrected.","section":"Results, Fig. 3a1,b1 references"},{"comment":"Equation (5) has garbled mathematical symbols in the typeset text; it should be rewritten with clear notation so that the DGI estimate is unambiguous.","section":"Eq. (5)"},{"comment":"There are several typographical errors, such as 'state-of-art' in the Results section and 'near-filed' in the Introduction; a careful proofread is recommended.","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"The central idea is promising and the experimental effort is substantial, but the λ/7 resolution claim currently overreaches the evidence: the evanescent-field attenuation calculation strongly suggests the recovered subwavelength features are prior-driven rather than measurement-driven. The authors should be asked to either provide control experiments that isolate the role of the measurements, or substantially reframe the claim as a demonstration of deep-image-prior super-resolution for known binary objects at known depth. With that reframing and the requested quantitative rigor, the paper could be publishable."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The short version: this is a real advance in integrating an untrained deep image prior with angular-spectrum backpropagation for THz single-pixel imaging, and the low-sampling-ratio results are promising. But the headline λ/7 resolution claim is not established. The stress-test arithmetic is correct: for 118 μm features at 0.36 THz, the evanescent component at the back of a 500-μm silicon wafer is suppressed by roughly ten orders of magnitude. The measured intensities carry essentially no information about that spatial frequency, so the clean three-slit reconstruction in Fig. 4f3 must be coming from the network's implicit prior, not from the data. That is fine if the paper is framed as a super-resolution demonstration for known binary targets, but it does not support a claim that the imaging system resolves λ/7. What is missing is a control: a non-periodic 118 μm feature, a shuffled-measurement test, or an ASP-off ablation with a quantitative contrast metric. Without one, the refocusing trend in the d_BP sweep could simply reflect the prior selecting the expected pattern.\n\nThe paper does several things well. The combination of DIP with ASP in an untrained network is new in THz SPI, and the hardware is simple: a 500-μm Si wafer, CW source, and a single-pixel detector. The comparison against HSPI, CS-TV, and a trained network at low SR is informative, and the authors are honest about the d_FP dependence and the artifacts at wrong backpropagation distances.\n\nThe soft spots beyond the resolution claim: the ASP equations are written for complex fields, but the network outputs an intensity distribution. The paper never states the phase assumption (presumably zero-phase amplitude object). That's an easy fix but important. Also, resolution is assessed visually from three slits, with no error bars, no repeated trials, and SSIM computed against HSPI rather than ground truth. For a paper claiming λ/7, that is thin.\n\nMy take: I would send this to peer review, not desk-reject it. The idea is good and the experimental data are real. The λ/7 claim needs to be either substantially tempered or backed by controls; a reviewer should ask for those experiments. This is a paper for THz imaging and computational imaging researchers, and with revision it could be a solid contribution. Right now, I would not cite the λ/7 number.","headline":"Promising untrained DIP + ASP combination for THz SPI, but the λ/7 resolution claim is likely prior-driven and lacks the controls to back it up.","tokens_in":18053,"tokens_out":4912,"would_cite":true,"duration_ms":58018,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper reports that an untrained neural network constrained by angular-spectrum backpropagation recovers 118-micron (λ/7) features at 0.36 THz through a 500-micron silicon wafer from far-field single-pixel measurements taken at a…","keywords":["terahertz imaging","sub-diffraction imaging","single-pixel imaging","untrained neural network","angular spectrum propagation","computational imaging","sampling ratio","near-field imaging"],"falsifier":"Run the same experiment with a grayscale object whose transmission varies continuously and contains λ/7 features, keeping $d_{\\mathrm{FP}} = 0.5$ mm and $d_{\\mathrm{BP}} = 0.5$ mm; if the reconstruction cannot resolve those features or produces artifacts, the claim that backpropagation recovers lost evanescent information is false for non-binary objects. Alternatively, deliberately mis-set $d_{\\mathrm{BP}}$ to 0.55 mm and observe whether the three-slit target begins to show the four-aperture artifact predicted at $d_{\\mathrm{BP}} = 1.0$ mm, which would confirm the method's sharp dependence on knowing $d_{\\mathrm{FP}}$.","tokens_in":16939,"feed_emoji":"🔬","tokens_out":6818,"duration_ms":69687,"temperature":0.7,"pith_summary":"This paper claims that a major practical obstacle in terahertz single-pixel imaging—the need for ultrathin photomodulators to capture evanescent, subwavelength information—can be removed by doing the 'thinness' in software. The authors describe an imaging system at 0.36 THz in which a 500-micron silicon wafer both encodes the object field and blurs it by diffraction, a far-field detector records only total intensities, and an untrained neural network iteratively reconstructs the object by simulating the diffraction and encoding process in reverse. They report recovering 118-micron features (about one seventh of the 833-micron wavelength) at a sampling ratio of 3.125%, from data that would normally be hopelessly undersampled and blurred. If correct, this would let subwavelength THz imaging use commercially practical thicker modulators and far fewer measurements, without any labeled training data.","feed_headline":"THz imaging at one-seventh the wavelength through a thick wafer","feed_subtitle":"An untrained network plus backpropagation recovers 118-micron features at 0.36 THz from 3.125% sampling.","key_machinery":"The central mechanism is the pair consisting of the untrained TSPIDL network and the angular-spectrum propagation (ASP) model. The network, a lightweight convolutional architecture, maps the detector-derived input to an estimated object intensity distribution; the ASP model (equations 2 and 3) then propagates that estimate forward by a distance $d_{\\mathrm{BP}}$, decomposing the field into homogeneous and evanescent components, and the result is compared with experimental single-pixel measurements. The load-bearing identity is setting the backpropagation distance $d_{\\mathrm{BP}}$ equal to the known forward distance $d_{\\mathrm{FP}}$ (0.5 mm, the silicon wafer thickness), which numerically reverses diffraction and recovers near-field information that the physical field lost through exponential evanescent decay.","core_discovery":"On its own terms, the paper's central discovery is that embedding the angular-spectrum propagation (ASP) model into the output layer of an untrained deep network turns single-pixel measurements of a diffracted THz field into a refocused, subwavelength image. The network takes a rough differential-ghost-imaging estimate of the blurred diffraction image as input and is constrained, through an MSE loss, to match the actual detector readouts after its output is numerically propagated by distance $d_{\\mathrm{BP}}$ and multiplied by the Hadamard masks. When $d_{\\mathrm{BP}}$ equals the true object-to-wafer distance $d_{\\mathrm{FP}} = 0.5$ mm, the reconstruction resolves three air slits separated by 118 $\\mu$m, i.e., $\\lambda_0/7$ at 0.36 THz; when $d_{\\mathrm{BP}}$ is smaller or larger, the image stays blurred or breaks into artifacts. The paper therefore claims that backpropagation through a thick silicon wafer can substitute for the ultrathin modulators previously thought necessary for near-field THz imaging, and that the untrained network's implicit prior supplies the missing high-spatial-frequency content for binary amplitude objects.","pith_inferences":["If the mechanism holds, the same 'propagate-then-untrained-prior' recipe should transfer to other spectral bands where evanescent fields matter, provided the forward model is known and the object is binary; the key comparison would be resolution against a known ground truth.","The demonstrated λ/7 figure is for high-contrast slit separations, not a general resolution metric; for grayscale or rough-surface objects the network cannot invent frequencies the wafer has already suppressed, so the practical resolution gain will likely be smaller.","A natural testable extension is to make $d_{\\mathrm{BP}}$ an adaptive parameter optimized jointly with the network weights, turning the known-depth requirement into an autofocus capability that would also reveal the method's sensitivity to distance error.","The reported 3-minute reconstruction time for a 64×64 image means the method is currently computation-bound; faster architectures or better initialization could make the approach practical for near-field microscopy where acquisition is fast."],"forward_implications":["Thick silicon wafers, 500 μm rather than a few micrometers, become sufficient for subwavelength THz imaging, eliminating the need for fragile ultrathin photomodulators.","The method reconstructs usable THz images at sampling ratios as low as 1.5625% and demonstrates its λ/7 result at 3.125%, a large reduction in acquisition time compared with the roughly 80% sampling used by earlier single-pixel THz imaging.","The same ASP-constrained network refocuses objects placed 2, 4, and 6 mm from the wafer, recovering detail that far-field diffraction had blurred.","Because the network is untrained, it generalizes to unseen objects without dataset-specific supervision, avoiding the retraining burden of supervised deep-learning THz imaging.","Reconstruction quality depends on knowing the object's position: with an incorrect backpropagation distance, artifacts such as four apertures appearing where three exist can emerge."],"supporting_citations":[{"why":"Provides the fully sampled Hadamard single-pixel reconstruction used as the reference baseline for SSIM comparisons.","marker":"[17]"},{"why":"Establishes noninvasive near-field THz single-pixel imaging with an ultrathin silicon modulator and supplies the angular-spectrum framework for evanescent fields.","marker":"[23]"},{"why":"Demonstrates compressed-sensing near-field THz imaging, the approach the paper compares against and surpasses in SNR at low sampling ratios.","marker":"[24]"},{"why":"Reports λ/100 THz resolution using an ultrathin VO2 film, representing the stringent modulator-thickness requirement this work aims to remove.","marker":"[26]"},{"why":"Provides the deep image prior concept that justifies using an untrained network as an implicit regularizer for inverse problems.","marker":"[36]"},{"why":"Supplies the phase-retrieval/angular-spectrum propagation formulation used in the paper's backpropagation output layer.","marker":"[42]"},{"why":"Serves as the standard reference for scalar diffraction theory and angular-spectrum propagation used in equations 2 and 3.","marker":"[43]"},{"why":"Presents differential ghost imaging, the algorithm used to form the network's low-quality input estimate from the raw measurements.","marker":"[45]"}],"fun_headline_variants":["THz backpropagation resolves λ/7 features","Sub-diffraction THz imaging with single pixel","Untrained net sharpens THz to λ/7","Thick wafer, thin detail: THz backpropagation","Backpropagation beats diffraction at 0.36 THz"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The paper assumes that numerically reversing diffraction with the correctly chosen distance can restore spatial frequencies that a 500-micron silicon slab has already attenuated exponentially, and that the untrained network's built-in prior can synthesize the missing subwavelength edges for binary amplitude objects.","fun_headline_variants_meta":{"raw":{"variants":["THz backpropagation resolves λ/7 features","Sub-diffraction THz imaging with single pixel","Untrained net sharpens THz to λ/7","Thick wafer, thin detail: THz backpropagation","Backpropagation beats diffraction at 0.36 THz"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000394,"raw_usage":{"total_tokens":2135,"prompt_tokens":1076,"completion_tokens":1059,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":692,"completion_tokens_details":{"reasoning_tokens":978}},"tokens_in":692,"tokens_out":1059,"duration_ms":11730,"temperature":1.0,"reasoning_tokens":978,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T00:49:34.066929+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same experiment with a grayscale object whose transmission varies continuously and contains λ/7 features, keeping $d_{\\mathrm{FP}} = 0.5$ mm and $d_{\\mathrm{BP}} = 0.5$ mm; if the reconstruction cannot resolve those features or produces artifacts, the claim that backpropagation recovers lost evanescent information is false for non-binary objects. Alternatively, deliberately mis-set $d_{\\mathrm{BP}}$ to 0.55 mm and observe whether the three-slit target begins to show the four-aperture artifact predicted at $d_{\\mathrm{BP}} = 1.0$ mm, which would confirm the method's sharp dependence on knowing $d_{\\mathrm{FP}}$.","supporting_citations":[],"review_version":1}