Pith. sign in

REVIEW 4 major objections 5 minor 13 references

Deep Learning Empowered Sub-Diffraction Terahertz Backpropagation Single-Pixel Imaging

T0 review · 4 major / 5 minor · reviewed 2026-08-16 · deepseek-v4-flash

Pith's one-line read The paper reports that an untrained neural network constrained by angular-spectrum backpropagation recovers 118-micron (λ/7) features at 0.36 THz through a 500-micron silicon wafer from far-field single-pixel measurements taken at a…

desk verdict Promising untrained DIP + ASP combination for THz SPI, but the λ/7 resolution claim is likely prior-driven and lacks the controls to back it up. read the letter →

arxiv 2505.07839 v2 pith:VPRS5JRU submitted 2025-05-05 eess.IV cs.AI

classification eess.IVcs.AI
keywords terahertzimagingsub-diffractionsingle-pixeluntrainedneuralnetworkangularspectrumpropagationcomputationalsamplingrationear-field
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper claims that a major practical obstacle in terahertz single-pixel imaging—the need for ultrathin photomodulators to capture evanescent, subwavelength information—can be removed by doing the 'thinness' in software. The authors describe an imaging system at 0.36 THz in which a 500-micron silicon wafer both encodes the object field and blurs it by diffraction, a far-field detector records only total intensities, and an untrained neural network iteratively reconstructs the object by simulating the diffraction and encoding process in reverse. They report recovering 118-micron features (about one seventh of the 833-micron wavelength) at a sampling ratio of 3.125%, from data that would normally be hopelessly undersampled and blurred. If correct, this would let subwavelength THz imaging use commercially practical thicker modulators and far fewer measurements, without any labeled training data.

What carries the argument

The central mechanism is the pair consisting of the untrained TSPIDL network and the angular-spectrum propagation (ASP) model. The network, a lightweight convolutional architecture, maps the detector-derived input to an estimated object intensity distribution; the ASP model (equations 2 and 3) then propagates that estimate forward by a distance $d_{\mathrm{BP}}$, decomposing the field into homogeneous and evanescent components, and the result is compared with experimental single-pixel measurements. The load-bearing identity is setting the backpropagation distance $d_{\mathrm{BP}}$ equal to the known forward distance $d_{\mathrm{FP}}$ (0.5 mm, the silicon wafer thickness), which numerically reverses diffraction and recovers near-field information that the physical field lost through exponential evanescent decay.

What would settle it

Run the same experiment with a grayscale object whose transmission varies continuously and contains λ/7 features, keeping $d_{\mathrm{FP}} = 0.5$ mm and $d_{\mathrm{BP}} = 0.5$ mm; if the reconstruction cannot resolve those features or produces artifacts, the claim that backpropagation recovers lost evanescent information is false for non-binary objects. Alternatively, deliberately mis-set $d_{\mathrm{BP}}$ to 0.55 mm and observe whether the three-slit target begins to show the four-aperture artifact predicted at $d_{\mathrm{BP}} = 1.0$ mm, which would confirm the method's sharp dependence on knowing $d_{\mathrm{FP}}$.

Watch

Extended reading notes

Core claim

On its own terms, the paper's central discovery is that embedding the angular-spectrum propagation (ASP) model into the output layer of an untrained deep network turns single-pixel measurements of a diffracted THz field into a refocused, subwavelength image. The network takes a rough differential-ghost-imaging estimate of the blurred diffraction image as input and is constrained, through an MSE loss, to match the actual detector readouts after its output is numerically propagated by distance $d_{\mathrm{BP}}$ and multiplied by the Hadamard masks. When $d_{\mathrm{BP}}$ equals the true object-to-wafer distance $d_{\mathrm{FP}} = 0.5$ mm, the reconstruction resolves three air slits separated by 118 $\mu$m, i.e., $\lambda_0/7$ at 0.36 THz; when $d_{\mathrm{BP}}$ is smaller or larger, the image stays blurred or breaks into artifacts. The paper therefore claims that backpropagation through a thick silicon wafer can substitute for the ultrathin modulators previously thought necessary for near-field THz imaging, and that the untrained network's implicit prior supplies the missing high-spatial-frequency content for binary amplitude objects.

Load-bearing premise

The paper assumes that numerically reversing diffraction with the correctly chosen distance can restore spatial frequencies that a 500-micron silicon slab has already attenuated exponentially, and that the untrained network's built-in prior can synthesize the missing subwavelength edges for binary amplitude objects.

Editorial extensions

If this is right

  • Thick silicon wafers, 500 μm rather than a few micrometers, become sufficient for subwavelength THz imaging, eliminating the need for fragile ultrathin photomodulators.
  • The method reconstructs usable THz images at sampling ratios as low as 1.5625% and demonstrates its λ/7 result at 3.125%, a large reduction in acquisition time compared with the roughly 80% sampling used by earlier single-pixel THz imaging.
  • The same ASP-constrained network refocuses objects placed 2, 4, and 6 mm from the wafer, recovering detail that far-field diffraction had blurred.
  • Because the network is untrained, it generalizes to unseen objects without dataset-specific supervision, avoiding the retraining burden of supervised deep-learning THz imaging.
  • Reconstruction quality depends on knowing the object's position: with an incorrect backpropagation distance, artifacts such as four apertures appearing where three exist can emerge.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the mechanism holds, the same 'propagate-then-untrained-prior' recipe should transfer to other spectral bands where evanescent fields matter, provided the forward model is known and the object is binary; the key comparison would be resolution against a known ground truth.
  • The demonstrated λ/7 figure is for high-contrast slit separations, not a general resolution metric; for grayscale or rough-surface objects the network cannot invent frequencies the wafer has already suppressed, so the practical resolution gain will likely be smaller.
  • A natural testable extension is to make $d_{\mathrm{BP}}$ an adaptive parameter optimized jointly with the network weights, turning the known-depth requirement into an autofocus capability that would also reveal the method's sensitivity to distance error.
  • The reported 3-minute reconstruction time for a 64×64 image means the method is currently computation-bound; faster architectures or better initialization could make the approach practical for near-field microscopy where acquisition is fast.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper proposes an untrained physics-constrained neural network, termed untrained TSPIDL, for terahertz single-pixel imaging (THz SPI). A 0.36 THz continuous-wave source illuminates an object, the transmitted field is encoded by a 500-μm-thick silicon photomodulator, and a far-field single-pixel detector records intensity. The network is optimized so that simulated single-pixel measurements from its output match actual detector readings, with an angular-spectrum propagation (ASP) model embedded in the output layer to backpropagate the diffracted field from the recording plane to the object plane. The authors claim a spatial resolution of 118 μm (about λ0/7) using this approach at a sampling ratio of 3.125%, and demonstrate refocusing of images at distances up to 6 mm. The central demonstration is a three-slit target with 118 μm separation recovered when the backpropagation distance is set to the known wafer thickness of 0.5 mm.

Significance. If the resolution claim were substantiated, this would be a valuable contribution: the method requires no training data, uses a thick silicon modulator rather than ultrathin photomodulators, and achieves very low sampling ratios. The integration of an untrained deep-image prior with a physical propagation model is a sensible and timely idea. However, the current evidence is not sufficient to establish that the recovered 118 μm features are information from the measurements rather than artifacts of the network's implicit prior. The paper also lacks quantitative rigor in the resolution assessment, with a single target, no error bars, and SSIM computed against HSPI reconstructions rather than ground truth.

major comments (4)
  1. [Methods, Eqs. (2)–(3); Results, Fig. 4f] The λ/7 resolution claim is not supported as a property of the imaging system because the 118 μm feature corresponds to a normalized spatial frequency u ≈ 7 in air (u ≈ 2 in silicon), and the evanescent transfer factor in Eq. (3), exp(−2πd/λ·√(u²−1)), is below 10⁻¹⁰ over d = 0.5 mm inside silicon. The measured far-field intensities therefore contain essentially no information about that gap, so the recovered contrast in Fig. 4f3 must come from the network's implicit deep-image prior. The authors need a control experiment—e.g., a non-periodic 118 μm feature, a shuffled-measurement control, or a no-ASP ablation—to distinguish measurement-driven recovery from prior-driven hallucination before claiming a spatial resolution of the imaging system.
  2. [Results, Fig. 4f1–f4 and Fig. 4g] The resolution claim rests on visual separation of three slits in a single target, with no error bars, no repeated trials, and no quantitative contrast metric. The cross-sectional profiles in Fig. 4g show a trend, but the authors should report repeated measurements, a resolution metric (e.g., modulation depth or a Rayleigh-type criterion), and ideally compare against the known optical mask rather than only against HSPI reference images. The current evidence is anecdotal and cannot support a quantitative λ/7 resolution statement.
  3. [Methods, Eqs. (2)–(3) and Eq. (6)] The ASP model is defined for complex field propagation, while the network output is described as an intensity distribution (Eq. (6): O0 = |E0|²). The manuscript never specifies how the missing phase is obtained or why applying the complex ASP propagator to an intensity-only estimate is physically valid. This missing step is load-bearing for the backpropagation claim, and the paper should clarify the field-to-intensity conversion and justify the intensity-only ASP approximation.
  4. [Results, paragraphs after Fig. 5 and Conclusion] The method requires prior knowledge of the forward propagation distance dFP, as the authors explicitly state, and the dBP sweep in Fig. 4f1–f4 shows that a wrong dBP produces artifacts (four apertures instead of three at dBP = 1.0 mm). This means the sub-diffraction result is conditional on feeding the correct object position into the algorithm. The paper should state this conditioning prominently in the abstract and conclusion, and it should discuss whether the method can be extended to unknown object depth or whether the claimed resolution only holds for known binary targets at a known distance.
minor comments (5)
  1. [Abstract and Fig. 4 caption] The abstract advertises an ultralow sampling ratio of 1.5625%, but the λ/7 result in Fig. 4 uses SR = 3.125%. Please clarify which sampling ratio belongs to which claim and avoid implying that the sub-diffraction result was obtained at the lowest stated ratio.
  2. [Fig. 3 caption] The caption refers to images 'at d = 0 mm' while the text states the object is at dFP = 0.5 mm from the back surface of the wafer; please make the distance definition consistent.
  3. [Results, Fig. 3a1,b1 references] The sentence 'Even at a low SR = 1.5625%, the SSIM can reach 0.67 and the reconstructed images are acceptable (Figure 3a1,b1)' appears to cite the full-sampling HSPI reference images rather than the low-SR reconstructions; the figure references should be corrected.
  4. [Eq. (5)] Equation (5) has garbled mathematical symbols in the typeset text; it should be rewritten with clear notation so that the DGI estimate is unambiguous.
  5. [Throughout] There are several typographical errors, such as 'state-of-art' in the Results section and 'near-filed' in the Introduction; a careful proofread is recommended.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the method is a physics-constrained untrained-network inverse problem fitted to measured single-pixel data, with no fitted parameter renamed as a prediction and no load-bearing self-citation chain.

full rationale

The derivation chain is self-contained rather than circular. The measured detector vector I and the known Hadamard patterns P are the inputs; the network output O_hat_0 is optimized so that angular-spectrum propagation followed by simulated pattern measurements reproduces I, via Equations (2)-(8). This is a standard inverse-problem formulation: the network parameters are not pre-trained on the target, the loss is the MSE against measured detector readings, and no parameter is fitted to the claimed 118-um (lambda/7) result. The three-slit structure is used only as a validation object; dBP is set equal to the separately known forward distance dFP=0.5 mm, and the paper explicitly acknowledges that wrong dBP produces artifacts ('our proposed algorithm relies on prior knowledge of the object position, specifically the forward propagation distance dFP... In the absence of this information, artifacts can appear'). The ASP equations are standard scalar diffraction theory (Born-Wolf), and the untrained-network prior is attributed to Ulyanov et al., an external reference, not to an unverified self-citation. Self-citations in the reference list are background and are not load-bearing. The serious concerns in the paper are validation limitations rather than constructional circularity: the phase specification for propagating the intensity output is not given, and no control experiment (e.g., non-periodic targets or measurement-shuffling) quantitatively separates data-driven recovery from deep-image-prior hallucination of binary slit edges. These affect whether the lambda/7 claim is convincingly established, but they do not make the derivation equivalent to its inputs by construction.

Assumptions & free parameters 2 free parameters · 4 assumptions · 0 invented entities

The central reconstruction depends on the scalar diffraction model being the correct forward operator, on the binary-amplitude object assumption, and on the known propagation distance dFP. The paper does not introduce new physical entities. The externally supplied numbers that steer the result are dBP and the regularization or hyperparameter settings; TV weight and network architecture details are not in the main text.

free parameters (2)
  • backpropagation distance dBP = 0.5 mm (scanned values 0.1, 0.3, 0.5, 1.0 mm)
    The λ/7 resolution is only reported at dBP=0.5 mm; dBP=1.0 mm produces artifacts and dBP=0.1 mm is blurred. The authors acknowledge requiring prior knowledge of dFP, so this is an externally supplied, hand-set parameter that the claim depends on.
  • TV regularization weight and network hyperparameters = Not specified in main text (Supporting Information)
    Deep-image-prior reconstructions are sensitive to the total-variation weight, learning rate, and architecture; without code or SI values these are hand-tuned choices that affect the reported SSIM and SNR.
assumptions (4)
  • domain assumption Scalar angular-spectrum propagation (eqs 2-3) accurately models the THz field from the object plane through the 500-um silicon wafer and free space to the detector.
    The whole backpropagation layer is built on this model; errors in thickness, refractive index, or scalar approximation would shift the refocusing distance.
  • domain assumption The object is a binary amplitude mask, and the neural-network prior plus TV regularization can synthesize subwavelength spatial frequencies lost by evanescent decay.
    Evanescent components with 118-um features decay strongly over 500 um, so the reconstructed sharp edges are supplied by the prior, not directly by measured data.
  • domain assumption The THz illumination is a normally incident plane wave (eq 1) and the single-pixel measurement is linear with additive noise (eqs 4 and 10).
    The experimental beam is collimated but not an ideal plane wave; detector nonlinearity and pattern misalignment are ignored in the model.
  • domain assumption Deep image prior provides a suitable natural-image prior for this inverse problem.
    The untrained network's convolutional architecture biases outputs toward natural images, and this is the mechanism that fills missing information; the paper does not validate the prior against other reconstruction priors.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Deep Learning Empowered Sub-Diffraction Terahertz Backpropagation Single-Pixel Imaging." pith.science (2026). https://pith.science/paper/VPRS5JRU

@misc{pith2026250507839,
  author       = {Pith},
  title        = {Pith review of: Deep Learning Empowered Sub-Diffraction Terahertz Backpropagation Single-Pixel Imaging},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/VPRS5JRU}},
  note         = {Machine review of arXiv:2505.07839}
}
read the original abstract

Terahertz single-pixel imaging (THz SPI) has garnered widespread attention for its potential to overcome challenges associated with THz focal plane arrays. However, the inherently long wavelength of THz waves limits imaging resolution, while achieving subwavelength resolution requires harsh experimental conditions and time-consuming processes. Here, we propose a sub-diffraction THz backpropagation SPI technique. We illuminate the object with continuous-wave 0.36-THz radiation ({\lambda}0 = 833.3 {\mu}m). The transmitted THz wave is modulated by prearranged patterns generated on a 500-{\mu}m-thick silicon wafer and subsequently recorded by a far-field single-pixel detector. An untrained neural network constrained with the physical SPI process iteratively reconstructs the THz images with an ultralow sampling ratio of 1.5625%, significantly reducing the long sampling times. To further suppress the THz diffraction-field effects, a backpropagation SPI from near field to far field is implemented by integrating with a THz physical propagation model into the output layer of the network. Notably, using the thick wafer where THz evanescent field cannot be fully recorded, we achieve a spatial resolution of 118 {\mu}m (~{\lambda}0/7) through backpropagation SPI, thus eliminating the need for ultrathin photomodulators. This approach provides an efficient solution for advancing THz microscopic imaging and addressing other inverse imaging challenges.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

13 extracted references · 13 canonical work pages

  1. [1]

    Y.; Zhu, L

    (1) Yan, Z. Y.; Zhu, L. G.; Meng, K.; Huang, W. X. & Shi, Q. W., THz Medical Imaging: From in Vitro to in Vivo. Trends Biotechnol. 2022, 40(7): 816–830. (2) Peng, Y.; Shi, C.; Zhu, Y.; Gu, M. & Zhuang, S., Terahertz Spectroscopy in Biomedical Field: A Review on Signal-to-Noise Ratio Improvement. PhotoniX 2020, 1(1):

  2. [6]

    I.; Phillips, D

    (24) Stantchev, R. I.; Phillips, D. B.; Hobson, P.; Hornett, S. M.; Padgett, M. J. & Hendry, E., Compressed Sensing with near-Field THz Radiation. Optica 2017, 4(8): 989–992. (25) Olivieri, L.; Totero Gongora, J. S.; Pasquazi, A. & Peccianti, M., Time -Resolved Nonlinear Ghost Imaging. ACS Photonics 2018, 5(8): 3379–3388. (26) Chen, S.; Du, L.; Meng, K.; ...

  3. [8]

    A., Se lf-Consistent Optical Parameters of Intrinsic Silicon at 300k Including Temperature Coefficients

    (48) Green, M. A., Se lf-Consistent Optical Parameters of Intrinsic Silicon at 300k Including Temperature Coefficients. Sol. Energy Mater. Sol. Cells 2008, 92(11): 1305–1310. (49) Hooper, I. R.; Grant, N. E.; Barr, L. E.; Hornett, S. M.; Murphy, J. D. & Hendry, E., High Efficiency Photomodulators for Millimeter Wave and THz Radiation. Sci. Rep. 2019, 9(1)...

  4. [11]

    & Qadir, J., Untrained Neural Network Priors for Inverse Imaging Problems: A Survey

    (39) Qayyum, A.; Ilahi, I.; S hamshad, F.; Boussaid, F.; Bennamoun, M. & Qadir, J., Untrained Neural Network Priors for Inverse Imaging Problems: A Survey. IEEE Trans. Pattern Anal. Mach. Intell. 2023, 45(5): 6511–6536. (40) Jin, X.; Zhao, J.; Wang, D.; Healy, J. J.; Rong, L.; Wang, Y. & Lin, S., Continuous -Wave Terahertz in -Line Holographic Diffraction...

  5. [12]

    A.; Aloma iny, A

    (3) Ren, A.; Zahid, A.; Fan, D.; Yang, X.; Imran, M. A.; Aloma iny, A. & Abbasi, Q. H., State - of-the-Art in Terahertz Sensing for Food and Water Security — a Comprehensive Review. Trends Food Sci. Technol. 2019, 85: 241–251. (4) Costa, F. B.; Machado, M. A.; Bonfait, G. J.; Vieira, P. & Santos, T. G., Continuous Wave Terahertz Imaging for NDT: Fundament...

  6. [55]

    & Boyd, R

    (31) Zhao, J.; Dai, J.; Braverman, B.; Zhang, X. & Boyd, R. W., Compressive Ultrafast Pulse Measurement via Time-Domain Single-Pixel Imaging. Optica 2021, 8(9): 1176–1185. (32) Song, K.; Bian, Y.; Wang, D.; Li, R.; Wu, K.; Liu, H.; Qin, C., et al., Advances and Challenges of Single-Pixel Imaging Based on Deep Learning. Laser Photon. Rev. 2024: 2401397. 33...

  7. [77]

    & Situ, G., Far -Field Super- Resolution Ghost Imaging with a Deep Neural Network Constraint

    (38) Wang, F.; Wang, C.; Chen, M.; Gong, W.; Zhang, Y.; Han, S. & Situ, G., Far -Field Super- Resolution Ghost Imaging with a Deep Neural Network Constraint. Light: Sci. Appl. 2022, 11(1):

  8. [99]

    (28) Olivieri, L.; Gongora, J. S. T.; Peters, L.; Cecconi, V.; Cutrona, A.; Tunesi, J.; Tucker, R., et al., Hyperspectral Terahertz Microscopy via Nonlinear Ghost Imaging. Optica 2020, 7(2): 186–191. (29)Olivieri, L.; Peters, L.; Cecconi, V.; Cutrona, A.; Rowley, M.; Totero Gongora, J. S.; Pasquazi, A., et al., Terahertz Nonlinear Ghost Imaging via Plane ...

Show all 13 references
  1. [416]

    (22) Liang, J.; Zhang, J.; Wang, Z.; Wang, R.; Yao, Z.; Singh, R.; Tian, Z., et al., Photoactive Polymer-Silicon Heterostructures for Terahertz Spat ial Light Modulation and Video -Rate Single-Pixel Compressive Imaging. Adv. Funct. Mater. 2025: 2422478. 32 (23) Stantchev, R. I...

  2. [2013]

    & Situ, G., Single -Pixel Imaging Using Physics Enhanced Deep Learning

    (44) Wang, F.; Wang, C.; Deng, C.; Han, S. & Situ, G., Single -Pixel Imaging Using Physics Enhanced Deep Learning. Photon. Res. 2022, 10(1): 104–110. (45) Ferri, F.; Magatti, D.; Lugiato, L. A. & Gatti, A., Differential Ghost Imaging. Phys. Rev. Lett. 2010, 104(25): 253603. (4...

  3. [2535]

    P.; Morandotti, R.; Piccoli, R

    (18) Zanotto, L.; Balistreri, G.; Rovere, A.; Kwon, O. P.; Morandotti, R.; Piccoli, R. & Razzari, L., Terahertz Scanless Hypertemporal Imaging. Laser Photon. Rev. 2023, 17(8): 2200936. (19) Chen, H.; Wang, X.; Liu, S.; Cao, Z.; Li, J.; Zhu, H.; Li, S., et al., Monolithic Multi...

  4. [2634]

    ACS Photonics 2024, 11(2): 362–368

    (10) Cecconi, V.; Kumar, V.; Bertolotti, J.; Peters , L.; Cutrona, A.; Olivieri, L.; Pasquazi, A., et al., Terahertz Spatiotemporal Wave Synthesis in Random Systems. ACS Photonics 2024, 11(2): 362–368. (11) Ismagilov, A.; Lappo-Danilevskaya, A.; Grachev, Y.; Nasedkin, B.; Zali...

  5. [2911]

    (9) Kumar, V.; Cecconi, V.; Peters, L.; Bertolotti, J.; Pasquazi, A.; Totero Gongora, J. S. & Peccianti, M., Deterministic Terahertz Wave Control in Scattering Media. ACS Photon. 2022, 9:

Pith tools

Reviewed August 16, 2026 · model on record in the stance chip above.