{"id":"dc1e9366-fe43-4a09-b15c-7124a1fbcac6","arxiv_id":"2502.06025","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Spatially incoherent diffractive processors can learn to create arbitrary 3D point spread functions, enabling all-optical snapshot 3D and multispectral imaging in numerical simulations.","lead":"This paper shows in simulations that layered light-processing surfaces can be trained to create any desired 3D blur pattern between a volume of light sources and a detector, a capability called universal 3D point spread function engineering. A generalist might care because it points toward snapshot 3D multispectral cameras that need no filters, scanning, or computer reconstruction.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Universality rests on an empirical DOF rule tested on single random targets; no proof or repeated trials.","rationale":"The reader's weakest_assumption correctly identified two premises: the independent-emitter linearity and the unproven DOF scaling rule. I agree that both are load-bearing, but the independent-emitter assumption is explicitly stated and qualified in the Discussion, making it an honest boundary condition rather than an internal weakness. The more serious gap for the paper's advertised 'universal' and 'rigorous' claims is the empirical nature of the N≈2N_iN_o threshold: it is derived from single random target matrices and lacks repeated trials, theoretical guarantees, or released code. A single counterexample—a physically admissible target matrix that cannot be approximated at the claimed DOF—would falsify the universality assertion as stated. The proposed check directly tests that risk by sweeping multiple random and structured targets and measuring error distributions. If the check passes, the conditional acceptance is appropriate; if it fails, the claim should be narrowed to 'demonstrated on random matrices' rather than 'arbitrary.' Since the reader already issued a CONDITIONAL verdict that conditions on tempering the claim and releasing code, my stress-test does not change the verdict; it sharpens the specific condition that should be satisfied.","tokens_in":8578,"tokens_out":7808,"duration_ms":87272,"concrete_test":"For fixed small N_i,N_o (e.g., 4 input planes × 6×6 voxels and matching output volume), train K=4 diffractive processors with N=2N_iN_o on at least 20 independent random nonnegative target matrices and on several structured targets: an identity-like permutation of voxels, sparse impulse PSFs, and a matrix whose column sums vary strongly. Report the distribution of normalized elementwise errors and compare with the error level in Fig. 2c. If any target fails to reach that error level, the universal claim at the stated DOF is unsupported; if all succeed, the empirical rule is substantially strengthened.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that a diffractive processor can synthesize an arbitrarily defined set of 3D PSFs given sufficient features—depends on the empirical threshold N≈2N_iN_o shown in Fig. 2b. Each plotted point is a separately trained processor for one random target matrix; there are no repeated trials, no error bars, and no theoretical guarantee. The factor-of-2 argument (phase-only features have one real DOF versus N_iN_o nonnegative matrix entries) is a heuristic lower-bound intuition, not a proof that the physical map from phase features to the intensity matrix is surjective onto the nonnegative orthant. Because the targets are i.i.d. uniform random matrices, they may not be representative of all physically plausible PSF sets; a structured or adversarial target could require more features or be unachievable. The paper also provides no code or data to independently reproduce Fig. 2c/2d. Thus the 'arbitrary' and 'rigorous' wording exceeds the evidence, and the universality claim remains an extrapolation from a few small-scale simulations. The linearity assumption (non-interacting emitters) is explicitly stated and is not the weakest link; the unproven DOF scaling is what carries the claim of universality.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a method for universal point spread function (PSF) engineering in three dimensions using spatially incoherent diffractive optical processors composed of cascaded transmissive surfaces. The central claim is that, given sufficient optimizable phase-only features N distributed over K≥4 surfaces, such a processor can approximate, with negligible error, any arbitrarily defined set of spatially varying 3D PSFs, corresponding to an arbitrary nonnegative linear intensity transformation between input and output volumes. The authors numerically demonstrate this by training diffractive networks to approximate random nonnegative target matrices, reporting an empirical threshold of N≈2N_iN_o for low error. They also present numerical demonstrations of snapshot 3D imaging and snapshot multispectral 3D imaging using the engineered PSFs. The physical assumptions are explicitly stated: the input emitters are independent and non-interacting, so the system is linear in intensity.","tokens_in":8805,"tokens_out":4656,"duration_ms":44690,"significance":"If the central claim holds, the work would substantially extend PSF engineering from 2D pupil-plane masks to arbitrary 3D spatially varying PSFs, enabling all-optical 3D and multispectral imaging without axial scanning, spectral filters, or digital reconstruction. This would be a meaningful advance for computational imaging, microscopy, and optical information processing. The numerical demonstrations in Figs. 4 and 5 show that the proposed approach can realize these applications at least at a small scale. However, the universality claim rests on an empirical, unproven design rule, and the paper provides no code, data, or repeated-trial statistics to support the generality of the results. The significance is therefore conditional on strengthening the numerical evidence and carefully qualifying the scope of the universality claim.","major_comments":[{"comment":"The manuscript reports only numerical simulations and contains no code, data, or data-availability statement. Because the universality claim rests entirely on the numerical results in Figs. 2–5, the absence of reproducible artifacts hinders independent verification of the convergence threshold and the reported error values. I request that the authors release the simulation code and the trained network parameters, or at minimum provide a detailed specification of the forward model and optimization hyperparameters in the main text or supplementary material.","section":"Results (Fig. 2b)"},{"comment":"The linearity assumption—independent, non-interacting emitters with no shadowing or re-excitation—is explicitly acknowledged in the Discussion, and the authors correctly note that if this assumption is violated the transformation becomes nonlinear and object-dependent. However, the abstract and introduction present the method as 'universal PSF engineering' without prominently stating this restriction. The universality claim should be qualified in the abstract and introduction to make clear that it applies only to spatially incoherent, non-interacting emitters, which is a meaningful constraint for many imaging scenarios.","section":"Discussion"},{"comment":"The application demonstrations use very small volumes: Fig. 4 uses 4 input planes each discretized into 6×6 pixels, and Fig. 5 uses 3 input planes at 3 wavelengths with 6×6 pixels per plane. The required number of features N≈2N_iN_o grows quadratically with the product of input and output voxel counts, which would become impractically large for realistic imaging volumes (e.g., 10^6 input and output voxels would require on the order of 2×10^12 features). The paper does not discuss this scalability constraint, yet it is central to the practical significance of the claimed universality. I ask the authors to address the scalability of the approach and to clarify the range of problem sizes for which the design rule is practically feasible.","section":"Figs. 4 and 5"}],"minor_comments":[{"comment":"There is a typo in the sentence 'see Fig. 5c) )'—an extra closing parenthesis appears after 'Fig. 5c'.","section":"Discussion (Fig. 5 paragraph)"},{"comment":"The phrase 'rigorously analyze' overstates the nature of the evidence: the paper provides numerical demonstrations, not a formal mathematical proof. Suggest rewording to 'numerically analyze' or 'characterize'.","section":"Abstract"},{"comment":"The main text references 'Methods' for the forward model and optimization details, but the Methods section is placed in the Supplementary Information. The main text should explicitly direct readers to the Supplementary Methods, especially since the Methods are not included in the main manuscript.","section":"Throughout"},{"comment":"For the diffraction-limit study, it would be informative to overlay the theoretical axial resolution limit (e.g., 2λ/NA²) on the error curves in Fig. 3a and 3b to allow a direct quantitative comparison between the numerical error onset and the expected diffraction-limited resolution.","section":"Fig. 3"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is from a group with a strong track record in diffractive optical networks, and the extension from 2D to 3D PSF engineering is a natural and potentially valuable step. However, the central universality claim is supported only by a small number of un-replicated numerical experiments, and the lack of code/data availability is a concern for a claim of this strength. The reviewers were unable to examine the supplementary Methods because it is not included in the arXiv text; I would encourage the editor to request a full version with supplementary materials before any final decision. The scale of the demonstrations is also a limiting factor, and I would not recommend acceptance until the numerical evidence is made more robust and the scalability limitations are discussed."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a solid numerical extension of the Ozcan group's earlier 2D universal linear transformation results to 3D volumes, with spatially varying 3D PSFs. The core feasibility claim is credible for small-scale cases; the 'universal' wording, however, outruns the evidence because the DOF threshold is empirical, not proven.\n\nWhat's genuinely new: prior work did 2D input/output planes; here they let the input and output be volumes, define target PSF columns for each input voxel, and train diffractive layers to approximate a random nonnegative target matrix. Using externally specified random targets is non-circular, and the error curves in Fig. 2 do drop to negligible values near N≈2N_iN_o. The axial-resolution analysis (Fig. 3) is a useful addition, and the snapshot 3D and multispectral demos (Figs. 4-5) show the concept working end-to-end in simulation. The authors also deserve credit for explicitly stating the non-interacting emitter assumption and discussing what happens if it is violated (object-dependent nonlinearity). That is more honest than most papers in this space.\n\nSoft spots, in proportion: (1) The load-bearing N≈2N_iN_o rule is read off interpolated curves from one or a few random target matrices per point, with no error bars over repeated training runs and no code or data released. That makes the 'arbitrary' and 'rigorous' language too strong. It is a plausible heuristic (phase-only features give one real DOF each) but not a surjectivity proof. (2) The imaging demos are small (6×6 pixels, 3-4 planes), which is acceptable for proof of concept but not a demonstration of scaling. (3) The linearity assumption is real, but the authors flag it clearly, so it's not a hidden flaw.\n\nI largely agree with the stress-test: the weak link is the unproven DOF scaling, not the linearity assumption. The paper would be stronger if the authors released code and repeated training runs with variance, and if they replaced 'universal' with 'can approximate arbitrary target matrices at the scales tested.'\n\nWho it's for: readers working on diffractive networks, incoherent computational imaging, or PSF engineering. It deserves a serious referee; the numerical results are clean and the topic is important. I'd send it to review but ask for data/code and tempered claims before acceptance.","headline":"Clean numerical extension of 2D universal diffractive transformations to 3D PSF engineering, but the universality claim floats on an empirical DOF threshold with no error bars or code.","tokens_in":9353,"tokens_out":2968,"would_cite":true,"duration_ms":30077,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A trained stack of thin passive surfaces can synthesize an arbitrary set of 3D point-spread functions for incoherent light, enabling snapshot 3D and multispectral imaging without digital reconstruction.","keywords":["point spread function engineering","diffractive optical processor","3D imaging","multispectral imaging","spatially incoherent light","optical information processing","computational imaging","phase-only optics"],"falsifier":"Fabricate the K=4, N approximately 2 Ni No processor from the numerical study and measure its full intensity transformation on random nonnegative test inputs; if the measured outputs deviate from A i beyond the simulation's negligible error, the claimed universal synthesis does not hold. A second decisive test is to put two scatterers in the input volume and check whether the output is the sum of their individual PSFs, since any shadowing-induced deviation would violate the linear model.","tokens_in":8326,"feed_emoji":"🔬","tokens_out":7641,"duration_ms":72491,"temperature":0.7,"pith_summary":"This paper claims that a stack of thin, passive, phase-only optical surfaces can be trained so that every voxel of a 3D input volume produces a prescribed diffraction-limited intensity pattern at an output plane. Because the patterns can be chosen independently for each input voxel and each wavelength, the same all-optical stack can implement essentially any nonnegative linear transformation between 3D intensity volumes. The authors demonstrate numerically that such a stack can image a 3D scene from a single detector snapshot, with axial planes separated by about 2.67 wavelengths, and can separate three emission wavelengths at once without spectral filters, axial scanning, or digital reconstruction. The significance, if true, is that complex 3D and multispectral imaging tasks could be performed passively and instantaneously by a fabricated optical element.","feed_headline":"A trained stack of thin optical surfaces creates any 3D point-spread function","feed_subtitle":"Diffractive layers turn a single camera snapshot into a 3D, multispectral image, with no filters, scanning, or reconstruction.","key_machinery":"The central object is the intensity transformation matrix A: its columns are the target 3D PSFs, so an input intensity vector i over the input voxels produces an output intensity o = A i, and the processor is trained to realize A-hat approximately equal to A. The processor itself is a cascade of K thin phase-only diffractive surfaces whose N feature values are optimized by backpropagation through a digital forward model. The load-bearing design rule is that roughly N approximately 2 Ni No optimizable features distributed over K at least 4 surfaces suffice to approximate any nonnegative target A with negligible error, with depth suppressing the larger errors seen for K = 2; the diffraction limit enters through the chosen voxel spacing and plane-to-plane distances, which set the axial resolution of the achievable PSFs.","core_discovery":"The central claim is that a spatially incoherent diffractive processor, built from K cascaded transmissive surfaces with N optimizable phase-only features, can synthesize an arbitrarily defined set of 3D diffraction-limited PSFs mapping Ni input voxels to No output voxels, provided the number of features is large enough. In numerical experiments with random nonnegative target matrices A, the approximation error between the all-optical transformation A-hat and A becomes negligible as N approaches 2 Ni No when K is at least 4, and the factor 2 is attributed to the phase-only nature of the features. The authors also show spectrally engineered 3D PSFs, using them to demonstrate snapshot 3D imaging and snapshot 3D multispectral imaging from a single output frame, with demultiplexing accomplished simply by rearranging detector pixel values according to the assigned input planes and wavelengths.","pith_inferences":["The paper's rule that N approximately 2 Ni No features suffice is inferred from random target matrices at modest sizes; scaling it to large volumes is an extrapolation, and a proof or a larger-scale test would be needed before treating the universality claim as quantitative.","Because the output pixels are multiplexed across input planes and wavelengths, the achievable field of view shrinks as the number of axial planes and colors grows; the paper does not address this information-capacity trade-off explicitly.","A direct experimental falsification would be to fabricate one of the optimized K=4 designs and measure the full transformation on random test inputs, since the paper's demonstrations are entirely numerical.","The same intensity-linear architecture suggests a route to depth-resolved fluorescence lifetime or polarimetric imaging if the programmed PSFs encode those additional degrees of freedom, but that extension is not explored here."],"forward_implications":["A single passive diffractive processor can perform any prescribed nonnegative linear intensity transformation between 3D volumes at the speed of light, with no digital processing step.","One output snapshot from a single detector array can recover the full 3D distribution of independent emitters, with axial planes separated by about 2.67 wavelengths, by pixel rearrangement alone.","Spectral information can be encoded in the same snapshot: different wavelengths are routed to different output pixels, so multispectral 3D imaging needs no filters, no scanning, and no reconstruction.","The optimized layer designs transfer directly to other wavelengths by rescaling physical dimensions, since the spectral engineering relies on free-space dispersion rather than material dispersion.","Shallow processors with only two surfaces are not sufficient for accurate 3D PSF synthesis, so structural depth is a necessary part of the design."],"supporting_citations":[{"why":"Showed that diffractive surfaces can synthesize an arbitrary linear transformation, the 2D predecessor of the 3D PSF synthesis claimed here.","marker":"[28]"},{"why":"Established wavelength-multiplexed universal linear transformations, which underpins the spectral engineering of 3D PSFs in this paper.","marker":"[29]"},{"why":"Introduced universal linear intensity transformations using spatially incoherent diffractive processors, directly supplying the intensity-linear forward model used here.","marker":"[30]"},{"why":"Analyzed the all-optical information-processing capacity of diffractive surfaces, providing the depth-dependent error background for the K comparison.","marker":"[27]"},{"why":"Describes nonlinear encoding in diffractive processing, which the paper invokes to define the regime in which its linear model fails.","marker":"[33]"},{"why":"Supplies the deep-learning optimization used to train the diffractive features in the numerical demonstrations.","marker":"[37]"},{"why":"Demonstrates complex-valued universal linear transformations with spatially incoherent diffractive networks, the contrast case for the phase-only design rule.","marker":"[34]"}],"fun_headline_variants":["Universal PSF engineering enables arbitrary 3D PSFs for snapshot imaging","Cascaded diffractive layers synthesize any diffraction-limited 3D PSF","One shot, any 3D PSF: diffractive processor enables snapshot multispectral imaging","Universal PSF engineering gives snapshot 3D multispectral imaging without filters","Any 3D PSF is possible with cascaded diffractive layers"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that each emitter in the input volume radiates independently and does not block, shadow, or re-excite its neighbors, so the output is exactly the sum of the individual point-spread functions; any emitter-to-emitter coupling makes the transformation nonlinear and object-dependent, which would void the claimed universality.","fun_headline_variants_meta":{"raw":{"variants":["Universal PSF engineering enables arbitrary 3D PSFs for snapshot imaging","Cascaded diffractive layers synthesize any diffraction-limited 3D PSF","One shot, any 3D PSF: diffractive processor enables snapshot multispectral imaging","Universal PSF engineering gives snapshot 3D multispectral imaging without filters","Any 3D PSF is possible with cascaded diffractive layers"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001184,"raw_usage":{"total_tokens":4856,"prompt_tokens":878,"completion_tokens":3978,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":494,"completion_tokens_details":{"reasoning_tokens":3877}},"tokens_in":494,"tokens_out":3978,"duration_ms":28684,"temperature":1.0,"reasoning_tokens":3877,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-08T16:58:50.620921+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Fabricate the K=4, N approximately 2 Ni No processor from the numerical study and measure its full intensity transformation on random nonnegative test inputs; if the measured outputs deviate from A i beyond the simulation's negligible error, the claimed universal synthesis does not hold. A second decisive test is to put two scatterers in the input volume and check whether the output is the sum of their individual PSFs, since any shadowing-induced deviation would violate the linear model.","supporting_citations":[{"cited_title":"& Ozcan, A","cited_arxiv_id":null,"evidence_quote":"Showed that diffractive surfaces can synthesize an arbitrary linear transformation, the 2D predecessor of the 3D PSF synthesis claimed here."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Established wavelength-multiplexed universal linear transformations, which underpins the spectral engineering of 3D PSFs in this paper."},{"cited_title":"& Ozcan, A","cited_arxiv_id":null,"evidence_quote":"Describes nonlinear encoding in diffractive processing, which the paper invokes to define the regime in which its linear model fails."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the deep-learning optimization used to train the diffractive features in the numerical demonstrations."}],"review_version":1}