{"id":"db50388c-f74b-4463-893d-d596c62e731e","arxiv_id":"2412.10300","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"A SPAD array captures the full first-order transient light transport matrix of a relay wall, and beamforming algorithms extract the second-order transient light transport matrix of the hidden scene, enabling relighting, direct/indirect separation, and dual photography.","lead":"By measuring every combination of laser and detector position on a relay wall with a 256-pixel single-photon camera, this paper computes a second-order light transport matrix that describes how light moves within a hidden scene. The result lets researchers relight hidden objects, separate direct from multi-bounce light, and take dual photographs around corners.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Eq. (3) is asserted to extract the true hidden-scene TLTM-2, but the output is never compared to any ground-truth transport matrix; the demonstrated effects are qualitative and could be PSF/side-lobe artifacts of the virtual focusing system.","rationale":"The reader's weakest assumption was the static-scene requirement during sequential acquisition. That is a real practical limitation, but it is standard for NLOS proof-of-concept demonstrations and does not test the mathematical core of the paper. The load-bearing point is instead the unvalidated equivalence between the linear transform f in Eq. (3) and a true second-order transport matrix. The paper provides qualitative demonstrations with internal controls, and the phasor-field framework is established, which gives the claim credibility; however, no quantitative ground truth exists for any extracted H^(2) entry. A simulation with a known hidden scene would directly settle whether f(H) recovers true transport or merely a PSF-blurred projection of it. This concern does not require rejecting the paper; it is an addressable validation gap, so the reader's conditional verdict remains appropriate. I therefore recommend no change to the verdict.","tokens_in":24002,"tokens_out":14471,"duration_ms":125869,"concrete_test":"Render a small synthetic hidden scene (e.g., two Lambertian patches plus one mirror) with a transient path tracer: (a) compute the true hidden-scene transport H_true^(2)(x_p^(2), x_c^(2), t) directly between surface points; (b) simulate the relay-wall measurement H(xp, xc, t) for the same scene including third- and higher-bounce paths and the SPAD temporal response; (c) apply Eq. (3) and the supplement's fast implementation to H to obtain H_extracted^(2). Compare H_extracted^(2) to H_true^(2) after convolving the latter with the known P(omega) band-limit and the predicted focusing/imaging PSF. If the residual after this comparison is at the noise level, the central claim is supported; if systematic spurious arrivals or wrong relative amplitudes appear, Eq. (3) has a hidden assumption that needs revision.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Eq. (3) defines TLTM-2 as f(H) by beamforming the illumination over xp and phasor-field imaging over xc. For the central claim to hold, f(H) must be a faithful representation of the spatiotemporal transport between points in the hidden scene. This is not established. Both constituent operations are lossy: the projector P(omega) band-limits the temporal response, the finite aperture blurs the focus by the PSF of Eq. (5), and the imaging step is an ill-posed diffraction inversion with known missing-cone ambiguities and side lobes, as the paper acknowledges in Sec. 6.1. The validations in Fig. 4B, Fig. S10, and Fig. S11 are qualitative and are consistent with, but do not uniquely require, genuine hidden-scene transport; the same observations could plausibly arise from the known PSF and residual illumination of the virtual focusing system. No code, data, or synthetic test is provided to compare an extracted H^(2) entry against a known value. Since the headline claim is precisely that TLTM-1 contains sufficient information to compute TLTM-2, this unvalidated equivalence is the weakest load-bearing point.","agreement_with_reader":"disagree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a non-line-of-sight (NLOS) imaging system that measures the full first-order transient light transport matrix (TLTM-1) on a relay surface using a gated 16x16 SPAD array and a scanned laser. The central methodological claim is that a linear operator applied to TLTM-1, given by Eq. (3), extracts a second-order TLTM, H^(2)(x_p^(2), x_c^(2), t), for surfaces in the hidden scene by computationally focusing virtual illumination and imaging the response. The authors demonstrate three applications: scene relighting with novel illumination, separation of direct and indirect light transport, and dual photography. They also report an FFT-accelerated algorithm with complexity O(k N^3 log N) for a single column of TLTM-2, and an experiment showing that increasing the number of SPAD pixels can substitute for longer exposure times.","tokens_in":24216,"tokens_out":6787,"duration_ms":62779,"significance":"If the extracted TLTM-2 is a faithful representation of hidden-scene light transport, this work would be a valuable step toward higher-order NLOS light transport analysis and multi-corner imaging. The FFT-based complexity reduction relative to prior virtual light transport matrix work is a concrete algorithmic contribution, and the use of a full 16x16 SPAD array to capture the complete TLTM-1 is a useful hardware demonstration. The supplementary shadow and water-versus-milk experiments are falsifiable checks that go beyond simple 3D reconstruction. However, the central equivalence between Eq. (3) and true hidden-scene transport is not validated against any ground-truth matrix, and the experimental demonstrations are largely qualitative. No code, data, or synthetic test is provided to support the headline claim. The paper is therefore promising but not yet established.","major_comments":[{"comment":"The central claim that TLTM-1 contains sufficient information to compute the full TLTM-2 is not validated. Equation (3) is a composition of a band-limiting projector P(omega), a finite-aperture beamformer, and a diffraction-based imaging step; each of these operations is lossy, as the paper itself acknowledges through the missing-cone discussion and the residual illumination observed in Sec. 6.1. The demonstrations in Fig. 4B and Supp. S6 are qualitative and are consistent with, but do not uniquely require, genuine hidden-scene transport; the same observations could plausibly arise from the PSF and side lobes of the virtual focusing system. I recommend adding a synthetic or laboratory validation in which a hidden scene with known transport properties is simulated or directly measured, and the extracted H^(2) is compared quantitatively against a ground-truth TLTM-2 entry or column.","section":"Eq. (3), Sec. 6.1, Fig. 4B, Supp. S6"},{"comment":"There is an algebraic error in the resolution derivation. From Eq. (5) with lambda_s = lambda_c/2 and D = N * lambda_s, one obtains Delta_x = 1.22 * lambda_c * z / (N * lambda_c/2) = 2.44 z / N, not 1.22 z / (2 N) as printed in Eq. (6). The Discussion's example (z = 2 m, N = 100, approximately 5 cm resolution) corresponds to Delta_x = 2.44 z / N ≈ 4.9 cm, not to the printed Eq. (6). Please correct Eq. (6) and ensure that all subsequent quantitative statements use the corrected formula.","section":"Sec. 3.1, Eq. (6), Sec. 7"},{"comment":"The extraction of TLTM-2 assumes that all rows and columns of the measured TLTM-1 are mutually consistent, meaning that the hidden scene and the imaging system remain static over the full sequential acquisition, which takes minutes per dataset as described in Sec. 5. This assumption should be stated explicitly as a limitation and, ideally, validated with a stability check such as repeated measurement of a reference row or column. If the scene or system drifts during acquisition, the linear combination in Eq. (3) does not represent any single light transport state and the resulting TLTM-2 would be physically meaningless.","section":"Sec. 5, Eq. (3)"}],"minor_comments":[{"comment":"The sentence 'Helmholtz reciprocity is used to treat each SPAD pixel xp as a virtual illumination source, and each laser position xc is a point on the virtual aperture' appears to reverse the physical roles of detector and illumination; please clarify whether xp and xc are used consistently with their definitions in Eq. (3).","section":"Sec. 4.3"},{"comment":"The PSNR growth values are inconsistent between the text (0.186 vs. 0.184) and the figure caption (10.806 vs. 10.236), and no error bars or statistical tests are reported; please unify the numbers and quantify the uncertainty.","section":"Supp. S5, Fig. S7"},{"comment":"The focusing delay is written as Delta t(x_p^(2)) in Eq. (3), but in Supplement S2.2.2 it depends on both the relay-surface position x_p and the focus point; please make this dependence explicit in the notation.","section":"Eq. (3), Supp. S2.2.2"},{"comment":"The sentence 'the test set is used to generate the projector function ... while the test set is used to generate the reconstruction' uses 'test set' twice; one of these should almost certainly be 'training set'.","section":"Sec. 6.2.2"}],"recommendation":"major_revision","confidential_remarks":"This paper is heavily reliant on self-authored prior work and does not ship code or data. For a claim as strong as 'TLTM-1 contains sufficient information to compute TLTM-2', I would expect at least one synthetic validation with a known ground-truth transport matrix. The incorrect Eq. (6) is a red flag that should be caught in a careful revision. The paper fits the journal's scope, but the central validation gap and the resolution-analysis error need to be addressed before publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"What you should know: this is not incremental tinkering. They build an NLOS system that captures the full first-order TLTM on the relay surface with a 2D SPAD array, then use phasor-field beamforming to computationally focus virtual illumination inside the hidden scene and read out a second-order TLTM. The three demonstrations—relighting, direct/indirect separation, dual photography—are new for hidden scenes, and the water/milk and shadow experiments show genuine transport effects, not just reconstruction noise.\n\nThe math is sound and the contributions are real. Equation (3) is a linear integral transform that focuses and images, and the FFT acceleration in Eq. (4) is a genuine step forward in complexity. The beamforming resolution analysis is clean, and the side-lobe suppression with train/test split shows care. The hardware description is detailed, and the SNR equivalence between pixel count and exposure time is useful, even if the PSNR fits lack error bars.\n\nThe stress-test is right about the load-bearing gap: no ground-truth validation that the extracted H^(2) actually equals the true transport matrix. The demonstrations are qualitative. That said, the shadow appearing and disappearing with the vase, and the water/milk scattering differences, are hard to explain as pure PSF artifacts—they are genuinely physical. What is missing is a synthetic test: simulate a known hidden scene, compute H, extract H^(2), and compare against the known matrix. That would settle whether Eq. (3) is a faithful inverse or just a useful beamformer. Also, 'full' TLTM-1 is qualified by non-functional pixels and non-uniform spacing, and the static-scene assumption during long sequential scans is unaddressed. These are addressable.\n\nWho is this for? Anyone working on NLOS imaging, computational light transport, or time-of-flight sensors. It deserves a serious referee pass—the central claim is plausible and the new capability is important. I would recommend conditional acceptance, with a synthetic validation and uncertainty quantification as required revisions.","headline":"A genuinely new NLOS capability—extracting second-order transient transport from a full relay-surface measurement—backed by plausible math but missing the quantitative ground-truth check that would make the central claim bulletproof.","tokens_in":24825,"tokens_out":2112,"would_cite":true,"duration_ms":583429,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Full measurement of the relay-surface transient light transport matrix lets you compute a second-order matrix for the hidden scene, turning the relay wall into a remote active imaging system.","keywords":["non-line-of-sight imaging","transient light transport matrix","SPAD array","beamforming","phasor field","time-of-flight","dual photography","scene relighting"],"falsifier":"Place a moving object, such as a rotating fan, in the hidden scene while the relay surface is scanned, and check whether the computed second-order matrix contains time-inconsistent shadows, ghosting, or artifacts that vary with focus position; any such inconsistency would confirm that the static-scene assumption is the load-bearing premise. A more direct check is to put a physical detector at a computed hidden-surface location and compare its measured transient response with the corresponding TLTM-2 column.","tokens_in":23774,"feed_emoji":"🔦","tokens_out":9395,"duration_ms":81796,"temperature":0.7,"pith_summary":"This paper claims that a single full measurement of the transient light transport matrix (TLTM) of a relay surface—recording, for every illumination position, every detection position, and every time bin, how light returns from a hidden scene—contains enough information to compute a second-order TLTM describing light transport among the hidden surfaces themselves. If true, the relay surface becomes a virtual active imaging system: the hidden scene can be relit under synthesized illumination, direct bounces can be separated from indirect ones, and dual-photography views can be obtained without new hardware. The authors build a 16x16 gated SPAD array system that samples the full first-order TLTM, then use beamforming ideas from phased arrays to focus virtual illumination at chosen hidden-scene points and image the response as a function of time. They demonstrate the three applications experimentally and argue that the same iteration can be pushed to higher-order TLTMs as SPAD arrays grow.","feed_headline":"Full relay-surface light matrix reveals hidden-scene light transport","feed_subtitle":"Beamforming the measured matrix yields hidden-scene light transport, enabling relighting, direct/indirect separation, and dual photography.","key_machinery":"The load-bearing object is the second-order TLTM and the linear focus-and-image operator that produces it. Equation (3) defines this operator: a projector function $P(\\omega)$ carrying a focusing delay $e^{-j\\omega\\Delta t(\\mathbf{x}_p^{(2)})}$ is multiplied into the measured $H(\\mathbf{x}_p,\\mathbf{x}_c,\\omega)$, summed over $\\mathbf{x}_p$ to form $P_F(\\mathbf{x}_c,\\omega)$, then propagated with the spherical phase $e^{-j\\omega|\\mathbf{x}_c^{(2)}-\\mathbf{x}_c|/c}$ and integrated over $\\mathbf{x}_c$ and $\\omega$. When $\\mathbf{x}_c$ lies on a regular planar grid, the inner propagation becomes a convolution and is evaluated with 2D FFTs, giving $O(kN^3\\log N)$ per TLTM-2 column. Resolution is set by the Rayleigh criterion $\\Delta x = 1.22\\lambda_c z/D$, which under Nyquist sampling becomes $\\Delta x = 1.22z/(2N)$, so the number of SPAD grid points $N$ ultimately limits how sharply the virtual illumination can be focused. Grid interpolation and a ridge-regression side-lobe-suppression step are used to sharpen the focus with the available array.","core_discovery":"The central claim is that the first-order TLTM $H(\\mathbf{x}_p,\\mathbf{x}_c,t)$ measured on the relay surface can be transformed by a linear operator $f$ into the second-order TLTM $H^{(2)}(\\mathbf{x}_p^{(2)},\\mathbf{x}_c^{(2)},t)=f(H(\\mathbf{x}_p,\\mathbf{x}_c,t))$, where $\\mathbf{x}_p^{(2)}$ and $\\mathbf{x}_c^{(2)}$ are illumination and detection points in the hidden scene. Equation (3) realizes $f$ as a delay-and-sum beamformer in the temporal-frequency domain: it applies a focusing time shift $\\Delta t(\\mathbf{x}_p^{(2)})$ to the projector, integrates over relay-surface illumination and detection positions with a spherical imaging phase, and inverse-transforms in time. The extracted TLTM-2 preserves complex light transport—shadows, subsurface scattering, caustics, specular reflections—and supports scene relighting, direct/indirect separation by time gating, and dual photography through Helmholtz reciprocity. The experimental system captures the full TLTM-1 with a 16x16 gated SPAD array and a scanned laser, and the paper reports that increasing the number of SPAD pixels is equivalent to increasing exposure time for signal-to-noise ratio.","pith_inferences":["The linearity of $f$ means the same beamforming operator could be reapplied to TLTM-2 to get TLTM-3, but each iteration compounds the depth-dependent resolution loss $\\Delta x = 1.22z/(2N)$, so two-corner reconstruction would need substantially larger arrays or compressed sensing.","The static-scene requirement is a hard practical limit: because the galvanometer scans the relay surface sequentially, any motion in the hidden scene during the multi-minute acquisition corrupts the matrix combination, so dynamic scenes would need motion compensation or simultaneous multi-point illumination.","The demonstrated equivalence between pixel count and exposure time suggests acquisition time can be traded against array size, but the data-rate bottleneck will then shift to processing, pointing toward compression of the TLTM as the next constraint.","A direct validation of the core claim would be to place a physical detector or light source at a computed hidden-scene point and compare its measured transient response with the corresponding TLTM-2 column, which the paper does not yet report."],"forward_implications":["The full first-order TLTM of a relay surface is sufficient to compute the complete second-order TLTM of the hidden scene, so no additional hardware beyond a multi-pixel time-of-flight array is needed for relighting, separation, or dual photography.","Hidden scenes can be relit from arbitrary directions by treating individual SPAD pixels as virtual illumination sources, and arbitrary patterns can be projected onto hidden surfaces by optimizing the projector function over space and frequency to create incoherent point sources.","Time-gating TLTM-2 separates the direct third-bounce component from indirect higher-bounce components, revealing multi-bounce paths, shadows, and subsurface scattering in the hidden scene.","Primal and dual images of the hidden scene are equivalent under Helmholtz reciprocity, and multi-pixel acquisition acquires them faster because more photons are collected in parallel, matching the quality of a longer single-pixel exposure.","As SPAD arrays grow, TLTM-2 resolution improves linearly, and the same iteration procedure becomes feasible for extracting TLTM-3, enabling reconstruction around two corners."],"supporting_citations":[{"why":"Supplies the 16x16 gated SPAD array hardware that makes full TLTM-1 sampling feasible.","marker":"[1]"},{"why":"Introduces the ultrafast time-of-flight backprojection approach that underlies the reconstruction operator.","marker":"[6]"},{"why":"Provides the fast phasor-field diffraction reconstruction whose FFT convolution accelerates TLTM-2 computation.","marker":"[9]"},{"why":"Establishes the phasor-field virtual wave model that turns the relay surface into a virtual line-of-sight imaging system.","marker":"[16]"},{"why":"Gives the Huygens-like phasor-field wave propagation model used for focusing and imaging.","marker":"[22]"},{"why":"Introduces dual photography and Helmholtz reciprocity, the basis for the dual-photography application.","marker":"[36]"},{"why":"Supplies the delay-and-sum beamforming principles used to focus virtual illumination.","marker":"[41]"},{"why":"Is the prior virtual light transport matrix method whose computational complexity and missing temporal dimension the paper improves on.","marker":"[46]"}],"fun_headline_variants":["Iterating light matrix turns corners into cameras","Second-order light matrix unlocks hidden-scene relighting","Beamforming the light matrix sees around corners","Full TLTM iteration yields dual photography and relighting","SPAD array accelerates NLOS imaging via full light matrix"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The hidden scene must remain completely static during the full sequential scan of the relay surface, which takes minutes, because the entire first-order matrix is assembled from rows measured at different times; any motion in the hidden scene invalidates the matrix combination used to extract TLTM-2.","fun_headline_variants_meta":{"raw":{"variants":["Iterating light matrix turns corners into cameras","Second-order light matrix unlocks hidden-scene relighting","Beamforming the light matrix sees around corners","Full TLTM iteration yields dual photography and relighting","SPAD array accelerates NLOS imaging via full light matrix"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000441,"raw_usage":{"total_tokens":2324,"prompt_tokens":1119,"completion_tokens":1205,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":735,"completion_tokens_details":{"reasoning_tokens":1131}},"tokens_in":735,"tokens_out":1205,"duration_ms":12125,"temperature":1.0,"reasoning_tokens":1131,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T15:58:51.946275+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Place a moving object, such as a rotating fan, in the hidden scene while the relay surface is scanned, and check whether the computed second-order matrix contains time-inconsistent shadows, ghosting, or artifacts that vary with focus position; any such inconsistency would confirm that the static-scene assumption is the load-bearing premise. A more direct check is to put a physical detector at a computed hidden-surface location and compare its measured transient response with the corresponding TLTM-2 column.","supporting_citations":[{"cited_title":"Fast-Gated 16 16 SPAD Array with 16 on-chip 6 ps Time-to-Digital Converters for Non-Line-of-Sight Imaging,","cited_arxiv_id":null,"evidence_quote":"Supplies the 16x16 gated SPAD array hardware that makes full TLTM-1 sampling feasible."},{"cited_title":"Recovering three-dimensional shape around a corner using ultrafast time-of-flight imaging,","cited_arxiv_id":null,"evidence_quote":"Introduces the ultrafast time-of-flight backprojection approach that underlies the reconstruction operator."},{"cited_title":"Phasor field diffraction based reconstruction for fast non-line-of-sight imaging systems,","cited_arxiv_id":null,"evidence_quote":"Provides the fast phasor-field diffraction reconstruction whose FFT convolution accelerates TLTM-2 computation."},{"cited_title":"Non-line-of-sight imaging using phasor-field virtual wave optics,","cited_arxiv_id":null,"evidence_quote":"Establishes the phasor-field virtual wave model that turns the relay surface into a virtual line-of-sight imaging system."},{"cited_title":"Phasor field waves: A Huygens-like light transport model for non-line-of-sight imaging applications,","cited_arxiv_id":null,"evidence_quote":"Gives the Huygens-like phasor-field wave propagation model used for focusing and imaging."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the delay-and-sum beamforming principles used to focus virtual illumination."},{"cited_title":"Virtual Light Transport Matrices for Non-Line-of-Sight Imaging,","cited_arxiv_id":null,"evidence_quote":"Is the prior virtual light transport matrix method whose computational complexity and missing temporal dimension the paper improves on."}],"review_version":1}