{"id":"8c578a83-a987-4d68-a16f-a0fb56520f9d","arxiv_id":"2506.08163","paper_version":2,"verdict":"REJECT","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"SpINRv2 combines a differentiable closed-form frequency-domain forward model with an implicit neural representation for 3D FMCW radar reconstruction, beating baselines on synthetic data but tested only up to 5 GHz start frequency.","lead":"SpINRv2 is a neural method that reconstructs 3D scenes from FMCW radar data by matching a closed-form model of the radar's frequency-domain response. The paper reports better geometry and image metrics than several baselines, but only on synthetic data that share the model's own simplifying assumptions.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'high start frequency' claim is tested only up to f0=5 GHz (f0/B≈1.4), whereas the motivating AWR1843BOOST radar operates near 77 GHz (f0/B≈21.5); the phase-aliasing regime central to the paper is extrapolated by an order of magnitude without evidence.","rationale":"The paper's strongest claim is explicitly about high start frequencies, where phase aliasing and sub-bin ambiguity become prominent. The experiments, however, never leave the f0/B ≈ 1.4 regime, while the motivating hardware is at f0/B ≈ 21.5. The number of phase wraps per range bin, f0/B, is the natural dimensionless quantity controlling the ambiguity the method claims to resolve, and the tested range falls far short of the claimed operating point. This is not merely an unreported detail: it is the difference between a mildly ambiguous regime and a severely aliased one, so the central claim is an extrapolation. The reader's weakest_assumption identified exactly this gap. The additional synthetic-data circularity mentioned in the reader's rationale is real but secondary; the f0 discrepancy alone is enough to reject or at least heavily condition the headline claim. Therefore the reader's REJECT verdict should stand, and my concrete test would settle whether the high-frequency claim survives at the actual operating frequency.","tokens_in":13277,"tokens_out":10532,"duration_ms":138419,"concrete_test":"Re-run the Section 6.5.3 start-frequency ablation at f0 = 77 GHz, keeping B = 3.585 GHz, the cylindrical aperture, the INR architecture, and the regularization weights fixed, and compare SpINRv2 against the strongest baseline (TF-SS or coherent backprojection) using IoU, Chamfer distance, and convergence curves. If SpINRv2's IoU drops by more than 2x relative to its f0 = 5 GHz result, or if a baseline comes within 10% relative error of SpINRv2's accuracy, the paper's high-frequency claim is contradicted.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The abstract and Section 6.5.3 rest the central contribution on behavior 'under high start frequencies—where phase aliasing and sub-bin ambiguity become prominent.' The ablation fixes B=3.585 GHz and varies f0 only from 1 to 5 GHz, so the dimensionless aliasing severity f0/B runs from 0.28 to 1.39. Because the phase phi = 2*pi*f0*tau advances by 2*pi*(f0/B) across one range bin of width c/(2B), the tested configuration reaches at most ~1.4 phase wraps per bin. The TI AWR1843BOOST listed in Section 5 operates at 76–81 GHz, giving f0/B ≈ 21.5 and lambda/4 ≈ 1 mm versus a 41.8 mm range bin—about 21 phase wraps per bin. That is a qualitatively different, much more ambiguous regime. In addition, the synthetic measurements are generated from the same noiseless, perfectly phase-coherent point-scatterer chirp model that the forward model implements, so the experiments do not establish physical fidelity at any frequency. The claim that SpINRv2 'establishes a new benchmark' for high-frequency radar imaging is therefore unsupported by the reported evidence.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes SpINRv2, an implicit neural representation framework for volumetric FMCW radar reconstruction. It derives a closed-form frequency-domain forward model from the DFT of a dechirped beat signal, supervises the complex spectrum directly, and adds smoothness and sparsity regularizers. The method is evaluated on synthetic cylindrical-aperture data against time-domain, range-quantized, and backprojection baselines, with reported gains in IoU, Chamfer distance, Hausdorff distance, PSNR, SSIM, LPIPS, and runtime. The central claim is that the approach works under high start frequencies, where phase aliasing and sub-bin ambiguity become prominent.","tokens_in":13539,"tokens_out":8190,"duration_ms":97030,"significance":"The closed-form differentiable spectral synthesis is a useful idea that, if correct and properly validated, could improve both the stability and efficiency of neural radar reconstruction relative to time-domain simulation plus FFT. Explicitly modeling spectral leakage and using sparsity/smoothness priors are sensible design choices. However, the current evidence does not establish the paper's stronger claims: all measurements are simulated under the same ideal point-scatterer, noiseless, occlusion-free assumptions embodied by the forward model, and the 'high start frequency' regime is tested at carrier frequencies an order of magnitude below the motivating hardware. The absence of error bars, code, or real data further limits the claim of a new benchmark.","major_comments":[{"comment":"The DFT synthesis formula is dimensionally inconsistent as written. For the sampled beat signal b(nTs)=exp(j2π f0 τ) exp(j2π S τ n Ts), the per-sample phase increment is 2π S τ / Fs, where Fs is the sampling rate. The display equation in Section 4.2 instead uses 2π S τ(x) in the numerator and subtracts β_k = 2πk/N from it. This mixes a frequency in Hz with an angle in radians and is not a valid DFT of the sampled chirp. Since the synthetic measurements in Section 5 are generated by a time-domain simulation followed by DFT, the reported experiments can be internally self-consistent even if this equation is not physically correct. Please correct the formula (including the 1/Fs or Ts factor) and confirm that the implementation matches the corrected expression.","section":"Section 4.2"},{"comment":"The central 'high start frequency' claim is not tested in the regime that motivates it. The ablation fixes B=3.585 GHz and varies f0 only from 1 GHz to 5 GHz, so the dimensionless parameter f0/B reaches at most 1.39 and λ/4 is no smaller than about 15 mm. The TI AWR1843BOOST hardware cited in Section 5 operates at 76–81 GHz, for which f0/B ≈ 21.5 and λ/4 ≈ 1 mm against a range bin of about 41.8 mm. Phase aliasing is therefore an order of magnitude more severe than anything evaluated, and the paper gives no scaling argument for why behavior at 1–5 GHz carries over to 77 GHz. The sentence 'SpINRv2 works under high start frequencies—where phase aliasing and sub-bin ambiguity become prominent' is consequently unsupported by the reported experiments.","section":"Section 6.5.3 and Abstract"},{"comment":"The evaluation is entirely synthetic and is generated from the same ideal point-scatterer assumptions that the forward model encodes. Section 5 describes a noiseless simulation of 691,200 multistatic measurements followed by a multistatic-to-monostatic transformation, but the transformation is not specified and is not validated against true multistatic or real-world data. There is no test with noise, clutter, occlusion, multipath, non-point scatterers, or the antenna patterns of actual mmWave hardware. Tables 1 and 2 report point estimates without error bars or significance tests, so the statement that SpINRv2 'significantly outperforms' the baselines is not statistically grounded. A real-radar validation (even limited) or a clearly stated restriction of the claims to the simulated ideal scenario is needed before the 'new benchmark' claim can be accepted.","section":"Section 5 and Section 6"},{"comment":"The optimization procedure is incompletely specified. Section 4.3 defines a fixed total loss with weights β and γ, but Section 6.5.5 introduces a staged supervision schedule in which the abs-term is used alone for the first 10% of transmitter locations and the complex terms are added later. This schedule is not part of the loss definition in Section 4, and the manuscript does not report values for β, γ, ε, the stage duration, or the switch criterion. Because these choices directly affect all reported reconstructions, they are load-bearing for reproducibility.","section":"Sections 4.3 and 6.5.5"}],"minor_comments":[{"comment":"The word 'commertial' should be 'commercial'.","section":"Section 5"},{"comment":"The notation 'm(t) ∗ ˆm(t)' should be multiplication, not convolution; the beat-signal derivation is otherwise described in words as mixing.","section":"Section 3.2"},{"comment":"The text says f0 is varied from 1 GHz to 5 GHz, but Figure 10 caption lists only 2, 3, and 4 GHz; please reconcile the experimental range and the figure.","section":"Section 6.5.3"},{"comment":"The sentence 'In the next section we show how regularization helps with higher frequencies' is inaccurate: the next section is the high-frequency sheet benchmark, while regularization is the topic of Section 6.5.1.","section":"Section 6.5.3"},{"comment":"The PSNR entry for TF-SS is given as 13.5134 with excessive and inconsistent decimal places; use uniform formatting across all entries.","section":"Table 1"},{"comment":"The baseline list says 'All methods are trained' but coherent backprojection is not trained; please rephrase to distinguish learned baselines from the classical method.","section":"Section 6"}],"recommendation":"major_revision","confidential_remarks":"The paper is a preprint without code or data release. The main risks are that the forward-model equation in Section 4.2 appears dimensionally incorrect as written, and that the experimental claims rest on a synthetic dataset generated under the same ideal assumptions as the forward model, with the 'high-frequency' regime extrapolated from f0/B ≈ 1.4 to f0/B ≈ 21.5 without evidence. These issues are addressable in a revision by correcting the formula, adding a real mmWave experiment or at least a noise/clutter sensitivity study, and narrowing the claims. As it stands, the evidence is not yet sufficient for a journal acceptance."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: the core idea is legitimate and the forward-model math is clean, but the paper's central claim is the one thing it never actually tests. The frequency-domain synthesis is a textbook DFT of a delayed chirp, but putting it inside an INR with spectral supervision on only the relevant bins is a sensible and modestly new combination, and the efficiency argument (skip full time-domain simulation and FFT) holds in the synthetic setup. I credit the authors for the closed-form leakage treatment and for ablating the regularization terms.\n\nThe soft spots are load-bearing. First, the evaluation is entirely generated with the same noiseless isotropic point-scatterer model the forward model encodes, so it is a round-trip simulation test, not evidence about physical fidelity. No noise, no error bars, no hardware. Second, the \"high start frequencies\" claim is tested only up to f0 = 5 GHz with B = 3.585 GHz, so f0/B is at most about 1.4, while the motivating AWR1843BOOST operates near 77 GHz with f0/B around 21. The phase-aliasing regime they say is central is extrapolated by an order of magnitude, and their own sheet benchmark (Fig. 13) shows degradation at fine detail, which does not make the extrapolation reassuring. Third, the multistatic-to-monostatic transformation is asserted but never specified, and the data normalization is vague. Fourth, they cite their prior SpINR in the abstract but omit it from the references, so the delta over prior work cannot be checked. That last one is sloppy but fixable.\n\nThe paper is not a waste of time. If the authors run the method at true mmWave parameters, add noise and clutter, specify the transformation, and release code, it could become a credible benchmark. As is, I would not accept the strong claims, but I would still send it to peer review: the method is coherent, the math is internally consistent, and the evaluation can be fixed in a major revision. My guess is reviewers will want realistic or real data before believing the \"new benchmark\" line.","headline":"A sensible frequency-domain forward model for FMCW radar INR, but the headline claim about high-frequency performance is exactly the part that goes untested.","tokens_in":14099,"tokens_out":2021,"would_cite":false,"duration_ms":24733,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"SpINRv2 claims that FMCW radar volume reconstruction should be supervised directly on the complex frequency spectrum through a differentiable closed-form forward model, with sparsity and smoothness priors to resolve sub-bin phase ambiguity.","keywords":["FMCW radar","implicit neural representation","frequency-domain forward model","volumetric reconstruction","spectral leakage","phase aliasing","sparsity regularization","mmWave radar"],"falsifier":"Run the identical reconstruction pipeline with the chirp start frequency set to 77 GHz, or an equivalent simulation where $\\lambda/4$ is about a millimetre and well below the bin resolution, and measure IoU and chamfer distance against the same baselines; if the reported gap closes or shell artifacts persist despite regularization, the high-frequency claim does not cover the intended mmWave regime.","tokens_in":13011,"feed_emoji":"📡","tokens_out":5390,"duration_ms":55640,"temperature":0.7,"pith_summary":"SpINRv2 claims that high-fidelity 3D volume reconstruction from FMCW radar is best done by supervising an implicit neural field directly in the frequency domain, using a closed-form differentiable model of the complex beat spectrum instead of simulating time-domain waveforms. The paper argues that time-domain supervision is ill-conditioned because tiny range shifts scramble the oscillating beat signal, and that naive frequency-domain shortcuts such as range quantization throw away spectral leakage and phase information. SpINRv2 synthesizes only the DFT bins that correspond to valid scene ranges, adds sparsity and smoothness priors to resolve sub-bin phase ambiguities, and reports better IoU, chamfer distance, PSNR, SSIM, and LPIPS than backprojection and learning-based baselines on synthetic cylindrical-aperture scenes. The claim matters because if it holds, neural radar imaging can move from coarse voxel grids to continuous scattering fields with less computation and better high-frequency behavior.","feed_headline":"Frequency-domain neural field sharpens FMCW radar 3D reconstructions","feed_subtitle":"A differentiable spectral forward model beats backprojection and time-domain baselines in simulated volumetric tests.","key_machinery":"The load-bearing object is a closed-form spectral synthesis formula: for a scatterer at point $x$ with round-trip delay $\\tau(x)$, the complex response at DFT bin $k$ is an integral over the scene of $\\sigma(x)$ times a propagation factor and a Dirichlet-kernel leakage term, with phase $\\phi(x) = 2\\pi f_0\\tau(x)$ set by the chirp start frequency. This formula lets the model compute only the $K$ frequency bins that correspond to valid scene delays, avoiding full time-domain simulation and FFT, and it keeps the forward map fully differentiable. Paired with an implicit neural representation for the scattering field $\\sigma(x)$, plus $\\ell^1$ sparsity and local smoothness regularizers, the mechanism turns FMCW spectral measurements into a continuous volumetric reconstruction.","core_discovery":"On its own terms, the paper's central discovery is that the FMCW beat signal's DFT has a closed form in terms of round-trip delay and chirp parameters, so the radar forward model can be evaluated directly at relevant frequency bins and supervised against complex measurements. The resulting system, SpINRv2, represents the scene as an implicit neural field predicting scatterer intensity and trains it with a magnitude and complex-component spectral loss plus smoothness and sparsity regularization. The paper reports that this combination outperforms time-domain simulation, range quantization, and coherent backprojection across all six metrics it evaluates, and degrades gracefully as bandwidth shrinks from 4 GHz to 40 MHz. It also reports that the high start-frequency regime is where the gap is largest, provided the regularizers are present.","pith_inferences":["A direct test the paper leaves open is whether the same gains appear at true automotive mmWave frequencies of 76-81 GHz: the experiments stop at 5 GHz, where $\\lambda/4$ is roughly 15 mm, while at 77 GHz it is roughly 1 mm, a far harsher aliasing regime.","The frequency-domain forward model could be lifted to other coherent modalities that produce complex spectra, such as stepped-frequency radar or FMCW LiDAR, wherever the measurement is a Fourier transform over delay.","The staged loss schedule, magnitude first and then real and imaginary components, hints at a coarse-to-fine spectral supervision strategy that may benefit other inverse problems in sonar or ultrasound.","Because the experimental setup is a synthetic aperture, extending SpINRv2 with motion estimation could open the way to dynamic scenes, a direction the paper lists as future work but does not explore."],"forward_implications":["Frequency-domain supervision replaces time-domain MSE, which should make neural radar reconstruction trainable where time-domain losses plateau or explode.","Because only the bins inside the scene's delay range are computed, forward passes become cheaper than full 256-sample simulation plus FFT; the paper reports a runtime advantage that grows with scene size.","The closed-form leakage model captures sub-bin energy spill, so scatterers no longer need to align with bin centers for accurate reconstruction.","Smoothness and sparsity regularizers suppress the shell artifacts that appear when $\\lambda/4 < c/(2B)$, extending the usable start-frequency range.","The continuous INR representation degrades gracefully under low bandwidth down to 40 MHz, unlike voxel-based backprojection which blurs sharply."],"supporting_citations":[{"why":"Defines coherent backprojection, the classical baseline that SpINRv2 is compared against and whose voxel-grid limitations motivate the continuous representation.","marker":"Duersch (2013)"},{"why":"Supplies the range-Doppler processing baseline for FMCW systems that SpINRv2 contrasts with its native frequency-domain model.","marker":"Wagner et al. (2013)"},{"why":"Provides the FMCW automotive radar signal model, including the beat-frequency mapping and dechirping operation that the forward model is built on.","marker":"Patole et al. (2017)"},{"why":"Justifies dropping the residual video phase term to reach the simplified beat-signal equation used throughout the derivation.","marker":"Wang et al. (2014)"},{"why":"Establishes the implicit neural representation paradigm that SpINRv2 adapts from view synthesis to radar volume reconstruction.","marker":"Mildenhall et al. (2021)"},{"why":"Presents the earlier pulse-echo neural volumetric reconstruction framework that SpINRv2 extends by moving from time-domain to frequency-domain supervision for FMCW.","marker":"Reed et al. (2023)"},{"why":"Supports the use of sinusoidal activations and the gradient-stability argument that motivates the frequency-domain forward model's shallower backpropagation path.","marker":"Sitzmann et al. (2020)"},{"why":"Provides the AirSAS cylindrical-aperture data-generation setup that the paper's experimental geometry draws on.","marker":"Cowen et al. (2021)"}],"fun_headline_variants":["Passband FMCW 3D improved by spectral-supervised neural field","Differentiable radar forward model wins high-frequency 3D tests","Neural INR with spectral loss outperforms time-domain baselines","Closed-form DFT supervises neural field for passband FMCW 3D","High-start-frequency FMCW 3D improved by neural spectral model"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The paper tests start frequencies only from 1 to 5 GHz, far below the 76-81 GHz band of the mmWave sensor it mimics, and gives no reason the results would carry over to that more severe phase-wrapping regime.","fun_headline_variants_meta":{"raw":{"variants":["Passband FMCW 3D improved by spectral-supervised neural field","Differentiable radar forward model wins high-frequency 3D tests","Neural INR with spectral loss outperforms time-domain baselines","Closed-form DFT supervises neural field for passband FMCW 3D","High-start-frequency FMCW 3D improved by neural spectral model"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000935,"raw_usage":{"total_tokens":3961,"prompt_tokens":870,"completion_tokens":3091,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":486,"completion_tokens_details":{"reasoning_tokens":2996}},"tokens_in":486,"tokens_out":3091,"duration_ms":27526,"temperature":1.0,"reasoning_tokens":2996,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T05:19:11.650545+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the identical reconstruction pipeline with the chirp start frequency set to 77 GHz, or an equivalent simulation where $\\lambda/4$ is about a millimetre and well below the bin resolution, and measure IoU and chamfer distance against the same baselines; if the reported gap closes or shell artifacts persist despite regularization, the high-frequency claim does not cover the intended mmWave regime.","supporting_citations":[{"cited_title":"Backprojection for synthetic aperture radar","cited_arxiv_id":null,"evidence_quote":"Defines coherent backprojection, the classical baseline that SpINRv2 is compared against and whose voxel-grid limitations motivate the continuous representation."},{"cited_title":"Wide-band range-doppler processing for fmcw systems","cited_arxiv_id":null,"evidence_quote":"Supplies the range-Doppler processing baseline for FMCW systems that SpINRv2 contrasts with its native frequency-domain model."},{"cited_title":"Application of linear-frequency-modulated continuous-wave (lfmcw) radars for tracking of vital signs","cited_arxiv_id":null,"evidence_quote":"Justifies dropping the residual video phase term to reach the simplified beat-signal equation used throughout the derivation."},{"cited_title":"Neural volumetric reconstruction for coherent synthetic aperture sonar","cited_arxiv_id":null,"evidence_quote":"Presents the earlier pulse-echo neural volumetric reconstruction framework that SpINRv2 extends by moving from time-domain to frequency-domain supervision for FMCW."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supports the use of sinusoidal activations and the gradient-stability argument that motivates the frequency-domain forward model's shallower backpropagation path."},{"cited_title":"Airsas: Controlled dataset generation for physics-informed machine learning","cited_arxiv_id":null,"evidence_quote":"Provides the AirSAS cylindrical-aperture data-generation setup that the paper's experimental geometry draws on."}],"review_version":1}