{"id":"48c6810f-8d44-433c-86de-6c1f8c6cc148","arxiv_id":"2511.14019","paper_version":4,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A single static mmWave radar, using human-motion multipath ghosts and a diffusion prior, reconstructs indoor wall layouts (16 cm Chamfer distance) and detects furniture (58% IoU).","lead":"RISE shows that a single stationary millimeter-wave radar can reconstruct indoor room layouts and detect furniture by exploiting the multipath reflections created by a person walking through the room. If it holds up, this would let existing home routers or access points act as privacy-preserving indoor sensors without cameras.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"SRHD's simulated partial-layout distribution is not validated against real BAME+inversion outputs; the diffusion prior that drives the largest reported gain (32→20 cm) may be fitting synthetic artifacts, and no quantitative domain-gap measure is given.","rationale":"The paper's central claim is conditional on the diffusion prior transferring from simulated partial layouts to real radar-derived observations. The ablation table shows the diffusion step is responsible for the largest single reduction in Chamfer distance (32.31 to 19.82 cm), so the transfer assumption is load-bearing. The paper provides no quantitative domain-gap metric: the simulator uses first-intersection ray casting plus random angular deletions, rotations, and scalings, whereas real inputs are generated by BAME 3D CFAR ghost detection and multipath inversion. These two processes have very different failure modes; random augmentation is a weak proxy for CFAR missed detections and spurious clusters. I therefore agree with the reader's assessment. I do not see an internal inconsistency in the geometric derivation (Eqs. 10-16 check out), and the multipath-enhancement motivation is plausible. The missing elements are external validation: no released code/data, no error bars, and no distributional comparison. My concrete test—computing a distribution distance between simulated and real initial layouts, or retraining on real initial layouts—would settle whether the concern lands. Because the reader already conditioned acceptance on releasing artifacts and re-benchmarking, my recommendation is UNCHANGED.","tokens_in":17889,"tokens_out":6793,"duration_ms":71059,"concrete_test":"Hold out a subset of real trajectories (e.g., 20 of 100) and compute a distribution-distance metric—such as sliced-Wasserstein or FID on 2D occupancy grids—between (i) simulated partial layouts from the released simulator and (ii) real BAME+inversion initial layouts, stratified by scene and after the same coordinate normalization. If per-scene distances are comparable to or smaller than within-scene variation across real trajectories, the sim2real claim is supported. If they are substantially larger, retrain SRHD on the real initial-layout distribution (or calibrate the simulator to match the empirical missing-segment/clutter statistics) and check whether the 19.82 cm ablation entry is reproduced; if not, the reported diffusion gain is an artifact of the domain gap.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim depends on the diffusion prior trained by SRHD transferring from simulation to real radar. The simulator (Section 4.4) generates partial layouts by ray-casting that keeps only the first surface intersection per ray, then applies hand-designed angular deletion, rotation and scaling augmentations (Eqs. 5-6). The real inputs to the same diffusion are not direct ray-cast maps; they are reflector estimates produced by BAME ghost detection and multipath inversion (Sections 4.2-4.3), which contain missed detections, spurious clusters, and angle errors whose statistics are unlikely to be captured by random missing/rotation/scaling. The paper calls the simulator 'high-fidelity' but gives no quantitative comparison of the synthetic partial-layout distribution to the real initial-layout distribution. This is load-bearing because Table 2 attributes the bulk of the improvement to the diffusion model (32.31->19.82 cm, and then 19.82->16.32 cm with reverse optimization). If the synthetic distribution is shifted relative to real observations, the reported 16 cm Chamfer distance and 58% IoU could reflect the prior filling in simulator-like geometry rather than actual radar evidence, leaving the sim-to-real transfer claim unsupported.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"RISE proposes a single-static mmWave radar framework and benchmark for joint 2D wall layout reconstruction and object detection. The method models multipath ghosts generated by a moving human, derives reflector positions from ghost geometry (Eq. 12), recovers off-diagonal multipath reflections using a Bi-Angular Multipath Enhancement (BAME), and then uses a simulation-trained hierarchical diffusion model (SRHD) plus inference-time overlap optimization to complete fragmented layouts and detect furniture. The paper reports 50,000 frames across 100 trajectories in 11 scenes, a Chamfer distance of 16.03 cm and F1 of 83.63 for layout, and 57.78 IoU / 69.34 Dice for object detection, versus 39.06 cm Chamfer for the baseline EMT [9].","tokens_in":18263,"tokens_out":5439,"duration_ms":58540,"significance":"The central idea—exploiting human-induced multipath ghosts for static-radar indoor mapping—is timely and potentially impactful. The geometric derivation of the first-order reflector position is not fitted and provides a principled signal-processing basis, and the modular ablation in Table 2 gives evidence that each component contributes. The proposed benchmark, if actually released, would be a valuable community asset. However, the main sim-to-real claim underlying the diffusion prior, the validity of the state-of-the-art comparison, and the statistical support for the headline numbers are not yet at the standard the paper's claims require.","major_comments":[{"comment":"The simulator-to-real transfer claim is load-bearing but unsupported. Table 2 attributes the largest gain to the diffusion model (Chamfer 32.31→19.82 cm, with reverse optimization adding 19.82→16.32 cm), yet no quantitative comparison is provided between the synthetic partial-layout distribution and the real BAME+inversion output distribution. The simulator retains only first surface intersections per ray and applies hand-designed missing, rotation, and scaling augmentations (Eqs. 5–6), whereas real inputs are reflector estimates containing missed detections, spurious clusters, and angle errors whose statistics are unlikely to match random augmentations. The paper calls the simulator \"high-fidelity\" but gives no domain-gap measure. Please add a quantitative distribution comparison (e.g., coverage statistics, noise/error statistics, or a distribution-distance metric on initial partial lay","section":"Section 4.4 and Table 2"},{"comment":"The baseline used for the headline comparison appears to be mischaracterized. The paper refers to EMT [9] as \"the multi-frame radar layout reconstruction method\" and \"the state of the art in mmWave layout reconstruction,\" but reference [9] is titled \"Environment-aware Multi-person Tracking in Indoor Environments with mmWave Radars\" and appears to be a tracking method, not a layout-reconstruction method. If EMT does not reconstruct layouts from raw radar, the comparison in Table 2 and Figure 7 (16.03 vs. 39.06 cm Chamfer) is not a valid state-of-the-art comparison. The authors must clarify what EMT outputs, and either compare against actual mmWave layout reconstruction baselines or remove the state-of-the-art claim.","section":"Section 5.1 and reference [9]"},{"comment":"The central quantitative claims are reported only as single averages over 100 trajectories, without error bars, standard deviations, confidence intervals, or significance tests. Table 1 shows very large scene-to-scene variation (IoU from 29.48 to 96.02), so the average IoU of 57.78 alone is not informative. Similarly, the layout comparison in Figure 7 and the trajectory-length analysis in Figure 9 would benefit from per-trajectory distributions and statistical tests. Please report variance and, where appropriate, paired comparisons to support the claimed improvements.","section":"Sections 5.2, 5.3, Tables 1–2"},{"comment":"The paper is framed as a benchmark contribution and states \"Our website and code are available at https://rise-cvpr.github.io,\" but no dataset, annotations, or evaluation code are included or referenced in the submission; Section 3 says \"We'll release\" in the future. For a benchmark paper, releasing the dataset and evaluation protocol is central to reproducibility and verifiability. Please include the release, a stable link with an anonymized version, or a clear statement of availability and access conditions.","section":"Abstract and Section 3"}],"minor_comments":[{"comment":"The proof of Eq. (12) is only sketched: the path-length equation (10) and the cosine-law substitution (11) are stated, but the algebra leading to Eq. (12) is omitted, and the definitions of θ_s^1 and θ_s^2 are not fully clear. Given that Eq. (12) is the geometric foundation for reflector estimation, please provide a complete derivation and a notation table.","section":"Section 8.2.1"},{"comment":"The bi-angular beamforming equation is introduced without a number, and then Step 2 says \"Apply Eq. 4 separately,\" where Eq. (4) is the standard virtual-array formula that the paper argues suppresses multipath. Number the AOA–AOD equation and refer to it explicitly so the reader can follow the BAME pipeline.","section":"Section 4.3"},{"comment":"The header \"A VG\" should read \"AVG,\" and the \"Metric Type\" entries S and M are not defined in the text. Please clarify what these types mean and how the averages are computed.","section":"Table 1"},{"comment":"The manuscript contains numerous typos, incomplete sentences, and informal phrases (e.g., \"Could be found in our Appendix,\" \"more details could be found,\" and the abstract's \"layouts reconstruction\"). A thorough language edit would improve readability.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"This is an interesting and promising system, and I do not see a fatal flaw in the geometric derivation. However, the sim-to-real transfer of the diffusion prior is the key scientific risk: the largest reported improvements come from a model trained on a simulator that has not been validated against real BAME/inversion outputs. The apparent mismatch between the claimed layout-reconstruction baseline EMT [9] and the actual title of [9] is also concerning and must be resolved. The benchmark and code release are essential for a paper whose main contribution is a new dataset. I would be willing to consider a revised version that addresses these items."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThis is the first single-static-radar layout-plus-object system I've seen, and the core signal-processing idea — BAME — is genuinely clever. Conventional virtual-array beamforming implicitly assumes AOA≈AOD, which suppresses multipath ghosts that bounce off walls at asymmetric angles. By keeping the two angles separate and running 3D CFAR on the range-AOA-AOD cube, they recover some of those ghosts. The geometric derivation leading to Eq. 12 is consistent, and the ablation in Table 2 is honest: each module adds something. The dataset (50k frames, 100 trajectories, 11 rooms) is a real resource if they actually release it.\n\nThe soft spots are real, and one is load-bearing. The diffusion prior carries most of the improvement in Table 2 (32→20→16 cm Chamfer), but the simulator trains on first-hit ray-cast partial layouts with hand-designed random missing/rotation/scaling, while the real inputs are reflector estimates from BAME + inversion, which have their own clutter, missed detections, and angle biases. There is no quantitative domain-gap measure anywhere in the paper. So the 16 cm result could partly be the prior filling in simulator-like walls rather than reading actual radar geometry. That is not a proven fatal flaw, but it is an unproven assumption doing heavy lifting.\n\nAlso: no error bars over the 100 trajectories; the abstract promises a website and code, but the full text says they \"will release\" the dataset — that's a gap between promise and artifact. The layout baseline EMT [9] is cited as SOTA, but the reference title is about multi-person tracking; if it isn't a native layout reconstructor, the 40 cm baseline may be weak and the \"60% reduction\" inflated. And \"first mmWave-based object detection\" is too broad — there are mmWave detectors in other settings. To their credit, the discussion section is candid about needing human motion and about the 2D-only output.\n\nBottom line: the paper is coherent, the geometric core is sound, and BAME deserves a serious look. But the current evidence is not independent-checkable enough to trust the headline numbers. I'd send it to peer review, not desk reject, and I'd want the revision to include released artifacts, error bars, and a direct comparison between the simulated partial-layout distribution and the real BAME+inversion output.\n\nBest,\n[your name]","headline":"RISE is the first single-static-radar layout-plus-object system I've seen, and the BAME idea is genuinely clever, but the headline numbers lean on an unvalidated sim-to-real transfer and no artifacts are out yet.","tokens_in":18669,"tokens_out":2731,"would_cite":false,"duration_ms":30514,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A single, stationary millimetre-wave radar can reconstruct room walls and detect furniture by decoding multipath 'ghost' reflections, cutting layout error to 16 cm.","keywords":["mmWave radar","multipath ghost","indoor layout reconstruction","object detection","diffusion model","sim-to-real","angle of arrival/departure","privacy-preserving sensing"],"falsifier":"Measure the distribution distance (e.g., Wasserstein or FID) between simulated partial layouts and real BAME reflector estimates on matched floor plans. If the distance is large, or if SRHD trained purely on simulation fails to improve Chamfer distance over the non-diffusion baseline on real data, the claim that the diffusion model understands radar geometry collapses.","tokens_in":17826,"feed_emoji":"📡","tokens_out":3775,"duration_ms":36978,"temperature":0.7,"pith_summary":"This paper tries to establish that a single static mmWave radar—no robot, no camera—can understand an indoor scene, provided a person is moving through it. The key move is to treat multipath reflections, normally discarded as noise, as geometric signals: a person's motion creates 'ghost' targets that reveal walls that are not in the radar's direct line of sight. The authors build a pipeline (BAME) that recovers these ghosts by separating angle-of-arrival from angle-of-departure, then a diffusion model (SRHD) trained on simulated partial layouts completes the fragmented wall map and locates furniture. If true, this would make privacy-preserving layout mapping and object detection available from a device no more invasive than a WiFi router. The reported numbers are 16 cm average Chamfer distance and 58% IoU, roughly a 60% error reduction over the prior static-radar method.","feed_headline":"Radar 'ghost' reflections map rooms to 16 cm","feed_subtitle":"By mining multipath ghosts from a person's motion, a single stationary mmWave sensor reconstructs walls and finds objects without a camera.","key_machinery":"Bi-Angular Multipath Enhancement (BAME): a beamforming step that keeps receiver (AOA) and transmitter (AOD) angles separate rather than merging them into a single virtual array, so off-diagonal ghost reflections survive CFAR detection and are reintegrated into the range–angle map. Multipath inversion then converts ghost ranges and angles into wall reflector points via a derived closed-form relation, and a Sim2Real Hierarchical Diffusion (SRHD) uses two diffusion stages—object mask first, then wall layout—trained on a ray-cast simulator that keeps only the first surface hit per ray, with random missing sectors, rotation, and radial scaling augmentations.","core_discovery":"RISE's central claim is that the multipath ghosts induced by a walking human are not noise but a usable geometric channel. By computing a full range–angle-of-arrival–angle-of-departure cube instead of collapsing the two angles into one, the system recovers first-order ghost reflections that standard beamforming suppresses, and from their geometry it estimates reflector points on walls and objects. A hierarchical diffusion model, pretrained on 35,000 synthetic floor plans with randomized missing sectors, rotations, and scalings, then fills in occluded regions and predicts furniture masks, with a reverse optimization that forces the output to be consistent with the human's free-space trajector","pith_inferences":["The claim that moving reflectors generate usable geometry implies that other moving objects—pets, ceiling fans, robot vacuums—might substitute for human motion; this is a directly testable extension the paper does not explore.","Because the diffusion prior is trained on simulated first-return ray casts, the method's real-world ceiling is likely set by how faithfully that simulation matches real multipath statistics; a future domain-adaptation step could replace the hand-designed augmentations.","The 58% IoU is measured on coarse axis-aligned bounding boxes for a handful of furniture types; extending to per-instance segmentation or 3D boxes is a natural next step, and the privacy argument would strengthen if ground truth did not rely on a depth camera."],"forward_implications":["If correct, existing mmWave access points or routers already deployed in homes and offices could map rooms and detect furniture without cameras, sidestepping occlusion and privacy concerns.","The separation of AOA from AOD offers a general recipe for other MIMO radar tasks where multipath components are currently filtered out as noise.","The hierarchical diffusion completes a room map from a single 30-second walk, and even with only 40% of the trajectory it beats the prior full-trajectory baseline in Chamfer distance.","The system outputs 2D top-down geometry, which is enough for navigation, elder-care monitoring, fall detection, and safety analysis in privacy-sensitive environments.","The new benchmark—50,000 frames, 100 trajectories, 11 scenes—provides the first standardized testbed for single-static-radar indoor scene understanding."],"fun_headline_variants":["Single radar maps rooms via multipath ghosts","Ghost reflections turn one radar into a scene mapper","Radar multipath ghosts reveal indoor layouts","One static radar reconstructs rooms to 16 cm"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The load-bearing premise is that the distribution of partial layouts produced by the simulator (first-surface ray casts with random missing sectors, rotations, and scalings) matches the distribution of reflector estimates produced by BAME on real radar data; the paper does not quantify this domain gap, and the diffusion model's large reported improvement (32.3 → 19.8 cm in ablation) could be fitting synthetic artifacts if the gap is large.","fun_headline_variants_meta":{"raw":{"variants":["Single radar maps rooms via multipath ghosts","Ghost reflections turn one radar into a scene mapper","Radar multipath ghosts reveal indoor layouts","One static radar reconstructs rooms to 16 cm"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000188,"raw_usage":{"total_tokens":1204,"prompt_tokens":815,"completion_tokens":389,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":559,"completion_tokens_details":{"reasoning_tokens":331}},"tokens_in":559,"tokens_out":389,"duration_ms":4403,"temperature":1.0,"reasoning_tokens":331,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-03T21:39:50.458274+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Measure the distribution distance (e.g., Wasserstein or FID) between simulated partial layouts and real BAME reflector estimates on matched floor plans. If the distance is large, or if SRHD trained purely on simulation fails to improve Chamfer distance over the non-diffusion baseline on real data, the claim that the diffusion model understands radar geometry collapses.","supporting_citations":[],"review_version":1}