{"id":"4b742898-e72f-4425-ac0d-1628105dafb5","arxiv_id":"2607.25569","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":7,"one_line_summary":"CORF-GS reconstructs wireless radiance fields online by sharing Gaussian geometry between optical images and RF spectra, cutting reconstruction time ~6.4x versus RF-3DGS while improving RF spectrum PSNR.","lead":"CORF-GS is an online method that reconstructs a room's radio signal map in minutes from sequential camera images and RF measurements, using shared 3D Gaussian blobs for both light and radio. It is a step toward live wireless channel maps for 6G digital twins and sensing.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'real-time' and 'SOTA' claims are not yet substantiated: only total 2m18s is reported, not per-keyframe latency, and the fast GSpaRC baseline is omitted from all comparisons.","rationale":"The reader's weakest assumption—known keyframe poses—is a legitimate limitation, but it is explicitly acknowledged in Sec. IV-A and is common in Gaussian-splatting and SLAM papers; it does not directly threaten the reconstruction-quality comparison. The more load-bearing gap for the advertised contribution is that the 'real-time' property and the 'state-of-the-art' label are not empirically pinned down. A 2:18 total is a speed improvement over offline baselines, not itself a real-time guarantee; without per-keyframe timing versus the capture interval, the inferred 'online update' capability is unsupported. The omission of GSpaRC is especially concerning because the paper itself identifies GSpaRC as a fast WRF-reconstruction method, so a claim of SOTA speed or quality cannot be evaluated without it. Both deficiencies are fixable with one focused experimental pass: instrument per-keyframe latency and add GSpaRC to the benchmark. Therefore the reader's CONDITIONAL verdict is appropriate and should remain unchanged rather than being hardened to REJECT or loosened to ACCEPT.","tokens_in":11681,"tokens_out":6242,"duration_ms":62376,"concrete_test":"Run CORF-GS on the NIST-lobby sequence while logging timestamps per keyframe for the LoG sampling (Eqs. 12–14), depth estimation/back-projection (Eqs. 15–17), and the 75 coupled-optimization iterations; compare the average per-keyframe latency with the dataset's keyframe arrival interval. In the same harness, run GSpaRC [19] on the identical train/test split and report PSNR_rf/SSIM_rf/LPIPS_rf and total/per-keyframe time. If CORF-GS cannot keep pace with the capture interval, or GSpaRC matches or exceeds its quality-time trade-off, the 'real-time' and 'state-of-the-art' statements in the abstract need to be weakened.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Sec. IV-A defines the update operator U(·) in Eq. (8) as a per-keyframe recursion, and Table I reports the headline reconstruction time of 2 min 18 s. But the paper never reports the per-keyframe wall-clock time of U(·), nor the keyframe capture interval of the NIST-lobby trajectory. With 147 keyframes, 2:18 implies an average of roughly 0.94 s per keyframe; at any typical camera or measurement rate, that is not self-evidently 'real-time', and the paper does not demonstrate that the sequential update completes before the next keyframe arrives. The optional offline fine-tuning in Table III further muddies what the 2:18 figure actually covers. Independently, the related-work section cites GSpaRC [19] as a fast WRF-reconstruction framework, yet GSpaRC is absent from Tables I/II and the qualitative comparison. The central 'state-of-the-art' and speed claims therefore rest on an incomplete baseline set: if GSpaRC is faster or comparably accurate, CORF-GS is an incremental improvement rather than the claimed real-time SOTA. This is an evidence gap rather than an internal inconsistency, but it is exactly the gap that must be closed before the central claim can be accepted.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes CORF-GS, a 3D Gaussian Splatting framework for incremental reconstruction of wireless radiance fields (WRFs) from sequential optical-RF keyframes. The method shares Gaussian geometry between optical and RF modalities while maintaining modality-specific appearance. For each incoming keyframe, CORF-GS performs optical-guided Gaussian expansion based on Laplacian-of-Gaussian saliency, initializes new primitives via estimated depth, and then jointly optimizes the shared geometry and modality-specific appearance under a combined optical-RF loss. Experiments on the synthetic NIST-lobby optical-RF dataset report PSNR_rf = 18.52 dB, SSIM_rf = 0.852, LPIPS_rf = 0.239 and a total reconstruction time of 2 min 18 s, a 6.4x speedup over RF-3DGS. Ablations show that both the optical-guided expansion and coupled optimization contribute to the reported gains.","tokens_in":12098,"tokens_out":3377,"duration_ms":37362,"significance":"If the claims are substantiated, CORF-GS would be a meaningful step toward online WRF reconstruction, where spatial channel maps are updated as new measurements arrive rather than being rebuilt from batch data. The design of a coupled optical-RF Gaussian representation is sensible and the ablation study is internally consistent, giving evidence that optical-guided expansion and joint optimization are each beneficial. The use of 19 held-out keyframes for quantitative evaluation is a genuine test-set measurement, and no obvious circularity in the evaluation protocol was found. However, the significance is currently limited by the absence of a fast sequential baseline, the lack of per-keyframe latency evidence for the 'real-time' claim, and the reliance on a single synthetic scene with no variance reporting. These issues are fixable and do not contradict the internal logic of the method.","major_comments":[{"comment":"The central 'real-time' claim is not quantified at the per-keyframe level. With 147 keyframes and a total time of 2 min 18 s, the average time per keyframe is about 0.94 s, but the paper never reports the wall-clock latency of the update operator U(·), the keyframe capture interval, or the fraction of the 138 s spent on optical-guided sampling, depth initialization, coupled optimization, and I/O. Without showing that U(·) completes before the next keyframe arrives, the real-time claim is not substantiated. Additionally, Table III includes an optional offline fine-tuning stage; the manuscript should state explicitly whether the 2:18 figure in Table I excludes fine-tuning and whether any fine-tuning is needed for the reported main results.","section":"Sec. IV-A, Eq. (8); Table I; Table III"},{"comment":"GSpaRC [19] is cited in the related work as a 'fast WRF reconstruction framework', yet it is absent from all quantitative comparisons and qualitative results. The headline claims of state-of-the-art synthesis quality and 6.4x speedup are therefore established only against offline baselines (NeRF2, WRF-GS+, RF-3DGS). If GSpaRC is a sequential or fast baseline, it must be included in Tables I, II, and Fig. 3 (or a concrete justification for its exclusion must be given). Otherwise, the paper's central comparison set is incomplete and the 'state-of-the-art' and speed claims are overstated.","section":"Sec. II-A, Ref. [19]; Tables I-II; Fig. 3"},{"comment":"The evaluation is conducted on a single synthetic scene (NIST-lobby replica). All quantitative results are reported as averages over ten repetitions, but no standard deviations or confidence intervals are provided. The margin that supports the 'state-of-the-art' claim is modest: PSNR_rf of 18.52 dB vs. 18.09 dB for WRF-GS+ at 30k iterations. Without error bars or a second scene, it is difficult to assess whether this margin is statistically meaningful. Adding per-repetition variance, or at least reporting standard deviations, and ideally evaluating on a second scene, would materially strengthen the claim. This is a load-bearing point for the generality of the headline result.","section":"Sec. V-B, Table I; Sec. V-A"},{"comment":"The real-time pipeline assumes keyframe poses are directly available from device metadata, and the parameter initialization step uses monocular depth estimation (Depth-Anything-V2) followed by multi-view matching. The manuscript does not include the computational cost of these steps in the reported 2:18 reconstruction time, nor does it provide the per-keyframe breakdown. If the intended deployment is a robot or handheld platform, pose estimation is not free, and the depth/alignment step may dominate the update operator. The real-time claim should either be scoped explicitly to the reconstruction-update operator under known poses, or the pose and depth estimation costs should be included in the latency budget. This is not a fundamental flaw but needs to be stated precisely.","section":"Sec. IV-A; Sec. IV-D2"}],"minor_comments":[{"comment":"The text refers to 'Appendix A1' and 'Appendix A2' for sensitivity analysis and dataset details, but no appendix appears in the arXiv version. Either include the referenced appendices or remove the citations.","section":"Sec. V-A"},{"comment":"In Fig. 4, the x-axis mixes different iteration counts for Ours (50-125) and WRF-GS+/RF-3DGS (in thousands). The figure is readable but would benefit from a logarithmic or annotated axis to avoid implying comparable iteration budgets.","section":"Sec. V-B, Fig. 4"},{"comment":"The RF-3DGS row labels use parentheses like '(15,10) ×10^3'; this is concise but easy to misread. Add a single clarifying line in the caption, e.g., '(N_opt, N_rf)'.","section":"Table II"},{"comment":"The ablation 'w/o optical-guided exp.' replaces the proposed sampling with uniform sampling at probability 0.1. The paper reports that this yields about 1.0M Gaussians vs. 0.2M for the full model. It would be useful to also ablate with a tuned uniform sampling probability to separate the effect of sampling strategy from the effect of primitive count.","section":"Sec. V-C, Table IV"},{"comment":"The phrase 'state-of-the-art' in the abstract and conclusion should be qualified as 'state-of-the-art on the NIST-lobby synthetic dataset' until broader evaluation supports a general claim.","section":"Global"},{"comment":"Minor grammatical issues: 'existing WRF reconstruction methods' vs. 'existing methods' is used inconsistently, and Eq. (20) is missing a period after the displayed equation.","section":"Sec. VI"}],"recommendation":"major_revision","confidential_remarks":"The paper is technically sound in its internal logic and the ablations support the proposed method's components. The main issues are evidentiary rather than conceptual: the real-time claim lacks per-keyframe timing, the closest fast baseline is missing, and the evaluation is single-scene with no variance information. I recommend major revision rather than rejection because these gaps can be closed within the manuscript's scope. I also note that no code is released; given the empirical nature of the claims, a code/data release or a detailed reproducibility description would substantially increase confidence."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a real engineering contribution, not a repackaging. The authors take three known pieces—3DGS for RF, on-the-fly Gaussian expansion for SLAM-style mapping, and optical-RF shared geometry—and combine them in a way that is new and appears to work in their experiments. The coupled optimization is the key move: RF spectra get to adjust the shared Gaussian geometry rather than being forced onto a frozen optical scaffold. Their ablations make that case cleanly: uniform sampling hurts a lot, and freezing geometry drops PSNR_rf from 18.52 to 17.11. The held-out keyframes are a genuine test set, and the comparison to WRF-GS+ and RF-3DGS is fairly configured. I believe the central mechanism.\n\nThe soft spots are evidence gaps rather than errors. First, the evaluation is one synthetic NIST lobby scene. Ten repetitions are averaged but no error bars appear, so we do not know if a 0.3 dB edge over WRF-GS+ is meaningful. Second, GSpaRC [19] is cited as a fast WRF-reconstruction framework but is absent from every comparison table. That is the closest baseline to what they claim to beat, and leaving it out weakens the word “state-of-the-art.” Third, the real-time label: the paper reports 2:18 total for 147 keyframes, which averages to about 0.94 s per keyframe, but never gives per-keyframe latency or the keyframe capture interval. The optional offline fine-tuning in Table III also muddies which number actually covers the real-time stage. For an online pipeline, the relevant question is whether U(·) completes before the next keyframe arrives; that is not answered. Fourth, known-pose is assumed. The paper is upfront about it, and it is a reasonable simplification, but it means the system is really “real-time reconstruction given poses,” not a full SLAM pipeline.\n\nNone of this is load-bearing enough to make me distrust the method. The writing is direct, the design decisions are motivated, and the limitations are at least partially acknowledged. I would send this to peer review. The authors should add a second scene, report error bars, include GSpaRC, and either provide per-keyframe timings or soften the real-time claim. With those changes, this could be a solid IEEE-level contribution. A serious referee will help them calibrate the claims; the work deserves that scrutiny.","headline":"Credible online WRF reconstruction system with a sensible design and honest ablations, but the real-time and SOTA claims rest on thin evidence (no per-keyframe timings, one scene, missing GSpaRC baseline).","tokens_in":12544,"tokens_out":2049,"would_cite":true,"duration_ms":23209,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"CORF-GS reconstructs a wireless radiance field in real time from sequential optical-RF keyframes, achieving state-of-the-art RF spectrum quality (PSNR_rf 18.52 dB) in 2 minutes 18 seconds, a 6.4x speedup over prior offline methods, by coupl","keywords":["wireless radiance field","3D Gaussian splatting","real-time reconstruction","RF spectrum synthesis","optical-RF fusion","channel modeling","incremental mapping","coupled optimization"],"falsifier":"A concrete test: on the simulated lobby dataset, delete the optical images for the last 30 keyframes and feed only RF spectra; if the coupled optimization still maintains RF PSNR within 1 dB, the optical guide is not load-bearing. Alternatively, run the same coupled optimization on a scene where keyframe poses are estimated by a SLAM system instead of given, and measure whether the 2-minute budget and RF PSNR hold; a scene with a strong non-line-of-sight diffracted path, where no visible surface anchors the peak, would also stress the shared-surface assumption.","tokens_in":11634,"feed_emoji":"📡","tokens_out":4991,"duration_ms":49375,"temperature":0.7,"pith_summary":"The paper proposes CORF-GS, a framework that reconstructs a wireless radiance field (WRF) — a dense representation of how radio power arrives from different directions at any point — in real time as a camera and antenna array move through a scene. Instead of collecting all observations first and optimizing offline, CORF-GS processes keyframes one by one: optical images provide structural hints for where to add new Gaussian primitives, and a shared Gaussian geometry with separate optical and RF appearance is refined jointly under both image and radio-spectrum supervision. This design is meant to overcome the mismatch between how light and radio waves interact with surfaces, which causes purely optical geometry to miss RF-informative regions. The claimed payoff: a WRF that can be queried for synthesized RF spectra at new positions, built in 2 minutes 18 seconds with RF PSNR of 18.52 dB, outperforming offline baselines while being 6.4x faster.","feed_headline":"Radio field maps in 2 minutes, not hours","feed_subtitle":"Sharing one Gaussian geometry for light and radio and refining it jointly beats offline methods while streaming keyframes.","key_machinery":"The central object is the coupled Gaussian primitive G_i = {μ_i, Σ_i, o_opt_i, c_opt_i, o_rf_i, c_rf_i}: shared spatial center and covariance, but separately learned opacity and color for optical and RF. New primitives are seeded by comparing Laplacian-of-Gaussian responses of the current optical frame and the rendered optical view, which highlights under-represented high-frequency regions. The update operator U(·) runs a coarse-to-fine optimization of L = L_opt + λ_rf L_rf, where L_rf is an L1 plus SSIM loss on rendered versus observed RF spectra. This coupled loss is what lets RF supervision reshape the shared Gaussian geometry.","core_discovery":"The central claim is that optical images and RF spatial spectra can share a single 3D Gaussian geometry, where each primitive carries both an optical color/opacity and an RF color/opacity, and that this coupled representation can be built incrementally from streaming keyframes. The key move is to use the optical image to decide where to place new Gaussians via Laplacian-of-Gaussian sampling, then to optimize the shared geometry under both optical and RF losses rather than freezing the optical geometry. The paper reports that this joint refinement improves RF spectrum synthesis by about 1.4 dB PSNR over the frozen-geometry alternative and by over 1 dB over prior offline methods, while cutting","pith_inferences":["If poses become available from a SLAM front-end instead of metadata, the same update operator could run in true real-time SLAM-style settings; the paper does not time pose estimation, so that extension remains untested.","The shared-surface assumption suggests the method will perform best where reflection and specular scattering dominate; scenes with strong diffraction or non-line-of-sight contributions that do not attach to visible surfaces may require a separate mechanism, which the current formulation does not include.","One testable extension is to replace the optical-guided sampling with RF-guided or hybrid sampling for environments without usable optical images (e.g., dark or textureless indoor spaces) and measure how much RF quality degrades; the ablation with uniform sampling suggests the optical guide is worth roughly 3 dB.","The fidelity of the monocular depth prior and its alignment to triangulated depth is likely to dominate geometric accuracy for newly spawned Gaussians; failures of such priors in repetitive or low-texture scenes may become the practical bottleneck."],"forward_implications":["Wireless channel maps can be updated online as a device moves, enabling adaptive beamforming, coverage prediction, and digital-twin updates without batch re-optimization.","RF-informative regions that are optically unremarkable can still be captured, because RF losses can move or add Gaussians beyond the optical sampling.","The 6.4x speedup means a 147-keyframe trajectory is reconstructed in about the time it takes to capture it, making near-real-time channel acquisition practical on a single GPU.","Because the representation remains a standard 3DGS renderer, any query pose can synthesize both an optical view and an RF spatial spectrum from the same compact primitive set.","The optional offline fine-tuning stage gives a quality-versus-latency knob: 75 extra iterations bring RF quality above all baselines while still costing about a third of the prior two-stage method's time."],"fun_headline_variants":["Real-time radio maps via shared optical-RF Gaussians","6x faster wireless channel modeling with CORF-GS","Coupling light and radio for streaming radiance fields","Joint optical-RF splatting for real-time radio mapping","CORF-GS: 6x faster real-time radio field reconstruction"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The pipeline assumes keyframe poses come free from device metadata rather than being estimated, and that radio radiance arises from the same surface-representing Gaussians that carry optical radiance; if either fails, the reported real-time budget and the accuracy against non-surface propagation both collapse.","fun_headline_variants_meta":{"raw":{"variants":["Real-time radio maps via shared optical-RF Gaussians","6x faster wireless channel modeling with CORF-GS","Coupling light and radio for streaming radiance fields","Joint optical-RF splatting for real-time radio mapping","CORF-GS: 6x faster real-time radio field reconstruction"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000179,"raw_usage":{"total_tokens":1157,"prompt_tokens":788,"completion_tokens":369,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":532,"completion_tokens_details":{"reasoning_tokens":285}},"tokens_in":532,"tokens_out":369,"duration_ms":4810,"temperature":1.0,"reasoning_tokens":285,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T02:02:17.928388+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A concrete test: on the simulated lobby dataset, delete the optical images for the last 30 keyframes and feed only RF spectra; if the coupled optimization still maintains RF PSNR within 1 dB, the optical guide is not load-bearing. Alternatively, run the same coupled optimization on a scene where keyframe poses are estimated by a SLAM system instead of given, and measure whether the 2-minute budget and RF PSNR hold; a scene with a strong non-line-of-sight diffracted path, where no visible surface anchors the peak, would also stress the shared-surface assumption.","supporting_citations":[],"review_version":1}