{"id":"51b94c26-fad0-4f12-b81d-b4227b4c9eed","arxiv_id":"2608.02304","paper_version":3,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":9,"one_line_summary":"TRACE replaces greedy next-best-view selection in Gaussian-splatting active reconstruction with ergodic trajectory optimization, improving reconstruction PSNR by 1.5 dB on Replica.","lead":"This paper presents TRACE, a planner that turns robotic 3D scene reconstruction into a continuous coverage problem: the robot's path is optimized so it spends more time looking at regions the current map says are uncertain. The authors report a 1.5 dB PSNR gain over a next-best-view baseline on eight indoor scenes and direct execution on a quadruped robot.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The online information map (Eq. 2) is never ablated; if replacing it with a uniform target leaves the +1.5 dB PSNR gain intact, the central causal mechanism is misattributed to ergodic matching of the map.","rationale":"With the reader, I read the paper as a coherent reformulation whose main empirical claim is plausible. My strongest concern is not the kernel-ergodic math or the real-robot section but the unexamined bridge between map state and trajectory objective. The method is only as good as Eq. (2); the paper's ablations isolate trajectory-level mechanisms but never vary the target distribution. The uniform-map ablation is the minimal experiment that would show whether φ_t carries the causal load. If it does not, the claim that dwell time is proportional to information density is not established even though the PSNR numbers might still hold. This supports the reader's CONDITIONAL verdict and adds a specific required check; I do not see an internal inconsistency that would warrant rejection. The missing map validation and hyperparameter disclosure are fixable with a moderate revision, so the verdict remains conditional rather than accept or reject.","tokens_in":14551,"tokens_out":24818,"duration_ms":307807,"concrete_test":"Run TRACE on the eight Replica scenes with Eq. (2) replaced by a uniform distribution over the same masked free space, both in the depletion-aware kernel-ergodic term (Eq. 6) and in the gaze reward (Eq. 5), keeping all other components, budgets, and evaluation settings identical. If the PSNR advantage over ActiveGS persists within roughly 0.5 dB, the information map is not load-bearing and the central mechanism should be reattributed; if the gain collapses, the map is necessary. As a complementary check, log per-horizon φ_t values together with the actual per-voxel reduction in rendered PSNR/depth error after observing each voxel, and report the rank correlation on two scenes.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central causal story is that the map-derived target distribution φ_t (Eq. 2) guides the ergodic trajectory to the right places, producing the +1.5 dB gain. This link is the least secure part of the argument. The ablations in Table 2 remove the footprint depletion/gaze mechanism (Ours-Kernel ES) and densify the baseline's sampling, but no experiment removes or perturbs the information map itself. This matters because ActiveGS, the baseline, also uses the same per-Gaussian confidence and frontier signals; the comparison therefore does not isolate whether φ_t's specific weighting (α_u, α_f, α_b, β), the low-confidence gate 1_low, or the diffusion/masking projection is responsible for the improvement. If the map mis-ranks regions, the ergodic cost spends motion on the wrong places; the reported gain could instead come from footprint depletion, continuous-pose sampling, or the gaze reward. Eq. (2) is the sole bridge from map state to trajectory objective, and the paper provides no validation that high-φ voxels are the voxels whose observation reduces reconstruction error. The parameters of Eq. (2) are also unreported, so the map's behavior cannot be inspected or reproduced.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"TRACE replaces greedy next-best-view selection in active 2D Gaussian-splatting reconstruction with ergodic trajectory optimization. The method builds an online target information distribution phi_t from per-Gaussian rendering confidence, frontier/unexplored/unbuilt voxel states, and a low-confidence gate (Eq. 2); it then optimizes a kernel-ergodic horizon cost with footprint-overlap depletion and a gaze reward (Eqs. 3-7). The paper reports +1.5 dB average PSNR over ActiveGS on eight Replica scenes, ablations showing the footprint mechanism is worth 2.1 dB and that densifying the baseline's sampling does not close the gap, and direct trajectory execution on a Unitree Go2 and Franka FR3 without an intermediate path planner.","tokens_in":14949,"tokens_out":4071,"duration_ms":42168,"significance":"If the claims hold, the paper makes a useful contribution: it moves active reconstruction from discrete viewpoint selection to trajectory-level ergodic optimization, and the ablation in Table 2 gives concrete evidence that footprint depletion matters (2.1 dB) and that the gain is not explained by sampling density. The real-robot demonstrations, while somewhat under-quantified, indicate a practical payoff of trajectory-level planning. The main technical risk is the unvalidated information map: phi_t is the sole bridge from map state to trajectory objective, yet the paper never tests whether its ranking of informative regions actually reduces reconstruction error. The reader's stress-test concern lands on this point, and it is the main reason the paper needs revision.","major_comments":[{"comment":"The information map phi_t is never ablated. Table 2 removes the footprint-depletion mechanism and densifies the baseline's sampling, but no experiment replaces phi_t with a uniform target, drops the low-confidence gate 1_low(v), or varies the weights (alpha_u, alpha_f, alpha_b, beta). Because ActiveGS also uses per-Gaussian confidence and frontier signals, the +1.5 dB PSNR gain cannot be attributed to phi_t's specific ranking of informative regions; it could instead be carried by footprint depletion, continuous-pose optimization, or the gaze reward. The paper should validate that high-phi voxels are the voxels whose observation reduces reconstruction error, and it must report the parameter values (alpha_u, alpha_f, alpha_b, beta, eta, sigma_fp, sigma, lambda_g, lambda_s, lambda_r, kappa, D, depth samples, box-filter width, well-built threshold). This is load-bearing for the central causal story.","section":"§3.2, Eq. (2); §4.4"},{"comment":"The table reports PSNR, SSIM, and LPIPS averaged over five independent runs but gives no standard deviations or confidence intervals. Several per-scene PSNR gains are under 1 dB (e.g., Of4: 34.92 vs 34.08; R0: 31.55 vs 29.93). Without variance, the reader cannot judge whether the mean +1.5 dB advantage and the per-scene gains are statistically significant. Please report error bars or paired per-run comparisons.","section":"Table 1, §4.1"},{"comment":"NARUTO and FisherRF numbers are cited from Jin et al. (2025) rather than re-run under the same mapper, budget, and evaluation protocol. Since ActiveGS is the central baseline this is not fatal, but the text should clearly distinguish re-run results from cited results and state that the NARUTO/FisherRF comparison inherits the original protocol. This matters for the claim that ActiveGS is the strongest NBV baseline.","section":"Table 1, §4.1"},{"comment":"The main text asserts a 100% success rate and that every planned waypoint is dynamically consistent, but the quantitative collision-free evidence in the supplementary is a simulation result on the FR3 (Table 3), not a hardware measurement, and the real Go2 deployment is described qualitatively for a single 42 m^2 scene. Please clarify what was measured on the physical platforms (tracking error, contact detection, reconstruction metrics, number of executed horizons/trials) and separate simulated collision counts from real-robot observations. Also, the Go2 planner uses single-integrator motion, so 'dynamically consistent by construction' is stronger than what the model actually enforces.","section":"§4.5 and supplementary Table 3"}],"minor_comments":[{"comment":"The labels 'H' and 'L' in the left panel are not defined in the caption; please define them as high- and low-information regions.","section":"Figure 2 caption"},{"comment":"The 'well-built' threshold used in 1_low(v) is never specified, so the low-confidence gate is not reproducible. Please give the criterion or the threshold value.","section":"§3.2, Eq. (2)"},{"comment":"Several references have broken line breaks (e.g., the ActiveNeRF entry) and the table captions use inconsistent highlighting descriptions; please clean these up at revision.","section":"References and formatting"}],"recommendation":"major_revision","confidential_remarks":"The central mechanism is plausible and the footprint-depletion ablation is a real strength, but the missing information-map ablation and the unreported parameters are the key weaknesses. I would encourage the editor to require the authors to add an ablation that perturbs or replaces phi_t, and to report error bars on Table 1, before publication. The self-citation to Zheng et al. (2025) is used appropriately to distinguish the footprint mechanism and is not a concern."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nQuick take on TRACE (arXiv:2608.02304). The genuinely new thing is replacing greedy next-best-view selection in active Gaussian-splatting reconstruction with a trajectory-level ergodic objective, and the paper makes that concrete: an online information map from voxel states and per-Gaussian confidence, a kernel-ergodic cost with footprint depletion, and a gaze reward. The ablations are honest: removing the footprint mechanism costs 2.1 dB, densifying the baseline does not close the gap, and the random-view baseline is basically noise. That is real evidence that trajectory-level planning, not sampling density, drives the gain. The real-robot section is a nice practical payoff, even if thin.\n\nThe soft spots are real but not fatal. Table 1 reports five-run averages without variance, so the +1.5 dB could wobble. The NARUTO/FisherRF numbers are cited from the ActiveGS paper rather than re-run under the same protocol. The information map of Eq. (2) carries a lot of weight but none of its weights or the low-confidence gate are reported, and the map itself is never ablated. The stress-test concern is fair: if a uniform target distribution produced the same PSNR gain, the central causal story would be misattributed. That experiment is missing. The real-robot claim of 100% success for TRACE and \"minor contact\" for the baseline is undercooked—no trial counts, no definition of success, no quantitative comparison.\n\nNone of this sinks the paper. The core idea is sound, the math follows Sun et al. and is correctly adapted (free-space projection and footprint depletion are sensible solutions to the surface-vs-traversable-space mismatch), and the joint-space FR3 extension is thoughtful. The PSNR gain is measured externally, so there is no circularity problem. I'd send this to a serious referee and expect major revision: add error bars, disclose hyperparameters, and run at least one experiment that perturbs or removes the information map. I'd cite it if I worked in active reconstruction.\n\nWorth a full review.","headline":"Solid integration of ergodic search into active GS reconstruction with honest ablations, but missing variance and a map-ablation leaves the central mechanism only partially supported.","tokens_in":15415,"tokens_out":1848,"would_cite":true,"duration_ms":46604,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Ergodic trajectory optimization lifts active 3D reconstruction quality by 1.5 dB PSNR, and the resulting paths run directly on physical robots.","keywords":["active 3D reconstruction","ergodic search","Gaussian splatting","trajectory optimization","next-best-view planning","information map","mobile manipulation","coverage planning"],"falsifier":"On the eight Replica scenes, run TRACE with the information map of Eq. (2) replaced by a uniform distribution over the same free-space mask (all weights equal), keeping every other parameter fixed; if the PSNR gain over ActiveGS persists, the claim that map-guided dwell time drives the improvement is falsified.","tokens_in":14370,"feed_emoji":"🤖","tokens_out":7211,"duration_ms":79393,"temperature":0.7,"pith_summary":"TRACE proposes that the bottleneck in active 3D reconstruction is the planner's decision unit: greedy next-best-view selection wastes the motion between views and ignores the robot's dynamics. The paper replaces it with ergodic trajectory optimization, in which the time the sensor spends in each region is made proportional to an online information map built from per-Gaussian confidence and voxel states. The key technical moves are deriving that map from the live Gaussian-splatting map, diffusing and masking its mass from surfaces into free space so real trajectories can match it, and adding a footprint-depletion term so the target distribution decays as regions are re-observed within a horizon. The reported payoff is a 1.5 dB average PSNR gain over the strongest NBV baseline on eight Replica scenes under matched conditions, plus trajectories that a quadruped and a manipulator can execute without a separate path planner. If true, this shifts active reconstruction toward a single, differentiable objective that couples information gathering with control feasibility.","feed_headline":"Ergodic paths beat next-best-view planning by 1.5 dB","feed_subtitle":"Continuous trajectory optimization produces higher-fidelity maps and robot-ready motions.","key_machinery":"The load-bearing machinery is a differentiable, horizon-level trajectory objective $J_t(u_t)$ (Eq. 7) whose core is the depletion-aware kernel-ergodic metric $E_{\\mathrm{dep}}^{\\mathrm{kernel}}$ (Eqs. 3 and 6): the first term rewards placing waypoints where the time-discounted information map $\\phi_t^k$ is high, and the second term spreads waypoints out with a pairwise Gaussian repulsion. That metric acts on an information map $\\phi_t$ (Eq. 2) built online at each horizon from voxel classes—unexplored, frontier, unbuilt, and low-confidence-Gaussian voxels—then box-filtered and masked to collision-free space so the volume the sensor body can occupy carries the information mass that actually lives on surfaces. Footprint-overlap depletion (Eqs. 3–4) discounts already-covered surface samples within the horizon, and a gaze reward (Eq. 5) aims the camera at high-information surfaces while the position stays governed by the ergodic metric. All terms are differentiable in the control sequence, which is optimized by Adam with a warm start from the previous horizon.","core_discovery":"The paper's central claim is that active 3D reconstruction with Gaussian-splatting maps is better posed as an ergodic coverage problem than as a sequence of next-best-view selections: instead of committing to the most informative discrete viewpoint and routing between such views, the agent should optimize a continuous trajectory whose time-averaged spatial statistics match a target information distribution. TRACE derives that distribution online from the current 2DGS map and voxel map, using a weighted sum of unexplored, frontier, unbuilt, and low-confidence regions (Eq. 2), then projects it into traversable free space by diffusion and masking. The trajectory is obtained by gradient-descent optimization of a kernel-ergodic objective with a footprint-overlap depletion term that suppresses re-coverage within the horizon and a gaze reward that orients the camera at under-covered surfaces. Under a matched mapper, budget, and evaluation protocol on eight Replica scenes, the author reports a mean +1.5 dB PSNR improvement over the ActiveGS NBV baseline, with better SSIM and LPIPS on every scene, and trajectories that execute directly on a Unitree Go2 and a Franka FR3 without an intermediate path planner.","pith_inferences":["A natural next test is whether the hand-crafted weights in Eq. (2) are near-optimal: replacing them with a learned map trained to predict per-voxel reduction in reconstruction error could either improve the PSNR gain or reveal a ceiling set by the ergodic objective itself.","The direct-execution result suggests that the information loss between discrete viewpoints, not just viewpoint selection, is a major source of inefficiency in NBV pipelines; the same ergodic objective could be applied to TSDF or NeRF mappers by substituting their uncertainty proxies.","The depletion mechanism is local to the current horizon; a persistent coverage memory across horizons might extend the approach to long-horizon missions where the map updates alone are too slow to suppress re-coverage.","One could also test whether ergodic trajectories generalize across scenes and robots without retuning: the paper reports weights but not their values, so the sensitivity of the 1.5 dB gain to $\\alpha_u, \\alpha_f, \\alpha_b, \\beta$ is an open empirical question."],"forward_implications":["Under the same mapper, sensor, and time budget, trajectory-level ergodic planning beats greedy NBV on eight Replica scenes by 1.5 dB average PSNR, with SSIM and LPIPS improving on every scene.","Because the planned trajectory is the optimization variable, the output is dynamically feasible by construction and can be sent directly to a robot controller, eliminating the separate path-planner step of NBV pipelines.","Simply adding more observations along the NBV path—random or uniform interpolation—does not close the gap to TRACE, so the gain comes from where the trajectory dwells, not from sheer frame count.","The footprint-depletion mechanism is necessary: without it, the kernel-ergodic version loses 2.1 dB and falls below the NBV baseline on six of eight scenes, whereas with it the planner deliberately re-visits under-confident surfaces.","The same objective transfers to joint-space planning on a manipulator, where forward kinematics map the information map into the joint trajectory, keeping every planned view reachable."],"supporting_citations":[{"why":"Supplies the ActiveGS NBV baseline, the per-primitive confidence score, and the matched mapper/budget/evaluation protocol that TRACE uses for comparison.","marker":"Jin et al. 2025"},{"why":"Provides the kernel-ergodic metric that TRACE adapts with footprint depletion; removing it yields the Ours-Kernel ES ablation.","marker":"Sun et al. 2025"},{"why":"Foundation of ergodic-search metrics that equate time-averaged trajectory statistics with a target distribution, the core formulation adopted here.","marker":"Mathew and Mezić 2011"},{"why":"Defines 2D Gaussian Splatting (2DGS), the surface-aligned map representation whose per-Gaussian confidence drives the information map.","marker":"Huang et al. 2024"},{"why":"OctoMap provides the voxel map with occupancy probability, updated on the fly and used for the information map's class indicators.","marker":"Hornung et al. 2013"},{"why":"Supplies the eight Replica scenes used for the main quantitative evaluation.","marker":"Straub et al. 2019"},{"why":"Habitat simulator provides the RGB-D sensor model and rendering pipeline for the experiments.","marker":"Savva et al. 2019"}],"fun_headline_variants":["Ergodic trajectory planning beats next-best-view by 1.5 dB PSNR","Active reconstruction via ergodic coverage improves PSNR by 1.5 dB","TRACE: ergodic optimization for active reconstruction yields +1.5 dB","Replace greedy next-best-view with ergodic trajectories for better maps"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the online information map of Eq. (2), after diffusion and masking, faithfully ranks where new observations will most improve reconstruction quality; if that ranking is wrong, the ergodic trajectory will dwell in the wrong places and the PSNR gain will not transfer to other scenes or settings.","fun_headline_variants_meta":{"raw":{"variants":["Ergodic trajectory planning beats next-best-view by 1.5 dB PSNR","Active reconstruction via ergodic coverage improves PSNR by 1.5 dB","TRACE: ergodic optimization for active reconstruction yields +1.5 dB","Replace greedy next-best-view with ergodic trajectories for better maps"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000354,"raw_usage":{"total_tokens":1925,"prompt_tokens":947,"completion_tokens":978,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":563,"completion_tokens_details":{"reasoning_tokens":894}},"tokens_in":563,"tokens_out":978,"duration_ms":11529,"temperature":1.0,"reasoning_tokens":894,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T00:54:01.362650+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"On the eight Replica scenes, run TRACE with the information map of Eq. (2) replaced by a uniform distribution over the same free-space mask (all weights equal), keeping every other parameter fixed; if the PSNR gain over ActiveGS persists, the claim that map-guided dwell time drives the improvement is falsified.","supporting_citations":[],"review_version":3}