{"id":"389203c2-1b4a-44ea-8b41-936b51118229","arxiv_id":"2608.08531","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"ERF-GS improves dynamic 3D Gaussian splatting by using simulated event streams to supervise motion and densify moving regions, yielding higher dynamic PSNR on blurred and disjoint-event versions of Neu3D and Nvidia.","lead":"This paper introduces two event-based components for dynamic 3D Gaussian splatting, a regularized loss and a densification strategy, that improve reconstruction of fast-moving scenes from sparse, blurry RGB frames and disjoint event viewpoints. It reports consistent gains in dynamic PSNR over a strong RGB-only baseline on simulated event versions of Neu3D and Nvidia benchmarks.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Simulated-event proxy is the load-bearing weakness: v2e events derive from the same RGB frames used for evaluation, so real-hardware transfer of the >0.9 dB gain is not established.","rationale":"The reader's weakest_assumption correctly identifies the simulated-event proxy as the central vulnerability. I agree with that identification and sharpen it: the proxy is not merely a question of simulation realism, but of information leakage, because v2e synthesizes events from the very high-frame-rate RGB videos that also supply the evaluation ground truth. This gives ERF-GS access, at intermediate timesteps, to exact log-intensity differences between frames that the subsampled RGB inputs deliberately hide. Such leakage could inflate DPSNR gains relative to what a real event camera, with its own noise, latency, and unknown pose, would provide. The paper's own Sec. 5.3 explicitly concedes this limitation. Consequently, the central real-world applicability claim should remain conditional until validated on real event hardware or on a simulation pipeline that does not derive events from the evaluation frames. The internal ablations and the favorable results on -ts, -mb, and -dv variants are credible as evidence that the method works under the paper's protocol, but they do not by themselves establish transfer to real event sensors. I therefore leave the reader's CONDITIONAL verdict unchanged: acceptance should require the proposed real-data or independent-simulation validation.","tokens_in":19274,"tokens_out":5556,"duration_ms":66703,"concrete_test":"Settle by evaluating ERF-GS and the 4DGS baseline on a real multiview event-RGB dataset with fast motion and a disjoint event/RGB camera configuration (at minimum 10 static cameras, e.g., a DAVIS346 rig). If such data is unavailable, use a synthetic scene where events are generated by a physically-based simulator from a render pass independent of the evaluation frames and ground truth, so event streams are not derived from the same RGB videos used for testing. Rerun the exact Sec. 4.2 hyperparameters and report DPSNR deltas on the same -dv style split; if the >0.9 dB gain does not reproduce, the headline claim is conditional on the simulation proxy.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim is that fusing events into 4DGS improves dynamic NVS by >0.9 dB DPSNR in realistic disjoint-view setups. The load-bearing assumption is that the v2e event streams created in Sec. 4.1 faithfully represent real event-camera observations. This is not merely a domain-gap caveat: in every experiment, events are generated from the same original high-frame-rate RGB videos that later define the evaluation ground truth and the blurred training inputs. At an intermediate timestep, an event frame is essentially the log-intensity difference between two original frames, i.e., exactly the supervision needed to recover the motion that subsampled RGB hides. A real event camera at a disjoint viewpoint would produce asynchronous, noisy, low-resolution events with different optics, latency, and pose relative to RGB; Sec. 5.3 concedes 'the domain gap between simulated and real-captured events is unavoidable.' Since EARL (Eq. 6) and EDS (Sec. 3.3) are optimized against these simulated events, the measured gains can overstate what is achievable with real hardware. The paper notes that no public real-world multiview event-RGB dataset exists, so the reported numbers are internally consistent but not externally validated.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes ERF-GS, a dynamic 3D Gaussian splatting framework that fuses event streams into both the optimization and densification stages of a 4DGS backbone. The two proposed components are EARL, an event-assisted regularized loss that supervises rendered images against simulated event frames from possibly disjoint viewpoints, and EDS, an event-guided densification strategy that lifts event pixels to 3D candidates and cross-validates them across views. Experiments are conducted on temporally subsampled and motion-blurred variants of the Neu3D and Nvidia datasets, with v2e-generated events, and the paper reports average DPSNR improvements of 0.91 dB over 4DGS on Neu3D-dv and 1.44 dB on Nvidia-dv, plus a comparison with the E-D3DGS baseline. The paper also releases code and processed datasets.","tokens_in":19540,"tokens_out":4317,"duration_ms":48619,"significance":"If the reported gains transfer to real event-camera captures, the paper would be a useful contribution to event-assisted dynamic scene reconstruction: it addresses realistic constraints such as disjoint event-RGB viewpoints, does not require ground-truth RGB at intermediate times, and is designed as a plug-in for existing dynamic Gaussian splatting methods. The ablations in Table 5 support the usefulness of both EARL and EDS, and the release of code and processed datasets is a concrete strength. However, the significance is conditional on external validity: every experiment uses simulated events generated from the same RGB videos that supply both training supervision and evaluation ground truth, and the paper itself concedes the simulated-to-real domain gap. The central claim of a >0.9 dB improvement is therefore not yet validated for real hardware, which is the main risk in assessing this work.","major_comments":[{"comment":"The evaluation's only event modality is v2e events generated from the same high-frame-rate RGB videos that define the evaluation ground truth and the blurred training inputs. An event frame is essentially a log-intensity difference between original frames, so at intermediate timesteps it encodes exactly the motion that subsampled RGB hides. Because EARL in Eq. (6) and EDS in Sec. 3.3 are optimized against these same-data-derived events, the reported 0.91 dB and 1.44 dB DPSNR gains are internally consistent but do not establish that the method improves reconstruction from real event sensors. The paper's own statement in Sec. 5.3 that 'the domain gap between simulated and real-captured events is unavoidable' makes this a load-bearing validity gap, not a cosmetic caveat. The authors should either validate on real multiview event-RGB data or provide a simulation protocol in which events are not derived from the exact frames used for supervision and evaluation.","section":"Sec. 4.1, Sec. 5.3, Eq. (6)"},{"comment":"The comparison with E-D3DGS is not on equal footing. The paper states that E-D3DGS had to be run with 2x2 downsampled inputs 'essentially reducing the reconstruction difficulty,' while ERF-GS runs at full resolution. The 10.30 dB DPSNR gap on Neu3D-dv (16.35 vs. 26.65) could be largely attributable to resolution rather than method quality. Please retrain E-D3DGS at full resolution under comparable memory settings, or clearly label it as a lower-resolution baseline and do not use the margin to support the claim of superiority.","section":"Sec. 4.5, Table 8"},{"comment":"The headline average improvements are not supported by error bars or multiple runs. Per-scene results include meaningful regressions, for example 'cut roasted beef' drops from 27.17 to 26.09 DPSNR in Table 1 and 'Dynamicface' is worse in all metrics in Table 2, while 'Balloon2' is excluded post hoc in Sec. 4.1. Since the central claim is an average gain of more than 0.9 dB, the authors should report standard deviations over several seeds, state the per-scene success rate, and either include Balloon2 or justify its exclusion with a pre-specified criterion rather than a post hoc observation of poor quality.","section":"Sec. 4.1, Tables 2 and 3"}],"minor_comments":[{"comment":"The section title says 'confidence-weighed event loss' but 'weighed' should be 'weighted'.","section":"Sec. 3.2.1"},{"comment":"The abstract's phrase 'event-based learning detached from RGB inputs' is overstated: Eq. (5) still uses rendered RGB-to-luma images, and Eqs. (7) and (8) anchor intermediate states to RGB-supervised Gaussian positions and colors. Suggest rewording to 'detached from ground-truth RGB frames at intermediate timesteps.'","section":"Abstract and Sec. 3.2"},{"comment":"The event-frame visualizations are small and hard to read; please annotate the timestamps, viewpoints, and failure regions explicitly so the qualitative claims can be verified.","section":"Fig. 9 and Fig. 12"},{"comment":"The '-mb' variants are created by averaging each subsampled frame with its two original neighbors, and the paper itself notes this can leak information from adjacent frames; this makes the -mb results difficult to interpret as a test of robustness to genuine motion blur.","section":"Sec. 4.3, Table 3"}],"recommendation":"major_revision","confidential_remarks":"The paper is a solid engineering contribution with reproducible code and a clean ablation story, but the simulated-event proxy is the key validity risk. I would like to see either a real-event experiment or a serious reframing of the claims as simulation-only, plus a fairer E-D3DGS comparison and error bars, before publication in a top venue."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague, the genuinely new pieces here are EARL and EDS: event information goes into both the loss and the densification of dynamic 3DGS, and the evaluation protocol is harder than what most event-assisted reconstruction papers use, with temporally subsampled, motion-blurred RGB and disjoint event-RGB viewpoints. The paper is also honest that events are simulated, and the ablation study shows each component contributes. Code and processed datasets are released, which is real work.\n\nThe central internal result — about 0.9–1.4 dB DPSNR gain over the 4DGS baseline on the -dv variants — holds in their simulated setting. The numbers are consistent with the claim that event-based supervision recovers fast motion that subsampled RGB hides. The trajectory and color regularizers address a known failure mode of event-supervised splatting, and the failure analysis in Sec. 5.2 is refreshingly direct.\n\nSoft spots, in order. First, the load-bearing assumption is that v2e events generated from the same RGB videos are a faithful proxy for real event cameras at disjoint viewpoints. This is not a minor caveat: at intermediate timesteps, an event frame is essentially a log-intensity difference between two original frames, so it carries exactly the motion information that subsampled RGB drops. Sec. 5.3 concedes the domain gap. That makes the quantitative gains credible only for simulated events until real-hardware validation appears. Second, the E-D3DGS comparison is weakened by running that baseline at 2x downsampled resolution; the gap is large, but the comparison reads as more about protocol than method. Third, no error bars, one run per setting, and Balloon2 excluded post hoc. None of these are fatal for the internal claim, but they cap how strongly the paper can claim real-world transfer.\n\nCitation pattern is normal for the area; nothing looks off. Who this is for: people working on event-aided 3D reconstruction or fast-motion capture. They will get a clear method description and a reusable benchmark protocol. That deserves a serious referee. I would ask the authors to add variance across seeds, restore or justify Balloon2, and either run E-D3DGS at full resolution or drop the direct comparison. The central internal result is intact; the external claim is not yet supported.","headline":"Solid incremental work on event-RGB dynamic Gaussian splatting with a useful new benchmark protocol, but the headline gain rests on simulated events derived from the same RGB supervision.","tokens_in":20095,"tokens_out":1936,"would_cite":true,"duration_ms":20822,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Fusing event-camera data into Gaussian splatting lifts dynamic view synthesis by about 1 dB on moving pixels.","keywords":["dynamic scene reconstruction","event camera","3D Gaussian splatting","multimodal fusion","novel view synthesis","motion blur","event simulation","densification"],"falsifier":"Run the same Neu3D- and Nvidia-style scenes with physical event cameras in a disjoint-view capture rig; if the dynamic-PSNR gain over the RGB-only baseline drops to within noise, or color-shift artifacts grow as they do on simulated reflective and dark regions, the central transfer claim fails.","tokens_in":19070,"feed_emoji":"⚡","tokens_out":7069,"duration_ms":67807,"temperature":0.7,"pith_summary":"The paper claims that the missing ingredient for reconstructing fast-moving scenes is not a better RGB model but a second, faster modality: the asynchronous event stream from an event camera. It proposes ERF-GS, which attaches two event-driven components to an existing dynamic Gaussian splatting pipeline, in which a dynamic scene is stored as thousands of moving 3D Gaussian blobs. The first component, EARL, supervises the representation between RGB frames with a confidence-weighted event loss plus trajectory and color regularizers, using no ground-truth RGB frames. The second, EDS, places new Gaussians exactly where event data indicates motion, after cross-validating candidates across multiple event viewpoints. On temporally subsampled, motion-blurred natural videos with disjoint RGB and event viewpoints, the paper reports that ERF-GS improves dynamic PSNR over the RGB-only baseline by 0.91–1.44 dB; the point is to make event-assisted reconstruction work without aligned RGB-event viewpoints or clean RGB frames.","feed_headline":"Event-camera data adds ~1 dB to fast-motion 3D view synthesis","feed_subtitle":"Fusing sparse event streams into Gaussian splatting recovers moving detail that blurred, disjoint RGB views miss.","key_machinery":"The load-bearing machinery is two attached components rather than a new representation. EARL (event-assisted regularized loss) treats an event frame as the difference in log-luma between two rendered images, weights each pixel by the inverse of its event count so that noisy spikes are downweighted, and adds linear-trajectory and color-stability regularizers to stop Gaussians from zigzagging or shifting hue while fitting sparse events. EDS (event-guided densification) unprojects nonzero event pixels into depth-sampled 3D candidates, re-projects them into several other event viewpoints for cross-validation, and filters the survivors by depth uniqueness, spatial coherence with existing Gaussians, and deformation consistency before adding them. Together these components give the Gaussian cloud the high-temporal-resolution motion evidence that RGB-only training lacks.","core_discovery":"ERF-GS establishes that event information can be injected into both optimization and densification of a dynamic 3D Gaussian representation without requiring ground-truth RGB frames, color-channel alignment, or shared viewpoints between event and RGB cameras. Using the v2e simulator to turn natural multiview videos into event streams, then degrading RGB by temporal subsampling and motion blur and holding out event-only viewpoints, the method reports that the 4DGS backbone gains 0.91 dB in dynamic PSNR on Neu3D-dv and 1.44 dB on Nvidia-dv, with average PSNR gains of 0.17 dB and 0.57 dB. The central discovery is that sparse monochrome events, interpreted as relative log-intensity changes between two rendered predictions, can both supervise Gaussian motion and guide where new Gaussians are placed, recovering fast-moving content that blurred RGB frames alone do not resolve.","pith_inferences":["A natural next experiment is to run the same pipeline on real event streams from a multiview rig; if the 1 dB margin persists, the simulation-to-real gap the paper acknowledges would be largely bridged.","The robustness of EDS to the number of disjoint event viewpoints, tested with two to four, suggests a practical capture design with only two or three event cameras supplementing a larger RGB rig.","The failure cases around reflective surfaces and dark regions indicate that future event-assisted methods will need to distinguish motion-triggered events from lighting- and specularity-triggered events before deployment in uncontrolled scenes.","A testable extension would combine EARL and EDS with simultaneous camera-pose optimization for moving event cameras, since the paper's static-camera assumption is an acknowledged simplification."],"forward_implications":["If the result holds, an existing dynamic Gaussian splatting method can be upgraded with event fusion without redesigning its Gaussian representation or deformation network.","The method removes the usual requirement that event and RGB cameras share a viewpoint, so a capture rig can combine a sparse RGB camera set with event-only cameras at different positions.","The reported improvements concentrate on moving pixels, rising by 0.91–1.44 dB in dynamic PSNR while overall PSNR rises by 0.17–0.57 dB, indicating the event stream's value is specific to fast scene dynamics.","Training-set metrics at RGB-supervised timesteps also improve slightly, implying the event-driven losses regularize the representation even where ordinary RGB supervision exists."],"supporting_citations":[{"why":"Provides the dynamic Gaussian splatting backbone and the baseline that ERF-GS extends and compares against.","marker":"[56]"},{"why":"Generates the simulated event streams from RGB videos that all experiments rely on.","marker":"[18]"},{"why":"Defines the Gaussian splatting representation and differentiable rasterization pipeline that the backbone builds on.","marker":"[23]"},{"why":"The concurrent event-boosted deformable Gaussian baseline compared on the disjoint-view datasets.","marker":"[63]"},{"why":"One of the two multiview natural-video benchmarks used to construct the event-RGB variants.","marker":"[24]"},{"why":"The other multiview benchmark, with more complex dynamics, used for the -ts, -mb, and -dv variants.","marker":"[66]"},{"why":"Generates the dynamic masks used to compute the dynamic PSNR metric on moving pixels.","marker":"[44]"}],"fun_headline_variants":["Event-RGB Gaussian splatting sharpens fast-motion 3D","Event streams lift dynamic 3D quality by ~1 dB","Disjoint event and RGB views still rebuild motion","ERF-GS fuses events into Gaussian densification","No RGB alignment needed for event-supervised splatting"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that event streams synthesized by the v2e simulator from RGB video are a faithful stand-in for real event-camera recordings, so the measured gains would transfer to actual hardware.","fun_headline_variants_meta":{"raw":{"variants":["Event-RGB Gaussian splatting sharpens fast-motion 3D","Event streams lift dynamic 3D quality by ~1 dB","Disjoint event and RGB views still rebuild motion","ERF-GS fuses events into Gaussian densification","No RGB alignment needed for event-supervised splatting"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000261,"raw_usage":{"total_tokens":1601,"prompt_tokens":959,"completion_tokens":642,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":575,"completion_tokens_details":{"reasoning_tokens":558}},"tokens_in":575,"tokens_out":642,"duration_ms":5938,"temperature":1.0,"reasoning_tokens":558,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T04:32:47.338683+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same Neu3D- and Nvidia-style scenes with physical event cameras in a disjoint-view capture rig; if the dynamic-PSNR gain over the RGB-only baseline drops to within noise, or color-shift artifacts grow as they do on simulated reflective and dark regions, the central transfer claim fails.","supporting_citations":[{"cited_title":"4d gaussian splatting for real-time dynamic scene rendering","cited_arxiv_id":null,"evidence_quote":"Provides the dynamic Gaussian splatting backbone and the baseline that ERF-GS extends and compares against."},{"cited_title":"3D gaussian splatting for real-time radiance field rendering.ACM Trans","cited_arxiv_id":null,"evidence_quote":"Defines the Gaussian splatting representation and differentiable rasterization pipeline that the backbone builds on."},{"cited_title":"Event-boosted deformable 3d gaussians for dynamic scene reconstruction","cited_arxiv_id":null,"evidence_quote":"The concurrent event-boosted deformable Gaussian baseline compared on the disjoint-view datasets."},{"cited_title":"Novel view synthesis of dynamic scenes with globally coherent depths from a monocular camera","cited_arxiv_id":null,"evidence_quote":"The other multiview benchmark, with more complex dynamics, used for the -ts, -mb, and -dv variants."}],"review_version":1}