{"id":"8be5ea45-027f-4f82-8200-603bad99245c","arxiv_id":"2412.20720","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"4D Gaussian Splatting optimizes a set of 4D Gaussian primitives to render photorealistic novel views of dynamic scenes in real time.","lead":"Dynamic 3D scenes are represented as a cloud of 4D Gaussian ellipsoids spanning space and time, optimized directly from video for real-time novel-view rendering. The method reports top quality scores on dynamic-scene benchmarks and adds compact, driving-scene, generation, and segmentation extensions.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The abstract's 'first real-time high-resolution photorealistic dynamic novel view synthesis' claim is contradicted by Table 1, where 4D-Rotor-Gaussian achieves 277 FPS with comparable PSNR; the priority claim is false as stated.","rationale":"The reader's verdict of CONDITIONAL is sound. The paper's strongest claim—being the first to achieve real-time photorealistic dynamic novel view synthesis—is contradicted by its own Table 1: 4D-Rotor-Gaussian reaches 277 FPS on the same benchmark. This is a directly checkable factual issue, not a matter of interpretation, and it requires softening the abstract and Section 1. The reader's weakest_assumption focuses on the linear-conditional-mean property of 4D Gaussians. That concern is legitimate but less load-bearing than the priority overclaim, because the paper explicitly frames its motion model as piecewise-linear (Section 4.1) and the experiments demonstrate that it suffices on the tested benchmarks. The method is mathematically consistent: the conditional Gaussian factorisation is correct, the unnormalised-Gaussian proof holds, and the 4DSH basis is a standard Fourier/spherical product basis. No internal contradiction undermines the technical construction. The practical concerns—no released code, no error bars, and a typo in Table 2's LPIPS column for 4DGSC—strengthen the CONDITIONAL recommendation but do not change the core risk. The verdict should remain CONDITIONAL pending revision of the 'first real-time' claim and the listed clarifications.","tokens_in":25205,"tokens_out":11422,"duration_ms":107745,"concrete_test":"Run the official 4D-Rotor-Gaussian [44] code on the Plenoptic Video benchmark under the same hardware, resolution, and evaluation protocol used for Table 1, and measure FPS and PSNR. If 4D-Rotor-Gaussian exceeds 30 FPS with PSNR within about 0.5 dB of 4DGS, the 'first real-time' claim is false. Also report the exact GPU, resolution, and measurement methodology for Table 1 to assess the 'high-resolution' qualifier.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim in the abstract and Section 1 is that 4DGS 'has been the first solution to achieve real-time rendering of high-resolution, photorealistic novel views for complex dynamic scenes.' This claim fails on the paper's own evidence: Table 1 lists 4D-Rotor-Gaussian [44] at 277 FPS with PSNR 31.62, versus 4DGS at 114 FPS with PSNR 32.01 on the Plenoptic Video benchmark. Since 4D-Rotor-Gaussian is a prior published method (SIGGRAPH 2024) that also renders dynamic scenes in real time at comparable quality, the 'first' qualifier is demonstrably false. Furthermore, the FPS numbers in Table 1 are measured on Plenoptic Video at 2x downsampled resolution (Section 8.1.1), so the 'high-resolution' part of the claim is not established for the real-time figure. The reader's weakest_assumption about linear conditional motion is a real modeling limitation, but it is explicitly acknowledged in Section 4.1 and is not internally inconsistent; the priority overclaim is a factual error in the paper's headline contribution. This concern is load-bearing because it invalidates the strongest claim as written, even though the technical method and experiments remain valuable.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces 4D Gaussian Splatting (4DGS), a representation for dynamic scenes built from native 4D Gaussian primitives. Each primitive has a 4D mean and covariance, parameterized through 4D rotations and scales; at render time, the conditional 3D Gaussian at time t and the marginal temporal distribution are derived from the full 4D Gaussian, and appearance is modeled with a proposed 4D Spherindrical Harmonics basis. The representation is trained end-to-end with photometric supervision and is extended with compact variants (R-VQ, Huffman encoding, mask pruning), urban-scene adaptations with LiDAR and motion regularizations, video-to-4D generation with diffusion priors, and SAM-based 4D segmentation. Experiments cover Plenoptic Video, Technicolor, D-NeRF, Consistent4D, and Waymo dynamic scenes, reporting state-of-the-art or competitive quality and high rendering speed.","tokens_in":25523,"tokens_out":5725,"duration_ms":51800,"significance":"The technical core of the paper is sound: the derivation in Section 3.2 is standard multivariate-Gaussian conditioning, and the provided proof that the unnormalized Gaussian factorizes into conditional and marginal unnormalized Gaussians is correct. This yields a clean explicit primitive that can model appearance and disappearance without global tracking, and the paper demonstrates considerable breadth across reconstruction, compression, generation, and segmentation. The experiments are extensive and the compression results are practically valuable. However, the paper's headline claim of being the 'first solution to achieve real-time rendering of high-resolution, photorealistic novel views for complex dynamic scenes' is contradicted by the paper's own Table 1, which lists the prior 4D-Rotor-Gaussian method at 277 FPS with comparable PSNR. This overclaim is load-bearing because it is the stated central contribution in the abstract and introduction, and it must be corrected. The underlying representation and pipeline remain a useful contribution once the priority and speed claims are stated accurately.","major_comments":[{"comment":"The paper's claim that 4DGS 'has been the first solution to achieve real-time rendering of high-resolution, photorealistic novel views for complex dynamic scenes' is contradicted by Table 1 on the paper's own evidence. 4D-Rotor-Gaussian [44], a prior SIGGRAPH 2024 method, is listed at 277 FPS with PSNR 31.62, while 4DGS is listed at 114 FPS with PSNR 32.01. The same table also makes the statement in Section 8.1.2 that 4DGS is 'the sole method capable of real-time rendering while delivering high-quality dynamic novel view synthesis' false as written. Additionally, the FPS figures in Table 1 are measured at 2x downsampled resolution (Section 8.1.1), so the 'high-resolution' component of the real-time claim is not established by the reported measurements. The abstract, introduction, and Section 8.1.2 should be revised to state the method's actual comparative position rather than a priority claim.","section":"Abstract, Section 1, Section 8.1.2, Table 1"},{"comment":"Table 2 reports LPIPS 0.865 for 4DGSC, while 4DGS is 0.084 and the surrounding text states that the compact variant achieves 'significant improvement in compactness without significant compromise in quality.' An LPIPS of 0.865 is an order-of-magnitude outlier and appears to be a typo, likely 0.086 or 0.084. As printed, the table directly undermines the compression-quality claim, so the entry must be corrected and verified against the actual evaluation outputs.","section":"Table 2"},{"comment":"The text states that a higher opacity threshold 'eliminates floaters in the scene, leading to a higher PSNR (33.46 vs. 30.48)', but Table 7 reports 33.81 vs. 33.53 for Sear Steak and 34.02 vs. 33.87 for Cut Beef. The numbers in the text appear in neither row, and the value 30.48 does not occur in the table. Since the section's conclusion about the pruning threshold depends on this comparison, the text and table need to be reconciled.","section":"Section 8.3.2, Table 7"}],"minor_comments":[{"comment":"The equation numbered (32) in Section 7 appears to be a placeholder: the displayed formula for sM contains the text '(32)' and no right-hand side, while Eq. (32) in Section 6 is the generative loss. The scale formula must be completed and renumbered to avoid the collision.","section":"Section 7, Eq. (32)"},{"comment":"The row labeled 'No-Time split' is not defined in the text; it is not clear which component of the densification or rendering pipeline is removed. Please define this ablation.","section":"Table 3"},{"comment":"The sentence 'suggest that the proposed general representation can also work well in such an ill-posed task' is a sentence fragment with a subject-verb disagreement; it should be attached to the preceding sentence about initialization.","section":"Section 8.2.1"},{"comment":"There are numerous typographical issues, including 'T raining', 'T echnicolor', 'Storge', 'T able', and 'sckit-image'; these should be corrected in a final pass.","section":"Throughout"},{"comment":"The caption does not explain the color coding of the rendered optical flow; please add a brief description of the flow color convention.","section":"Figure 6"},{"comment":"The table header 'Deform-GS C-DyNeRF' appears to concatenate two separate baseline entries; clarify the column structure so each method is identified individually.","section":"Table 5"},{"comment":"The piecewise-linear interpretation of motion is a useful and honest discussion, but the associated trade-off for very long videos is not quantified; a sentence on how the number of Gaussians grows with video length would strengthen this section.","section":"Section 4.1"}],"recommendation":"major_revision","confidential_remarks":"The technical derivation and experimental breadth are solid, and the paper should not be rejected on technical grounds. The main issue is the overstated priority and real-time claim, which is contradicted by the authors' own Table 1. I also found an apparent LPIPS typo in Table 2 and an inconsistency between the text and Table 7. These are all fixable within the manuscript's scope, so I recommend major revision rather than rejection. The editor may also wish to check whether the paper adequately discloses the relationship to the authors' prior ICLR work [13], although the current text does mention it explicitly."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Read this one before you assign a referee to it. The key thing to know: the abstract's 'first solution to achieve real-time rendering of high-resolution, photorealistic novel views for complex dynamic scenes' is contradicted by the paper's own Table 1, where 4D-Rotor-Gaussian hits 277 FPS at 31.62 PSNR versus 114 FPS at 32.01 for this method. The FPS numbers are also measured on 2x downsampled Plenoptic Video, so 'high-resolution' is doing work the experiments don't back. The headline claim needs to be cut or qualified.\n\nThat said, the technical content is solid and the paper is more than a repackaging. The Gaussian conditional/marginal factorization in Section 3.2 is correct; the proof of the unnormalized Gaussian property is clean. The authors openly describe the piecewise-linear motion assumption in Section 4.1, which is the main modeling limitation, and they show it still works across several benchmarks. The extensions beyond the preliminary ICLR paper are real: compact variants with R-VQ and mask pruning bring Cut Roasted Beef from 1183 MB to 56.7 MB with a 0.1 dB drop; the urban-scene adaptation with LiDAR supervision and the segmentation via scale-gated features are non-trivial and sensibly evaluated. The generative pipeline is more standard but the comparison is fair.\n\nSoft spots, in order: the priority overclaim is load-bearing in the abstract and introduction, and it should be fixed before publication. The 4DSH basis is asserted to be orthonormal without proof or supporting reference; that's a minor gap but worth addressing. There are no error bars on any of the reported numbers, and no code release, which makes the FPS comparisons hard to verify. Table 2 has an LPIPS typo for 4DGSC (0.865 presumably 0.0865). None of these undercut the central method; they're housekeeping plus one framing error.\n\nIf you work on dynamic NVS, Gaussian splatting, or 4D generation, this is worth your time, mostly for the application extensions and the compression study. I'd send it to peer review with a required revision that deflates the 'first' claim, clarifies the 4DSH basis, and ideally adds error bars or a code link.","headline":"Solid extension of the authors' ICLR 4DGS work, with useful applications and compression, but the abstract's 'first real-time high-resolution' claim is contradicted by its own Table 1.","tokens_in":26100,"tokens_out":2740,"would_cite":true,"duration_ms":25128,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper introduces native 4D Gaussian primitives—Gaussians defined in space and time—as a single explicit representation for dynamic scenes, claiming the first real-time, high-resolution, photorealistic novel view synthesis on complex…","keywords":["4D Gaussian splatting","dynamic scene representation","novel view synthesis","real-time rendering","spatiotemporal 4D volume","4D spherindrical harmonics","4D segmentation","Gaussian compression"],"falsifier":"Render a scene containing a small object rotating rapidly around its own center (e.g., a spinning wheel or a twirling baton) and track the number of 4D Gaussians assigned to that object as render quality reaches a target PSNR: if the count grows steeply with rotation speed while rendering frame rate drops below real-time, the piecewise-linear trajectory assumption is the cause. A more direct check is to compute the model's implied optical flow (from conditional means) in a high-acceleration region and compare it with ground-truth flow, looking for systematic error at high curvature.","tokens_in":24929,"feed_emoji":"🎥","tokens_out":4925,"duration_ms":43614,"temperature":0.7,"pith_summary":"The paper's central claim is that a dynamic scene can be modeled directly as a collection of 4D Gaussian primitives, each carrying a mean and covariance in (x, y, z, t), and that optimizing these primitives under photometric loss alone yields real-time, high-resolution, photorealistic novel views of complex dynamic scenes, which the authors say is a first. The significance is that a single explicit spatiotemporal representation could serve rendering, 4D content generation, and 4D scene understanding without the interference or tracking ambiguities of previous implicit or deformation-based approaches. The authors also show that memory cost can be cut by roughly an order of magnitude with vector quantization and pruning, producing compact variants that retain most of the quality.","feed_headline":"Native 4D Gaussians render dynamic scenes in real time","feed_subtitle":"One explicit spatiotemporal primitive set handles novel views, 4D generation, and 4D segmentation at interactive frame rates.","key_machinery":"The central object is the 4D Gaussian with mean µ and covariance Σ = R S S^T R^T, where S is a 4D scale matrix and R is a 4D rotation encoded by a pair of quaternions (left and right isotropic rotations). Its work is to make time a coordinate rather than a conditioning input: given a query timestamp t, equation (9) produces a conditional 3D Gaussian for splatting and a marginal 1D Gaussian p(t) for temporal gating, so the existing 3D Gaussian rasterizer can be adapted with negligible overhead. The 4D Spherindrical Harmonics carry appearance evolution in time and direction.","core_discovery":"The discovery is the native 4D Gaussian primitive: a 4D anisotropic Gaussian distribution over space and time, parameterized by a 4D mean, a diagonal scale, and a rotation built from a pair of quaternions. Because the conditional distribution of any multivariate Gaussian is Gaussian, at each time t the primitive yields a conditional 3D Gaussian—mean linear in t and covariance independent of t—and a marginal 1D Gaussian in time; the paper proves an unnormalized-Gaussian factorization that makes this splitting exact. This turns dynamic scene rendering into a 3D Gaussian splatting operation with an extra time-marginal weighting, and a 4D extension of spherical harmonics adds time-varying, view-dependent color. Trained end-to-end on rendering loss with densification in space and time, the representation produces state-of-the-art or competitive quality on the Plenoptic Video and Technicolor benchmarks while rendering around 100 fps, and the same primitives are shown to support video-to-4D generation, urban driving scene reconstruction, and 4D segmentation.","pith_inferences":["A testable consequence the paper leaves implicit: if the piecewise-linear trajectory assumption holds, the number of primitives needed to represent a moving object should grow with the curvature of its motion, so a scene with many fast multi-axis rotations should require proportionally more 4D Gaussians than the same scene with translational motion.","The conditional-mean linearity suggests a natural diagnostic: compare the rendered optical flow (which the paper extracts from conditional means) against ground-truth flow in regions of high acceleration; systematic underestimation of curvature would confirm the piecewise-linear limitation.","One could extend the representation to higher-order motion (e.g., adding velocity or acceleration states per primitive) while keeping the same rendering pipeline; the paper's derivation suggests this would be a drop-in change to the conditional distribution formula."],"forward_implications":["If the claim is right, real-time dynamic novel view synthesis no longer requires per-frame reconstruction or deformation networks: one explicit 4D primitive set is trained end-to-end from multi-view video and renders at interactive rates.","The same 4D Gaussians can drive downstream tasks—4D generation from monocular video with score distillation, dynamic urban scene reconstruction with LiDAR initialization, and multi-granularity 4D segmentation—without changing the underlying representation.","Compact variants (residual vector quantization of shape and color attributes, mask-based pruning of insignificant Gaussians, half-precision positions, 8-bit opacity) shrink a scene from over 1 GB to about 57 MB with roughly 0.1 dB PSNR loss, making storage less prohibitive.","Because rendering time depends mainly on how many Gaussians are active at a given time (filtered by p(t) < 0.05), rendering speed stays roughly constant as video length grows."],"supporting_citations":[{"why":"Supplies the differentiable rasterizer, optimization defaults, and anisotropic Gaussian parameterization that the 4D representation extends.","marker":"[23]"},{"why":"Provides the Plenoptic Video benchmark and the neural volume baseline that the method must beat.","marker":"[4]"},{"why":"Defines the monocular dynamic scene setting and dataset used in the video-to-4D reconstruction experiments.","marker":"[9]"},{"why":"The closest Gaussian-based baseline for dynamic scenes, showing the prior state of 4D Gaussian approaches.","marker":"[41]"},{"why":"The preliminary version of this work, establishing the initial 4DGS formulation that this paper extends.","marker":"[13]"},{"why":"Supplies residual vector quantization and mask pruning strategies adopted for the compact variants.","marker":"[34]"},{"why":"Provides the score distillation sampling (SDS) objective used in the generative 4D pipeline.","marker":"[65]"},{"why":"Supplies the Waymo-NOTR Dynamic-32 split and the emergent decomposition baseline for urban scene experiments.","marker":"[61]"}],"fun_headline_variants":["4D Gaussian splatting: real-time dynamic scenes","Native 4D primitives render dynamic scenes at 100 fps","Real-time 4D Gaussian rendering for dynamic scenes","Explicit 4D Gaussians enable real-time dynamic scenes","One 4D primitive set does real-time dynamic rendering"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing assumption is that each 4D Gaussian's spatial mean moves along a straight line in time (its conditional mean is linear in t), so all real motion must be approximated piecewise by many primitives; if a scene's dynamics cannot be captured this way without exploding the primitive count, the real-time and quality claims weaken.","fun_headline_variants_meta":{"raw":{"variants":["4D Gaussian splatting: real-time dynamic scenes","Native 4D primitives render dynamic scenes at 100 fps","Real-time 4D Gaussian rendering for dynamic scenes","Explicit 4D Gaussians enable real-time dynamic scenes","One 4D primitive set does real-time dynamic rendering"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001116,"raw_usage":{"total_tokens":4670,"prompt_tokens":994,"completion_tokens":3676,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":610,"completion_tokens_details":{"reasoning_tokens":3593}},"tokens_in":610,"tokens_out":3676,"duration_ms":24122,"temperature":1.0,"reasoning_tokens":3593,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T23:13:35.242893+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Render a scene containing a small object rotating rapidly around its own center (e.g., a spinning wheel or a twirling baton) and track the number of 4D Gaussians assigned to that object as render quality reaches a target PSNR: if the count grows steeply with rotation speed while rendering frame rate drops below real-time, the piecewise-linear trajectory assumption is the cause. A more direct check is to compute the model's implied optical flow (from conditional means) in a high-acceleration region and compare it with ground-truth flow, looking for systematic error at high curvature.","supporting_citations":[{"cited_title":"In: IEEE Conference on Computer Vision and Pattern Recognition (2024)","cited_arxiv_id":null,"evidence_quote":"The closest Gaussian-based baseline for dynamic scenes, showing the prior state of 4D Gaussian approaches."},{"cited_title":"In: IEEE Conference on Computer Vision and Pattern Recognition (2023)","cited_arxiv_id":null,"evidence_quote":"Provides the score distillation sampling (SDS) objective used in the generative 4D pipeline."},{"cited_title":": Emernerf: Emer- gent spatial-temporal scene decomposition via self-supervision","cited_arxiv_id":null,"evidence_quote":"Supplies the Waymo-NOTR Dynamic-32 split and the emergent decomposition baseline for urban scene experiments."}],"review_version":1}