{"id":"c3b8e34a-80d2-4cab-b603-bbf8d0f47cef","arxiv_id":"2505.10049","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":3.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A survey that categorizes dynamic scene reconstruction methods from NeRF to 3D Gaussian splatting into a unified framework based on motion type and representation paradigm.","lead":"This paper reviews more than 200 studies on modeling moving 3D scenes with neural radiance fields and 3D Gaussian splatting. It organizes the field by motion type and representation style, and proposes a unified framework for comparing dynamic scene reconstruction methods.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Unified framework omits factorization-based motion representation, so the central claim of a single organizing scheme is internally incomplete.","rationale":"The most load-bearing part of the central claim is the 'unified representational framework.' If the framework actually covered all reviewed methods, the survey would deliver on its main promise. But Sec. 4.6's reference-frame spectrum omits factorization, a paradigm Sec. 3.2.5 defines and Table 3 populates. This is a direct, demonstrable gap in the paper's own analytical structure, not a speculative external critique. It means a reader cannot use the framework to locate a large class of methods (all plane- and basis-decomposition approaches), undermining the 'definitive reference' aspiration. The reader's CONDITIONAL verdict remains appropriate: the survey is useful and mostly accurate, but its central organizational claim is overfull. A secondary concern, the 'over 200 papers' count, is also unsupported because the reference list includes many background citations, but we treat the framework gap as primary because it is internal and checkable. No change to the verdict is needed; the authors should either extend the framework to cover factorization or revise the claim to a narrower scope.","tokens_in":36875,"tokens_out":6953,"duration_ms":62805,"concrete_test":"Derive from Sec. 3.2.5 the set of factorization-based methods in Table 3 (NPGs, FPO, Tensor4D, Hexplane, K-Planes, 4K4D, 4D GS, DeformGS, DynMF). For each, attempt to place it on the reference-frame spectrum of Fig. 5 and Sec. 4.6 (canonical, multi-keyframe, two-frame flow, per-frame volume, per-point trajectory). If none of the five positions can accommodate the factorization mechanism described—e.g., basis-driven motion does not use a reference frame at all—the unified framework omits an entire paradigm, confirming the internal inconsistency.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The survey's central contribution, stated in the Abstract and Sec. 1.3, is that it 'organize[s] diverse methodological approaches under a unified representational framework.' That framework is made concrete in Fig. 5 and Sec. 4.6: any dynamic scene is treated as a set of static reference frames plus transformations, and methods are placed along a spectrum by the number of reference frames (single canonical space, multiple keyframes, two-frame flow fields, per-frame 4D spacetime, per-point tracking). This spectrum, however, leaves out one of the five motion-representation paradigms the paper itself defines in Sec. 3.2: factorization (Sec. 3.2.5). Factorization appears in two forms—hyperplane-based (e.g., Hexplane [15], K-Planes [14], 4K4D [142]) and basis-driven (e.g., NPGs [173], DynMF [175])—and Table 3 dedicates a ten-row block to it. Sec. 4.6 does not assign factorization any position on the reference-frame spectrum; it is never mentioned in the unified-framework discussion. The claim that the framework can 'encapsulate these methods' and that the survey organizes all approaches is therefore internally inconsistent: a substantial category of the surveyed literature is outside the proposed organizing axis. This is not a matter of external corpus completeness but a flaw visible within the paper's own structure.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This survey reviews over 200 papers on dynamic scene representation and reconstruction using neural radiance fields and 3D Gaussian splatting. It proposes a taxonomy of motion types (rigid, articulated, non-rigid, hybrid) and of motion representation paradigms (4D spacetime, canonical space with deformation fields, flow fields, point tracking, and factorization), and organizes reconstruction methods into tables by motion type. It further discusses auxiliary information, regularization, and future trends, and claims to provide a unified representational framework (Sec. 4.6, Fig. 5) that relates methods by the number of reference frames used. The paper also maintains an active online repository of methods and implementations.","tokens_in":37119,"tokens_out":2486,"duration_ms":25278,"significance":"If the survey's claims hold, it would serve as a valuable entry point and reference for researchers in dynamic scene reconstruction, bridging implicit neural fields and explicit Gaussian primitives. The paper's strengths include a broad corpus of methods, a structured taxonomy of motion types and representation paradigms, useful summary tables (Tables 1–4), and an explicit discussion of regularization and auxiliary information. However, the central claim of a \"unified representational framework\" is weakened by the omission of one of the paper's own categories (factorization) from that framework, and the lack of a systematic search protocol reduces confidence in the survey's completeness and in its \"definitive reference\" claim. The paper's equations (1)–(20) are standard formulations from the cited literature, and the taxonomy is internally consistent apart from the noted gap.","major_comments":[{"comment":"The unified framework proposed in Sec. 4.6 and illustrated in Fig. 5 organizes methods along a spectrum defined by the number of reference frames (single canonical space, multiple keyframes, two-frame flow fields, per-frame 4D spacetime, per-point tracking). This framework omits factorization, which is one of the five motion-representation paradigms defined in Sec. 3.2.5 and which occupies a substantial block in Table 3 (e.g., Hexplane, K-Planes, 4K4D, NPGs, DynMF). Consequently, the claim in the Abstract and Sec. 1.3 that the survey organizes diverse methodological approaches under a unified representational framework is internally incomplete: a major category of the surveyed literature is not placed on the proposed organizing axis. The authors should either extend the framework to include factorization (e.g., as a complementary dimension concerned with how the motion field is decomposed rather than how many reference frames are used) or explicitly explain how factorization relates to the reference-frame spectrum.","section":"Sec. 4.6, Fig. 5, Sec. 3.2.5, Table 3"},{"comment":"The paper claims to provide a \"systematic analysis of over 200 papers\" and to be a \"definitive reference,\" but it reports no systematic search protocol, no inclusion or exclusion criteria, no database or time-window specification, and no audit trail linking each table entry to search decisions. Without such methodology, the completeness and representativeness of the corpus cannot be assessed, and the assertion of a definitive reference is not supported. The authors should describe their literature collection process in the paper or in a supplementary document, including how papers were selected and how the tables were populated.","section":"Abstract, Sec. 1.3"},{"comment":"Equation (9), which formalizes the decomposition of hybrid motion into a coarse global transformation and a fine non-rigid residual, is corrupted in the manuscript: it contains a long run of repeated tokens (e.g., \"rl\", \"mo\", \"r\") that makes the formula unreadable. Because hybrid motion is a core concept in the taxonomy and the equation is meant to provide the mathematical basis for the discussion that follows, this corruption is a load-bearing presentation error and must be fixed.","section":"Sec. 3.1.4, Eq. (9)"}],"minor_comments":[{"comment":"The heading \"Disscussion\" should be \"Discussion\".","section":"Sec. 3.3 heading"},{"comment":"The column header \"Auxilary\" is misspelled; it should be \"Auxiliary\".","section":"Tables 1–4"},{"comment":"The author affiliation contains \"T ao\" where \"Tao\" is intended.","section":"Author affiliation, Abstract footnote"},{"comment":"The term \"Lidar\" is inconsistently capitalized; use \"LiDAR\" throughout.","section":"Sec. 2.1.1"},{"comment":"The phrase \"casual video capture\" is used where \"casually captured video\" or \"casual capture\" would be clearer; the meaning is understandable but the phrasing is informal for a survey.","section":"Sec. 2.1.2 and elsewhere"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is within the journal's scope and addresses a timely topic. The main technical weakness is the gap between the claimed unified framework and the actual coverage of factorization; this is fixable by extending the framework or by reframing the claim. The corrupted equation and missing methodology are also fixable. I recommend major revision rather than rejection because the survey's core organizing ideas are useful and the corpus is broad, but the central claim of a unified framework needs to be made accurate and auditable."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's the quick take. This is a broad, well-organized survey of dynamic scene reconstruction from NeRF to 3DGS. It deserves to be used as an entry point, but the central claim of a unified representational framework is overstated. The reference-frame spectrum in Sec. 4.6 is a useful pedagogical device, but it leaves out factorization, one of the five motion-representation paradigms defined in Sec. 3.2.5. Fig. 5 and the discussion place canonical space, keyframes, flow fields, 4D volumes, and point tracking on a spectrum of reference-frame granularity, but factorization gets no slot. That's an internal inconsistency, not a coverage quibble: the paper's own taxonomy is wider than its unifying axis.\n\nWhat the paper does well: the motion-type taxonomy (rigid, articulated, non-rigid, hybrid) is sensible, the tables (especially Tables 2 and 3) provide a dense map of citations, and the discussion of auxiliary information and regularization is practical. The roadmap figures will genuinely help a newcomer. The survey covers a lot of ground and the reference list looks representative and current.\n\nThe soft spots are what you'd expect, plus one notable editorial failure. There's no systematic search protocol or inclusion criteria, so the 'over 200 papers' claim isn't auditable; and there are no quantitative comparisons—acceptable for a survey, but the abstract's 'definitive reference' phrase is too strong. The corrupted Eq. (9) in Sec. 3.1.4 looks like a production error and should be fixed before anyone cites it. The manuscript also has an unusual number of typos ('Disscussion', 'Auxilary', 'segentic'), which suggests it wasn't carefully proofread. None of these undermine the core survey value, but they dent the polish.\n\nWho's this for? A newcomer to dynamic scene reconstruction who wants a structured map and a bibliography. An experienced researcher will find the tables useful for checking coverage. The stress-test concern about factorization is real; it should be addressed in revision, either by adding factorization to the unified-framework discussion or by explicitly stating that the reference-frame spectrum covers only one axis and factorization is a complementary one.\n\nI'd send this to peer review. It's a genuinely useful survey with one structural gap and a set of fixable editorial problems. A serious referee could help the authors close the gap and clean it up.","headline":"A useful, broad survey whose unified-framework claim omits factorization (one of its own categories); worth refereeing after fixes.","tokens_in":37613,"tokens_out":4309,"would_cite":true,"duration_ms":39799,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This survey of over 200 papers argues that every method for reconstructing moving scenes fits one framework: a static reference space plus a motion model, with the number of reference frames setting the trade-off between detail and cost.","keywords":["motion representation","dynamic scenes","neural radiance field","3D gaussian splatting","deformation field","scene flow","novel view synthesis","4D reconstruction"],"falsifier":"Take a random sample of roughly twenty recent dynamic-scene papers and check each against the taxonomy: if a well-known method resists placement in one of the four motion types and one of the five representation paradigms, or if a table entry contradicts what the method's own paper claims, the unified framework would be incomplete rather than universal.","tokens_in":36694,"feed_emoji":"🎥","tokens_out":18314,"duration_ms":132077,"temperature":0.7,"pith_summary":"This survey of more than 200 papers tries to establish that the literature on reconstructing moving scenes with radiance fields is one coherent design space rather than an unstructured collection of approaches. Its central claim is that every method, from implicit neural radiance fields to explicit Gaussian primitives, can be read as a static reference space plus a motion model, with approaches differing mainly in how many reference frames anchor the motion. The paper organizes motion into four types (rigid, articulated, non-rigid, hybrid) and five representation paradigms, and argues that the shift to explicit Gaussian primitives is what unlocked real-time rendering, dense tracking, and editing. If the framework holds, it gives newcomers a single map of a fast-moving field and shows experienced researchers where design combinations remain unfilled.","feed_headline":"200+ moving-scene methods collapse into one design spectrum","feed_subtitle":"A survey of 200+ papers shows every 4D reconstruction method is a static reference plus a motion model.","key_machinery":"The organizing mechanism is the reference-frame spectrum introduced in the paper's unified framework, formalized by the point-motion map $x_t = T_\\theta(x_{t-1}; \\pi(t))$, which sends any 3D point to its position at a later time through a transformation conditioned on a temporal code. The spectrum spans two base representations: NeRF, the implicit neural field mapping position and view direction to color and density, and 3D Gaussian Splatting (3DGS), explicit collections of anisotropic Gaussian primitives rendered by splatting. The framework classifies every method by how many static reference frames anchor this transformation - one canonical space, multiple keyframes, consecutive-frame flow, per-frame 4D spacetime, or per-point trajectories - and by a motion-type taxonomy (rigid, articulated, non-rigid, hybrid). The survey's tables then classify each of the 200+ papers along these axes together with the auxiliary information (depth, segmentation, optical flow, data-driven priors) and regularizers (smoothness, rigidity, volume preservation) they employ.","core_discovery":"The paper's central organizational claim is that any dynamic scene can be conceptualized as a static reference space coupled with an appropriate motion representation, and that all reconstruction methods differ in how many reference frames they use. A single canonical space with a learned deformation field covers rigid, articulated, and mildly non-rigid scenes; several keyframes serve as local references when one global space fails; reducing the window to two frames yields frame-to-frame flow fields; letting each frame be its own reference gives full 4D spacetime optimization, where the fourth dimension is time; and at the finest granularity, per-point tracking builds continuous trajectories across the whole sequence. The survey claims that finer granularity captures more temporal detail at higher computational cost, and that hybrid methods - structured coarse motion plus a neural residual - tend to win on real scenes with mixed motion. Alongside this spectrum, the paper maps the field's trajectory from implicit MLP fields to explicit 3D Gaussian primitives and credits that shift with enabling real-time rendering, dense long-term tracking, and object-level editing.","pith_inferences":["The spectrum predicts that an adaptive method choosing its reference-frame granularity region by region - canonical where motion is small, flow where it is large - should beat any fixed-paradigm method; the paper argues for the spectrum's usefulness but does not test this construction.","The framework makes canonical-space and flow-field methods endpoints of one continuum rather than rivals; a published method that interpolates between them as sequence length grows would directly test the framework's predictive value.","By the paper's own logic, the analogue of the SMPL template for arbitrary object categories is a foundation-model semantic field: data-driven semantic features should supply the priors that hand-built kinematic trees supplied for humans, enabling template-free articulated reconstruction at scale.","The taxonomy yields a quantitative corollary the survey does not state: hybrid-classified methods in the tables should outperform single-paradigm methods on benchmarks with mixed rigid and non-rigid motion, which a reader could verify by aggregating the reported metrics of the cited papers."],"forward_implications":["A reader encountering a new dynamic-scene method can place it on the reference-frame spectrum and immediately read off the expected trade-off between temporal detail and computational cost.","Hybrid representations that layer a structured coarse motion - rigid or articulated - over a neural residual field are presented as the most effective pattern for real scenes with mixed motion, because they keep interpretability while capturing fine deformation.","The shift from implicit MLP radiance fields to explicit Gaussian primitives is framed as the enabler of real-time rendering, dense long-term tracking, and part-level editing, at the price of higher memory use.","Auxiliary information such as depth, segmentation, optical flow, and data-driven priors is treated as a load-bearing component of monocular reconstruction, without which the motion and appearance ambiguities are underdetermined.","The open challenges the survey names - editing, scalability to long videos and large spaces, reconstruction by generation, and LLM-guided semantic priors - are the directions where it expects the field's next advances."],"supporting_citations":[{"why":"Supplies the base representation: NeRF, the implicit radiance field that all dynamic extensions and the survey's canonical-space and 4D spacetime categories build on.","marker":"[1]"},{"why":"Supplies the explicit alternative: 3D Gaussian Splatting, the primitive representation whose real-time splatting defines the second half of the survey's implicit-to-explicit trajectory.","marker":"[2]"},{"why":"Establishes the canonical-space plus deformation-field approach for non-rigid scenes, one of the survey's five motion representation paradigms.","marker":"[9]"},{"why":"Defines the time-conditioned deformation approach for monocular dynamic scenes, a defining example of the canonical-space row in the taxonomy tables.","marker":"[8]"},{"why":"Introduces bidirectional neural scene flow fields, the load-bearing example of the frame-to-frame flow representation paradigm.","marker":"[13]"},{"why":"Provides the point-tracking paradigm: dense long-term 3D trajectories, the finest granularity on the survey's reference-frame spectrum.","marker":"[18]"},{"why":"Supplies the factorization paradigm via six hyperplanes; the survey treats it as the template for efficient 4D feature-grid methods.","marker":"[15]"},{"why":"Provides the SMPL parametric body model, the template prior that anchors the survey's articulated-motion and human-avatar categories.","marker":"[33]"},{"why":"Represents time-dependent deformation of 3D Gaussian primitives, the direct Gaussian counterpart to NeRF-style deformation fields.","marker":"[12]"},{"why":"Establishes canonical-space reconstruction of articulated humans with neural skinning, a foundational example of the articulated-motion category.","marker":"[30]"}],"fun_headline_variants":["All dynamic-scene methods: static reference plus a motion model","200+ papers: every 4D method is a static base plus motion","From implicit MLPs to Gaussian splats: survey of dynamic scenes","Reference-frame spectrum: how dynamic scenes are modeled in 200+ works","Unified view: static reference + motion model covers all 4D reconstruction"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The survey's map of the field depends on its corpus of over 200 papers being complete and representative and on its taxonomy being applied consistently, but it reports no auditable search protocol or inclusion criteria.","fun_headline_variants_meta":{"raw":{"variants":["All dynamic-scene methods: static reference plus a motion model","200+ papers: every 4D method is a static base plus motion","From implicit MLPs to Gaussian splats: survey of dynamic scenes","Reference-frame spectrum: how dynamic scenes are modeled in 200+ works","Unified view: static reference + motion model covers all 4D reconstruction"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000768,"raw_usage":{"total_tokens":3412,"prompt_tokens":962,"completion_tokens":2450,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":578,"completion_tokens_details":{"reasoning_tokens":2355}},"tokens_in":578,"tokens_out":2450,"duration_ms":17999,"temperature":1.0,"reasoning_tokens":2355,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T21:16:40.309964+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a random sample of roughly twenty recent dynamic-scene papers and check each against the taxonomy: if a well-known method resists placement in one of the four motion types and one of the five representation paradigms, or if a table entry contradicts what the method's own paper claims, the unified framework would be incomplete rather than universal.","supporting_citations":[],"review_version":1}