{"id":"07a08c40-6043-4132-985d-730131547c3d","arxiv_id":"2607.18539","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"AniGS animates a static 3D Gaussian Splatting scene by iteratively distilling video-diffusion motion into a time-conditioned deformation field while keeping static regions fixed.","lead":"AniGS adds living motion — swaying leaves, moving grass — to still 3D reconstructions of large outdoor scenes, so viewers can walk around a world that feels alive. It uses a pretrained video generator to teach a 3D Gaussian Splatting scene how to move, iteratively expanding across camera viewpoints.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Quantitative claims rely on FVD against diffusion-generated references; FVD-LTX is circular, and Cosmos/DynamiCrafter are also generative, so 'natural dynamics' lacks a real-world benchmark.","rationale":"The reader's weakest assumption—that FVD against diffusion-generated references is a valid proxy for real motion realism—is also the single most load-bearing concern I find. The method itself is internally coherent: a canonical 3DGS plus DCT-based deformation field, supervised by diffusion-refined pseudo-targets, with ablations supporting the design choices. However, the quantitative evaluation does not demonstrate proximity to real ambient dynamics; it demonstrates proximity to video-diffusion-model outputs, one of which is the training teacher. The circularity of FVD-LTX is direct, and using two other generators (Cosmos, DynamiCrafter) only shows generalization across generators, not across the simulation-to-reality gap. A real-video FVD test or a user study against real captures would settle whether the claimed 'natural' motion holds. Since this concern is already the basis of the reader's CONDITIONAL verdict, my stress-test does not change the verdict recommendation.","tokens_in":15181,"tokens_out":7265,"duration_ms":84315,"concrete_test":"Recompute the FVD evaluation using a held-out set of real outdoor videos containing natural ambient motion (e.g., tripod-captured foliage, trees, and grass under wind) as the reference distribution, with the same I3D backbone and 256x256 preprocessing, for AniGS, Gaussians2Life, and PhysDreamer. If the ranking changes or AniGS no longer significantly outperforms the baselines on real-video FVD, the diffusion-generated FVD proxy is invalid and the quantitative claim of 'natural dynamics' collapses.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central quantitative evidence for 'natural ambient dynamics' rests on Fréchet Video Distance computed against reference videos generated by pretrained video diffusion models (Sec. 4.3, Tables 1–2). This is a load-bearing evaluation weakness. Since the pseudo-training targets in Secs. 3.4–3.5 are produced by LTX (Sec. 3.6), FVD-LTX is a self-comparison to the method's own teacher: a low score may only indicate faithful mimicry of LTX's motion distribution, not real-world realism. The two additional references, Cosmos and DynamiCrafter, are also video generators, so they do not ground the metric in actual scene dynamics. The paper explicitly acknowledges that ground-truth dynamic videos are unavailable (Sec. 4.3), but this remains the linchpin: without a real-video reference, the FVD numbers cannot substantiate the abstract's claim of 'natural' motion. The user study provides some support, but its reported percentages lack participant counts and significance tests, and it compares against other generative baselines, not against real captures. This is a correctness risk in the evaluation, not an internal inconsistency.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes AniGS, a method for animating static 3D Gaussian Splatting reconstructions of large, cluttered outdoor scenes. The method combines a canonical 3DGS with a time-conditioned deformation field parameterized by DCT coefficients, and uses a pretrained flow-based video diffusion model (LTX) as a motion prior. Training is bootstrapped through an iterative dataset–model update: a dataset of camera-fixed rendered clips is progressively expanded and refined by a composed video-to-video (ComposedV2V) scheme that keeps static regions fixed, and the model is optimized against these pseudo-targets with a photometric loss plus canonical regularization and an SDS loss. Experiments on five collected scenes and DL3DV scenes compare against Gaussians2Life and PhysDreamer using FVD against diffusion-generated references, a multiview-consistency PSNR check, and a 2AFC user study. The authors claim natural ambient dynamics and high-quality novel-view videos at scene scale without ground-truth dynamic supervision.","tokens_in":15528,"tokens_out":2837,"duration_ms":32907,"significance":"If validated, AniGS addresses a real gap: adding plausible ambient motion to large, navigable 3D scene reconstructions without synchronized multi-camera dynamic capture. The proposed iterative viewpoint expansion and ComposedV2V static-preservation mechanism are sensible and potentially useful to the 4D-scene community. The multiview-consistency PSNR experiment (Table 3) is a good positive check, and the qualitative space-time slices illustrate the intended behavior. However, the quantitative evaluation currently relies on Fréchet Video Distance computed against reference videos generated by the same or similar video diffusion models. Since LTX is both the teacher that produces the pseudo-training targets (Sec. 3.4/3.5) and the generator of the reference distribution for FVD-LTX (Sec. 4.3), the primary numerical evidence is circular and does not establish that the synthesized motion is natural or realistic relative to real-world dynamics. This is a load-bearing weakness that requires a non-circular evaluation before the central claim can be accepted.","major_comments":[{"comment":"FVD-LTX in Tables 1, 2, and 5 is a self-comparison. The pseudo-training targets are produced by LTX via ComposedV2V (Eqs. 9–12 and Sec. 3.6), and the reference distribution for FVD-LTX is also generated by LTX (Sec. 4.3). A low score therefore measures how well the model imitates its own teacher distribution, not how natural the motion is with respect to real scenes. Because the abstract's central claim of 'natural ambient dynamics' is supported primarily by these FVD numbers, an independent reference—for example real captured dynamic video, a non-teacher video model, or a metric grounded in human judgment—is needed. The ablations in Table 5 inherit this issue and cannot validate the design choices.","section":"Sec. 4.3; Sec. 3.4/3.6"},{"comment":"FVD-Cosmos and FVD-DynamiCrafter are also computed against reference videos generated by pretrained video diffusion models. While these models are not the teacher for AniGS, they are still generative priors rather than ground-truth observations of the actual scenes. The paper acknowledges that ground-truth scene-level dynamic videos are unavailable (Sec. 4.3), but that does not justify treating diffusion-generated distributions as a gold standard for 'realistic motion'. The metric can reward matching a generic generator's motion prior while failing to capture the true dynamics of the scene. A concrete test would be to evaluate on a real dynamic outdoor capture (even with a single moving camera or limited multi-camera data) or to report human agreement against real footage.","section":"Sec. 4.3, Tables 1–2"},{"comment":"The user study reports preference percentages (e.g., 91.7% in Fireplace) but does not report the number of participants, the number of comparison pairs, the variance across participants, or any statistical significance test. Without this information, the percentages cannot be distinguished from noise, especially when the numbers are close for some scenes (e.g., 74.5% in Trail). Please report sample sizes, confidence intervals, and a significance test (e.g., paired bootstrap or Wilcoxon).","section":"Sec. 4.5, Table 4"}],"minor_comments":[{"comment":"Typo: 'modeling the the synthesized motion' should read 'modeling the synthesized motion'.","section":"Sec. 1"},{"comment":"Hyperparameters such as DCT basis size K, loss weights λ_cano and λ_SDS, and the SDS clip-length parameter m are not specified. Please provide these values for reproducibility.","section":"Sec. 3.6"},{"comment":"The 'Original capture' PSNR row includes DL3DV scenes (Gazebo, Courtyard) but the caption does not clarify whether these are from the same capture protocol as the collected scenes. Please specify the source and processing for each scene.","section":"Sec. 4.1, Table 3"},{"comment":"The baseline name is inconsistently spelled as 'Gaussians2Life' and 'Gaussians2life' in Tables 1 and 4 and in the text. Please unify.","section":"General"},{"comment":"Reference [38] contains an editorial note ('Use this entry for both ... bullets'), which should be removed.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The core technical idea is plausible and the qualitative results are encouraging, but the quantitative evaluation needs to be rethought. If the authors can provide a non-circular metric or at least a rigorous user study with real-video grounding, the contribution would be much stronger. I would not reject on the current evidence, but the manuscript is not yet ready in its present form."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: AniGS is a real systems contribution. It targets scene-level animation of static 3DGS reconstructions — large, cluttered, walkable scenes — which is distinct from the object-centric 4D generation work. The iterative dataset–model update with incremental viewpoint expansion and the composed video-to-video refinement are sensible engineering choices, and the qualitative results show natural ambient motion on foliage while keeping structures stable. The ablations support the design decisions. The multiview consistency check (training a 3DGS at a novel timestep and measuring PSNR) is a good idea and gives the central claim some non-circular support.\n\nThe soft spot is the quantitative evaluation. FVD-LTX is nearly circular: LTX generates the pseudo-training targets in Sec 3.4–3.5, and LTX also generates the reference distribution in Sec 4.3. A low FVD-LTX mostly says the model mimics its own teacher. FVD-Cosmos and FVD-DynamiCrafter are less circular but still measure distance to a generative prior, not to real scene dynamics. The paper acknowledges that ground-truth dynamic videos don't exist, so this is an honest limitation rather than a hidden flaw — but it means the abstract's claim of 'natural' motion rests on qualitative videos and a user study with no participant counts or significance tests. That's a moderate weakness, not a fatal one. The core idea does not depend on the FVD numbers; the qualitative evidence and the multiview PSNR check are more convincing.\n\nMinor points: no code or data release, and the free parameters (DCT basis size, loss weights, number of cameras, SDS clip length) are standard but not sensitivity-analyzed. The limitation section is candid about failure on large non-reversible dynamics.\n\nWho this is for: anyone working on animating 3D reconstructions, VR/AR immersive content, or leveraging video diffusion priors for 3D. It deserves a serious referee; the evaluation should be strengthened but the contribution is real. I'd send it to review.","headline":"Genuine systems contribution for scene-level 3DGS animation, but the headline FVD metric is partly circular; the qualitative and multiview-consistency evidence carries the load.","tokens_in":16003,"tokens_out":1451,"would_cite":true,"duration_ms":16415,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"AniGS animates a static 3D scene reconstruction into a free-viewpoint video with natural ambient motion, while keeping rigid structures still.","keywords":["scene animation","3D Gaussian splatting","video diffusion model","ambient dynamics","deformation field","novel view synthesis","dataset-model update","free-viewpoint video"],"falsifier":"A concrete test: record or obtain real dynamic video of one of the same outdoor scenes (real wind-driven foliage) with a synchronized multi-camera rig, render AniGS's animation from matching viewpoints, and run a two-alternative forced choice comparing AniGS's motion against the real captured motion, or compute FVD against this independent reference. If AniGS is not preferred or its FVD is not lower than a static baseline, the claim of natural ambient dynamics would be falsified.","tokens_in":15098,"feed_emoji":"🌿","tokens_out":3052,"duration_ms":31583,"temperature":0.7,"pith_summary":"The paper tries to establish that a static 3D Gaussian Splatting reconstruction of a large outdoor scene can be turned into a dynamic, free-viewpoint video with plausible ambient motion (vegetation swaying, leaves moving) without any ground-truth dynamic supervision. It does so by coupling the static representation with a time-conditioned deformation field and a pretrained video diffusion model that supplies motion cues, using an iterative dataset–model update loop to bootstrap temporally consistent multi-view supervision across the whole scene. The central claim is that this yields natural, scene-wide dynamics while rigid structures remain stable, supported by FVD scores, a multiview-consistency reconstruction check, and a user study on five real-world scenes.","feed_headline":"Static 3D reconstructions become animated scenes with natural sway","feed_subtitle":"A diffusion prior plus iterative refinement adds wind-like motion to large outdoor reconstructions while keeping buildings still.","key_machinery":"The time-conditioned deformation field: each Gaussian's position and rotation are offset by a discrete cosine transform (DCT) basis expansion of time, which is storage-efficient and smooth. The iterative dataset–model update loop with incremental viewpoint expansion and composed video-to-video refinement (ComposedV2V) is the mechanism that generates multi-view consistent pseudo-targets; ComposedV2V pastes the static region from the first frame into all frames to prevent drift in static areas.","core_discovery":"The central discovery is that scene-level animation of 3DGS reconstructions can be bootstrapped from a static canonical model plus a video diffusion prior, via an iterative process: render camera-fixed clips from the current model, refine them with a composed video-to-video diffusion step that freezes static regions, add these as training targets, then optimize a DCT-parameterized deformation field over the canonical Gaussians. The key insight is that by iterating dataset updates and model updates while expanding viewpoints, the method avoids the need for synchronized multi-camera dynamic capture, producing temporally coherent ambient motion across a large navigable scene.","pith_inferences":["A testable extension: applying the same loop to indoor scenes with different animatable objects (the paper shows one such result) could verify whether the method's reliance on the diffusion model's prior limits it to vegetation-like oscillatory motion.","The FVD-LTX metric is partly self-referential because LTX is also the teacher; an independent evaluation against real captured dynamic footage or a blind comparison with a held-out scene would strengthen or weaken the claim.","The method may inherit biases from the video diffusion model (e.g., a tendency for certain motion patterns or a preference for particular lighting), and the iterative updates could amplify these biases; this is an open question not addressed in the paper.","The deformation field with DCT basis limits the motion to reversible, small deformations, so the claim naturally stops short of state changes; extending to non-reversible dynamics would require a different representation."],"forward_implications":["If correct, reconstructed 3D environments from casual captures can be rendered as animated free-viewpoint videos, making VR/AR walkthroughs more immersive.","The method establishes a template for lifting video diffusion priors into 3D scene-level motion without per-scene dynamic capture, potentially extensible to other generative priors.","The iterative dataset–model update could serve as a general recipe for distilling 2D generative models into consistent 3D representations.","The composed refinement shows a way to constrain generative video models to move only desired regions, which may transfer to video editing tasks.","The multiview consistency of generated motion suggests the deformation field learns a true 3D motion, enabling novel-view synthesis of the animation."],"fun_headline_variants":["AniGS: Turn static 3D reconstructions into animated scenes","Diffusion prior brings natural sway to large 3D scene reconstructions","Static 3D to living scenes: AniGS uses video diffusion","AniGS: Add ambient motion to static 3D scenes"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The load-bearing assumption is that FVD computed against videos generated by pretrained video diffusion models (especially LTX, which also generates the training targets) is a valid measure of motion realism; if that proxy does not reflect genuine scene dynamics, the central quantitative evidence loses its footing.","fun_headline_variants_meta":{"raw":{"variants":["AniGS: Turn static 3D reconstructions into animated scenes","Diffusion prior brings natural sway to large 3D scene reconstructions","Static 3D to living scenes: AniGS uses video diffusion","AniGS: Add ambient motion to static 3D scenes"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001308,"raw_usage":{"total_tokens":5162,"prompt_tokens":730,"completion_tokens":4432,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":474,"completion_tokens_details":{"reasoning_tokens":4354}},"tokens_in":474,"tokens_out":4432,"duration_ms":34261,"temperature":1.0,"reasoning_tokens":4354,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T15:05:29.545732+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A concrete test: record or obtain real dynamic video of one of the same outdoor scenes (real wind-driven foliage) with a synchronized multi-camera rig, render AniGS's animation from matching viewpoints, and run a two-alternative forced choice comparing AniGS's motion against the real captured motion, or compute FVD against this independent reference. If AniGS is not preferred or its FVD is not lower than a static baseline, the claim of natural ambient dynamics would be falsified.","supporting_citations":[],"review_version":1}