{"id":"37492c55-b946-4973-abac-46a401176ca5","arxiv_id":"2509.07653","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":8,"one_line_summary":"TaoGS introduces a two-layer Gaussian representation with lifespan-aware 2D lookup-table compression for tracking and rendering human performances through topological changes at up to 40x compression.","lead":"TaoGS is a new dynamic 3D Gaussian method for volumetric human videos that keeps tracking people through topological changes such as removing clothes or drawing a sword. It also packs the Gaussians into a lifespan-aware 2D layout so standard video codecs can shrink the stream by up to 40x.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Compression operating point in Table 1 is unreproducible from Table 5: claimed 36.697 PSNR at 1.335 MB/frame is better than the best QP sweep point (36.507 PSNR at 2.47 MB/frame), so the headline 40x claim lacks a specified codec setting.","rationale":"The paper is a credible systems contribution: the motion/appearance decomposition, CoTracker/RoMa-guided densification, GLUT, and Morton/lifespan sorting are coherent, and the ablations (Tables 3-5) support the design choices qualitatively. I am not objecting to the method's conceptual validity or to using a private dataset, which is common in performance capture. The load-bearing problem is that the single most prominent quantitative claim—after-compression quality at 1.335 MB/frame—does not match any operating point in the authors' own QP sweep. This is an internal inconsistency, not a disagreement with field consensus, and it directly affects the strongest claim quoted from Section 4.1. The reader's chosen weakest assumption (tracker/segmentation failure in textureless regions) is acknowledged by the authors and would degrade specific scenarios; it does not undermine the reported numbers themselves. My concern is more fundamental: without the codec operating point, the headline number is unverifiable even from the paper's own tables. I therefore recommend CONDITIONAL rather than REJECT, because a clarifying supplement (exact QP/encoder settings, per-sequence rate-distortion table, or corrected Tables 1/2) could plausibly resolve the inconsistency. The condition should be that the after-compression row is reproduced on the Table 5 curve. I agree partially with the reader: we share the compression inconsistency as a concern, but their primary 'weakest assumption' is the tracker dependency.","tokens_in":17182,"tokens_out":4706,"duration_ms":38781,"concrete_test":"Request the exact compression configuration for Table 1's 'After Compression' row (codec, preset, GOP length, fixed QP or rate-control target, and whether QP varies per frame/sequence). Then rerun the Table 5 QP sweep under that same configuration on the same three 200-frame sequences and plot (PSNR, storage) for QP=5,15,25 plus the claimed point. If the claimed point does not lie on or within measurement error of the resulting rate-distortion curve, the after-compression improvement is unsupported. Alternatively, decode the QP5 stream (2529.28 KB, 36.507 PSNR) and the QP15 stream (1351.68 KB, 36.059 PSNR): neither reaches 36.697 PSNR, so any claim exceeding both requires a stated per-sequence bit allocation.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The strongest claim—40x compression at 36.697 PSNR—is not internally reproducible. Table 1 reports Ours(After Compression) at 36.697 PSNR/0.9881 SSIM/0.0214 LPIPS, and Table 2 reports 1.335 MB/frame. Table 5 gives the only compression operating-point data: combined sorting at QP5 yields PSNR 36.507 with 2529.28 KB/frame (2.47 MB), QP15 yields 36.059 with 1351.68 KB (1.32 MB), QP25 yields 31.837 with 783.36 KB. No QP row simultaneously achieves the Table 1 PSNR (36.697) and storage (1.335 MB); the claimed point lies 0.19 dB above QP5 while using half its storage. Unless a per-sequence adaptive QP, a different codec preset, or a different evaluation subset is being reported, the numbers cannot come from the same pipeline. Because the after-compression row is the main evidence for the headline 'up to 40x' claim, this inconsistency is load-bearing. The omitted compression configuration (codec, preset, GOP, QP/rate control) is not a minor detail; it is the definition of the reported operating point. A reader cannot verify whether the method actually beats V3/HiFi4G at matched bitrates. This concern is independent of the acknowledged tracker limitations, which are honest but secondary.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents TaoGS, a dynamic 3D Gaussian representation for human-centric volumetric video that separates motion and appearance. Sparse motion Gaussians are tracked under an ARAP regularizer and extended by new candidate Gaussians that are inserted based on a spatio-temporal tracker and photometric filtering, enabling topological changes. Each motion Gaussian anchors a set of appearance Gaussians activated for fine texture, and a Global Gaussian Lookup Table packs attributes into a lifespan-aware 2D layout for video-codec compression. Experiments on a private 81-camera dataset with three 200-frame sequences report higher PSNR/SSIM/LPIPS than five baselines, faster training than several of them, and roughly 40x compression after encoding.","tokens_in":17552,"tokens_out":4745,"duration_ms":41416,"significance":"If the claims are reproducible, the paper makes a solid practical advance: it directly targets topological changes in dynamic scenes, a known weakness of tracking-based Gaussian methods, and it integrates compression-friendly layout into the representation rather than treating compression as a post-hoc add-on. The ablations cover candidate filtering, initialization strategies, sorting schemes, and QP values, and the limitations section is unusually honest about dependence on pretrained trackers and masks. However, the evaluation rests entirely on a private dataset, all numbers are single runs without error bars, and the main after-compression operating point is internally inconsistent with the reported QP sweep, so the headline claims are currently conditional on unresolved details.","major_comments":[{"comment":"The 'Ours(After Compression)' row of Table 1 (PSNR 36.697, SSIM 0.9881, LPIPS 0.0214) together with the 1.335 MB/frame storage in Table 2 cannot be reproduced from the QP sweep in Table 5. For combined sorting, Table 5 reports QP5: PSNR 36.507 at 2529.28 KB, QP15: PSNR 36.059 at 1351.68 KB, and QP25: PSNR 31.837 at 783.36 KB. No row simultaneously gives a PSNR of 36.697 and a storage of 1.335 MB; the closest QP5 is 0.19 dB lower while using about 1.9x the storage, and QP15 has similar storage but 0.64 dB lower PSNR. Please specify the exact compression configuration (codec, preset, GOP structure, per-sequence QP or rate control, and the frames/sequences used) that produces the Table 1 after-compression numbers, or align Table 1 with an existing row of Table 5. Without this, the headline 40x compression claim and the storage comparison against V3 and HiFi4G are not verifiable.","section":"Tables 1, 2, and 5"},{"comment":"All quantitative claims are based on a private dataset with no error bars, significance tests, code release, or data release. The numbers appear to be single-run averages over three 200-frame sequences; the method has many tuned hyperparameters (epsilon=0.46, lambda_smooth values, K=9, densification schedules), and the reported differences of roughly 0.2-2 dB against HiFi4G/V3 could be within run-to-run and sequence-to-sequence variation. Please provide per-sequence results, repeated-run statistics, or a public release of data/code so that the reported improvements can be independently confirmed.","section":"Section 4.1, Tables 1-5"},{"comment":"The paper states that 'we benchmark all methods on three sequences of 200 frames', but the dataset overview in Figure 6 appears to show more sequences and the gallery figures show multiple different scenes. It is unclear whether Tables 1-5 use the same three sequences as the qualitative comparisons, and whether the numbers are per-sequence or pooled. State the exact evaluation subset and, if different sequences were used for qualitative illustrations, clarify this explicitly.","section":"Section 4.1, Figure 6"}],"minor_comments":[{"comment":"The text says the photometric error is used to 'identify poorly reconstructed regions' but then discards candidates when the error w2 falls below epsilon. Low color error means the region is already well reconstructed and is likely an occlusion false-positive, so the behavior is correct, but the wording is confusing; please rephrase to say that low photometric error flags already-reconstructed regions that are then discarded.","section":"Section 3.1, Eq. (7)"},{"comment":"In the appearance Gaussian initialization ablation, '9x (random)' has LPIPS 0.0154, which is slightly lower (better) than '9x (edge-based)' at 0.0156. The text concludes that edge-based 9x achieves a good trade-off, but this LPIPS result mildly favors random; please acknowledge and explain the discrepancy or correct the conclusion.","section":"Table 4"},{"comment":"The abstract says 'supports up to 40 compression' with a missing multiplication sign; it should read '40x compression'.","section":"Abstract"},{"comment":"The caption states 'Our method achieves the highest rendering quality' before presenting the quantitative results; consider toning this to 'Our method produces the highest rendering quality in this qualitative comparison'.","section":"Figure 4 caption"},{"comment":"The integration of 'Taming 3DGS' is mentioned only in the experiments paragraph; please describe in the method section how this technique is adapted to the motion-appearance framework, since it affects training time and convergence behavior.","section":"Section 4"}],"recommendation":"major_revision","confidential_remarks":"The paper is within the scope of the journal and the method is plausible, but the irreproducible compression operating point is a serious issue because it sits at the center of the paper's practical claim. I would also urge the editor to consider whether the private dataset and lack of release are acceptable for a methods paper of this type; at minimum, per-sequence numbers and error bars are needed."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: TaoGS is a genuinely new dynamic-Gaussian system with solid engineering and strong reported pre-compression results. The headline 40x compression number, however, is internally inconsistent with the paper's own QP sweep. That unresolved operating point is the main blocker, not the method.\n\nWhat's new: sparse motion Gaussians with tracker-driven candidate insertion for topology changes, plus dense appearance Gaussians anchored per motion Gaussian and a lifespan-aware global lookup table that packs attributes into 2D streams for off-the-shelf video codecs. That combination is not in the cited literature. The ablations on initialization, candidate filtering, and sorting strategy are reasonable, and the reported pre-compression quality (37.264 PSNR) beats the five baselines on the authors' 81-camera capture. Training at 2 min/frame with real-time rendering is a real achievement.\n\nThe problem is Table 1 vs Table 5. After Compression is listed as 36.697 PSNR at 1.335 MB/frame. The QP sweep's best point (QP5) hits 36.507 at 2.47 MB; QP15 sits at 36.059 at 1.32 MB. No listed QP reaches 36.697 at that storage. The claimed point is half the storage of QP5 while being 0.19 dB better. Unless they ran per-sequence adaptive QP or a different codec preset/GOP, these numbers can't come from the same pipeline. Since the after-compression row is the evidence for 'up to 40x', this is load-bearing. The right fix is straightforward: specify codec, preset, rate control, and include a rate-distortion curve with the reported point on it.\n\nSecondary soft spots: all metrics are on a private dataset with no error bars or significance tests, no code/data release, and heavy reuse of the authors' own DualGS/HiFi4G/V3 regularizers. The limitations section honestly admits dependence on CoTracker and 2D masks and the indoor studio assumption, which tempers generalization claims. The stress-test's tracker concern is real but secondary; the compression mismatch is the one that matters.\n\nThis paper deserves peer review, not desk rejection. It's a serious contribution that needs a scrutiny pass on the numbers. I'd bring it to reading group for that discussion. Recommendation: send it to a serious referee, but the referee must require the compression operating point to be explained and ideally code/data release. As it stands, I'd cite the representation and the compression layout idea, but not the 40x number.","headline":"TaoGS is a genuinely new dynamic-Gaussian system with strong pre-compression results, but its headline 40x compression figure is internally inconsistent with its own QP sweep — that's the main blocker.","tokens_in":18141,"tokens_out":3748,"would_cite":true,"duration_ms":31130,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A dynamic Gaussian representation that handles topological changes and compresses volumetric video to a 40x smaller stream.","keywords":["volumetric video","3D Gaussian splatting","topology change","dynamic scene reconstruction","human performance capture","motion-appearance disentanglement","video codec compression","real-time rendering"],"falsifier":"A decisive test would be to annotate the ground-truth moments of surface emergence and disappearance in a held-out sequence, such as a person removing a plain, textureless shirt, and compare the Gaussian insertion locations against that annotation; if the tracker-driven insertion misses the newly exposed torso because the mask or photometric error is wrong, the rendered sequence will show ghosting or blur on exactly those frames, and PSNR on that segment should fall below the keyframe-based baselines.","tokens_in":16934,"feed_emoji":"🎥","tokens_out":6055,"duration_ms":51945,"temperature":0.7,"pith_summary":"TaoGS proposes a dynamic 3D Gaussian representation that treats scene motion and surface texture as two separate layers, so a video of a person changing clothes, drawing a sword, or interacting with objects can be rendered in real time while the set of Gaussians grows and shrinks as new surfaces appear and old ones vanish. The paper's central claim is that by anchoring dense appearance Gaussians to a sparse set of tracked motion Gaussians, and by recording each motion Gaussian's lifespan in a global lookup table, one can handle topological changes without the keyframe resets or strong rigidity constraints that limit prior dynamic-Gaussian methods. If the claim holds, volumetric video capture moves from short, topology-fixed clips to long sequences with genuine topological variation, at a storage cost low enough for standard video codecs: the paper reports 37.264 PSNR before compression, 36.697 after, and 1.335 MB per frame after a 40x compression on its own dataset.","feed_headline":"Changing-topology humans rendered in real time at 40x compression","feed_subtitle":"Sparse motion Gaussians anchor appearance details, so cloth removal and sword draws keep tracking without keyframe resets.","key_machinery":"The load-bearing mechanism is the motion-to-appearance Gaussian hierarchy together with the global Gaussian lookup table (GLUT). Motion Gaussians are sparse (around 20,000) and connected into a deformation graph; they are optimized with an as-rigid-as-possible regularizer, and new candidates are inserted only where a temporal tracker marks pixels as newly observed and the photometric error at the triangulated pixel is above a threshold. Appearance Gaussians are created along graph edges, nine per anchor by default, inherit the anchor's attributes, and are warped by it each frame, which lets the appearance layer converge quickly. The GLUT assigns every motion Gaussian a unique index, records its lifespan, and orders the associated appearance Gaussians so that persistent ones appear in a spatially coherent 2D layout while transient ones line up by activation time; that ordering is what lets standard video codecs remove redundancy and reach the reported compression.","core_discovery":"The paper's central discovery is that topological adaptation and long-range tracking can coexist in a Gaussian representation if motion and appearance are decoupled. A sparse set of motion Gaussians carries the deformation field, regularized by an as-rigid-as-possible constraint, and is extended on the fly: a pretrained temporal tracker identifies pixels with no correspondence to the previous frame, and a photometric-error filter removes candidates in occluded regions, so newly emerged surfaces trigger insertion of new motion Gaussians while stable regions keep their track. Each motion Gaussian then anchors a fixed set of appearance Gaussians sampled along its nearest-neighbor edges; the appearances are non-rigidly warped by the anchor, giving a strong initialization that cuts training to about two minutes per frame on a single GPU. A global lookup table records each Gaussian's lifespan and sorts persistent Gaussians by a spatial bit-interleaving order and transient Gaussians by activation time, which packs attributes into a 2D stream that standard video codecs compress efficiently. On the authors' 81-camera dataset, TaoGS reports higher PSNR, SSIM, and lower LPIPS than the compared dynamic-Gaussian methods, with real-time rendering and a 40x compression ratio.","pith_inferences":["Because the candidate-insertion filter relies on tracker masks and photometric error, the same decoupling could support an uncertainty-aware variant that weights the photometric threshold by tracker confidence, which would likely reduce artifacts in textureless regions the paper concedes.","A natural testable extension is to let deactivated motion Gaussians reactivate when their region reappears, which the paper lists as a limitation; semantic or feature-based reuse could cut redundancy in long sequences.","The approach is demonstrated on indoor multi-view studio capture, so applying the motion-to-appearance design to monocular or outdoor footage would require coupling with inverse rendering to handle lighting variation; the paper's own limitations discussion points in that direction.","The GLUT ordering principle could transfer to other primitive-based dynamic representations, such as tetrahedral or mesh-based ones, because it only needs a per-primitive lifespan and a spatial ordering key rather than Gaussian-specific attributes."],"forward_implications":["Dynamic scenes with true topological change, such as a jacket coming off, are modeled without splitting the sequence into keyframe segments, removing a costly step in earlier keyframe-based pipelines.","Because attributes are packed into a lifespan-aware 2D stream, the same compressed representation can be decoded with standard video codecs, making mobile and VR playback of such scenes feasible.","The reported training time of about two minutes per frame on one RTX 4090 brings per-sequence training into a practical range for production capture.","Rendering runs at 125 FPS before compression and 65 FPS after, so the fidelity gains do not come at the cost of interactivity.","The edge-based derivation of appearance Gaussians gives a controlled quality-versus-count trade-off: the paper's 9x scheme matches the visual quality of a 16x scheme with far fewer Gaussians."],"supporting_citations":[{"why":"Provides the base 3D Gaussian splatting rendering and optimization pipeline on which TaoGS is built.","marker":"[Kerbl et al. 2023]"},{"why":"Supplies the temporal tracker that outputs the dense warp field and occlusion mask used to detect newly observed regions.","marker":"[Karaev et al. 2024]"},{"why":"Supplies the dense matcher used to triangulate initial point clouds and to generate spatial candidate Gaussians.","marker":"[Edstedt et al. 2024]"},{"why":"Provides the dense-correspondence initialization strategy that avoids SfM and mesh-based initialization.","marker":"[Kotovenko et al. 2025]"},{"why":"Introduces the as-rigid-as-possible regularization that the paper adapts to constrain motion Gaussians.","marker":"[Luiten et al. 2024]"},{"why":"Contributes the isotropic and size losses used during initialization and serves as a keyframe-free baseline comparison.","marker":"[Jiang et al. 2024a]"},{"why":"Supplies the temporal regularizer applied to appearance Gaussian attributes and serves as a keyframe-based baseline.","marker":"[Jiang et al. 2024b]"},{"why":"Demonstrates packing dynamic Gaussians into 2D streams for codec compression, a comparison baseline and a compression reference.","marker":"[Wang et al. 2024b]"},{"why":"Provides the Morton ordering used to map persistent Gaussian positions to a spatially consistent 2D layout.","marker":"[Morton 1966]"}],"fun_headline_variants":["Sparse motion Gaussians track topology changes for real-time volumetric video","TaoGS: decouple motion and appearance for compressible human 3D video","Topology-aware dynamic Gaussians with 40x compression and real-time rendering","Handling cloth changes in Gaussian splatting for volumetric human capture","Fast topology adaptation for dynamic Gaussian rendering with long-term tracking"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the pretrained temporal tracker and dense matcher return correct occlusion masks and pixel correspondences wherever new surfaces appear, because new motion Gaussians are inserted only where those signals fire; in textureless or thin-structure regions the tracker can err, and the paper concedes this can produce visible artifacts.","fun_headline_variants_meta":{"raw":{"variants":["Sparse motion Gaussians track topology changes for real-time volumetric video","TaoGS: decouple motion and appearance for compressible human 3D video","Topology-aware dynamic Gaussians with 40x compression and real-time rendering","Handling cloth changes in Gaussian splatting for volumetric human capture","Fast topology adaptation for dynamic Gaussian rendering with long-term tracking"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.0019,"raw_usage":{"total_tokens":7506,"prompt_tokens":1063,"completion_tokens":6443,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":679,"completion_tokens_details":{"reasoning_tokens":6348}},"tokens_in":679,"tokens_out":6443,"duration_ms":36888,"temperature":1.0,"reasoning_tokens":6348,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T16:12:09.311194+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A decisive test would be to annotate the ground-truth moments of surface emergence and disappearance in a held-out sequence, such as a person removing a plain, textureless shirt, and compare the Gaussian insertion locations against that annotation; if the tracker-driven insertion misses the newly exposed torso because the mask or photometric error is wrong, the rendered sequence will show ghosting or blur on exactly those frames, and PSNR on that segment should fall below the keyframe-based baselines.","supporting_citations":[],"review_version":2}