GMOS grounds moving object segmentation in 3D space and time from RGB video, introduces the GMOS-2K dataset and MOS-I protocol, and reports state-of-the-art results on MOS and unsupervised VOS benchmarks with faster runtime.
arXiv preprint arXiv:2412.19584 (2024)
9 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
fields
cs.CV 9roles
background 2polarities
background 2representative citing papers
A two-stage diversity-plus-entropy token selection framework speeds up visual geometry transformers by over 85% on 500-image scenes while preserving baseline accuracy.
The paper proposes a problem-driven taxonomy for feed-forward 3D scene modeling that groups methods by five core challenges: feature enhancement, geometry awareness, model efficiency, augmentation strategies, and temporal-aware modeling.
The Robust 4D Visual Geometry Transformer with Uncertainty-Aware Priors outperforms prior methods on dynamic benchmarks by cutting Mean Accuracy error 13.43% and raising segmentation F-measure 10.49% via three uncertainty mechanisms while keeping feed-forward speed.
GA-GS uses motion segmentation, diffusion-based inpainting for pseudo-ground-truth, and per-Gaussian authenticity scalars to achieve SOTA static scene reconstruction from videos with dynamic occlusions.
Ground4D reconstructs dynamic 4D scenes from monocular video by initializing dynamic Gaussians from VGGT geometry and refining them with multi-view depth consistency at observed and virtual viewpoints.
A training-free two-pass adaptation of VGGT, with attention-based motion masking and inverse-variance depth fusion, improves dynamic-scene point-cloud reconstruction on DyCheck.
citing papers explorer
-
GMOS: Grounding Moving Object Segmentation in 3D Space and Time
GMOS grounds moving object segmentation in 3D space and time from RGB video, introduces the GMOS-2K dataset and MOS-I protocol, and reports state-of-the-art results on MOS and unsupervised VOS benchmarks with faster runtime.
-
Good Token Hunting: A Hitchhiker's Guide to Token Selection for Visual Geometry Transformers
A two-stage diversity-plus-entropy token selection framework speeds up visual geometry transformers by over 85% on 500-image scenes while preserving baseline accuracy.
-
Feed-Forward 3D Scene Modeling: A Problem-Driven Perspective
The paper proposes a problem-driven taxonomy for feed-forward 3D scene modeling that groups methods by five core challenges: feature enhancement, geometry awareness, model efficiency, augmentation strategies, and temporal-aware modeling.
-
Robust 4D Visual Geometry Transformer with Uncertainty-Aware Priors
The Robust 4D Visual Geometry Transformer with Uncertainty-Aware Priors outperforms prior methods on dynamic benchmarks by cutting Mean Accuracy error 13.43% and raising segmentation F-measure 10.49% via three uncertainty mechanisms while keeping feed-forward speed.
-
GA-GS: Generation-Assisted Gaussian Splatting for Static Scene Reconstruction
GA-GS uses motion segmentation, diffusion-based inpainting for pseudo-ground-truth, and per-Gaussian authenticity scalars to achieve SOTA static scene reconstruction from videos with dynamic occlusions.
-
Ground4D: Consistency-Aware 4D Reconstruction from Monocular Video
Ground4D reconstructs dynamic 4D scenes from monocular video by initializing dynamic Gaussians from VGGT geometry and refining them with multi-view depth consistency at observed and virtual viewpoints.
-
4DVGGT-D: 4D Visual Geometry Transformer with Improved Dynamic Depth Estimation
A training-free two-pass adaptation of VGGT, with attention-based motion masking and inverse-variance depth fusion, improves dynamic-scene point-cloud reconstruction on DyCheck.
- GEM-4D: Geometry-Enhanced Video World Models for Robot Manipulation
- PAGE-4D: Disentangled pose and geometry estimation for vggt-4d perception