Pith. sign in

REVIEW 32 cited by

VGGT: Visual Geometry Grounded Transformer

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2503.11651 v1 pith:55L4SHDZ submitted 2025-03-14 cs.CV

VGGT: Visual Geometry Grounded Transformer

classification cs.CV
keywords pointvggttaskscameradepthestimationfeed-forwardgeometry
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

We present VGGT, a feed-forward neural network that directly infers all key 3D attributes of a scene, including camera parameters, point maps, depth maps, and 3D point tracks, from one, a few, or hundreds of its views. This approach is a step forward in 3D computer vision, where models have typically been constrained to and specialized for single tasks. It is also simple and efficient, reconstructing images in under one second, and still outperforming alternatives that require post-processing with visual geometry optimization techniques. The network achieves state-of-the-art results in multiple 3D tasks, including camera parameter estimation, multi-view depth estimation, dense point cloud reconstruction, and 3D point tracking. We also show that using pretrained VGGT as a feature backbone significantly enhances downstream tasks, such as non-rigid point tracking and feed-forward novel view synthesis. Code and models are publicly available at https://github.com/facebookresearch/vggt.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 32 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. GenRecon: Bridging Generative Priors for Multi-View 3D Scene Reconstruction

    cs.CV 2026-05 unverdicted novelty 7.0

    GenRecon lifts object-level generative priors to scene-scale reconstruction by chunking scenes and using projection-based conditioning on multi-view features, claiming 16% better results than prior methods.

  2. No Pose, No Problem in 4D: Feed-Forward Dynamic Gaussians from Unposed Multi-View Videos

    cs.CV 2026-05 unverdicted novelty 7.0

    NoPo4D is the first feed-forward system for dynamic 4D Gaussian splatting from unposed multi-view videos, using velocity decomposition supervised by optical flow and a bidirectional motion encoder.

  3. Stream3D: Sequential Multi-View 3D Generation via Evidential Memory

    cs.CV 2026-05 unverdicted novelty 7.0

    Stream3D is a training-free method that maintains temporal consistency in 3D generation from monocular streams by dynamically caching a fixed number of informative historical frames using an evidence score.

  4. CRePE: Curved Ray Expectation Positional Encoding for Unified-Camera-Controlled Video Generation

    cs.CV 2026-05 unverdicted novelty 7.0

    CRePE supplies depth-aware positional distributions along curved rays for stable unified-camera control in frozen video DiT models.

  5. OA-WAM: Object-Addressable World Action Model for Robust Robot Manipulation

    cs.RO 2026-05 unverdicted novelty 7.0

    OA-WAM uses persistent address vectors and dynamic content vectors in object slots to enable addressable world-action prediction, improving robustness on manipulation benchmarks under scene changes.

  6. Towards Realistic and Consistent Orbital Video Generation via 3D Foundation Priors

    cs.CV 2026-04 unverdicted novelty 7.0

    A video generation approach conditions a base model with multi-scale 3D latent features and a cross-attention adapter to produce geometrically realistic and consistent orbital videos from one image.

  7. $\pi^3$: Permutation-Equivariant Visual Geometry Learning

    cs.CV 2025-07 conditional novelty 7.0

    π³ is a feed-forward network with full permutation equivariance that outputs affine-invariant poses and scale-invariant local point maps without reference frames, reaching state-of-the-art on camera pose, depth, and d...

  8. VGGT-SLAM: Dense RGB SLAM Optimized on the SL(4) Manifold

    cs.CV 2025-05 unverdicted novelty 7.0

    VGGT-SLAM aligns VGGT submaps via SL(4) manifold optimization of 15-DoF homographies to enable consistent dense RGB SLAM on long uncalibrated monocular videos.

  9. RayOcc: Occlusion-Aware Ray Occupancy Estimation via Gaussian Mixture Intensity

    cs.CV 2026-07 conditional novelty 6.0

    RayOcc models each camera ray as a non-normalized Gaussian mixture with Poisson-based occupancy probabilities, allowing multiple depth hypotheses per ray and improving Gaussian-initialized 3D occupancy prediction on nuScenes.

  10. PACE: Polar Axis-Conditioned Estimation for PairUAV Relative Localization

    cs.CV 2026-07 conditional novelty 6.0

    A shared image-pair network beats a single-head baseline by giving heading and range their own decoder readouts—PACE's raw model scores 0.002460 on the PairUAV hidden test.

  11. X-Lens: Real-Time Metric Depth Estimation with Heterogeneous Cameras

    cs.CV 2026-07 conditional novelty 6.0

    X-Lens fuses arbitrary calibrated fisheye and pinhole views into real-time metric depth at 41 FPS with a 0.04B-parameter model and a new 266K-frame synthetic dataset.

  12. X-Lens: Real-Time Metric Depth Estimation with Heterogeneous Cameras

    cs.CV 2026-07 unverdicted novelty 6.0

    A 0.04B-parameter feed-forward model estimates metric depth from variable calibrated fisheye and pinhole views using calibration tokens and Jacobian distortion bias, with a new multi-view synthetic dataset.

  13. HandsOnWorld: Unconstrained Egocentric Video Generation with Camera-Disentangled Hand Control

    cs.CV 2026-07 unverdicted novelty 6.0

    HandsOnWorld creates a hand-controlled egocentric video generator from unconstrained monocular video via a new EgoVid-Pro dataset from monocular reconstruction and a Plücker Hand Map that disentangles camera and hand motion.

  14. Error-Conditioned Neural Solvers

    cs.LG 2026-06 unverdicted novelty 6.0

    Error-Conditioned Neural Solvers improve PDE prediction accuracy by using the residual field as network input for learned corrections, outperforming residual-minimization methods by up to 10x on turbulent flows and ge...

  15. Lighting-Consistent Object Transfer Across Radiance Fields

    cs.GR 2026-06 unverdicted novelty 6.0

    Diffusion-based per-view harmonization for lighting-consistent object transfer between 3DGS scenes, using heterogeneous training data and final 3D consolidation.

  16. S-Agent: Spatial Tool-Use Elicits Reasoning for Spatial Intelligence

    cs.CV 2026-06 unverdicted novelty 6.0

    S-Agent augments VLMs with spatial tools, scene and agent memory for evidence accumulation on multi-view and video tasks, and produces an 8B model via SFT on its own trajectories that beats same-scale baselines.

  17. Unpaired RGB-Thermal Gaussian-Splatting Using Visual Geometric Transformers

    cs.CV 2026-06 unverdicted novelty 6.0

    Framework for unpaired RGB-thermal novel view synthesis via VGGT-based independent pose estimation, Procrustes alignment with cross-modal matcher, multi-modal 3D Gaussian Splatting, and a new benchmarking framework fo...

  18. TriSplat: Simulation-Ready Feed-Forward 3D Scene Reconstruction

    cs.CV 2026-05 unverdicted novelty 6.0

    TriSplat predicts oriented triangle primitives from images in one forward pass to produce simulation-ready 3D meshes with competitive rendering quality.

  19. Stabilizing Streaming Video Geometry via Dynamic Feature Normalization

    cs.CV 2026-05 unverdicted novelty 6.0

    DyFN is a lightweight recurrent module that dynamically normalizes latent feature statistics to remove scale-shift drift and achieve state-of-the-art temporal consistency in streaming monocular geometry estimation whi...

  20. Stream3D: Sequential Multi-View 3D Generation via Evidential Memory

    cs.CV 2026-05 unverdicted novelty 6.0

    Stream3D is a training-free method that maintains a fixed-size evidential memory of past frames to convert frozen view-conditioned 3D generators into consistent streaming generators.

  21. ROAR-3D: Routing Arbitrary Views for High-Fidelity 3D Generation

    cs.CV 2026-05 unverdicted novelty 6.0

    ROAR-3D adds a token-wise view router and dual-stream attention to pretrained single-view 3D generators so they can use arbitrary unposed images for higher-fidelity output.

  22. Feed-Forward 3D Scene Modeling: A Problem-Driven Perspective

    cs.CV 2026-04 unverdicted novelty 6.0

    The paper proposes a problem-driven taxonomy for feed-forward 3D scene modeling that groups methods by five core challenges: feature enhancement, geometry awareness, model efficiency, augmentation strategies, and temp...

  23. WARPED: Wrist-Aligned Rendering for Robot Policy Learning from Egocentric Human Demonstrations

    cs.RO 2026-04 unverdicted novelty 6.0

    WARPED synthesizes realistic wrist-view observations from monocular egocentric human videos via foundation models, hand-object tracking, retargeting, and Gaussian Splatting to train visuomotor policies that match tele...

  24. Diff4Splat: Controllable 4D Scene Generation with Latent Dynamic Reconstruction Models

    cs.CV 2025-11 unverdicted novelty 6.0

    A feed-forward video latent transformer that predicts time-varying 3D Gaussian primitives from one image to produce controllable 4D scenes with appearance, geometry, and motion.

  25. VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction

    cs.CV 2025-05 unverdicted novelty 6.0

    VLM-3R augments VLMs with implicit 3D tokens from monocular video via geometry encoding and 200K+ 3D reconstructive QA pairs, plus a new 138K-pair temporal benchmark, to support spatial and embodied reasoning.

  26. SalientGS: Unified SfM-to-3DGS with Importance-Guided MCMC Gaussian Allocation

    cs.CV 2026-07 conditional novelty 5.5

    Importance-guided MCMC reallocates 3D Gaussians toward multi-view underfit regions, enabling a unified SfM-to-3DGS pipeline that finishes in ~15 minutes with SOTA perceptual quality.

  27. VIGOR: VIdeo Geometry-Oriented Reward for Temporal Generative Alignment

    cs.CV 2026-03 conditional novelty 5.5

    A VGGT-based pointwise reprojection reward with geometry-aware sampling improves video geometric consistency via SFT/DPO and causal test-time search.

  28. SalientGS: Unified SfM-to-3DGS with Importance-Guided MCMC Gaussian Allocation

    cs.CV 2026-07 conditional novelty 5.0

    SalientGS integrates fast first-order SfM, joint pose refinement, and importance-guided MCMC Gaussian birth/relocation to reach 27.65 dB macro-average PSNR at 1.5M Gaussians in about 10 minutes end-to-end.

  29. S-Agent: Spatial Tool-Use Elicits Reasoning for Spatial Intelligence

    cs.CV 2026-06 unverdicted novelty 5.0

    S-Agent improves VLMs on spatial reasoning benchmarks via tool-based 3D evidence accumulation and dual memory, and fine-tuning on its generated trajectories produces an 8B model that surpasses similar-scale baselines ...

  30. EigeNet: Geometry-Informed Multi-Modal Learning for Few-shot Novel View RIR Prediction

    cs.SD 2026-05 unverdicted novelty 5.0

    EigeNet applies a cross-view alternate-attention transformer with geometry modulation for few-shot novel-view RIR prediction, reporting SOTA results on simulated and real data.

  31. StemVLA:An Open-Source Vision-Language-Action Model with Future 3D Spatial Geometry Knowledge and 4D Historical Representation

    cs.RO 2026-02 reject novelty 4.0

    StemVLA supervises a GPT-2-based VLA with predicted future 3D-geometry features (VGGT) and temporally aggregated history, reporting 86.0% on LIBERO-Long - but its CALVIN results and equations are placeholders.

  32. Understanding the Impact of Geometric Foundation Models on Vision-Language-Action Models

    cs.CV 2026-05 unverdicted novelty 3.0

    The paper quantifies the geometric gap in current VLAs via linear probing and compares three architectures for injecting geometry from GFMs while analyzing impacts of data, cameras, and reconstruction quality.