Pith. sign in

REVIEW 29 cited by

Continuous 3D Perception Model with Persistent State

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2501.12387 v1 pith:HFDRWFJ4 submitted 2025-01-21 cs.CV

Continuous 3D Perception Model with Persistent State

classification cs.CV
keywords imagesmodelpointmapsstatecontinuouscut3rmethodreconstruction
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

We present a unified framework capable of solving a broad range of 3D tasks. Our approach features a stateful recurrent model that continuously updates its state representation with each new observation. Given a stream of images, this evolving state can be used to generate metric-scale pointmaps (per-pixel 3D points) for each new input in an online fashion. These pointmaps reside within a common coordinate system, and can be accumulated into a coherent, dense scene reconstruction that updates as new images arrive. Our model, called CUT3R (Continuous Updating Transformer for 3D Reconstruction), captures rich priors of real-world scenes: not only can it predict accurate pointmaps from image observations, but it can also infer unseen regions of the scene by probing at virtual, unobserved views. Our method is simple yet highly flexible, naturally accepting varying lengths of images that may be either video streams or unordered photo collections, containing both static and dynamic content. We evaluate our method on various 3D/4D tasks and demonstrate competitive or state-of-the-art performance in each. Project Page: https://cut3r.github.io/

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 29 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. GenRecon: Bridging Generative Priors for Multi-View 3D Scene Reconstruction

    cs.CV 2026-05 unverdicted novelty 7.0

    GenRecon lifts object-level generative priors to scene-scale reconstruction by chunking scenes and using projection-based conditioning on multi-view features, claiming 16% better results than prior methods.

  2. EvoScene-VLA: Evolving Scene Beliefs Inside the Action Decoder for Chunked Robot Control

    cs.RO 2026-05 conditional novelty 7.0

    EvoScene-VLA maintains an action-updated scene prior across control chunks in VLA policies, raising success rates on RoboTwin tasks from 87.2% to 89.1% fixed and 86.1% to 88.5% randomized while outperforming baselines...

  3. Trust It or Not: Evidential Uncertainty for Feed-Forward 3D Reconstruction with Trust3R

    cs.CV 2026-05 unverdicted novelty 7.0

    Trust3R introduces a gated residual refinement plus Normal-Inverse-Wishart evidential head that produces closed-form multivariate Student-t uncertainty for per-point geometry in feed-forward 3D reconstruction and impr...

  4. PaceVGGT: Pre-Alternating-Attention Token Pruning for Visual Geometry Transformers

    cs.CV 2026-05 unverdicted novelty 7.0

    PaceVGGT reduces VGGT inference latency by up to 5.1x on ScanNet-50 via pre-AA token pruning with a distilled Token Scorer, per-frame keep budgets, adaptive merge/prune, and feature-guided restoration, while preservin...

  5. VGGT-SLAM: Dense RGB SLAM Optimized on the SL(4) Manifold

    cs.CV 2025-05 unverdicted novelty 7.0

    VGGT-SLAM aligns VGGT submaps via SL(4) manifold optimization of 15-DoF homographies to enable consistent dense RGB SLAM on long uncalibrated monocular videos.

  6. Robust Dreamer: Deviation-Aware Latent Gaussian Memory for Action-Controlled AR Video Generation

    cs.CV 2026-05 unverdicted novelty 6.0

    Robust Dreamer uses Latent Gaussian Memory anchored to diffusion latents and Deviation Learning with a Dynamic Deviation Archive to reduce drift in long-horizon action-controlled image-to-video generation, reporting S...

  7. Stabilizing Streaming Video Geometry via Dynamic Feature Normalization

    cs.CV 2026-05 unverdicted novelty 6.0

    DyFN is a lightweight recurrent module that dynamically normalizes latent feature statistics to remove scale-shift drift and achieve state-of-the-art temporal consistency in streaming monocular geometry estimation whi...

  8. Cross-View Splatter: Feed-Forward View Synthesis with Georeferenced Images

    cs.CV 2026-05 unverdicted novelty 6.0

    A feed-forward model aligns ground and satellite features to predict Gaussian splats for improved novel-view synthesis on georeferenced outdoor scenes.

  9. SpaceMind++: Toward Allocentric Cognitive Maps for Spatially Grounded Video MLLMs

    cs.CV 2026-05 unverdicted novelty 6.0

    SpaceMind++ adds an explicit voxelized allocentric cognitive map and coordinate-guided fusion to video MLLMs, claiming SOTA on VSI-Bench and improved out-of-distribution generalization on three other 3D benchmarks.

  10. Sat3R: Satellite DSM Reconstruction via RPC-Aware Depth Fine-tuning

    cs.CV 2026-05 unverdicted novelty 6.0

    Sat3R adapts Depth Anything V2 via RPC-aware metric depth fine-tuning to deliver satellite DSM reconstruction with 38% lower MAE than zero-shot baselines and over 300x speedup versus optimization methods.

  11. Syn4D: A Multiview Synthetic 4D Dataset

    cs.CV 2026-05 conditional novelty 6.0

    Syn4D supplies multiview synthetic dynamic scenes with dense geometric, tracking and pose ground truth that lets any pixel be unprojected to any time and camera.

  12. HD-VGGT: High-Resolution Visual Geometry Transformer

    cs.CV 2026-03 unverdicted novelty 6.0

    HD-VGGT achieves state-of-the-art high-resolution 3D reconstruction from image collections via a dual-branch architecture that predicts coarse geometry at low resolution and refines details at high resolution while mo...

  13. OVGGT: O(1) Constant-Cost Streaming Visual Geometry Transformer

    cs.CV 2026-03 conditional novelty 6.0

    OVGGT achieves constant O(1) memory and compute for streaming 3D geometry reconstruction by using FFN-residual-based KV cache compression and dynamic anchor protection, matching state-of-the-art accuracy on long sequences.

  14. Co-Me: Confidence-Guided Token Merging for Visual Geometric Transformers

    cs.CV 2025-11 unverdicted novelty 6.0

    Co-Me distills a confidence predictor to selectively merge low-confidence tokens in visual geometric transformers, delivering up to 21.5x speedup on VGGT and 20.4x on Pi3 while preserving spatial coverage and performance.

  15. HERO: Hierarchical Extrapolation and Refresh for Efficient World Models

    cs.CV 2025-08 unverdicted novelty 6.0

    HERO accelerates world model inference 1.73x via hierarchical patch-wise refresh in shallow layers and linear extrapolation in deeper layers with minimal quality loss.

  16. LongSplat: Robust Unposed 3D Gaussian Splatting for Casual Long Videos

    cs.CV 2025-08 conditional novelty 6.0

    An incremental 3D Gaussian Splatting pipeline that jointly optimizes camera poses and scene geometry using MASt3R priors and density-adaptive octree anchors achieves state-of-the-art novel view synthesis on casual lon...

  17. STream3R: Scalable Sequential 3D Reconstruction with Causal Transformer

    cs.CV 2025-08 conditional novelty 6.0

    A decoder-only Transformer with causal attention and cached past-frame features performs incremental 3D reconstruction from streaming images, beating the RNN-based CUT3R on several benchmark metrics.

  18. VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction

    cs.CV 2025-05 unverdicted novelty 6.0

    VLM-3R augments VLMs with implicit 3D tokens from monocular video via geometry encoding and 200K+ 3D reconstructive QA pairs, plus a new 138K-pair temporal benchmark, to support spatial and embodied reasoning.

  19. Multi-SpatialMLLM: Multi-Frame Spatial Understanding with Multi-Modal Large Language Models

    cs.CV 2025-05 unverdicted novelty 6.0

    Multi-SpatialMLLM integrates depth perception, visual correspondence, and dynamic perception into MLLMs via a 27M-sample MultiSPA dataset and benchmark, yielding gains on multi-frame spatial tasks.

  20. RoadVGGT: Road-Structure-Aware Feed-Forward Road Surface Reconstruction

    cs.CV 2026-07 conditional novelty 5.5

    A feed-forward Gaussian head on OmniVGGT plus road-plane grid fusion and structure-aware grouping reconstructs compact road surfaces that beat RoGS and AnySplat on Waymo and zero-shot nuScenes.

  21. ESAM++: Efficient Online 3D Perception on the Edge

    cs.CV 2026-05 unverdicted novelty 5.0

    ESAM++ introduces a 3D Sparse Feature Pyramid Network for efficient online 3D scene perception on edge devices, claiming competitive accuracy with up to 3x faster inference and 2x smaller model size than ESAM on four ...

  22. HorizonStream: Long-Horizon Attention for Streaming 3D Reconstruction

    cs.CV 2026-05 unverdicted novelty 5.0

    HorizonStream is a long-horizon Transformer that factorizes geometric evidence influence into channel-wise linear attention for long-range temporal propagation and local spatiotemporal attention for short-range matchi...

  23. Syn4D: A Multiview Synthetic 4D Dataset

    cs.CV 2026-05 unverdicted novelty 5.0

    Syn4D is a new multiview synthetic 4D dataset supplying dense ground-truth annotations for dynamic scene reconstruction, tracking, and human pose estimation.

  24. Context Unrolling in Omni Models

    cs.CV 2026-04 unverdicted novelty 5.0

    Omni is a multimodal model whose native training on diverse data types enables context unrolling, allowing explicit reasoning across modalities to better approximate shared knowledge and improve downstream performance.

  25. MonoEM-GS: Monocular Expectation-Maximization Gaussian Splatting SLAM

    cs.RO 2026-04 unverdicted novelty 5.0

    MonoEM-GS stabilizes view-dependent geometry from foundation models inside a global Gaussian Splatting representation via EM and adds multi-modal features for in-place open-set segmentation.

  26. FF3R: Feedforward Feature 3D Reconstruction from Unconstrained views

    cs.CV 2026-04 unverdicted novelty 5.0

    FF3R unifies geometric and semantic 3D reconstruction in a single annotation-free feed-forward network trained solely via RGB and feature rendering supervision.

  27. InstantSfM: Towards GPU-Native SfM for the Deep Learning Era

    cs.CV 2025-10 conditional novelty 5.0

    A fully GPU-native, PyTorch-based global Structure-from-Motion pipeline using sparse-aware Levenberg-Marquardt with optional metric depth priors reports ~8-40× speedups over COLMAP at comparable accuracy on several be...

  28. ViPE: Video Pose Engine for 3D Geometric Perception

    cs.CV 2025-08 unverdicted novelty 5.0

    ViPE estimates camera intrinsics, motion, and dense near-metric depth from uncalibrated videos, outperforming baselines on TUM and KITTI while releasing annotations for 96M frames across real and generated videos.

  29. VGGT-Long: Chunk it, Loop it, Align it -- Pushing VGGT's Limits on Kilometer-scale Long RGB Sequences

    cs.CV 2025-07 conditional novelty 4.0

    VGGT-Long extends VGGT with chunking, overlap alignment, and loop closure to produce consistent kilometer-scale 3D reconstructions from monocular RGB sequences without retraining or extra supervision.