Pith. sign in

REVIEW 13 cited by

WinT3R: Window-Based Streaming Reconstruction with Camera Token Pool

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2509.05296 v1 pith:6PKVHVIP submitted 2025-09-05 cs.CV cs.AI

classification cs.CVcs.AI
keywords camerareconstructionwint3rqualityestimationmodelonlineperformance
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present WinT3R, a feed-forward reconstruction model capable of online prediction of precise camera poses and high-quality point maps. Previous methods suffer from a trade-off between reconstruction quality and real-time performance. To address this, we first introduce a sliding window mechanism that ensures sufficient information exchange among frames within the window, thereby improving the quality of geometric predictions without large computation. In addition, we leverage a compact representation of cameras and maintain a global camera token pool, which enhances the reliability of camera pose estimation without sacrificing efficiency. These designs enable WinT3R to achieve state-of-the-art performance in terms of online reconstruction quality, camera pose estimation, and reconstruction speed, as validated by extensive experiments on diverse datasets. Code and model are publicly available at https://github.com/LiZizun/WinT3R.

Discussion (0). Sign in to comment.

Forward citations

Cited by 13 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. SpatialBench: Is Your Spatial Foundation Model an All-Round Player?

    cs.CV 2026-05 unverdicted novelty 8.0 of 10

    SpatialBench evaluates 41 spatial foundation models across 6 paradigms and 5 task suites, finds they are not all-round players, and introduces the DA-Next-5M dataset plus DA-Next baseline model.

  2. Geo-Align: Video Generation Alignment via Metric Geometry Reward

    cs.CV 2026-05 unverdicted novelty 7.0 of 10

    Geo-Align applies RL with a perceptual reward derived from 3D camera trajectory estimation to improve controllability and fidelity in video generation without paired training data.

  3. FrameVGGT: Coherence-Preserving Memory for Bounded Streaming Geometry

    cs.CV 2026-03 unverdicted novelty 7.0 of 10

    FrameVGGT replaces token-level KV retention with frame-level segments and prototypes to bound memory while preserving geometric coherence in streaming VGGT.

  4. Glob3R: Global Structure-from-Motion with 3D Foundation Models

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A frozen Pi3X backbone plus dense warping tracks and keyframe sliding-window global optimization yields more accurate, scalable SfM than feed-forward or classical baselines alone.

  5. Anchor3R: Streaming 3D Reconstruction with Transient Anchors for Long-Horizon Visual Mapping

    cs.CV 2026-06 unverdicted novelty 6.0 of 10

    Anchor3R reframes feed-forward 3D reconstruction as current-centric local measurement prediction, using loop-closure and motion averaging to produce coherent global maps from visual streams.

  6. Rethinking the State Update Gate for Long-Sequence Recurrent 3D Reconstruction

    cs.CV 2026-05 unverdicted novelty 6.0 of 10

    A closed-form scalar frame-level gate α_t derived from internal feature changes extends effective memory in recurrent 3D reconstruction and improves accuracy on long sequences up to 4541 frames.

  7. Attention Itself Could Retrieve.RetrieveVGGT: Training-Free Long Context Streaming 3D Reconstruction via Query-Key Similarity Retrieval

    cs.CV 2026-05 unverdicted novelty 6.0 of 10

    RetrieveVGGT enables constant-memory long-context streaming 3D reconstruction by retrieving relevant frames via query-key similarities in VGGT's first attention layer, outperforming StreamVGGT and others.

  8. Feed-Forward 3D Scene Modeling: A Problem-Driven Perspective

    cs.CV 2026-04 unverdicted novelty 6.0 of 10

    The paper proposes a problem-driven taxonomy for feed-forward 3D scene modeling that groups methods by five core challenges: feature enhancement, geometry awareness, model efficiency, augmentation strategies, and temp...

  9. TTSA3R: Training-Free Temporal-Spatial Adaptive Persistent State for Streaming 3D Reconstruction

    cs.CV 2026-01 conditional novelty 6.0 of 10

    TTSA3R adds temporal and spatial adaptive gating to the CUT3R persistent-state update rule, cutting error growth on long sequences from >4x to 1.33x over 50-250 frames.

  10. 2K Retrofit: Entropy-Guided Efficient Sparse Refinement for High-Resolution 3D Geometry Prediction

    cs.CV 2026-03 conditional novelty 5.5 of 10

    Entropy-guided sparse refinement upgrades frozen low-resolution geometric foundation models to accurate 2K depth and pointmap outputs at a fraction of full-resolution cost.

  11. IGGT4D: Streaming 4D Instance-Grounded Geometry Transformer

    cs.CV 2026-07 conditional novelty 5.0 of 10

    A causal streaming transformer that jointly predicts camera motion, 3D geometry, and persistent object-instance features from video, trained on a new 147K-sequence 4D dataset.

  12. SANA-WM: Efficient Minute-Scale World Modeling with Hybrid Linear Diffusion Transformer

    cs.CV 2026-05 unverdicted novelty 5.0 of 10

    SANA-WM is a 2.6B-parameter efficient world model that synthesizes minute-scale 720p videos with 6-DoF camera control, trained on 213K public clips in 15 days on 64 H100s and runnable on single GPUs at 36x higher thro...

  13. FrameVGGT: Coherence-Preserving Memory for Bounded Streaming Geometry

    cs.CV 2026-03 unverdicted novelty 5.0 of 10

    FrameVGGT maintains stable long-horizon 3D reconstruction, depth, and pose under fixed memory by organizing history as complementary frame-wise KV prototypes plus sparse anchors.

Pith tools