Pith. sign in

REVIEW 9 cited by

Learning Temporally Consistent Video Depth from Video Diffusion Priors

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.01493 v4 pith:HCFYXUNL submitted 2024-06-03 cs.CV

Learning Temporally Consistent Video Depth from Video Diffusion Priors

classification cs.CV
keywords clipdepthframesstrategytrainingvideochronodepthclips
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

This work addresses the challenge of streamed video depth estimation, which expects not only per-frame accuracy but, more importantly, cross-frame consistency. We argue that sharing contextual information between frames or clips is pivotal in fostering temporal consistency. Therefore, we reformulate depth prediction into a conditional generation problem to provide contextual information within a clip and across clips. Specifically, we propose a consistent context-aware training and inference strategy for arbitrarily long videos to provide cross-clip context. We sample independent noise levels for each frame within a clip during training while using a sliding window strategy and initializing overlapping frames with previously predicted frames without adding noise. Moreover, we design an effective training strategy to provide context within a clip. Extensive experimental results validate our design choices and demonstrate the superiority of our approach, dubbed ChronoDepth. Project page: https://xdimlab.github.io/ChronoDepth/.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Forget, Anticipate and Adapt: Test Time Training for Long Videos

    cs.CV 2026-06 unverdicted novelty 7.0

    FFN enables efficient TTT for long videos by operating on three frames and using a surprise-based adaptive window, shown on a new dataset of up to 3-hour videos for segmentation and classification tasks.

  2. Hallo4D: Multi-Modal Hallucination Mitigation for Consistent Spatio-Temporal Generation

    cs.CV 2026-07 conditional novelty 6.0

    Hallo4D uses vision-language models to detect and correct spatial and temporal mistakes in AI-generated 3D and 4D content, improving consistency without retraining the base generators.

  3. Forget, Anticipate and Adapt: Test Time Training for Long Videos

    cs.CV 2026-06 unverdicted novelty 6.0

    FFN performs TTT on multi-hour videos by restricting updates to three frames and using a surprise metric for adaptive window sizing, plus a new EpicTours dataset.

  4. Forget, Anticipate and Adapt: Test Time Training for Long Videos

    cs.CV 2026-06 conditional novelty 6.0

    FFN performs efficient test-time training on multi-hour videos by forgetting the exiting frame, anticipating the next, and adapting only when a surprise metric exceeds a dynamic threshold.

  5. Towards Consistent Video Geometry Estimation

    cs.CV 2026-05 conditional novelty 6.0

    One transformer, trained with random-sized temporal attention chunks, unifies offline, streaming, and long-video depth, normal, and point-map estimation and reports new best numbers on five public benchmarks.

  6. Hallo4D: Multi-Modal Hallucination Mitigation for Consistent Spatio-Temporal Generation

    cs.CV 2026-07 conditional novelty 5.0

    Hallo4D mitigates 3D/4D generation hallucinations via LMM-based detection, multi-model voting correction, and motion-aware optimization without retraining base generators.

  7. Towards Consistent Video Geometry Estimation

    cs.CV 2026-05 unverdicted novelty 5.0

    ViGeo is a feed-forward transformer for video geometry that introduces dynamic chunking attention and a completion-based data refinement framework to achieve SOTA on depth, normals, and point map estimation.

  8. Geometry-aware 4D Video Generation for Robot Manipulation

    cs.CV 2025-07 unverdicted novelty 5.0

    A geometry-aware 4D video generation model trained with cross-view pointmap alignment to produce spatio-temporally consistent future videos from novel viewpoints for robot manipulation.

  9. MonST3R: A Simple Approach for Estimating Geometry in the Presence of Motion

    cs.CV 2024-10 unverdicted novelty 5.0

    By fine-tuning DUST3R to output per-timestep pointmaps on scarce dynamic video datasets, MonST3R achieves stronger video depth and pose estimation without explicit motion modeling.