REVIEW 3 major objections 4 minor 3 cited by
Track4Gen: Teaching Video Diffusion Models to Track Points Improves Video Generation
T0 review · 3 major / 4 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read Point-tracking supervision on diffusion features reduces video appearance drift.
desk verdict Track4Gen has a genuinely new idea with a clean implementation, but the missing no-Lcorr control means the reported generation gains are not causally pinned to tracking supervision. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The machinery is the pair of a trainable refiner module $R_\phi$ and a zero-convolution gate $\zeta_\psi$ attached to the upsampler layer of the third decoder block of a Stable Video Diffusion U-Net. $R_\phi$, eight stacked 2D convolution layers initialized as identity, transforms raw diffusion features $h^{1:N}$ into refined features $\tilde{h}^{1:N}$; the correspondence loss computes cosine-similarity cost volumes between a query point's feature and a target frame's feature map, applies a differentiable soft-argmax over a radius-limited window to predict the target position, and penalizes prediction error with a Huber loss. The refined features are fed into the next U-Net block only through $\zeta_\psi$, with gradients detached before the refiner, so $L_\text{corr}$ trains the refiner and the temporal transformer blocks while the diffusion loss continues to train the whole generator. This design is what lets tracking supervision reshape the feature space without destroying the pretrained generation prior.
What would settle it
Train Track4Gen with the same architecture but replace the correspondence labels by random point pairs within each video; if appearance-drift scores on VBench stay high and tracking accuracy stays low, the claimed causal role of the tracking supervision would be refuted.
Extended reading notes
Core claim
Track4Gen's central claim is that explicit spatial-correspondence supervision at the feature level is the missing ingredient for temporally consistent video generation. The paper demonstrates this by training a single network to minimize both the video diffusion denoising loss and a correspondence loss: raw U-Net features from the third decoder block's upsampler layer are projected by an identity-initialized refiner into a correspondence-rich space, cosine-similarity cost volumes with soft-argmax predict point tracks, and a Huber loss supervises those predictions against pseudo-ground-truth trajectories. The refined features are also routed back into the generation backbone through a zero convolution, so the generation path inherits spatial awareness while preserving the pretrained model's prior. On VBench, DAVIS, and BADJA, the paper finds meaningfully improved subject consistency, reduced flickering, and tracking accuracy that approaches dedicated optical-flow chaining, and concludes that video generation and point tracking can be unified in one architecture.
Load-bearing premise
The load-bearing premise is that the pseudo-ground-truth point trajectories—generated by chaining optical-flow estimates from a small set of 567 short video clips and keeping only cycle-consistent matches—are accurate and diverse enough to teach a generalizable spatial-correspondence prior that transfers to other videos.
Editorial extensions
If this is right
- Video generators trained this way should keep a subject's identity stable across many frames, eliminating the object-mutation and object-replacement failures typical of baseline Stable Video Diffusion.
- The same checkpoint can double as a point tracker: zero-shot feature matching with Track4Gen features approaches RAFT optical-flow chaining accuracy on DAVIS and BADJA.
- Plugging Track4Gen features into a test-time-optimization tracker yields long-term tracking accuracy comparable to dedicated supervised trackers.
- Because only a small refiner and the temporal transformer blocks are finetuned, the recipe should transfer to other U-Net video diffusion models with minimal changes.
- Standard generation-quality metrics (FID, FVD, motion smoothness, image quality) do not degrade, so the drift reduction is not bought at the cost of overall video fidelity.
Reading between the lines
- If the causal mechanism is the tracking supervision itself, then scaling the 567-clip training set with automatically annotated real videos should further reduce drift; the small data scale makes this an untested consequence of the paper's hypothesis.
- The same refiner-plus-tracking-loss design should apply to transformer-based (DiT) video generators, but the paper only tests U-Net architectures, so this extension is speculative.
- The paper's acknowledged trade-off of reduced camera motion suggests an extension that supervises global motion separately from point-level correspondences could recover dynamism without sacrificing identity stability.
- Because feature tracking still fails on fast motion, occlusions, and semantically similar objects, adding an explicit occlusion-prediction term to the correspondence loss is a natural next step and is flagged by the authors as future work.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes Track4Gen, a method to add point-tracking supervision to a video diffusion model (Stable Video Diffusion) in order to reduce appearance drift. The method attaches a trainable refiner module to a selected U-Net block, supervises the refined features with a correspondence loss computed against RAFT-generated point tracks, and injects the refined features back into the backbone through a zero convolution initialized to identity. The authors claim that this joint training unifies video generation and point tracking, and report quantitative improvements on VBench metrics, FID/FVD, and user studies, as well as improved zero-shot feature-based tracking on TAP-Vid benchmarks.
Significance. If the central causal claim is correct, Track4Gen would be a valuable and simple recipe for improving spatial coherence of video generators by adding correspondence supervision. The paper is clearly written, the method is well motivated by a feature analysis (Sec. 3.2), and the evaluation is broad, including tracking benchmarks and a user study. The authors are honest about limitations (reduced camera motion, failure cases). However, the headline claim that tracking supervision specifically improves generation is not directly tested, because no ablation removes the correspondence loss while keeping the architecture identical. The quantitative results also lack error bars and significance tests, which is important given the small evaluation set and the modest metric gaps.
major comments (3)
- [Sec. 4.2, Table 4 and Eq. (4)] The experimental design does not isolate the correspondence loss Lcorr from the architectural change. Track4Gen differs from the finetuned-SVD baseline in two coupled ways: the refiner module Rphi is supervised by Lcorr, and the zero-convolution zeta_psi injects Rphi's output into the backbone. The ablation 'Track4Gen w/o refiner' removes both the refiner and the injection path, so the improved scores in Table 1 could come from the added residual-adapter capacity rather than from tracking supervision. Please add a variant trained with the full Track4Gen architecture (refiner Rphi plus zero-conv zeta_psi) but with Lcorr removed (lambda = 0 in the joint loss), so the effect of the correspondence loss is directly measurable. This control is necessary to support the paper's central claim.
- [Table 1 and Fig. 8] Quantitative generation results are reported without error bars, confidence intervals, or significance tests. Several VBench differences between Track4Gen and finetuned SVD are small (e.g., Temporal Flickering 0.9806 vs 0.9800, Motion Smoothness 0.9921 vs 0.9909), and the VBench evaluation uses only 355 images. The FID/FVD values are also single numbers without variance; the user study reports only aggregate preference percentages. Without variance estimates or at least paired significance tests, it is hard to assess whether the reported improvements are robust. Please report results over multiple seeds or bootstrap confidence intervals, and include the user-study per-participant agreement or a significance test.
- [Sec. 4.1 and Supplementary A.1] The training data is only 567 short video clips, with pseudo-ground-truth trajectories generated by chaining RAFT optical flow and filtering via cycle consistency. This is a small and potentially noisy supervision source. Since the method finetunes the temporal transformer blocks on this data, the observed reduction in appearance drift could partly reflect adaptation to the specific training distribution (including reduced camera motion, as the authors acknowledge in the Conclusion) rather than a generalizable correspondence prior. Please report statistics on the quality and coverage of the generated tracklets (e.g., number of surviving tracks, distribution of motion magnitudes), and consider evaluating generation on a broader set of prompts to demonstrate generalization beyond the training distribution.
minor comments (4)
- [Sec. 3.3, Eq. (4)] The weighting factor lambda in the joint loss Ldiff + lambda * Lcorr is not defined until Sec. 4.1; it would be helpful to introduce it at the point of the loss definition in Sec. 3.3.
- [Sec. 4.2 and Table 2] The formatting of Table 2 is dense and the column alignment is difficult to read; please reformat it for clarity, separating generation and tracking metrics more clearly.
- [Sec. 3.2 and Sec. 4.3.1] The cosine similarity threshold of 0.6 is used both for pruning correspondences in feature analysis and for occlusion prediction in tracking evaluation; this dual use should be clarified, as the two tasks may benefit from different thresholds.
- [Supplementary A.2] The description of the refiner network states that the last layer has no BatchNorm, while the first seven layers do; please confirm that this design choice is intentional and report whether BatchNorm in the last layer was tried.
Circularity Check
No significant circularity: training targets and evaluation benchmarks are external, and the central generation claim is independently measured.
full rationale
Track4Gen's claimed derivation chain is self-contained and externally anchored. The correspondence supervision targets are pseudo-labels produced by chaining RAFT optical flow with cycle-consistency filtering (Sec. 4.1 and Supp. A.1), not by the model's own outputs or by the metrics being reported. The central generation claim is evaluated on the external VBench-I2V benchmark with DINO-based Subject Consistency, Temporal Flickering, Motion Smoothness, Image Quality, Video-Image Alignment, plus FID/FVD and a user study (Sec. 4.2); none of these metrics is optimized during training. The tracking auxiliary claim is tested on TAP-Vid DAVIS and BADJA with human-annotated ground truth (Sec. 4.3), so measuring how closely the refiner imitates RAFT is a transfer result, not a definitional identity. The training corpus draws on public video-segmentation datasets that include DAVIS, which is a potential train/test overlap for the tracking benchmark, but the benchmark labels are human-annotated TAP-Vid ground truth rather than the RAFT pseudo-labels used for training, so the tracking evaluation does not reduce to the training objective by construction. The only self-citation is Ref. [32] in Future Work, pointing to the authors' concurrent camera-pose unification paper; it is not load-bearing evidence. The absence of an ablation that removes Lcorr while keeping the zero-convolution injection path is a legitimate causal-attribution confound, but it is an experimental-design gap, not a reduction-by-construction of the paper's claims. Supplementary Sec. E candidly lists limitations (reduced camera motion, artifacts on faces and hands, tracking failures on fast-moving or ambiguous objects), and none of these indicate circularity. Therefore no circular step is present.
Assumptions & free parameters
free parameters (6)
- lambda_Lcorr =
8
- R_window_radius =
35
- cosine_sim_threshold =
0.6
- training_steps_lr_batch =
20K steps, lr 1e-5, batch 4
- feature_block_selection =
upsampler of 3rd decoder block
- training_dataset_composition =
567 video-trajectory pairs
assumptions (4)
- domain assumption Pre-trained Stable Video Diffusion provides a strong video prior and its internal features contain usable correspondences.
- domain assumption RAFT optical flow, chained across frames and filtered by cycle consistency, provides sufficiently accurate point trajectories for supervision.
- standard math Identity initialization of the refiner and zero-initialized convolution preserve the base model prior at the start of finetuning.
- ad hoc to paper Adding correspondence supervision on the chosen block improves appearance consistency without degrading video quality.
invented entities (1)
-
Refiner module R_phi
independent evidence
Cite this review
Pith. "Pith review of Track4Gen: Teaching Video Diffusion Models to Track Points Improves Video Generation." pith.science (2026). https://pith.science/paper/XIDZT3DH
@misc{pith2026241206016,
author = {Pith},
title = {Pith review of: Track4Gen: Teaching Video Diffusion Models to Track Points Improves Video Generation},
year = {2026},
howpublished = {\url{https://pith.science/paper/XIDZT3DH}},
note = {Machine review of arXiv:2412.06016}
}
read the original abstract
While recent foundational video generators produce visually rich output, they still struggle with appearance drift, where objects gradually degrade or change inconsistently across frames, breaking visual coherence. We hypothesize that this is because there is no explicit supervision in terms of spatial tracking at the feature level. We propose Track4Gen, a spatially aware video generator that combines video diffusion loss with point tracking across frames, providing enhanced spatial supervision on the diffusion features. Track4Gen merges the video generation and point tracking tasks into a single network by making minimal changes to existing video generation architectures. Using Stable Video Diffusion as a backbone, Track4Gen demonstrates that it is possible to unify video generation and point tracking, which are typically handled as separate tasks. Our extensive evaluations show that Track4Gen effectively reduces appearance drift, resulting in temporally stable and visually coherent video generation. Project page: hyeonho99.github.io/track4gen
Figures
Figures from the paper (18 more)
Forward citations
Cited by 3 Pith papers
-
SpatialTrackerV2: 3D Point Tracking Made Easy
A single feed-forward model jointly estimates video depth, camera poses, and 3D point trajectories from monocular video, setting a new state of the art on TAPVid-3D.
-
Diffusion as Shader: 3D-aware Video Diffusion for Versatile Video Generation Control
A single video diffusion model conditioned on colored 3D point trajectories performs camera control, motion transfer, mesh-to-video, and object manipulation with improved temporal consistency.
-
Improving Video Diffusion Transformer Training by Multi-Feature Fusion and Alignment from Self-Supervised Vision Encoders
Matching video diffusion transformer tokens to concatenated DINOv2 and SAM2 features improves FVD/FID and speeds convergence, e.g., 400K-step fusion beats 1M-step baseline on UCF-101.
Reference graph
Works this paper leans on
-
[1]
Deep vit features as dense visual descriptors
Shir Amir, Yossi Gandelsman, Shai Bagon, and Tali Dekel. Deep vit features as dense visual descriptors. arXiv preprint arXiv:2112.05814, 2(3):4, 2021. 14
arXiv 2021
-
[2]
Can visual foundation models achieve long-term point tracking? arXiv preprint arXiv:2408.13575, 2024
G ¨orkay Aydemir, Weidi Xie, and Fatma G ¨uney. Can visual foundation models achieve long-term point tracking? arXiv preprint arXiv:2408.13575, 2024. 2, 14
arXiv 2024
-
[3]
Context-PIPs: Persistent Independent Particles Demands Spatial Context Features
Weikang Bian, Zhaoyang Huang, Xiaoyu Shi, Yitong Dong, Yijin Li, and Hongsheng Li. Context-tap: Tracking any point demands spatial context features. arXiv preprint arXiv:2306.02000, 3, 2023. 2
work page Pith review arXiv 2023
-
[4]
Creatures great and smal: Recovering the shape and motion of animals from video
Benjamin Biggs, Thomas Roddick, Andrew Fitzgibbon, and Roberto Cipolla. Creatures great and smal: Recovering the shape and motion of animals from video. In Computer Vision–ACCV 2018: 14th Asian Conference on Computer Vision, Perth, Australia, December 2–6, 2018, Revised Se- lected Papers, Part V 14, pages 3–19. Springer, 2019. 7
2018
-
[5]
Stable video diffusion: Scaling latent video diffusion models to large datasets
Andreas Blattmann, Tim Dockhorn, Sumith Kulal, Daniel Mendelevitch, Maciej Kilian, Dominik Lorenz, Yam Levi, Zion English, Vikram V oleti, Adam Letts, et al. Stable video diffusion: Scaling latent video diffusion models to large datasets. arXiv preprint arXiv:2311.15127 , 2023. 1, 2, 3, 4, 6, 7, 13, 14
arXiv 2023
-
[6]
Align your latents: High-resolution video synthesis with la- tent diffusion models
Andreas Blattmann, Robin Rombach, Huan Ling, Tim Dock- horn, Seung Wook Kim, Sanja Fidler, and Karsten Kreis. Align your latents: High-resolution video synthesis with la- tent diffusion models. In Proceedings of the IEEE/CVF Con- ference on Computer Vision and Pattern Recognition, pages 22563–22575, 2023. 2
2023
-
[7]
Video generation models as world simulators
Tim Brooks, Bill Peebles, Connor Holmes, Will DePue, Yufei Guo, Li Jing, David Schnurr, Joe Taylor, Troy Luh- man, Eric Luhman, Clarence Ng, Ricky Wang, and Aditya Ramesh. Video generation models as world simulators
-
[8]
The 2018 davis challenge on video object seg- mentation
Sergi Caelles, Alberto Montes, Kevis-Kokitsi Maninis, Yuhua Chen, Luc Van Gool, Federico Perazzi, and Jordi Pont-Tuset. The 2018 davis challenge on video object seg- mentation. arXiv preprint arXiv:1803.00557, 2018. 5
arXiv 2018
Show all 87 references
-
[9]
Emerg- ing properties in self-supervised vision transformers
Mathilde Caron, Hugo Touvron, Ishan Misra, Herv ´e J´egou, Julien Mairal, Piotr Bojanowski, and Armand Joulin. Emerg- ing properties in self-supervised vision transformers. In Pro- ceedings of the IEEE/CVF international conference on com- puter vision, pages 9650–9660, 2021. 14
2021
-
[10]
Videocrafter1: Open diffusion models for high-quality video generation
Haoxin Chen, Menghan Xia, Yingqing He, Yong Zhang, Xiaodong Cun, Shaoshu Yang, Jinbo Xing, Yaofang Liu, Qifeng Chen, Xintao Wang, et al. Videocrafter1: Open diffusion models for high-quality video generation. arXiv preprint arXiv:2310.19512, 2023. 2
-
[11]
Deconstructing denois- ing diffusion models for self-supervised learning
X Chen, Z Liu, S Xie, and K He. Deconstructing denois- ing diffusion models for self-supervised learning. arxiv 2024. arXiv preprint arXiv:2401.14404. 3
2024 arXiv
-
[12]
Local all-pair correspon- dence for point tracking
Seokju Cho, Jiahui Huang, Jisu Nam, Honggyu An, Seun- gryong Kim, and Joon-Young Lee. Local all-pair correspon- dence for point tracking. arXiv preprint arXiv:2407.15420,
-
[13]
Diffusion models beat gans on image synthesis
Prafulla Dhariwal and Alexander Nichol. Diffusion models beat gans on image synthesis. Advances in neural informa- tion processing systems, 34:8780–8794, 2021. 2
2021
-
[14]
Tap-vid: A benchmark for track- ing any point in a video
Carl Doersch, Ankush Gupta, Larisa Markeeva, Adria Re- casens, Lucas Smaira, Yusuf Aytar, Joao Carreira, Andrew Zisserman, and Yi Yang. Tap-vid: A benchmark for track- ing any point in a video. Advances in Neural Information Processing Systems, 35:13610–13626, 2022. 2, 5, 7, 8, 13
2022
-
[15]
Tapir: Tracking any point with per-frame initialization and temporal refinement
Carl Doersch, Yi Yang, Mel Vecerik, Dilara Gokay, Ankush Gupta, Yusuf Aytar, Joao Carreira, and Andrew Zisserman. Tapir: Tracking any point with per-frame initialization and temporal refinement. In Proceedings of the IEEE/CVF In- ternational Conference on Computer Vision , pag...
2023
-
[16]
Bootstap: Boot- strapped training for tracking-any-point
Carl Doersch, Yi Yang, Dilara Gokay, Pauline Luc, Skanda Koppula, Ankush Gupta, Joseph Heyward, Ross Goroshin, Jo˜ao Carreira, and Andrew Zisserman. Bootstap: Boot- strapped training for tracking-any-point. arXiv preprint arXiv:2402.00847, 2024. 8
2024 arXiv
-
[17]
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Syl- vain Gelly, Jakob Uszkoreit, and Neil Houlsby. An image is worth 16x16 words: Transformers for image recognition ...
2021
-
[18]
Niladri Shekhar Dutt, Sanjeev Muralikrishnan, and Niloy J. Mitra. Diffusion 3d features (diff3f): Decorating untextured shapes with distilled semantic features. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 4494–4504, 2024. 2
2024
-
[19]
Jumpcut: non-successive mask transfer and interpolation for video cutout
Qingnan Fan, Fan Zhong, Dani Lischinski, Daniel Cohen-Or, and Baoquan Chen. Jumpcut: non-successive mask transfer and interpolation for video cutout. ACM Trans. Graph., 34 (6):195–1, 2015. 5
2015
-
[20]
Brandt, Axel Feld- mann, Zhoutong Zhang, and William T
Stephanie Fu, Mark Hamilton, Laura E. Brandt, Axel Feld- mann, Zhoutong Zhang, and William T. Freeman. Featup: A model-agnostic framework for features at any resolution. In ICLR, 2024. 2
2024
-
[21]
Tokenflow: Consistent diffusion features for consistent video editing
Michal Geyer, Omer Bar-Tal, Shai Bagon, and Tali Dekel. Tokenflow: Consistent diffusion features for consistent video editing. arXiv preprint arXiv:2307.10373, 2023. 2
2023 arXiv
-
[22]
Kubric: A scalable dataset generator
Klaus Greff, Francois Belletti, Lucas Beyer, Carl Doersch, Yilun Du, Daniel Duckworth, David J Fleet, Dan Gnanapra- gasam, Florian Golemo, Charles Herrmann, et al. Kubric: A scalable dataset generator. In Proceedings of the IEEE/CVF conference on computer vision and pattern re...
2022
-
[23]
Videoswap: Customized video subject swapping with interactive semantic point cor- respondence
Yuchao Gu, Yipin Zhou, Bichen Wu, Licheng Yu, Jia-Wei Liu, Rui Zhao, Jay Zhangjie Wu, David Junhao Zhang, Mike Zheng Shou, and Kevin Tang. Videoswap: Customized video subject swapping with interactive semantic point cor- respondence. In Proceedings of the IEEE/CVF Conference o...
2024
-
[24]
Animatediff: Animate your personalized text- to-image diffusion models without specific tuning
Yuwei Guo, Ceyuan Yang, Anyi Rao, Zhengyang Liang, Yaohui Wang, Yu Qiao, Maneesh Agrawala, Dahua Lin, and Bo Dai. Animatediff: Animate your personalized text- to-image diffusion models without specific tuning. arXiv preprint arXiv:2307.04725, 2023. 2
2023 arXiv
-
[25]
Particle video revisited: Tracking through occlusions using point trajectories
Adam W Harley, Zhaoyuan Fang, and Katerina Fragkiadaki. Particle video revisited: Tracking through occlusions using point trajectories. In European Conference on Computer Vi- sion, pages 59–75. Springer, 2022. 2
2022
-
[26]
Latent video diffusion models for high-fidelity long video generation
Yingqing He, Tianyu Yang, Yong Zhang, Ying Shan, and Qifeng Chen. Latent video diffusion models for high-fidelity long video generation. arXiv preprint arXiv:2211.13221 ,
-
[27]
Unsupervised semantic correspondence using stable diffu- sion
Eric Hedlin, Gopal Sharma, Shweta Mahajan, Hossam Isack, Abhishek Kar, Andrea Tagliasacchi, and Kwang Moo Yi. Unsupervised semantic correspondence using stable diffu- sion. In NIPS, 2023. 2
2023
-
[28]
Gans trained by a two time-scale update rule converge to a local nash equilib- rium
Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. Gans trained by a two time-scale update rule converge to a local nash equilib- rium. Advances in neural information processing systems , 30, 2017. 2, 6
2017
-
[29]
Denoising dif- fusion probabilistic models
Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising dif- fusion probabilistic models. Advances in neural information processing systems, 33:6840–6851, 2020. 3
2020
-
[30]
Imagen video: High definition video generation with diffusion mod- els
Jonathan Ho, William Chan, Chitwan Saharia, Jay Whang, Ruiqi Gao, Alexey Gritsenko, Diederik P Kingma, Ben Poole, Mohammad Norouzi, David J Fleet, et al. Imagen video: High definition video generation with diffusion mod- els. arXiv preprint arXiv:2210.02303, 2022. 2
-
[31]
Video dif- fusion models
Jonathan Ho, Tim Salimans, Alexey Gritsenko, William Chan, Mohammad Norouzi, and David J Fleet. Video dif- fusion models. Advances in Neural Information Processing Systems, 35:8633–8646, 2022. 2
2022
-
[32]
Mitra, and Duygu Ceylan
Chun-Hao Paul Huang, Jae Shin Yoon, Hyeonho Jeong, Niloy J. Mitra, and Duygu Ceylan. On unifying video gen- eration and camera pose estimation, 2025. 8
2025
-
[33]
Vbench: Comprehensive bench- mark suite for video generative models
Ziqi Huang, Yinan He, Jiashuo Yu, Fan Zhang, Chenyang Si, Yuming Jiang, Yuanhan Zhang, Tianxing Wu, Qingyang Jin, Nattapol Chanpaisit, et al. Vbench: Comprehensive bench- mark suite for video generative models. In Proceedings of the IEEE/CVF Conference on Computer Vision and P...
2024
-
[34]
Robust estimation of a location parameter
Peter J Huber. Robust estimation of a location parameter. In Breakthroughs in statistics: Methodology and distribution , pages 492–518. Springer, 1992. 5
1992
-
[35]
Space-time correspondence as a contrastive random walk
Allan Jabri, Andrew Owens, and Alexei Efros. Space-time correspondence as a contrastive random walk. Advances in neural information processing systems, 33:19545–19560,
-
[36]
Co- tracker: It is better to track together
Nikita Karaev, Ignacio Rocco, Benjamin Graham, Natalia Neverova, Andrea Vedaldi, and Christian Rupprecht. Co- tracker: It is better to track together. arXiv preprint arXiv:2307.07635, 2023. 2
2023 arXiv
-
[37]
Co- tracker3: Simpler and better point tracking by pseudo- labelling real videos
Nikita Karaev, Iurii Makarov, Jianyuan Wang, Natalia Neverova, Andrea Vedaldi, and Christian Rupprecht. Co- tracker3: Simpler and better point tracking by pseudo- labelling real videos. arXiv preprint arXiv:2410.11831 ,
-
[38]
Elucidating the design space of diffusion-based generative models
Tero Karras, Miika Aittala, Timo Aila, and Samuli Laine. Elucidating the design space of diffusion-based generative models. Advances in neural information processing systems, 35:26565–26577, 2022. 3, 4, 6
2022
-
[39]
Musiq: Multi-scale image quality transformer
Junjie Ke, Qifei Wang, Yilin Wang, Peyman Milanfar, and Feng Yang. Musiq: Multi-scale image quality transformer. In Proceedings of the IEEE/CVF international conference on computer vision, pages 5148–5157, 2021. 6
2021
-
[40]
Auto-encoding variational bayes
Diederik P Kingma. Auto-encoding variational bayes. arXiv preprint arXiv:1312.6114, 2013. 3
2013 arXiv
-
[41]
Harivo: Harnessing text-to-image models for video generation
Mingi Kwon, Seoung Wug Oh, Yang Zhou, Difan Liu, Joon-Young Lee, Haoran Cai, Baqiao Liu, Feng Liu, and Youngjung Uh. Harivo: Harnessing text-to-image models for video generation. arXiv preprint arXiv:2410.07763 , 2024. 15
2024 arXiv
-
[42]
Li, Mihir Prabhudesai, Shivam Duggal, Ellis Brown, and Deepak Pathak
Alexander C. Li, Mihir Prabhudesai, Shivam Duggal, Ellis Brown, and Deepak Pathak. Your diffusion model is secretly a zero-shot classifier. In ICCV, 2023. 2
2023
-
[43]
Video segmentation by tracking many figure- ground segments
Fuxin Li, Taeyoung Kim, Ahmad Humayun, David Tsai, and James M Rehg. Video segmentation by tracking many figure- ground segments. In Proceedings of the IEEE international conference on computer vision, pages 2192–2199, 2013. 5
2013
-
[44]
Taptr: Tracking any point with transformers as detection
Hongyang Li, Hao Zhang, Shilong Liu, Zhaoyang Zeng, Tianhe Ren, Feng Li, and Lei Zhang. Taptr: Tracking any point with transformers as detection. arXiv preprint arXiv:2403.13042, 2024. 2
2024 arXiv
-
[45]
Sd4match: Learning to prompt stable diffusion model for semantic matching
Xinghui Li, Jingyi Lu, Kai Han, and Victor Prisacariu. Sd4match: Learning to prompt stable diffusion model for semantic matching. In CVPR, 2023. 2
2023
-
[46]
Amt: All-pairs multi-field transforms for efficient frame interpolation
Zhen Li, Zuo-Liang Zhu, Ling-Hao Han, Qibin Hou, Chun- Le Guo, and Ming-Ming Cheng. Amt: All-pairs multi-field transforms for efficient frame interpolation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pat- tern Recognition, pages 9801–9810, 2023. 6
2023
-
[47]
Fixing weight decay regularization in adam
Ilya Loshchilov, Frank Hutter, et al. Fixing weight decay regularization in adam. arXiv preprint arXiv:1711.05101, 5,
-
[48]
Diffusion hyperfeatures: Searching through time and space for semantic correspondence
Grace Luo, Lisa Dunlap, Dong Huk Park, Aleksander Holyn- ski, and Trevor Darrell. Diffusion hyperfeatures: Searching through time and space for semantic correspondence. Ad- vances in Neural Information Processing Systems, 36, 2024. 3
2024
-
[49]
Dinov2: Learning robust visual features without supervision
Maxime Oquab, Timoth ´ee Darcet, Th ´eo Moutakanni, Huy V o, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez, Daniel Haziza, Francisco Massa, Alaaeldin El-Nouby, et al. Dinov2: Learning robust visual features without supervision. arXiv preprint arXiv:2304.07193, 2023. 2, 6, 8, 14
2023 arXiv
-
[50]
A benchmark dataset and evaluation methodology for video object segmentation
Federico Perazzi, Jordi Pont-Tuset, Brian McWilliams, Luc Van Gool, Markus Gross, and Alexander Sorkine-Hornung. A benchmark dataset and evaluation methodology for video object segmentation. In Proceedings of the IEEE conference on computer vision and pattern recognition, page...
-
[51]
Movie gen: A cast of media foundation models
Adam Polyak, Amit Zohar, Andrew Brown, Andros Tjandra, Animesh Sinha, Ann Lee, Apoorv Vyas, Bowen Shi, Chih- Yao Ma, Ching-Yao Chuang, et al. Movie gen: A cast of media foundation models. arXiv preprint arXiv:2410.13720,
-
[52]
The 2017 davis challenge on video object segmentation.arXiv preprint arXiv:1704.00675, 2017
Jordi Pont-Tuset, Federico Perazzi, Sergi Caelles, Pablo Ar- bel´aez, Alex Sorkine-Hornung, and Luc Van Gool. The 2017 davis challenge on video object segmentation.arXiv preprint arXiv:1704.00675, 2017. 5, 6
2017 arXiv
-
[53]
Learning transferable visual models from natural language supervi- sion
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervi- sion. In International conference on machine learning, ...
2021
-
[54]
Blattmann, Dominik Lorenz, Patrick Esser, and Bj ¨orn Ommer
Robin Rombach, A. Blattmann, Dominik Lorenz, Patrick Esser, and Bj ¨orn Ommer. High-resolution image synthesis with latent diffusion models. In CVPR, 2021. 2
2021
-
[55]
High-resolution image synthesis with latent diffusion models
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bj ¨orn Ommer. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 10684–10695, 2022. 2
2022
-
[56]
Progressive distillation for fast sampling of diffusion models
Tim Salimans and Jonathan Ho. Progressive distillation for fast sampling of diffusion models. arXiv preprint arXiv:2202.00512, 2022. 4
2022 arXiv
-
[57]
Particle video: Long-range mo- tion estimation using point trajectories.International journal of computer vision, 80:72–91, 2008
Peter Sand and Seth Teller. Particle video: Long-range mo- tion estimation using point trajectories.International journal of computer vision, 80:72–91, 2008. 2
2008
-
[58]
Make-a-video: Text-to-video generation without text-video data
Uriel Singer, Adam Polyak, Thomas Hayes, Xi Yin, Jie An, Songyang Zhang, Qiyuan Hu, Harry Yang, Oron Ashual, Oran Gafni, et al. Make-a-video: Text-to-video generation without text-video data. arXiv preprint arXiv:2209.14792 ,
-
[59]
Deep unsupervised learning using nonequilibrium thermodynamics
Jascha Sohl-Dickstein, Eric Weiss, Niru Maheswaranathan, and Surya Ganguli. Deep unsupervised learning using nonequilibrium thermodynamics. In International confer- ence on machine learning, pages 2256–2265. PMLR, 2015. 3
2015
-
[60]
Zeroscope, 2023
Spencer Sterling. Zeroscope, 2023. https : / / huggingface . co / cerspense / zeroscope _ v2 _ 576w. 3, 7, 14
2023
-
[61]
Tamir, Lucy Chai, Simon Kornblith, Trevor Darrell, and Phillip Isola
Shobhita Sundaram, Stephanie Fu, Lukas Muttenthaler, Ne- tanel Y . Tamir, Lucy Chai, Simon Kornblith, Trevor Darrell, and Phillip Isola. When does perceptual alignment benefit vision representations? In NIPS, 2024. 2
2024
-
[62]
Lift: A surprisingly simple lightweight feature transform for dense vit descriptors
Saksham Suri, Matthew Walmer, Kamal Gupta, and Abhinav Shrivastava. Lift: A surprisingly simple lightweight feature transform for dense vit descriptors. In ECCV, pages 110–
-
[63]
Emergent correspondence from image diffusion
Luming Tang, Menglin Jia, Qianqian Wang, Cheng Perng Phoo, and Bharath Hariharan. Emergent correspondence from image diffusion. Advances in Neural Information Pro- cessing Systems, 36:1363–1389, 2023. 3, 7
2023
-
[64]
Raft: Recurrent all-pairs field transforms for optical flow
Zachary Teed and Jia Deng. Raft: Recurrent all-pairs field transforms for optical flow. In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23– 28, 2020, Proceedings, Part II 16, pages 402–419. Springer,
2020
-
[65]
Plug-and-play diffusion features for text-driven image-to-image translation
Narek Tumanyan, Michal Geyer, Shai Bagon, and Tali Dekel. Plug-and-play diffusion features for text-driven image-to-image translation. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 1921–1930, 2023. 2
1921
-
[66]
Dino-tracker: Taming dino for self-supervised point tracking in a single video
Narek Tumanyan, Assaf Singer, Shai Bagon, and Tali Dekel. Dino-tracker: Taming dino for self-supervised point tracking in a single video. arXiv preprint arXiv:2403.14548, 2024. 2, 3, 7, 8, 13, 14, 20
2024 arXiv
-
[67]
To- wards accurate generative models of video: A new metric & challenges
Thomas Unterthiner, Sjoerd Van Steenkiste, Karol Kurach, Raphael Marinier, Marcin Michalski, and Sylvain Gelly. To- wards accurate generative models of video: A new metric & challenges. arXiv preprint arXiv:1812.01717, 2018. 2, 6
2018 arXiv
-
[68]
A connection between score matching and denoising autoencoders
Pascal Vincent. A connection between score matching and denoising autoencoders. Neural computation, 23(7):1661– 1674, 2011. 1
2011
-
[69]
Tracking everything everywhere all at once
Qianqian Wang, Yen-Yu Chang, Ruojin Cai, Zhengqi Li, Bharath Hariharan, Aleksander Holynski, and Noah Snavely. Tracking everything everywhere all at once. In Proceedings of the IEEE/CVF International Conference on Computer Vi- sion, pages 19795–19806, 2023. 2, 8, 13
2023
-
[70]
Zero-shot video semantic segmentation based on pre-trained diffusion models, 2024
Qian Wang, Abdelrahman Eldesokey, Mohit Mendiratta, Fangneng Zhan, Adam Kortylewski, Christian Theobalt, and Peter Wonka. Zero-shot video semantic segmentation based on pre-trained diffusion models, 2024. 2
2024
-
[71]
Lavie: High-quality video gener- ation with cascaded latent diffusion models
Yaohui Wang, Xinyuan Chen, Xin Ma, Shangchen Zhou, Ziqi Huang, Yi Wang, Ceyuan Yang, Yinan He, Jiashuo Yu, Peiqing Yang, et al. Lavie: High-quality video gener- ation with cascaded latent diffusion models. arXiv preprint arXiv:2309.15103, 2023. 2
2023 arXiv
-
[72]
Towards a better metric for text-to-video generation
Jay Zhangjie Wu, Guian Fang, Haoning Wu, Xintao Wang, Yixiao Ge, Xiaodong Cun, David Junhao Zhang, Jia-Wei Liu, Yuchao Gu, Rui Zhao, et al. Towards a better metric for text-to-video generation. arXiv preprint arXiv:2401.07781,
-
[73]
Denoising diffusion autoencoders are unified self-supervised learners
Weilai Xiang, Hongyu Yang, Di Huang, and Yunhong Wang. Denoising diffusion autoencoders are unified self-supervised learners. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 15802–15812, 2023. 3
2023
-
[74]
Spatialtracker: Tracking any 2d pixels in 3d space
Yuxi Xiao, Qianqian Wang, Shangzhan Zhang, Nan Xue, Sida Peng, Yujun Shen, and Xiaowei Zhou. Spatialtracker: Tracking any 2d pixels in 3d space. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 20406–20417, 2024. 2
2024
-
[75]
ODISE: Open-V ocabulary Panoptic Segmentation with Text-to-Image Diffusion Mod- els
Jiarui Xu, Sifei Liu, Arash Vahdat, Wonmin Byeon, Xiao- long Wang, and Shalini De Mello. ODISE: Open-V ocabulary Panoptic Segmentation with Text-to-Image Diffusion Mod- els. In CVPR, 2023. 2
2023
-
[76]
Pre-constancy vision in infants.Current Biology, 25(24):3209–3212, 2015
Jiale Yang, So Kanazawa, Masami K Yamaguchi, and Isamu Motoyoshi. Pre-constancy vision in infants.Current Biology, 25(24):3209–3212, 2015. 1
2015
-
[77]
Diffusion model as repre- sentation learner
Xingyi Yang and Xinchao Wang. Diffusion model as repre- sentation learner. In Proceedings of the IEEE/CVF Interna- tional Conference on Computer Vision, pages 18938–18949,
-
[78]
Representation alignment for generation: Training diffu- sion transformers is easier than you think
Sihyun Yu, Sangkyung Kwak, Huiwon Jang, Jongheon Jeong, Jonathan Huang, Jinwoo Shin, and Saining Xie. Representation alignment for generation: Training diffu- sion transformers is easier than you think. arXiv preprint arXiv:2410.06940, 2024. 2, 3
-
[79]
Improving 2D Feature Representa- tions by 3D-Aware Fine-Tuning
Yuanwen Yue, Anurag Das, Francis Engelmann, Siyu Tang, and Jan Eric Lenssen. Improving 2D Feature Representa- tions by 3D-Aware Fine-Tuning. In ECCV, 2024. 2
2024
-
[80]
Show-1: Marrying pixel and latent dif- fusion models for text-to-video generation
David Junhao Zhang, Jay Zhangjie Wu, Jia-Wei Liu, Rui Zhao, Lingmin Ran, Yuchao Gu, Difei Gao, and Mike Zheng Shou. Show-1: Marrying pixel and latent dif- fusion models for text-to-video generation. arXiv preprint arXiv:2309.15818, 2023. 2
2023 arXiv
-
[81]
A tale of two features: Stable diffusion complements dino for zero-shot semantic correspondence
Junyi Zhang, Charles Herrmann, Junhwa Hur, Luisa Pola- nia Cabrera, Varun Jampani, Deqing Sun, and Ming-Hsuan Yang. A tale of two features: Stable diffusion complements dino for zero-shot semantic correspondence. Advances in Neural Information Processing Systems, 36, 2024. 14
2024
-
[82]
Adding conditional control to text-to-image diffusion models
Lvmin Zhang, Anyi Rao, and Maneesh Agrawala. Adding conditional control to text-to-image diffusion models. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 3836–3847, 2023. 5
2023
-
[83]
The unreasonable effectiveness of deep features as a perceptual metric
Richard Zhang, Phillip Isola, Alexei A Efros, Eli Shecht- man, and Oliver Wang. The unreasonable effectiveness of deep features as a perceptual metric. In Proceedings of the IEEE conference on computer vision and pattern recogni- tion, pages 586–595, 2018. 14
2018
-
[84]
I2vgen-xl: High-quality image-to-video synthesis via cascaded diffusion models
Shiwei Zhang, Jiayu Wang, Yingya Zhang, Kang Zhao, Hangjie Yuan, Zhiwu Qin, Xiang Wang, Deli Zhao, and Jingren Zhou. I2vgen-xl: High-quality image-to-video synthesis via cascaded diffusion models. arXiv preprint arXiv:2311.04145, 2023. 2, 3
2023 arXiv
-
[85]
Pointodyssey: A large-scale synthetic dataset for long-term point tracking
Yang Zheng, Adam W Harley, Bokui Shen, Gordon Wet- zstein, and Leonidas J Guibas. Pointodyssey: A large-scale synthetic dataset for long-term point tracking. In Proceed- ings of the IEEE/CVF International Conference on Com- puter Vision, pages 19855–19865, 2023. 8
2023
-
[86]
Magicvideo: Efficient video generation with latent diffusion models
Daquan Zhou, Weimin Wang, Hanshu Yan, Weiwei Lv, Yizhe Zhu, and Jiashi Feng. Magicvideo: Efficient video generation with latent diffusion models. arXiv preprint arXiv:2211.11018, 2022. 2 Track4Gen: Teaching Video Diffusion Models to Track Points Improves Video Generation Suppl...
2022 arXiv
-
[87]
To better demonstrate the architecture of the baseline Track4Gen without Refiner , we provide a visualization in Fig
The first 7 layers follow the structure Conv2d → BatchNorm2d → ReLU, except for the last layer which consists of Conv2d → ReLU. To better demonstrate the architecture of the baseline Track4Gen without Refiner , we provide a visualization in Fig. 10. The figure compares the tra...
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.