REVIEW 2 major objections 4 minor 3 cited by
MegaFlow shows that frozen global Vision Transformer features plus light local refinement deliver state-of-the-art zero-shot large-displacement optical flow and transferable point tracking.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-13 18:00 UTC pith:LCAXW5TL
load-bearing objection Clean empirical win: frozen VGGT-style global features + matching + light multi-frame refinement give real zero-shot large-motion flow and free TAP-Vid transfer. the 2 major comments →
MegaFlow: Zero-Shot Large Displacement Optical Flow
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
By formulating optical flow as all-pairs global matching over multi-frame features from a frozen pre-trained Vision Transformer (augmented by a small trainable CNN fusion head) and then applying a lightweight recurrent local-refinement module, MegaFlow achieves state-of-the-art zero-shot accuracy on large-displacement optical flow while transferring directly to dense long-range point tracking.
What carries the argument
Global matching of pre-trained Vision Transformer features: each location is soft-matched against the entire next frame via correlation and expectation, yielding an initial flow that already captures large displacements; a few subsequent ConvNeXt-plus-temporal-attention iterations then restore sub-pixel accuracy and temporal consistency.
Load-bearing premise
The load-bearing premise is that geometric features learned only on static multi-view scenes already encode sufficiently accurate long-range correspondences for dynamic, large-displacement motion once a lightweight fusion and a handful of local refinements are added.
What would settle it
Train an identical architecture from scratch (no pre-trained Transformer weights) under the same curriculum and measure whether zero-shot Sintel Final EPE and s40+ large-motion error remain competitive with the pre-trained version; a large gap that cannot be closed by extra capacity would falsify the claim that the static priors are the essential ingredient.
If this is right
- Zero-shot optical-flow models can reach or surpass specialized fine-tuned systems on Sintel Final, Spring, and large-motion bins without any dataset-specific adaptation.
- The same architecture, used unchanged, yields competitive zero-shot dense point tracking on TAP-Vid and can become state-of-the-art after modest tracking fine-tuning.
- A unified global-matching-plus-local-refinement pipeline becomes a practical foundation for both short-range dense flow and long-horizon point trajectories.
- Multi-frame context (around four frames) measurably improves temporal consistency and occlusion handling without redesign of the network.
Where Pith is reading between the lines
- If static geometric priors already solve most of the large-displacement problem, further gains may come more from better multi-frame fusion and longer temporal windows than from ever-larger task-specific cost volumes.
- The same frozen-backbone recipe could be tested on related dense correspondence tasks such as stereo or scene flow with minimal architectural change.
- Because performance degrades when the Transformer is frozen or the CNN fusion is removed, hybrid fine-tuning strategies that keep the global prior while adapting only local layers may be a general pattern for motion foundation models.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. MegaFlow adapts frozen pre-trained global Vision Transformer features (DINOv2/VGGT-style alternating frame/global attention) plus a lightweight trainable CNN fusion head to cast optical flow as all-pairs global matching (Eqs. 1–3), followed by a few iterations of local correlation + ConvNeXt/temporal-attention refinement (Eqs. 4–5). The same architecture, without modification, is applied to dense long-range point tracking via sliding-window displacement fields (Sec. 3.5). After a standard multi-stage curriculum the model reports state-of-the-art zero-shot EPE on Sintel Final (1.83 with T=4), Spring, and the large-motion bin s40+, competitive KITTI numbers, and strong zero-shot TAP-Vid accuracy that becomes SOTA after a short Kubric fine-tune.
Significance. If the empirical claims hold, the work supplies a clean, reusable recipe for transferring static multi-view geometric priors into dynamic large-displacement motion estimation and shows that the same representation transfers to long-range point tracking. The ablations (Tables 6–7, S2) isolate the contribution of the frozen prior, feature fusion, and temporal attention; the evaluation suite covers the standard public zero-shot and benchmark protocols without dataset-specific fine-tuning. These results strengthen the case for foundation-model-based motion estimation and a unified dense-tracking paradigm.
major comments (2)
- [§4.4 / Table 5 / Table S4] Table 5 and the supplementary per-sequence breakdown (Table S4) show that the Sintel Final average is dominated by the Ambush 1 outlier (EPE 24). While the paper correctly flags this and reports the average without it, the main-text claim of “state-of-the-art zero-shot performance” on Sintel Final should be qualified more explicitly so that readers do not over-interpret the headline number.
- [§3.1–3.2 / Table 6] The central transfer assumption—that frozen DINOv2/VGGT features already supply sufficiently accurate long-range correspondences for dynamic scenes—is supported by the “w/o Pre-train” collapse in Table 6, yet the paper never quantifies how much of the residual large-motion error is still attributable to domain shift between static multi-view pre-training and optical-flow dynamics. A short diagnostic (e.g., matching accuracy of the raw global features before refinement on s40+ pixels) would make the claim more falsifiable.
minor comments (4)
- [Fig. 1 / Tables 1–5] Fig. 1 caption and several table headers use inconsistent abbreviations (EPE vs. Fl-epe vs. Fl-all); a single glossary would help.
- [Table 6] The latency and memory numbers in Table 6 are measured at 540×960 on an RTX 4090; stating the corresponding numbers for the default 4-frame 432×960 inference used in the main results would make the efficiency claims easier to compare.
- [§3.4 Eq. (6)] In Eq. (6) the smooth-L1 term is written with a double vertical bar that is not defined; a short note that it is the standard smooth-ℓ1 would remove ambiguity.
- [Abstract / throughout] A few typographical slips remain (e.g., “a few lightweight iterative refinement” in the abstract, missing spaces around citations). A final proof-reading pass is warranted.
Circularity Check
No significant circularity: empirical architecture + external public-benchmark evaluation; no prediction reduces to a fitted input or self-definition by construction.
full rationale
MegaFlow is an empirical computer-vision architecture paper. Its load-bearing claims are (1) that frozen DINOv2/VGGT global features plus all-pairs global matching (Eqs. 1–3) plus a few local iterative refinements (Eqs. 4–5) yield accurate large-displacement flow, and (2) that the same model transfers zero-shot to TAP-Vid point tracking. Both claims are validated exclusively on held-out public benchmarks (Sintel train/test, KITTI train/test, Spring, TAP-Vid) that never enter the training loss or hyper-parameter search. Training follows the standard multi-stage curriculum on FlyingChairs/TartanAir/FlyingThings/mixed data with ordinary smooth-L1 supervision (Eq. 6); no free parameter is fitted to a subset of a target metric and then re-reported as a “prediction.” Ablations (Tables 6–7, S2) isolate the contribution of the pre-trained prior, feature fusion and temporal attention by removing them and measuring the resulting degradation on the same external sets. Self-citations (GMFlow, SEA-RAFT, VGGT) supply architectural components or baselines, not uniqueness theorems or load-bearing uniqueness results that force the reported numbers. Consequently the derivation chain contains no self-definitional loop, no fitted-input-called-prediction, and no self-citation that collapses the central claim.
Axiom & Free-Parameter Ledger
free parameters (5)
- refinement iterations K =
8 (eval)
- loss weight gamma =
0.9
- Transformer depth L and feature dim D =
L=24, D=128
- local correlation radius r =
4
- training curriculum lengths and learning rates =
see Table S3
axioms (3)
- domain assumption Pre-trained DINOv2/VGGT features already contain usable long-range geometric correspondences for dynamic scenes.
- domain assumption Softmax expectation over all-pairs correlation yields a sufficiently accurate initial flow for subsequent local refinement.
- domain assumption Standard optical-flow evaluation metrics (EPE, Fl-all, δavg) and public benchmarks are faithful proxies for real-world large-displacement performance.
invented entities (1)
-
MegaFlow architecture
no independent evidence
read the original abstract
Accurate estimation of large displacement optical flow remains a critical challenge. Existing methods typically rely on iterative local search or/and domain-specific fine-tuning, which severely limits their performance in large displacement and zero-shot generalization scenarios. To overcome this, we introduce MegaFlow, a simple yet powerful model for zero-shot large displacement optical flow. Rather than relying on highly complex, task-specific architectural designs, MegaFlow adapts powerful pre-trained vision priors to produce temporally consistent motion fields. In particular, we formulate flow estimation as a global matching problem by leveraging pre-trained global Vision Transformer features, which naturally capture large displacements. This is followed by a few lightweight iterative refinements to further improve the sub-pixel accuracy. Extensive experiments demonstrate that MegaFlow achieves state-of-the-art zero-shot performance across multiple optical flow benchmarks. Moreover, our model also delivers highly competitive zero-shot performance on long-range point tracking benchmarks, demonstrating its robust transferability and suggesting a unified paradigm for generalizable motion estimation. Our project page is at: https://kristen-z.github.io/projects/megaflow.
Forward citations
Cited by 3 Pith papers
-
On the Real-World Generalisability of Optical Flow Models
Progress on Sintel, KITTI and Spring only weakly predicts real-world optical-flow accuracy; lighting and large displacements matter most, and extra synthetic data does not close the gap.
-
FlowPainter: Inpainting Optical Flow via Confidence-Guided Completion
Confidence-guided soft inpainting lets a lightweight flow prior stabilize and accelerate diffusion-based optical flow, yielding stronger results on Sintel, KITTI, and Spring with fewer training iterations.
-
Turning Video Models into Generalist Robot Policies
Decouples action-free video world models from embodiment-specific IDMs using Jacobian-based translation to achieve zero-shot cross-embodiment robot policies.
Reference graph
Works this paper leans on
-
[1]
In: The Thirteenth International Conference on Learning Representations (2025) 3
Aydemir, G., Cai, X., Xie, W., Güney, F.: Track-on: Transformer-based online point tracking with memory. In: The Thirteenth International Conference on Learning Representations (2025) 3
2025
-
[2]
Aydemir, G., Xie, W., Güney, F.: Can visual foundation models achieve long-term point tracking? arXiv preprint arXiv:2408.13575 (2024) 3
Pith/arXiv arXiv 2024
-
[3]
arXiv preprint arXiv:2509.19115 (2025) 3
Aydemir, G., Xie, W., Güney, F.: Track-on2: Enhancing online point tracking with memory. arXiv preprint arXiv:2509.19115 (2025) 3
arXiv 2025
-
[4]
arXiv preprint arXiv:2506.23151 (2025) 3, 7
Bargatin, V., Chistov, E., Yakovenko, A., Vatolin, D.: Memfof: High-resolution training for memory-efficient multi-frame optical flow estimation. arXiv preprint arXiv:2506.23151 (2025) 3, 7
Pith/arXiv arXiv 2025
-
[5]
In: ECCV
Butler, D.J., Wulff, J., Stanley, G.B., Black, M.J.: A naturalistic open source movie for optical flow evaluation. In: ECCV. pp. 611–625. Springer (2012) 7, 23, 24
2012
-
[6]
arXiv preprint arXiv:2407.15420 (2024) 3, 10
Cho, S., Huang, J., Nam, J., An, H., Kim, S., Lee, J.Y.: Local all-pair correspondence for point tracking. arXiv preprint arXiv:2407.15420 (2024) 3, 10
Pith/arXiv arXiv 2024
-
[7]
In: 2009 IEEE conference on computer vision and pattern recognition
Deng, J., Dong, W., Socher, R., Li, L.J., Li, K., Fei-Fei, L.: Imagenet: A large-scale hierarchical image database. In: 2009 IEEE conference on computer vision and pattern recognition. pp. 248–255. Ieee (2009) 7
2009
-
[8]
Advances in Neural Information Processing Systems35, 13610–13626 (2022) 3, 9, 10, 20, 21
Doersch, C., Gupta, A., Markeeva, L., Recasens, A., Smaira, L., Aytar, Y., Carreira, J., Zisserman, A., Yang, Y.: Tap-vid: A benchmark for tracking any point in a video. Advances in Neural Information Processing Systems35, 13610–13626 (2022) 3, 9, 10, 20, 21
2022
-
[9]
In: Proceedings of the Asian Conference on Computer Vision
Doersch, C., Luc, P., Yang, Y., Gokay, D., Koppula, S., Gupta, A., Heyward, J., Rocco, I., Goroshin, R., Carreira, J., et al.: Bootstap: Bootstrapped training for tracking-any-point. In: Proceedings of the Asian Conference on Computer Vision. pp. 3257–3274 (2024) 3, 10
2024
-
[10]
In: CVPR
Dong, Q., Fu, Y.: Memflow: Optical flow estimation and prediction with memory. In: CVPR. pp. 19068–19078 (2024) 3, 8, 9, 10, 11, 12, 20, 21, 24
2024
-
[11]
In: ICCV
Dosovitskiy, A., Fischer, P., Ilg, E., Hausser, P., Hazirbas, C., Golkov, V., Van Der Smagt, P., Cremers, D., Brox, T.: Flownet: Learning optical flow with convolu- tional networks. In: ICCV. pp. 2758–2766 (2015) 2, 3, 7, 23
2015
-
[12]
In: CVPR
Geiger, A., Lenz, P., Urtasun, R.: Are we ready for autonomous driving? the kitti vision benchmark suite. In: CVPR. pp. 3354–3361. IEEE (2012) 2, 7, 23, 24 16 D. Zhang et al
2012
-
[13]
In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition
Greff, K., Belletti, F., Beyer, L., Doersch, C., Du, Y., Duckworth, D., Fleet, D.J., Gnanapragasam, D., Golemo, F., Herrmann, C., et al.: Kubric: A scalable dataset generator. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. pp. 3749–3761 (2022) 10, 23
2022
-
[14]
In: ICCV (2025) 3, 6, 9, 10
Harley, A.W., You, Y., Sun, X., Zheng, Y., Raghuraman, N., Gu, Y., Liang, S., Chu, W.H., Dave, A., Tokmakov, P., You, S., Ambrus, R., Fragkiadaki, K., Guibas, L.J.: AllTracker: Efficient dense point tracking at high resolution. In: ICCV (2025) 3, 6, 9, 10
2025
-
[15]
In: CVPR
He, K., Zhang, X., Ren, S., Sun, J.: Deep residual learning for image recognition. In: CVPR. pp. 770–778 (2016) 4, 7
2016
-
[16]
Artificial intelligence17(1-3), 185–203 (1981) 2, 3
Horn, B.K., Schunck, B.G.: Determining optical flow. Artificial intelligence17(1-3), 185–203 (1981) 2, 3
1981
-
[17]
In: ECCV
Huang, Z., Shi, X., Zhang, C., Wang, Q., Cheung, K.C., Qin, H., Dai, J., Li, H.: Flowformer: A transformer architecture for optical flow. In: ECCV. pp. 668–685. Springer (2022) 2, 3, 8, 11, 12
2022
-
[18]
In: ICCV
Jiang, S., Campbell, D., Lu, Y., Li, H., Hartley, R.: Learning to estimate hidden motions with global motion aggregation. In: ICCV. pp. 9772–9781 (2021) 2, 3, 8, 11, 12
2021
-
[19]
In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recog- nition
Jung, H., Hui, Z., Luo, L., Yang, H., Liu, F., Yoo, S., Ranjan, R., Demandolx, D.: Anyflow: Arbitrary scale optical flow with implicit neural representation. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recog- nition. pp. 5455–5465 (2023) 8, 12
2023
-
[20]
Karaev, N., Makarov, I., Wang, J., Neverova, N., Vedaldi, A., Rupprecht, C.: CoTracker3: Simpler and better point tracking by pseudo-labelling real videos (2024) 3, 10
2024
-
[21]
In: European conference on computer vision
Karaev, N., Rocco, I., Graham, B., Neverova, N., Vedaldi, A., Rupprecht, C.: Cotracker: It is better to track together. In: European conference on computer vision. pp. 18–35. Springer (2024) 3, 10
2024
-
[22]
arXiv preprint arXiv:2509.13414 (2025) 2, 3, 22
Keetha, N., Müller, N., Schönberger, J., Porzi, L., Zhang, Y., Fischer, T., Knapitsch, A., Zauss, D., Weber, E., Antunes, N., et al.: Mapanything: Universal feed-forward metric 3d reconstruction. arXiv preprint arXiv:2509.13414 (2025) 2, 3, 22
Pith/arXiv arXiv 2025
-
[23]
In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops
Kondermann, D., Nair, R., Honauer, K., Krispin, K., Andrulis, J., Brock, A., Gussefeld, B., Rahimimoghaddam, M., Hofmann, S., Brenner, C., et al.: The hci benchmark suite: Stereo and flow ground truth with uncertainties for urban autonomous driving. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops. pp. 19–28 (201...
2016
-
[24]
arXiv preprint arXiv:2602.04877 (2026) 3
Lai, Z., Insafutdinov, E., Sucar, E., Vedaldi, A.: Cowtracker: Tracking by warping instead of correlation. arXiv preprint arXiv:2602.04877 (2026) 3
arXiv 2026
-
[25]
arXiv preprint arXiv:2507.10065 (2025) 2
Lin, C., Lin, Y., Pan, P., Yu, Y., Yan, H., Fragkiadaki, K., Mu, Y.: Movies: Motion- aware 4d dynamic view synthesis in one second. arXiv preprint arXiv:2507.10065 (2025) 2
arXiv 2025
-
[26]
In: The Fourteenth International Conference on Learning Representations 3
Liu, J., Liu, M., Zhu, S., Zhang, Y., Li, J., Yang, M.Y., Nex, F., Cheng, H., Wang, H.: Arflow: Auto-regressive optical flow estimation for arbitrary-length videos via progressive next-frame forecasting. In: The Fourteenth International Conference on Learning Representations 3
-
[27]
In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition
Liu, Z., Mao, H., Wu, C.Y., Feichtenhofer, C., Darrell, T., Xie, S.: A convnet for the 2020s. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. pp. 11976–11986 (2022) 7
2022
-
[28]
arXiv preprint arXiv:1711.05101 (2017) 7, 23 MegaFlow: Zero-Shot Large Displacement Optical Flow 17
Loshchilov, I., Hutter, F.: Decoupled weight decay regularization. arXiv preprint arXiv:1711.05101 (2017) 7, 23 MegaFlow: Zero-Shot Large Displacement Optical Flow 17
Pith/arXiv arXiv 2017
-
[29]
In: IJCAI’81: 7th international joint conference on Artificial intelligence
Lucas, B.D., Kanade, T.: An iterative image registration technique with an applica- tion to stereo vision. In: IJCAI’81: 7th international joint conference on Artificial intelligence. vol. 2, pp. 674–679 (1981) 2, 3
1981
-
[30]
In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
Luo, A., Li, X., Yang, F., Liu, J., Fan, H., Liu, S.: Flowdiffuser: Advancing optical flow estimation with diffusion models. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 19167–19176 (2024) 3, 8, 12
2024
-
[31]
In: European Conference on Computer Vision
Ma, Z., Teed, Z., Deng, J.: Multiview stereo with cascaded epipolar raft. In: European Conference on Computer Vision. pp. 734–750. Springer (2022) 2
2022
-
[32]
Mayer, N., Ilg, E., Häusser, P., Fischer, P., Cremers, D., Dosovitskiy, A., Brox, T.: A large dataset to train convolutional networks for disparity, optical flow, and scene flow estimation. In: IEEE International Conference on Computer Vision and Pattern Recognition (CVPR) (2016),http://lmb.informatik.uni-freiburg.de/ Publications/2016/MIFDB16, arXiv:1512...
Pith/arXiv arXiv 2016
-
[33]
In: CVPR
Mehl, L., Schmalfuss, J., Jahedi, A., Nalivayko, Y., Bruhn, A.: Spring: A high- resolution high-detail dataset and benchmark for scene flow, optical flow and stereo. In: CVPR. pp. 4981–4991 (2023) 2, 7
2023
-
[34]
In: CVPR
Menze, M., Geiger, A.: Object scene flow for autonomous vehicles. In: CVPR. pp. 3061–3070 (2015) 2
2015
-
[35]
In: AAAI
Morimitsu, H., Zhu, X., Ji, X., Yin, X.C.: Recurrent partial kernel network for efficient optical flow estimation. In: AAAI. vol. 38, pp. 4278–4286 (2024) 8, 11, 12
2024
-
[36]
arXiv preprint arXiv:2410.24211 (2024) 3
Ngo, T.D., Zhuang, P., Gan, C., Kalogerakis, E., Tulyakov, S., Lee, H.Y., Wang, C.: Delta: Dense efficient long-range 3d tracking for any video. arXiv preprint arXiv:2410.24211 (2024) 3
Pith/arXiv arXiv 2024
-
[37]
Oquab, M., Darcet, T., Moutakanni, T., Vo, H.V., Szafraniec, M., Khalidov, V., Fernandez, P., Haziza, D., Massa, F., El-Nouby, A., Howes, R., Huang, P.Y., Xu, H., Sharma, V., Li, S.W., Galuba, W., Rabbat, M., Assran, M., Ballas, N., Synnaeve, G., Misra, I., Jegou, H., Mairal, J., Labatut, P., Joulin, A., Bojanowski, P.: Dinov2: Learning robust visual feat...
2023
-
[38]
ArXiv (2019) 7
Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., Desmaison, A., Köpf, A., Yang, E., DeVito, Z., Raison, M., Tejani, A., Chilamkurthy, S., Steiner, B., Fang, L., Bai, J., Chintala, S.: Pytorch: An imperative style, high-performance deep learning library. ArXiv (2019) 7
2019
-
[39]
In: CVPR
Perazzi, F., Pont-Tuset, J., McWilliams, B., Van Gool, L., Gross, M., Sorkine- Hornung, A.: A benchmark dataset and evaluation methodology for video object segmentation. In: CVPR. pp. 724–732 (2016) 20, 21
2016
-
[40]
In: Proceedings of the IEEE/CVF international conference on computer vision
Ranftl, R., Bochkovskiy, A., Koltun, V.: Vision transformers for dense prediction. In: Proceedings of the IEEE/CVF international conference on computer vision. pp. 12179–12188 (2021) 4, 22
2021
-
[41]
In: CVPR
Ranjan, A., Black, M.J.: Optical flow estimation using a spatial pyramid network. In: CVPR. pp. 4161–4170 (2017) 2, 3
2017
-
[42]
In: ICCV
Richter, S.R., Hayder, Z., Koltun, V.: Playing for benchmarks. In: ICCV. pp. 2213–2222 (2017) 7
2017
-
[43]
International journal of computer vision80(1), 72–91 (2008) 3
Sand, P., Teller, S.: Particle video: Long-range motion estimation using point trajectories. International journal of computer vision80(1), 72–91 (2008) 3
2008
-
[44]
Saxena, S., Herrmann, C., Hur, J., Kar, A., Norouzi, M., Sun, D., Fleet, D.J.: The surprising effectiveness of diffusion models for optical flow and monocular depth estimation36, 39443–39469 (2023) 3, 12, 24
2023
-
[45]
Advances in Neural Information Processing Systems37, 68658–68685 (2024) 7, 22 18 D
Shah,J.,Bikshandi,G.,Zhang,Y.,Thakkar,V.,Ramani,P.,Dao,T.:Flashattention- 3: Fast and accurate attention with asynchrony and low-precision. Advances in Neural Information Processing Systems37, 68658–68685 (2024) 7, 22 18 D. Zhang et al
2024
-
[46]
In: Proceedings of the IEEE conference on computer vision and pattern recognition
Shi, W., Caballero, J., Huszár, F., Totz, J., Aitken, A.P., Bishop, R., Rueckert, D., Wang, Z.: Real-time single image and video super-resolution using an efficient sub-pixel convolutional neural network. In: Proceedings of the IEEE conference on computer vision and pattern recognition. pp. 1874–1883 (2016) 4, 23
2016
-
[47]
In: ICCV
Shi, X., Huang, Z., Bian, W., Li, D., Zhang, M., Cheung, K.C., See, S., Qin, H., Dai, J., Li, H.: Videoflow: Exploiting temporal cues for multi-frame optical flow estimation. In: ICCV. pp. 12469–12480 (2023) 2, 3, 8, 12
2023
-
[48]
In: CVPR
Shi, X., Huang, Z., Li, D., Zhang, M., Cheung, K.C., See, S., Qin, H., Dai, J., Li, H.: Flowformer++: Masked cost volume autoencoding for pretraining optical flow estimation. In: CVPR. pp. 1599–1610 (2023) 2, 3, 8
2023
-
[49]
Siméoni, O., Vo, H.V., Seitzer, M., Baldassarre, F., Oquab, M., Jose, C., Khalidov, V., Szafraniec, M., Yi, S., Ramamonjisoa, M., Massa, F., Haziza, D., Wehrstedt, L., Wang, J., Darcet, T., Moutakanni, T., Sentana, L., Roberts, C., Vedaldi, A., Tolan, J., Brandt, J., Couprie, C., Mairal, J., Jégou, H., Labatut, P., Bojanowski, P.: DINOv3 (2025),https://ar...
Pith/arXiv arXiv 2025
-
[50]
In: Artificial intelligence and machine learning for multi- domain operations applications
Smith, L.N., Topin, N.: Super-convergence: Very fast training of neural networks using large learning rates. In: Artificial intelligence and machine learning for multi- domain operations applications. vol. 11006, pp. 369–386. SPIE (2019) 23
2019
-
[51]
In: 2010 IEEE computer society conference on computer vision and pattern recognition
Sun, D., Roth, S., Black, M.J.: Secrets of optical flow estimation and their princi- ples. In: 2010 IEEE computer society conference on computer vision and pattern recognition. pp. 2432–2439. IEEE (2010) 2
2010
-
[52]
In: CVPR
Sun, D., Yang, X., Liu, M.Y., Kautz, J.: Pwc-net: Cnns for optical flow using pyramid, warping, and cost volume. In: CVPR. pp. 8934–8943 (2018) 2, 3, 8, 11, 12
2018
-
[53]
Sun, S., Liu, J., Li, H., Liu, G., Li, T., Gao, W.: Streamflow: streamlined multi-frame optical flow estimation for video sequences37, 9205–9228 (2025) 2, 3, 8, 11, 12
2025
-
[54]
In: European conference on computer vision
Teed, Z., Deng, J.: Raft: Recurrent all-pairs field transforms for optical flow. In: European conference on computer vision. pp. 402–419. Springer (2020) 2, 3, 6, 7, 8, 10, 11, 12, 21, 23
2020
-
[55]
In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition
Truong, P., Danelljan, M., Timofte, R.: Glu-net: Global-local universal network for dense flow and correspondences. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. pp. 6258–6268 (2020) 2
2020
-
[56]
In: 2024 IEEE International Conference on Robotics and Automation (ICRA)
Vecerik, M., Doersch, C., Yang, Y., Davchev, T., Aytar, Y., Zhou, G., Hadsell, R., Agapito, L., Scholz, J.: Robotap: Tracking arbitrary points for few-shot visual imitation. In: 2024 IEEE International Conference on Robotics and Automation (ICRA). pp. 5397–5403. IEEE (2024) 20, 21
2024
-
[57]
In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (2025) 2, 3, 4, 7, 22
Wang, J., Chen, M., Karaev, N., Vedaldi, A., Rupprecht, C., Novotny, D.: Vggt: Vi- sual geometry grounded transformer. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (2025) 2, 3, 4, 7, 22
2025
-
[58]
Wang, W., Zhu, D., Wang, X., Hu, Y., Qiu, Y., Wang, C., Hu, Y., Kapoor, A., Scherer, S.: Tartanair: A dataset to push the limits of visual slam (2020) 7, 21, 23
2020
-
[59]
Wang, Y., Zhou, J., Zhu, H., Chang, W., Zhou, Y., Li, Z., Chen, J., Pang, J., Shen, C., He, T.:π3: Scalable permutation-equivariant visual geometry learning (2025), https://arxiv.org/abs/2507.133472
Pith/arXiv arXiv 2025
-
[60]
arXiv preprint arXiv:2506.21526 (2025) 2, 3, 8, 9, 10, 11, 12, 13, 20, 21, 24
Wang, Y., Deng, J.: Waft: Warping-alone field transforms for optical flow. arXiv preprint arXiv:2506.21526 (2025) 2, 3, 8, 9, 10, 11, 12, 13, 20, 21, 24
arXiv 2025
-
[61]
arXiv preprint arXiv:2405.14793 (2024) 2, 3, 7, 8, 9, 10, 11, 12, 21, 23, 24
Wang, Y., Lipson, L., Deng, J.: Sea-raft: Simple, efficient, accurate raft for optical flow. arXiv preprint arXiv:2405.14793 (2024) 2, 3, 7, 8, 9, 10, 11, 12, 21, 23, 24
Pith/arXiv arXiv 2024
-
[62]
In: ICCV
Weinzaepfel, P., Lucas, T., Leroy, V., Cabon, Y., Arora, V., Brégier, R., Csurka, G., Antsfeld, L., Chidlovskii, B., Revaud, J.: Croco v2: Improved cross-view completion pre-training for stereo matching and optical flow. In: ICCV. pp. 17969–17980 (2023) 2, 3, 11, 12, 13 MegaFlow: Zero-Shot Large Displacement Optical Flow 19
2023
-
[63]
CVPR (2025) 3
Wen, B., Trepte, M., Aribido, J., Kautz, J., Gallo, O., Birchfield, S.: Foundation- stereo: Zero-shot stereo matching. CVPR (2025) 3
2025
-
[64]
In: Proceedings of the IEEE/CVF International Conference on Computer Vision
Wu, G., Liu, X., Luo, K., Liu, X., Zheng, Q., Liu, S., Jiang, X., Zhai, G., Wang, W.: Accflow: Backward accumulation for long-range optical flow. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. pp. 12119–12128 (2023) 3, 10, 21
2023
-
[65]
arXiv preprint arXiv:2312.03790 (2023) 3
Xu, G., Chen, S., Jia, H., Feng, M., Yang, X.: Memory-efficient optical flow via radius-distribution orthogonal cost volume. arXiv preprint arXiv:2312.03790 (2023) 3
Pith/arXiv arXiv 2023
-
[66]
In: CVPR
Xu, H., Zhang, J., Cai, J., Rezatofighi, H., Tao, D.: Gmflow: Learning optical flow via global matching. In: CVPR. pp. 8121–8130 (2022) 2, 3, 5, 7, 8
2022
-
[67]
IEEE Transactions on Pattern Analysis and Machine Intelligence (2023) 2, 3, 5, 7
Xu, H., Zhang, J., Cai, J., Rezatofighi, H., Yu, F., Tao, D., Geiger, A.: Unifying flow, stereo and depth estimation. IEEE Transactions on Pattern Analysis and Machine Intelligence (2023) 2, 3, 5, 7
2023
-
[68]
Yang, L., Kang, B., Huang, Z., Zhao, Z., Xu, X., Feng, J., Zhao, H.: Depth anything v2. arXiv:2406.09414 (2024) 3, 22
Pith/arXiv arXiv 2024
-
[69]
arXiv preprint arXiv:2506.09278 (2025) 2, 3, 8, 9
Zhang, Y., Keetha, N., Lyu, C., Jhamb, B., Chen, Y., Qiu, Y., Karhade, J., Jha, S., Hu, Y., Ramanan, D., et al.: Ufm: A simple path towards unified dense correspondence with flow. arXiv preprint arXiv:2506.09278 (2025) 2, 3, 8, 9
arXiv 2025
-
[70]
In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
Zhao, S., Zhao, L., Zhang, Z., Zhou, E., Metaxas, D.: Global matching with overlapping attention for optical flow estimation. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 17592–17601 (2022) 2
2022
-
[71]
In: ICCV (2023) 3, 10
Zheng, Y., Harley, A.W., Shen, B., Wetzstein, G., Guibas, L.J.: Pointodyssey: A large-scale synthetic dataset for long-term point tracking. In: ICCV (2023) 3, 10
2023
-
[72]
In: CVPR
Zheng, Z., Nie, N., Ling, Z., Xiong, P., Liu, J., Wang, H., Li, J.: Dip: Deep inverse patchmatch for high-resolution optical flow. In: CVPR. pp. 8925–8934 (2022) 3, 8
2022
-
[73]
In: Proceedings of the AAAI Conference on Artificial Intelligence
Zhou, S., He, R., Tan, W., Yan, B.: Samflow: Eliminating any fragmentation in optical flow with segment anything model. In: Proceedings of the AAAI Conference on Artificial Intelligence. vol. 38, pp. 7695–7703 (2024) 3, 8, 9, 12, 13, 24 20 D. Zhang et al. Appendix In this supplementary material, we provide additional quantitative and qualitative results, ...
2024
-
[74]
benchmarks in Tab. S1. Among models trained exclusively on optical flow datasets, MegaFlow establishes state-of-the-art average position accuracy (δavg). Notably, it significantly outperforms recent strong baselines, including WAFT [60] and MemFlow-T [10] on the zero-shot point tracking benchmarks. We provide additional qualitative long-range trajectory v...
-
[75]
We evaluate the average position accuracy (δavg) using an input resolution of384 × 512
benchmarks. We evaluate the average position accuracy (δavg) using an input resolution of384 × 512. Among models trained exclusively on optical flow datasets, MegaFlow establishes new state-of-the-art results. Method DAVIS↑Kinetics↑RGB-Stacking↑RoboTAP↑Mean↑ RAFT [54] 48.5 64.3 82.8 72.2 67.0 SEA-RAFT [61] 48.7 64.3 85.7 67.6 66.6 AccFlow [64] 23.5 38.8 6...
arXiv 2015
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.