REVIEW 2 major objections 4 minor 50 references
Gains on Sintel, KITTI and Spring stop predicting real-world optical-flow accuracy once models beat RAFT, and more synthetic data does not close the gap.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-14 11:26 UTC pith:7CDCDOBB
load-bearing objection Useful new factorized real-world benchmark and clear evidence that post-RAFT gains on Sintel/KITTI/Spring largely fail to transfer; sparse 5-point proxy is the main soft spot but not fatal. the 2 major comments →
On the Real-World Generalisability of Optical Flow Models
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
Progress on the standard optical-flow proxies (Sintel, KITTI, Spring) only weakly predicts accuracy on real-world video once models surpass RAFT; further proxy gains stagnate or reverse on a new real-world suite of 8,204 frame pairs, and simply scaling synthetic pre-training data does not close the gap.
What carries the argument
FlowFactor: a 1,000-pair, sparsely annotated real-world benchmark that isolates four classical failure factors (large displacements, repetitive textures, occlusions, lighting variation) so that each can be measured separately rather than averaged together.
Load-bearing premise
The claim that five hand-picked points per frame pair are enough to rank models the same way dense labels would, rests on a single Spearman check performed on two non-overlapping point sets drawn from TAP-Vid-DAVIS.
What would settle it
Re-annotate a non-trivial subset of FlowFactor or Slow Flow with dense or much denser labels and check whether model rankings and the reported factor-to-real-world correlations reverse or disappear.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper studies whether accuracy of modern optical flow models on standard synthetic and domain-specific benchmarks (Sintel, KITTI, Spring) predicts real-world performance. It introduces FlowFactor (1 000 sparsely annotated HD pairs factorised into large displacements, repetitive textures, occlusions and lighting variation), repurposes TAP-Vid-DAVIS into TAP-Flow, and combines both with Slow Flow into an 8 204-pair real-world suite. Using standard Things (C+T) checkpoints of a broad model set (after reproduction filtering), the authors report that post-RAFT gains on the proxies stagnate or reverse on the real-world suite, that lighting and large-displacement accuracy correlate most strongly with real-world scores, that large-motion optimisation can trade off against stationary photometric robustness, and that simply scaling synthetic pre-training data does not close the gap.
Significance. If the empirical findings hold, the work is a useful contribution to optical-flow evaluation practice. FlowFactor supplies a controlled, real-world diagnostic resource that isolates classical failure modes more cleanly than existing aggregate benchmarks; the public code and datasets, the transparent reproduction protocol against published Sintel/KITTI numbers, and the factor-wise Spearman analyses give the community concrete tools and evidence that continued proxy progress may not translate to real-world utility and that pure data scaling is insufficient. The paper therefore both documents a generalisation gap and supplies a practical instrument for measuring progress toward closing it.
major comments (2)
- [§3.1] The central claim that five sparsely chosen points constitute a reliable proxy for relative model ranking on an entire scene is validated only by a Spearman ρ=0.95 experiment on two non-overlapping 5-point subsets drawn from the first two frames of each TAP-Vid-DAVIS video (60 pairs total). That experiment does not cover FlowFactor’s deliberately extreme, single-factor regimes (identically zero-flow lighting changes, highly ambiguous repetitive textures, large displacements, occlusions). Because FlowFactor forms a substantial fraction of the 8 204-pair suite that underpins the post-RAFT stagnation plots (Fig. 7), the category–Slow-Flow correlations (Fig. 8) and the trade-off heat-maps (Fig. 9), the diagnostic and “weakly predicts” conclusions remain sensitive to point-selection bias. A modest additional validation—e.g., denser multi-annotator sampling or ranking-consistency checks on a F
- [§5.1, Fig. 7] The restriction to models that simultaneously outperform RAFT on the training splits of KITTI-15, Sintel-Final and Spring produces a small cloud of points; the OLS fits that are said to demonstrate stagnation or degradation (Fig. 7) are therefore sensitive both to the precise definition of “outperform” and to which models happen to clear all three thresholds. A continuous analysis (proxy score versus real-world score without hard thresholding) or a leave-one-out sensitivity check would make the claim that further proxy gains cease to predict real-world accuracy more convincing.
minor comments (4)
- [Fig. 6] Fig. 6 would be clearer with ground-truth flow or endpoint-error overlays so that the claimed superiority of GMA over SEA-RAFT under lighting and occlusion can be assessed quantitatively rather than by visual inspection alone.
- [§4.1] The exact number of models retained after the >2.0 EPE / Fl-all reproduction filter (§4.1) is never stated; a short sentence or a column in Appendix Table 1 would help readers gauge the breadth of the final comparison.
- [Abstract / §1] Several extracted passages contain concatenated words (“withreal-worldaccuracy”, “stronglywithreal-world”); these are presumably PDF-extraction artefacts but should be checked in the camera-ready source.
- [§3.1 Lighting Variation] The lighting category uses an identically zero ground-truth field; reporting the complementary mean EPE (or fraction of non-zero predictions) alongside Fl-all would make the photometric-robustness ranking easier to interpret under the 3 px / 5 % threshold.
Circularity Check
Purely empirical evaluation paper with no derivation chain that reduces predictions or claims to fitted inputs or self-definitional constructions.
full rationale
The paper constructs real-world benchmarks (FlowFactor with sparse manual annotations, TAP-Flow repurposed from TAP-Vid-DAVIS, Slow Flow) and evaluates standard Things (C+T) checkpoints of existing optical flow models using Fl-all. All reported results (OLS fits of real-world Fl-all vs. Sintel/KITTI/Spring, Spearman correlations of FlowFactor categories vs. Slow Flow/TAP-Flow, trade-offs among categories, training-set-size plots) are direct empirical measurements on held-out data never used to train or fit the evaluated models. No parameter is fitted to a subset and then re-presented as an independent prediction; no uniqueness theorem or ansatz is imported via self-citation to force a result; no quantity is defined in terms of the quantity it claims to predict. The 5-point proxy validation (Spearman ho=0.95 on TAP-Vid subsets) is an empirical check of ranking stability, not a circular definition. Self-citations are limited to standard datasets and annotation tooling and are not load-bearing for any central claim. The work is therefore free of the enumerated circularity patterns.
Axiom & Free-Parameter Ledger
free parameters (3)
- Fl-all correctness threshold =
3 px / 5 %
- large-displacement cutoff =
25 px
- points per scene =
5
axioms (3)
- domain assumption The four classical optical-flow failure modes (large displacement, occlusion, repetitive texture, lighting change) can be approximately isolated in real imagery so that each FlowFactor subset mainly varies only one factor.
- domain assumption Standard Things (C+T) checkpoints supplied by model authors constitute a fair, comparable measure of out-of-distribution generalisation.
- ad hoc to paper Five sparse points are a sufficient proxy for relative model ranking on an entire scene.
invented entities (1)
-
FlowFactor dataset
independent evidence
read the original abstract
Real-world deployment of vision models to broadly benefit society is arguably a main research objective. In optical flow, however, the difficulty to obtain the ground truth has focused research mainly on synthetic data and domain-specific benchmarks. Here, we investigate the severity of this mismatch. We study how well modern optical flow estimation models generalise to real-world video and question if accuracy on synthetic benchmark proxies actually predicts accuracy on real-world optical flow. To address this, we build a real-world evaluation benchmark and evaluate the real-world generalisability of a broad set of recent optical flow models using standard checkpoints. Our benchmark contains 8,204 frame pairs across TAP-Flow, Slow Flow, and our own dataset FlowFactor. FlowFactor is a manually annotated real-world benchmark of 1,000 HD frame pairs organised into four confounding factors: large displacements, repetitive textures, occlusions, and lighting variation. Each setting mainly varies only one factor, enabling diagnostic, confounder-specific analysis. Using FlowFactor, we reveal that performance on varying lighting and large displacements correlates most strongly with real-world accuracy, and that improvements on large-motion regimes can trade off against robustness in small-motion, stationary scenes. Our experiments show that progress on Sintel, KITTI and Spring only weakly predicts accuracy on real-world data, highlighting the need for a broad real-world optical flow benchmark. Interestingly, scaling up the amount of training data does not necessarily resolve the gap, calling for new innovative research instead of simply scaling data and compute.
Figures
Reference graph
Works this paper leans on
-
[1]
Robust vision challenge.http://www.robustvision.net/, accessed: 2025-11-13
2025
-
[2]
Transactions on Machine Learning Research (2025)
Agnihotri, S., Caspary, J.Y., Schwarz, L., Gao, X., Schmalfuss, J., Bruhn, A., Keu- per, M.: Flowbench: Benchmarking optical flow estimation methods for reliability and generalization. Transactions on Machine Learning Research (2025)
2025
-
[3]
Computer vision and image understanding 249, 104160 (2024)
Alfarano, A., Maiano, L., Papa, L., Amerini, I.: Estimating optical flow: A compre- hensive review of the state of the art. Computer vision and image understanding 249, 104160 (2024)
2024
-
[4]
International journal of computer vision92(1), 1–31 (2011)
Baker, S., Scharstein, D., Lewis, J.P., Roth, S., Black, M.J., Szeliski, R.: A database and evaluation methodology for optical flow. International journal of computer vision92(1), 1–31 (2011)
2011
-
[5]
International journal of computer vision12(1), 43–77 (1994)
Barron, J.L., Fleet, D.J., Beauchemin, S.S.: Performance of optical flow techniques. International journal of computer vision12(1), 43–77 (1994)
1994
-
[6]
Bauer, K., Bruhn, A., Schmalfuss, J.: On the generalization of optical flow: Quan- tifying robustness to dataset shifts
-
[7]
In: European conference on computer vision
Butler, D.J., Wulff, J., Stanley, G.B., Black, M.J.: A naturalistic open source movie for optical flow evaluation. In: European conference on computer vision. pp. 611–
-
[8]
Dahal, S.: Going against the flow: Evaluating optical flow estimation models on real-world non-rigid motion (2025),https://resolver.tudelft.nl/uuid: ab2db1e7-9122-43c3-9377-f188fceb30b6
2025
-
[9]
Advances in Neural Information Processing Systems35, 13610–13626 (2022)
Doersch,C.,Gupta,A.,Markeeva,L.,Recasens,A.,Smaira,L.,Aytar,Y.,Carreira, J.,Zisserman,A.,Yang,Y.:Tap-vid:Abenchmarkfortrackinganypointinavideo. Advances in Neural Information Processing Systems35, 13610–13626 (2022)
2022
-
[10]
In: Proceedings of the IEEE international conference on computer vision
Dosovitskiy, A., Fischer, P., Ilg, E., Hausser, P., Hazirbas, C., Golkov, V., Van Der Smagt, P., Cremers, D., Brox, T.: Flownet: Learning optical flow with convolu- tional networks. In: Proceedings of the IEEE international conference on computer vision. pp. 2758–2766 (2015)
2015
-
[11]
In: 2024 13th Iranian/3rd International Machine Vision and Image Processing Conference (MVIP)
Eslami, N., Arefi, F., Mansourian, A.M., Kasaei, S.: Rethinking raft for efficient optical flow. In: 2024 13th Iranian/3rd International Machine Vision and Image Processing Conference (MVIP). pp. 1–7. IEEE (2024)
2024
-
[12]
Ge, Z.: Real-world evaluation of optical flow with varying lighting condi- tions (2025),https://resolver.tudelft.nl/uuid:7ad41253-902a-4d78-a86a- c7bee94b9e0c
2025
-
[13]
In: 2012 IEEE conference on computer vision and pattern recognition
Geiger, A., Lenz, P., Urtasun, R.: Are we ready for autonomous driving? the kitti vision benchmark suite. In: 2012 IEEE conference on computer vision and pattern recognition. pp. 3354–3361. IEEE (2012)
2012
-
[14]
In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition
Greff, K., Belletti, F., Beyer, L., Doersch, C., Du, Y., Duckworth, D., Fleet, D.J., Gnanapragasam, D., Golemo, F., Herrmann, C., et al.: Kubric: A scalable dataset generator. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. pp. 3749–3761 (2022) 1 https://treesforall.nl/app/uploads/certificates/Certificaat_Trees_for_ ...
2022
-
[15]
Journal of the Optical Society of America A4(8), 1455–1471 (1987)
Heeger, D.J.: Model for the extraction of image flow. Journal of the Optical Society of America A4(8), 1455–1471 (1987)
1987
-
[16]
In: European conference on computer vision
Huang, Z., Shi, X., Zhang, C., Wang, Q., Cheung, K.C., Qin, H., Dai, J., Li, H.: Flowformer: A transformer architecture for optical flow. In: European conference on computer vision. pp. 668–685. Springer (2022)
2022
-
[17]
In: Proceedings of the IEEE conference on computer vision and pattern recognition
Ilg, E., Mayer, N., Saikia, T., Keuper, M., Dosovitskiy, A., Brox, T.: Flownet 2.0: Evolution of optical flow estimation with deep networks. In: Proceedings of the IEEE conference on computer vision and pattern recognition. pp. 2462–2470 (2017)
2017
-
[18]
In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition
Janai, J., Guney, F., Wulff, J., Black, M.J., Geiger, A.: Slow flow: Exploiting high- speed cameras for accurate and diverse optical flow reference data. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. pp. 3597– 3607 (2017)
2017
-
[19]
In: Proceedings of the IEEE/CVF international conference on computer vision
Jiang, S., Campbell, D., Lu, Y., Li, H., Hartley, R.: Learning to estimate hid- den motions with global motion aggregation. In: Proceedings of the IEEE/CVF international conference on computer vision. pp. 9772–9781 (2021)
2021
-
[20]
Klijnsma, J.B.: Real-world evaluation of optical flow on repetitive patterns (2025), https://resolver.tudelft.nl/uuid:dc709377-e116-48eb-91a5-4b6812902034
2025
-
[21]
In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops
Kondermann,D.,Nair,R.,Honauer,K.,Krispin,K.,Andrulis,J.,Brock,A.,Gusse- feld, B., Rahimimoghaddam, M., Hofmann, S., Brenner, C., et al.: The hci bench- mark suite: Stereo and flow ground truth with uncertainties for urban autonomous driving. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops. pp. 19–28 (2016)
2016
-
[22]
In: Proceedings of the IEEE/CVF international conference on computer vision
Li, H., Luo, K., Liu, S.: Gyroflow: Gyroscope-guided unsupervised optical flow learning. In: Proceedings of the IEEE/CVF international conference on computer vision. pp. 12869–12878 (2021)
2021
-
[23]
arXiv preprint arXiv:2506.07740 (2025)
Liang, Y., Fu, Y., Hu, Y., Shao, W., Liu, J., Zhang, D.: Flow-anything: Learn- ing real-world optical flow estimation from large-scale single-view images. arXiv preprint arXiv:2506.07740 (2025)
Pith/arXiv arXiv 2025
-
[24]
In: 2008 IEEE Conference on Computer Vision and Pattern Recognition
Liu, C., Freeman, W.T., Adelson, E.H., Weiss, Y.: Human-assisted motion anno- tation. In: 2008 IEEE Conference on Computer Vision and Pattern Recognition. pp. 1–8. IEEE (2008)
2008
-
[25]
Mayer, N., Ilg, E., Fischer, P., Hazirbas, C., Cremers, D., Dosovitskiy, A., Brox, T.: What makes good synthetic training data for learning disparity and optical flow estimation? International Journal of Computer Vision126(9), 942–960 (2018)
2018
-
[26]
In: Proceedings of the IEEE conference on computer vision and pattern recognition
Mayer, N., Ilg, E., Hausser, P., Fischer, P., Cremers, D., Dosovitskiy, A., Brox, T.: A large dataset to train convolutional networks for disparity, optical flow, and scene flow estimation. In: Proceedings of the IEEE conference on computer vision and pattern recognition. pp. 4040–4048 (2016)
2016
-
[27]
In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
Mehl, L., Schmalfuss, J., Jahedi, A., Nalivayko, Y., Bruhn, A.: Spring: A high- resolutionhigh-detaildatasetandbenchmarkforsceneflow,opticalflowandstereo. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 4981–4991 (2023)
2023
-
[28]
In: Proceedings of the IEEE conference on computer vision and pattern recognition
Menze, M., Geiger, A.: Object scene flow for autonomous vehicles. In: Proceedings of the IEEE conference on computer vision and pattern recognition. pp. 3061–3070 (2015)
2015
-
[29]
Morimitsu, H.: Pytorch lightning optical flow.https://github.com/hmorimitsu/ ptlflow(2021)
2021
-
[30]
Petre, I.A.: Performance of optical flow models on real-world occluded re- gions (2025),https://resolver.tudelft.nl/uuid:c25761a4-939d-4677-b790- b3fb9bc4bc28 On the Real-World Generalisability of Optical Flow Models 17
2025
-
[31]
arXiv preprint arXiv:2505.09368 (2025)
Schmalfuss, J., Oei, V., Mehl, L., Bartsch, M., Agnihotri, S., Keuper, M., Bruhn, A.: Robustspring: Benchmarking robustness to image corruptions for optical flow, scene flow and stereo. arXiv preprint arXiv:2505.09368 (2025)
Pith/arXiv arXiv 2025
-
[32]
In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition
Shi, X., Huang, Z., Li, D., Zhang, M., Cheung, K.C., See, S., Qin, H., Dai, J., Li, H.: Flowformer++: Masked cost volume autoencoding for pretraining optical flow estimation. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. pp. 1599–1610 (2023)
2023
-
[33]
Vision research 44(13), 1565–1573 (2004)
Stürzel, F., Spillmann, L.: Perceptual limits of common fate. Vision research 44(13), 1565–1573 (2004)
2004
-
[34]
In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
Sun, D., Vlasic, D., Herrmann, C., Jampani, V., Krainin, M., Chang, H., Zabih, R., Freeman, W.T., Liu, C.: Autoflow: Learning a better training set for optical flow. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 10093–10102 (2021)
2021
-
[35]
In: Proceedings of the IEEE conference on computer vision and pattern recognition
Sun, D., Yang, X., Liu, M.Y., Kautz, J.: Pwc-net: Cnns for optical flow using pyramid, warping, and cost volume. In: Proceedings of the IEEE conference on computer vision and pattern recognition. pp. 8934–8943 (2018)
2018
-
[36]
In: Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part II 16
Teed, Z., Deng, J.: Raft: Recurrent all-pairs field transforms for optical flow. In: Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part II 16. pp. 402–419. Springer (2020)
2020
-
[37]
nl/uuid:d7c4cc36-6582-47fc-94cd-cc0dc393725b
Timmerije, M.: Bridging the gap: A real-world dataset and evaluation of optical flow models in large displacement scenarios (2025),https://resolver.tudelft. nl/uuid:d7c4cc36-6582-47fc-94cd-cc0dc393725b
2025
-
[38]
In: 2020 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)
Wang, W., Zhu, D., Wang, X., Hu, Y., Qiu, Y., Wang, C., Hu, Y., Kapoor, A., Scherer, S.: Tartanair: A dataset to push the limits of visual slam. In: 2020 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS). pp. 4909–4916. IEEE (2020)
2020
-
[39]
arXiv preprint arXiv:2506.21526 (2025)
Wang, Y., Deng, J.: Waft: Warping-alone field transforms for optical flow. arXiv preprint arXiv:2506.21526 (2025)
arXiv 2025
-
[40]
In: European Conference on Computer Vision
Wang, Y., Lipson, L., Deng, J.: Sea-raft: Simple, efficient, accurate raft for optical flow. In: European Conference on Computer Vision. pp. 36–54. Springer (2024)
2024
-
[41]
Weinzaepfel, P., Revaud, J., Harchaoui, Z., Schmid, C.: Learning to detect motion boundaries.In:ProceedingsoftheIEEEconferenceoncomputervisionandpattern recognition. pp. 2578–2586 (2015)
2015
-
[42]
Electronics12(3), 682 (2023)
Weng, W., Zhu, X., Jing, L., Dong, M.: Attention mechanism trained with small datasets for biomedical image segmentation. Electronics12(3), 682 (2023)
2023
-
[43]
In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition
Xu, H., Zhang, J., Cai, J., Rezatofighi, H., Tao, D.: Gmflow: Learning optical flow via global matching. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. pp. 8121–8130 (2022)
2022
-
[44]
In: Proceedings of the IEEE/CVF conference on com- puter vision and pattern recognition
Yin, Z., Darrell, T., Yu, F.: Hierarchical discrete distribution decomposition for match density estimation. In: Proceedings of the IEEE/CVF conference on com- puter vision and pattern recognition. pp. 6044–6053 (2019)
2019
-
[45]
Pattern Recognition114, 107861 (2021)
Zhai, M., Xiang, X., Lv, N., Kong, X.: Optical flow and scene flow estimation: A survey. Pattern Recognition114, 107861 (2021)
2021
-
[46]
arXiv preprint arXiv:2603.25739 (2026)
Zhang, D., Wang, F., Pollefeys, M., Xu, H.: Megaflow: Zero-shot large displacement optical flow. arXiv preprint arXiv:2603.25739 (2026)
Pith/arXiv arXiv 2026
-
[47]
In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition
Zhao,S.,Sheng,Y.,Dong,Y.,Chang,E.I.,Xu,Y.,etal.:Maskflownet:Asymmetric feature matching with learnable occlusion mask. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. pp. 6278–6287 (2020)
2020
-
[48]
In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
Zhao, S., Zhao, L., Zhang, Z., Zhou, E., Metaxas, D.: Global matching with over- lapping attention for optical flow estimation. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 17592–17601 (2022) 18 P. Reijalt et al
2022
-
[49]
In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition
Zheng, Y., Zhang, M., Lu, F.: Optical flow in the dark. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. pp. 6749–6757 (2020)
2020
-
[50]
Zhou, H., Chang, Y., Shi, Z., Yan, W., Chen, G., Tian, Y., Yan, L.: Adverse weather optical flow: Cumulative homogeneous-heterogeneous adaptation. IEEE Transactions on Pattern Analysis and Machine Intelligence (2024) On the Real-World Generalisability of Optical Flow Models 19 Appendix A Reproducing Reported Checkpoint Results (Things) Table 1:Paper vs. o...
2024
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.