Pith. sign in

REVIEW 2 major objections 4 minor 50 references

Gains on Sintel, KITTI and Spring stop predicting real-world optical-flow accuracy once models beat RAFT, and more synthetic data does not close the gap.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-14 11:26 UTC pith:7CDCDOBB

load-bearing objection Useful new factorized real-world benchmark and clear evidence that post-RAFT gains on Sintel/KITTI/Spring largely fail to transfer; sparse 5-point proxy is the main soft spot but not fatal. the 2 major comments →

arxiv 2607.10470 v1 pith:7CDCDOBB submitted 2026-07-11 cs.CV eess.IV

On the Real-World Generalisability of Optical Flow Models

classification cs.CV eess.IV
keywords optical flowreal-world generalisationbenchmarkFlowFactorsynthetic-to-real gapSintelKITTIfactorised evaluation
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Optical flow models are trained almost entirely on synthetic data because dense real-world motion labels are nearly impossible to obtain. The field therefore treats accuracy on a handful of synthetic or domain-narrow proxies (Sintel, KITTI, Spring) as a stand-in for real-world performance. This paper shows that the stand-in has broken. After assembling an 8,204-pair real-world suite that includes a new factorised benchmark called FlowFactor, the authors find that further reductions on the usual proxies no longer translate into better real-world numbers; once models surpass the earlier RAFT baseline the real-world curve flattens or reverses. Simply adding larger synthetic pre-training sets does not fix the mismatch. The practical consequence is that the community needs a broad real-world evaluation suite if it wants optical flow that actually works outside the lab.

Core claim

Progress on the standard optical-flow proxies (Sintel, KITTI, Spring) only weakly predicts accuracy on real-world video once models surpass RAFT; further proxy gains stagnate or reverse on a new real-world suite of 8,204 frame pairs, and simply scaling synthetic pre-training data does not close the gap.

What carries the argument

FlowFactor: a 1,000-pair, sparsely annotated real-world benchmark that isolates four classical failure factors (large displacements, repetitive textures, occlusions, lighting variation) so that each can be measured separately rather than averaged together.

Load-bearing premise

The claim that five hand-picked points per frame pair are enough to rank models the same way dense labels would, rests on a single Spearman check performed on two non-overlapping point sets drawn from TAP-Vid-DAVIS.

What would settle it

Re-annotate a non-trivial subset of FlowFactor or Slow Flow with dense or much denser labels and check whether model rankings and the reported factor-to-real-world correlations reverse or disappear.

Watch this falsifier — get emailed when new claim-graph text bears on it.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 4 minor

Summary. The paper studies whether accuracy of modern optical flow models on standard synthetic and domain-specific benchmarks (Sintel, KITTI, Spring) predicts real-world performance. It introduces FlowFactor (1 000 sparsely annotated HD pairs factorised into large displacements, repetitive textures, occlusions and lighting variation), repurposes TAP-Vid-DAVIS into TAP-Flow, and combines both with Slow Flow into an 8 204-pair real-world suite. Using standard Things (C+T) checkpoints of a broad model set (after reproduction filtering), the authors report that post-RAFT gains on the proxies stagnate or reverse on the real-world suite, that lighting and large-displacement accuracy correlate most strongly with real-world scores, that large-motion optimisation can trade off against stationary photometric robustness, and that simply scaling synthetic pre-training data does not close the gap.

Significance. If the empirical findings hold, the work is a useful contribution to optical-flow evaluation practice. FlowFactor supplies a controlled, real-world diagnostic resource that isolates classical failure modes more cleanly than existing aggregate benchmarks; the public code and datasets, the transparent reproduction protocol against published Sintel/KITTI numbers, and the factor-wise Spearman analyses give the community concrete tools and evidence that continued proxy progress may not translate to real-world utility and that pure data scaling is insufficient. The paper therefore both documents a generalisation gap and supplies a practical instrument for measuring progress toward closing it.

major comments (2)
  1. [§3.1] The central claim that five sparsely chosen points constitute a reliable proxy for relative model ranking on an entire scene is validated only by a Spearman ρ=0.95 experiment on two non-overlapping 5-point subsets drawn from the first two frames of each TAP-Vid-DAVIS video (60 pairs total). That experiment does not cover FlowFactor’s deliberately extreme, single-factor regimes (identically zero-flow lighting changes, highly ambiguous repetitive textures, large displacements, occlusions). Because FlowFactor forms a substantial fraction of the 8 204-pair suite that underpins the post-RAFT stagnation plots (Fig. 7), the category–Slow-Flow correlations (Fig. 8) and the trade-off heat-maps (Fig. 9), the diagnostic and “weakly predicts” conclusions remain sensitive to point-selection bias. A modest additional validation—e.g., denser multi-annotator sampling or ranking-consistency checks on a F
  2. [§5.1, Fig. 7] The restriction to models that simultaneously outperform RAFT on the training splits of KITTI-15, Sintel-Final and Spring produces a small cloud of points; the OLS fits that are said to demonstrate stagnation or degradation (Fig. 7) are therefore sensitive both to the precise definition of “outperform” and to which models happen to clear all three thresholds. A continuous analysis (proxy score versus real-world score without hard thresholding) or a leave-one-out sensitivity check would make the claim that further proxy gains cease to predict real-world accuracy more convincing.
minor comments (4)
  1. [Fig. 6] Fig. 6 would be clearer with ground-truth flow or endpoint-error overlays so that the claimed superiority of GMA over SEA-RAFT under lighting and occlusion can be assessed quantitatively rather than by visual inspection alone.
  2. [§4.1] The exact number of models retained after the >2.0 EPE / Fl-all reproduction filter (§4.1) is never stated; a short sentence or a column in Appendix Table 1 would help readers gauge the breadth of the final comparison.
  3. [Abstract / §1] Several extracted passages contain concatenated words (“withreal-worldaccuracy”, “stronglywithreal-world”); these are presumably PDF-extraction artefacts but should be checked in the camera-ready source.
  4. [§3.1 Lighting Variation] The lighting category uses an identically zero ground-truth field; reporting the complementary mean EPE (or fraction of non-zero predictions) alongside Fl-all would make the photometric-robustness ranking easier to interpret under the 3 px / 5 % threshold.

Circularity Check

0 steps flagged

Purely empirical evaluation paper with no derivation chain that reduces predictions or claims to fitted inputs or self-definitional constructions.

full rationale

The paper constructs real-world benchmarks (FlowFactor with sparse manual annotations, TAP-Flow repurposed from TAP-Vid-DAVIS, Slow Flow) and evaluates standard Things (C+T) checkpoints of existing optical flow models using Fl-all. All reported results (OLS fits of real-world Fl-all vs. Sintel/KITTI/Spring, Spearman correlations of FlowFactor categories vs. Slow Flow/TAP-Flow, trade-offs among categories, training-set-size plots) are direct empirical measurements on held-out data never used to train or fit the evaluated models. No parameter is fitted to a subset and then re-presented as an independent prediction; no uniqueness theorem or ansatz is imported via self-citation to force a result; no quantity is defined in terms of the quantity it claims to predict. The 5-point proxy validation (Spearman ho=0.95 on TAP-Vid subsets) is an empirical check of ranking stability, not a circular definition. Self-citations are limited to standard datasets and annotation tooling and are not load-bearing for any central claim. The work is therefore free of the enumerated circularity patterns.

Axiom & Free-Parameter Ledger

3 free parameters · 3 axioms · 1 invented entities

The central claims rest on a small set of domain conventions (Fl-all threshold, Things checkpoint protocol, classical confounder list) plus one paper-specific modelling choice (5-point sparse proxy). No free parameters are fitted to produce the reported correlations; the invented entity is the FlowFactor collection itself, which is released and therefore independently usable.

free parameters (3)
  • Fl-all correctness threshold = 3 px / 5 %
    A pixel is counted correct if Euclidean error < 3 px or 5 % of ground-truth magnitude (KITTI convention). The 3 px figure is justified by the authors’ own re-annotation study (mean 0.58 EPE) but remains a conventional choice that affects absolute scores.
  • large-displacement cutoff = 25 px
    Pairs are labelled “large displacement” when Euclidean motion ≥ 25 px. The threshold is chosen by the authors to exceed typical Sintel/KITTI/Spring magnitudes; it is not fitted to maximise any reported correlation.
  • points per scene = 5
    Exactly five manually chosen points are used as the sparse ground truth for every FlowFactor pair. The number is a design decision validated only by a secondary Spearman experiment.
axioms (3)
  • domain assumption The four classical optical-flow failure modes (large displacement, occlusion, repetitive texture, lighting change) can be approximately isolated in real imagery so that each FlowFactor subset mainly varies only one factor.
    Stated in §3.1 and Fig. 2; underpins the diagnostic correlation analysis of Exp. 2 and Exp. 3.
  • domain assumption Standard Things (C+T) checkpoints supplied by model authors constitute a fair, comparable measure of out-of-distribution generalisation.
    Explicitly adopted in §4.2 because full re-training of every architecture is computationally infeasible; any systematic difference in training recipes remains unmeasured.
  • ad hoc to paper Five sparse points are a sufficient proxy for relative model ranking on an entire scene.
    Justified by a Spearman ρ=0.95 experiment on TAP-Vid subsets (§3.1); load-bearing for all FlowFactor-based claims.
invented entities (1)
  • FlowFactor dataset independent evidence
    purpose: Provide 1 000 real HD pairs organised into four approximately isolated confounder categories so that model failures can be diagnosed rather than merely averaged.
    Newly constructed and released by the authors; independent evidence is the public GitHub release and the annotation-consistency study.

pith-pipeline@v1.1.0-grok45 · 25169 in / 2980 out tokens · 52861 ms · 2026-07-14T11:26:46.078001+00:00 · methodology

0 comments
read the original abstract

Real-world deployment of vision models to broadly benefit society is arguably a main research objective. In optical flow, however, the difficulty to obtain the ground truth has focused research mainly on synthetic data and domain-specific benchmarks. Here, we investigate the severity of this mismatch. We study how well modern optical flow estimation models generalise to real-world video and question if accuracy on synthetic benchmark proxies actually predicts accuracy on real-world optical flow. To address this, we build a real-world evaluation benchmark and evaluate the real-world generalisability of a broad set of recent optical flow models using standard checkpoints. Our benchmark contains 8,204 frame pairs across TAP-Flow, Slow Flow, and our own dataset FlowFactor. FlowFactor is a manually annotated real-world benchmark of 1,000 HD frame pairs organised into four confounding factors: large displacements, repetitive textures, occlusions, and lighting variation. Each setting mainly varies only one factor, enabling diagnostic, confounder-specific analysis. Using FlowFactor, we reveal that performance on varying lighting and large displacements correlates most strongly with real-world accuracy, and that improvements on large-motion regimes can trade off against robustness in small-motion, stationary scenes. Our experiments show that progress on Sintel, KITTI and Spring only weakly predicts accuracy on real-world data, highlighting the need for a broad real-world optical flow benchmark. Interestingly, scaling up the amount of training data does not necessarily resolve the gap, calling for new innovative research instead of simply scaling data and compute.

Figures

Figures reproduced from arXiv: 2607.10470 by Jan van Gemert, Petter Reijalt, Rickard Karlsson, Sander Gielisse.

Figure 1
Figure 1. Figure 1: How well does optical flow quality on current benchmarks generalise [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Factorisation of Optical Flow failure modes. [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Example frame-pairs from the FlowFactor dataset. Annotated points are best viewed zoomed in. Closest to our work are efforts that analyse cross-dataset generalisation of modern deep optical flow models [2, 36, 43]. However, these datasets are often narrow in domain [7, 21] or synthetic in nature [7]. Moreover, these evaluations typically aggregate accuracy over entire datasets and do not systematically iso… view at source ↗
Figure 4
Figure 4. Figure 4: Feasibility and accuracy of hand-annotated points. [PITH_FULL_IMAGE:figures/full_fig_p008_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Correlation between real-world datasets and KITTI [28], Sintel-Final [PITH_FULL_IMAGE:figures/full_fig_p011_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: Qualitative analysis. The GMA model [19] demonstrates superior stability under conditions of extreme photometric volatility. Specifically, it maintains a near￾zero motion field when subjected to drastic lighting changes, whereas the SEA-RAFT small model [40] exhibits sensitivity to these variations, erroneously predicting motion. A similar performance gap is observed in the occlusion category. GMA [19]: Ro… view at source ↗
Figure 7
Figure 7. Figure 7: Correlation between real-world datasets and KITTI, Sintel-Final and [PITH_FULL_IMAGE:figures/full_fig_p012_7.png] view at source ↗
Figure 8
Figure 8. Figure 8: Correlation between Slow Flow and isolated categories in FlowFactor. [PITH_FULL_IMAGE:figures/full_fig_p013_8.png] view at source ↗
Figure 9
Figure 9. Figure 9: Correlation between rankings on individual categories of the Flow [PITH_FULL_IMAGE:figures/full_fig_p013_9.png] view at source ↗
Figure 10
Figure 10. Figure 10: Size training dataset vs. accuracy on FlowFactor. [PITH_FULL_IMAGE:figures/full_fig_p014_10.png] view at source ↗
Figure 11
Figure 11. Figure 11: Example of annotated pixels in the TAP-Flow dataset [9]. Just like the Flow￾Factor dataset the points were randomly chosen by the annotator. This image shows the ability to accurately annotate non-rigid motion [PITH_FULL_IMAGE:figures/full_fig_p023_11.png] view at source ↗
Figure 12
Figure 12. Figure 12: Example of annotated pixels in the Slow Flow dataset [18]. As opposed to the other datasets in our real-world benchmark this utilises a dense flow field. This image showcases the ability of Slow Flow [18] to account for occlusions as opposed to the TAP-Flow [9] dataset [PITH_FULL_IMAGE:figures/full_fig_p023_12.png] view at source ↗
Figure 13
Figure 13. Figure 13: A screenshot of the annotation tool used to annotate the FlowFactor dataset. It allows for setting a maximum of frame pairs, and the use of an offset to skip a certain amount of frames [PITH_FULL_IMAGE:figures/full_fig_p024_13.png] view at source ↗
Figure 14
Figure 14. Figure 14: Size training dataset vs. accuracy on TAP-Flow. [PITH_FULL_IMAGE:figures/full_fig_p025_14.png] view at source ↗
Figure 15
Figure 15. Figure 15: Size training dataset vs. accuracy on Slow Flow. [PITH_FULL_IMAGE:figures/full_fig_p025_15.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

50 extracted references · 3 linked inside Pith

  1. [1]

    Robust vision challenge.http://www.robustvision.net/, accessed: 2025-11-13

  2. [2]

    Transactions on Machine Learning Research (2025)

    Agnihotri, S., Caspary, J.Y., Schwarz, L., Gao, X., Schmalfuss, J., Bruhn, A., Keu- per, M.: Flowbench: Benchmarking optical flow estimation methods for reliability and generalization. Transactions on Machine Learning Research (2025)

  3. [3]

    Computer vision and image understanding 249, 104160 (2024)

    Alfarano, A., Maiano, L., Papa, L., Amerini, I.: Estimating optical flow: A compre- hensive review of the state of the art. Computer vision and image understanding 249, 104160 (2024)

  4. [4]

    International journal of computer vision92(1), 1–31 (2011)

    Baker, S., Scharstein, D., Lewis, J.P., Roth, S., Black, M.J., Szeliski, R.: A database and evaluation methodology for optical flow. International journal of computer vision92(1), 1–31 (2011)

  5. [5]

    International journal of computer vision12(1), 43–77 (1994)

    Barron, J.L., Fleet, D.J., Beauchemin, S.S.: Performance of optical flow techniques. International journal of computer vision12(1), 43–77 (1994)

  6. [6]

    Bauer, K., Bruhn, A., Schmalfuss, J.: On the generalization of optical flow: Quan- tifying robustness to dataset shifts

  7. [7]

    In: European conference on computer vision

    Butler, D.J., Wulff, J., Stanley, G.B., Black, M.J.: A naturalistic open source movie for optical flow evaluation. In: European conference on computer vision. pp. 611–

  8. [8]

    Dahal, S.: Going against the flow: Evaluating optical flow estimation models on real-world non-rigid motion (2025),https://resolver.tudelft.nl/uuid: ab2db1e7-9122-43c3-9377-f188fceb30b6

  9. [9]

    Advances in Neural Information Processing Systems35, 13610–13626 (2022)

    Doersch,C.,Gupta,A.,Markeeva,L.,Recasens,A.,Smaira,L.,Aytar,Y.,Carreira, J.,Zisserman,A.,Yang,Y.:Tap-vid:Abenchmarkfortrackinganypointinavideo. Advances in Neural Information Processing Systems35, 13610–13626 (2022)

  10. [10]

    In: Proceedings of the IEEE international conference on computer vision

    Dosovitskiy, A., Fischer, P., Ilg, E., Hausser, P., Hazirbas, C., Golkov, V., Van Der Smagt, P., Cremers, D., Brox, T.: Flownet: Learning optical flow with convolu- tional networks. In: Proceedings of the IEEE international conference on computer vision. pp. 2758–2766 (2015)

  11. [11]

    In: 2024 13th Iranian/3rd International Machine Vision and Image Processing Conference (MVIP)

    Eslami, N., Arefi, F., Mansourian, A.M., Kasaei, S.: Rethinking raft for efficient optical flow. In: 2024 13th Iranian/3rd International Machine Vision and Image Processing Conference (MVIP). pp. 1–7. IEEE (2024)

  12. [12]

    Ge, Z.: Real-world evaluation of optical flow with varying lighting condi- tions (2025),https://resolver.tudelft.nl/uuid:7ad41253-902a-4d78-a86a- c7bee94b9e0c

  13. [13]

    In: 2012 IEEE conference on computer vision and pattern recognition

    Geiger, A., Lenz, P., Urtasun, R.: Are we ready for autonomous driving? the kitti vision benchmark suite. In: 2012 IEEE conference on computer vision and pattern recognition. pp. 3354–3361. IEEE (2012)

  14. [14]

    In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition

    Greff, K., Belletti, F., Beyer, L., Doersch, C., Du, Y., Duckworth, D., Fleet, D.J., Gnanapragasam, D., Golemo, F., Herrmann, C., et al.: Kubric: A scalable dataset generator. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. pp. 3749–3761 (2022) 1 https://treesforall.nl/app/uploads/certificates/Certificaat_Trees_for_ ...

  15. [15]

    Journal of the Optical Society of America A4(8), 1455–1471 (1987)

    Heeger, D.J.: Model for the extraction of image flow. Journal of the Optical Society of America A4(8), 1455–1471 (1987)

  16. [16]

    In: European conference on computer vision

    Huang, Z., Shi, X., Zhang, C., Wang, Q., Cheung, K.C., Qin, H., Dai, J., Li, H.: Flowformer: A transformer architecture for optical flow. In: European conference on computer vision. pp. 668–685. Springer (2022)

  17. [17]

    In: Proceedings of the IEEE conference on computer vision and pattern recognition

    Ilg, E., Mayer, N., Saikia, T., Keuper, M., Dosovitskiy, A., Brox, T.: Flownet 2.0: Evolution of optical flow estimation with deep networks. In: Proceedings of the IEEE conference on computer vision and pattern recognition. pp. 2462–2470 (2017)

  18. [18]

    In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition

    Janai, J., Guney, F., Wulff, J., Black, M.J., Geiger, A.: Slow flow: Exploiting high- speed cameras for accurate and diverse optical flow reference data. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. pp. 3597– 3607 (2017)

  19. [19]

    In: Proceedings of the IEEE/CVF international conference on computer vision

    Jiang, S., Campbell, D., Lu, Y., Li, H., Hartley, R.: Learning to estimate hid- den motions with global motion aggregation. In: Proceedings of the IEEE/CVF international conference on computer vision. pp. 9772–9781 (2021)

  20. [20]

    Klijnsma, J.B.: Real-world evaluation of optical flow on repetitive patterns (2025), https://resolver.tudelft.nl/uuid:dc709377-e116-48eb-91a5-4b6812902034

  21. [21]

    In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops

    Kondermann,D.,Nair,R.,Honauer,K.,Krispin,K.,Andrulis,J.,Brock,A.,Gusse- feld, B., Rahimimoghaddam, M., Hofmann, S., Brenner, C., et al.: The hci bench- mark suite: Stereo and flow ground truth with uncertainties for urban autonomous driving. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops. pp. 19–28 (2016)

  22. [22]

    In: Proceedings of the IEEE/CVF international conference on computer vision

    Li, H., Luo, K., Liu, S.: Gyroflow: Gyroscope-guided unsupervised optical flow learning. In: Proceedings of the IEEE/CVF international conference on computer vision. pp. 12869–12878 (2021)

  23. [23]

    arXiv preprint arXiv:2506.07740 (2025)

    Liang, Y., Fu, Y., Hu, Y., Shao, W., Liu, J., Zhang, D.: Flow-anything: Learn- ing real-world optical flow estimation from large-scale single-view images. arXiv preprint arXiv:2506.07740 (2025)

  24. [24]

    In: 2008 IEEE Conference on Computer Vision and Pattern Recognition

    Liu, C., Freeman, W.T., Adelson, E.H., Weiss, Y.: Human-assisted motion anno- tation. In: 2008 IEEE Conference on Computer Vision and Pattern Recognition. pp. 1–8. IEEE (2008)

  25. [25]

    Mayer, N., Ilg, E., Fischer, P., Hazirbas, C., Cremers, D., Dosovitskiy, A., Brox, T.: What makes good synthetic training data for learning disparity and optical flow estimation? International Journal of Computer Vision126(9), 942–960 (2018)

  26. [26]

    In: Proceedings of the IEEE conference on computer vision and pattern recognition

    Mayer, N., Ilg, E., Hausser, P., Fischer, P., Cremers, D., Dosovitskiy, A., Brox, T.: A large dataset to train convolutional networks for disparity, optical flow, and scene flow estimation. In: Proceedings of the IEEE conference on computer vision and pattern recognition. pp. 4040–4048 (2016)

  27. [27]

    In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

    Mehl, L., Schmalfuss, J., Jahedi, A., Nalivayko, Y., Bruhn, A.: Spring: A high- resolutionhigh-detaildatasetandbenchmarkforsceneflow,opticalflowandstereo. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 4981–4991 (2023)

  28. [28]

    In: Proceedings of the IEEE conference on computer vision and pattern recognition

    Menze, M., Geiger, A.: Object scene flow for autonomous vehicles. In: Proceedings of the IEEE conference on computer vision and pattern recognition. pp. 3061–3070 (2015)

  29. [29]

    Morimitsu, H.: Pytorch lightning optical flow.https://github.com/hmorimitsu/ ptlflow(2021)

  30. [30]

    Petre, I.A.: Performance of optical flow models on real-world occluded re- gions (2025),https://resolver.tudelft.nl/uuid:c25761a4-939d-4677-b790- b3fb9bc4bc28 On the Real-World Generalisability of Optical Flow Models 17

  31. [31]

    arXiv preprint arXiv:2505.09368 (2025)

    Schmalfuss, J., Oei, V., Mehl, L., Bartsch, M., Agnihotri, S., Keuper, M., Bruhn, A.: Robustspring: Benchmarking robustness to image corruptions for optical flow, scene flow and stereo. arXiv preprint arXiv:2505.09368 (2025)

  32. [32]

    In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition

    Shi, X., Huang, Z., Li, D., Zhang, M., Cheung, K.C., See, S., Qin, H., Dai, J., Li, H.: Flowformer++: Masked cost volume autoencoding for pretraining optical flow estimation. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. pp. 1599–1610 (2023)

  33. [33]

    Vision research 44(13), 1565–1573 (2004)

    Stürzel, F., Spillmann, L.: Perceptual limits of common fate. Vision research 44(13), 1565–1573 (2004)

  34. [34]

    In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

    Sun, D., Vlasic, D., Herrmann, C., Jampani, V., Krainin, M., Chang, H., Zabih, R., Freeman, W.T., Liu, C.: Autoflow: Learning a better training set for optical flow. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 10093–10102 (2021)

  35. [35]

    In: Proceedings of the IEEE conference on computer vision and pattern recognition

    Sun, D., Yang, X., Liu, M.Y., Kautz, J.: Pwc-net: Cnns for optical flow using pyramid, warping, and cost volume. In: Proceedings of the IEEE conference on computer vision and pattern recognition. pp. 8934–8943 (2018)

  36. [36]

    In: Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part II 16

    Teed, Z., Deng, J.: Raft: Recurrent all-pairs field transforms for optical flow. In: Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part II 16. pp. 402–419. Springer (2020)

  37. [37]

    nl/uuid:d7c4cc36-6582-47fc-94cd-cc0dc393725b

    Timmerije, M.: Bridging the gap: A real-world dataset and evaluation of optical flow models in large displacement scenarios (2025),https://resolver.tudelft. nl/uuid:d7c4cc36-6582-47fc-94cd-cc0dc393725b

  38. [38]

    In: 2020 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)

    Wang, W., Zhu, D., Wang, X., Hu, Y., Qiu, Y., Wang, C., Hu, Y., Kapoor, A., Scherer, S.: Tartanair: A dataset to push the limits of visual slam. In: 2020 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS). pp. 4909–4916. IEEE (2020)

  39. [39]

    arXiv preprint arXiv:2506.21526 (2025)

    Wang, Y., Deng, J.: Waft: Warping-alone field transforms for optical flow. arXiv preprint arXiv:2506.21526 (2025)

  40. [40]

    In: European Conference on Computer Vision

    Wang, Y., Lipson, L., Deng, J.: Sea-raft: Simple, efficient, accurate raft for optical flow. In: European Conference on Computer Vision. pp. 36–54. Springer (2024)

  41. [41]

    Weinzaepfel, P., Revaud, J., Harchaoui, Z., Schmid, C.: Learning to detect motion boundaries.In:ProceedingsoftheIEEEconferenceoncomputervisionandpattern recognition. pp. 2578–2586 (2015)

  42. [42]

    Electronics12(3), 682 (2023)

    Weng, W., Zhu, X., Jing, L., Dong, M.: Attention mechanism trained with small datasets for biomedical image segmentation. Electronics12(3), 682 (2023)

  43. [43]

    In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition

    Xu, H., Zhang, J., Cai, J., Rezatofighi, H., Tao, D.: Gmflow: Learning optical flow via global matching. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. pp. 8121–8130 (2022)

  44. [44]

    In: Proceedings of the IEEE/CVF conference on com- puter vision and pattern recognition

    Yin, Z., Darrell, T., Yu, F.: Hierarchical discrete distribution decomposition for match density estimation. In: Proceedings of the IEEE/CVF conference on com- puter vision and pattern recognition. pp. 6044–6053 (2019)

  45. [45]

    Pattern Recognition114, 107861 (2021)

    Zhai, M., Xiang, X., Lv, N., Kong, X.: Optical flow and scene flow estimation: A survey. Pattern Recognition114, 107861 (2021)

  46. [46]

    arXiv preprint arXiv:2603.25739 (2026)

    Zhang, D., Wang, F., Pollefeys, M., Xu, H.: Megaflow: Zero-shot large displacement optical flow. arXiv preprint arXiv:2603.25739 (2026)

  47. [47]

    In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition

    Zhao,S.,Sheng,Y.,Dong,Y.,Chang,E.I.,Xu,Y.,etal.:Maskflownet:Asymmetric feature matching with learnable occlusion mask. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. pp. 6278–6287 (2020)

  48. [48]

    In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

    Zhao, S., Zhao, L., Zhang, Z., Zhou, E., Metaxas, D.: Global matching with over- lapping attention for optical flow estimation. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 17592–17601 (2022) 18 P. Reijalt et al

  49. [49]

    In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition

    Zheng, Y., Zhang, M., Lu, F.: Optical flow in the dark. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. pp. 6749–6757 (2020)

  50. [50]

    Zhou, H., Chang, Y., Shi, Z., Yan, W., Chen, G., Tian, Y., Yan, L.: Adverse weather optical flow: Cumulative homogeneous-heterogeneous adaptation. IEEE Transactions on Pattern Analysis and Machine Intelligence (2024) On the Real-World Generalisability of Optical Flow Models 19 Appendix A Reproducing Reported Checkpoint Results (Things) Table 1:Paper vs. o...