REVIEW 3 major objections 5 minor 72 references
CompactFlowNet: Efficient Real-time Optical Flow Estimation on Mobile Devices
T0 review · 3 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read CompactFlowNet runs in 40 ms per 512x512 frame on an iPhone 8 while keeping accuracy on par with lightweight baselines.
desk verdict A useful mobile optical flow system with a real speed/memory win, but the abstract's accuracy claim is contradicted by the paper's own test tables. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is architectural: the flow estimator, which dominates parameter count, FLOPs, and latency, is changed from a densely connected block to a sequentially connected block. Dense connections force every layer to receive concatenated maps from all earlier layers, keeping large tensors alive in memory and inflating both latency and the downstream refiner's input size; sequential connections remove that storage and compute burden. The other components are depthwise separable convolutions in the estimator and refiner, a shortened refiner with channels 128, 64, 32, and 2, and a MobileNetV3 backbone chosen by an on-device comparison of candidate backbones at 512x512 on an iPhone 8. Training adds a second mechanism: knowledge distillation from the heavier MobileNetV3-backed PWC-Net teacher, with loss $L = L_{\mathrm{sup}} + \gamma L_{\mathrm{dist}}$ at $\gamma = 0.1$, combined with an Autoflow pre-training dataset, extra augmentations, a OneCycle learning rate, and gradient clipping.
What would settle it
Reproduce the benchmark with the models converted under identical settings (same precision, same delegate, same iPhone OS) and check the 40 ms latency and 118 MB peak memory at 512x512 on an iPhone 8; if the margin over FastFlowNet disappears or the memory numbers grow past the reported values, the central claim fails.
Extended reading notes
Core claim
The paper argues that a PWC-Net-style coarse-to-fine optical flow model can be made mobile-friendly without sacrificing accuracy, by attacking the flow estimator block. Profiling on the iPhone 8 shows the estimator accounts for roughly 6.05 of the 8.75 million parameters and 131 of the 212 ms total latency of the original PWC-Net at 512x512. The authors replace the estimator's dense concatenation with a sequential connection pattern, adopt depthwise separable convolutions in the estimator and refiner, shrink the refiner from seven layers to four with channels 128, 64, 32, and 2, and swap the backbone for MobileNetV3 while keeping the full model under five million parameters. A heavy teacher model, the original architecture equipped with the MobileNetV3 backbone, is then used to distill CompactFlowNet, with a per-pixel L2 supervision loss plus a distillation term weighted at 0.1. The resulting model reaches 40 ms latency on an iPhone 8, 26 ms on an iPhone XR, 19 ms on an iPhone 12, and 13 ms on an iPhone 14 Pro at 512x512, with peak memory of 118 MB at that resolution; it also beats FastFlowNet and SpyNet on the Sintel test sets and runs faster than all compared lightweight models across three resolutions and four devices.
Load-bearing premise
The speed and memory comparisons assume that every model is converted to the mobile format with equally fair settings, so the measured differences on the phone reflect the architectures rather than conversion or precision artifacts.
Editorial extensions
If this is right
- Tasks that depend on optical flow, such as stabilization, tracking, restoration, frame interpolation, and generative video editing, can be executed on-device without uploading footage.
- The model runs at real-time speed (40 ms, 25 FPS) on an iPhone 8 and 13 ms on an iPhone 14 Pro at 512x512, with peak memory of 118 MB.
- At Full HD (1080x1920) input it reaches nearly 10 FPS on an iPhone 14 Pro, and outperforms FastFlowNet and SpyNet on the Sintel test sets.
- The under-10 MB float16 weight size and reduced memory footprint make the model compatible with the constraints the paper cites for mobile deployment.
Reading between the lines
- The profiling table implies that dense feature concatenation, not the backbone, is the main latency and memory bottleneck in lightweight flow models; a direct activation-memory profiler could verify this on other PWC variants.
- The accuracy gain from distillation (Sintel Final validation from 4.05 to 3.36, KITTI 2015 train from 2.14 to 1.91) suggests that a stronger teacher from a different model family might reduce the remaining error, although the paper reports trying such teachers without improvement.
- The KITTI 2015 test f1-all of 14.64 is higher than FDFlowNet's 9.38, so the 'superior or comparable' claim should be read as benchmark-dependent; the speed and memory story does not depend on that comparison.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. CompactFlowNet is a proposed mobile-optimized optical flow network built on a PWC-Net-style architecture with a MobileNetV3 backbone, a sequentially connected flow estimator that removes dense concatenation, depthwise separable convolutions, and a reduced-depth flow refiner. The paper reports 40 ms average latency on an iPhone 8 at 512x512 input, lower peak memory than several lightweight competitors, and claims 'superior or comparable performance' to state-of-the-art lightweight models on KITTI and Sintel. Training uses a distillation pipeline with the authors' own heavy PWC-Net+MobileNetV3 model as teacher, and results are reported on Sintel and KITTI benchmarks as well as on-device latency and memory across four iPhone models.
Significance. If the reported speed and memory numbers are reproducible, the paper makes a useful engineering contribution: it demonstrates that a full optical flow network can run at interactive rates on a 2017-era phone, and it provides a cross-device latency and memory comparison for several lightweight methods. The paper ships no code or weights, however, and the accuracy claim is contradicted by the paper's own public test-set numbers, which show CompactFlowNet trailing FDFlowNet on every benchmark. The speed advantage may be real, but the overall claim as stated is overstated and the benchmarking methodology is not sufficiently specified to assess it.
major comments (3)
- [Abstract; Tables 4 and 5; Section 4.3] The abstract and Section 1 claim 'superior or comparable performance' to state-of-the-art lightweight models, but the paper's own public test-set results contradict this. On Sintel Clean test, CompactFlowNet reports 4.43 AEPE versus FDFlowNet's 3.71; on Sintel Final test, 5.55 versus 5.11; and on KITTI 2015 test (f1-all), 14.64 versus 9.38. CompactFlowNet is worse than FDFlowNet on every public test benchmark, with a relative KITTI gap exceeding 50%. The Discussion's statement (Section 4.3) that CompactFlowNet 'aligns closely with' FDFlowNet is inconsistent with these numbers, which are the conventional basis for benchmark comparison. The claim must be revised or re-scoped, for example to 'competitive accuracy at a fraction of the latency,' and the Discussion should honestly state the accuracy trade-off.
- [Section 4.2; Tables 6 and 7] The on-device speed and memory comparison is under-specified. The text says all models were 'implemented in PyTorch and converted to TFLite,' but it does not report the conversion tool versions, whether weights were quantized or cast to float16, which TFLite delegate (GPU or CPU) was used, or any per-model conversion settings. If competitor models convert less favorably, for example due to unsupported operators falling back to CPU, the claimed speed and memory advantage of CompactFlowNet could be an artifact of the conversion pipeline rather than of the architecture. The paper also gives no variance or error bars for the latency measurements, despite averaging over 100 runs. This is load-bearing for the paper's central speed claim.
- [Section 4.1; Section 3.1] The paper provides no code, pretrained weights, or complete training configuration, and it refers to 'specific augmentation parameters in Supplementary Materials' that are not present in the arXiv submission. The training pipeline includes many undocumented choices, such as the exact OneCycle schedule details, batch size, augmentation ranges, and distillation teacher configuration, which would be needed to reproduce the models. For a paper whose contribution is an architecture plus a training recipe, this lack of reproducibility evidence weakens the verification of both the accuracy and the speed claims.
minor comments (5)
- [Section 1] The abstract states 'the first real-time mobile neural network for optical flow prediction,' while the introduction uses 'To our knowledge, this is the first compact and memory-efficient model...' These are different claims; the hedging should be consistent, and the novelty claim should be supported by a more thorough survey of mobile deployment efforts.
- [Table 3] The parenthesized values in Table 3 are not sufficiently explained; the caption says they 'represent the results of networks that were trained on the same dataset,' but the reader has to infer which dataset is meant. Please clarify in the caption.
- [Table 6] The bold font for real-time compatible latency is helpful, but the threshold for real-time, for example 25 FPS or 40 ms, is not defined in the text or caption; please state it explicitly.
- [Section 3.2] The description of depthwise separable convolutions is standard, but it would help to state whether they replace all standard convolutions in the flow estimator and refiner, or only some; the text says 'adopted in both the flow estimator and flow refiner blocks' without specifying the extent.
- [Section 2] Reference [54] is cited for GPU latency and memory of RAFT and PWC-Net, but the numbers given in the text, such as 107 ms for RAFT, should be attributed explicitly to that paper to avoid ambiguity.
Circularity Check
No significant circularity; results are empirical measurements against external benchmarks.
full rationale
CompactFlowNet is an empirical systems paper: the accuracy claims are evaluated on the public KITTI and Sintel test sets against ground-truth flow, and the latency/memory claims are direct on-device measurements of PyTorch-to-TFLite conversions of all compared models. The teacher-student distillation (Section 3.4) uses the authors' heavy PWC-Net+MobileNetV3 as teacher, but the student's final evaluation is against ground-truth benchmarks (Tables 4-5), not against the teacher, and the supervision loss L_sup is computed on ground truth, so the architecture is not defined in terms of its own outputs. The backbone selection (Section 3.3) uses external benchmarks to choose among published backbones (MobileNetV3, ReXNet, HardCoReNAS), all cited as prior work; no load-bearing self-citation or imported uniqueness theorem appears. There is an apparent internal inconsistency between the abstract's 'superior or comparable performance' claim and the test-set rows of Tables 4-5 (e.g., CompactFlowNet's KITTI test f1-all 14.64 vs FDFlowNet 9.38), but this is a correctness/consistency issue, not circularity: the numbers are still measured against external data rather than derived from assumptions. Similarly, the on-device comparison assumes equal conversion quality across models (Section 4.2), but that is an experimental-condition concern, not a definitional reduction. No fitted parameter is relabeled as a prediction, and no claim is justified solely by a self-citation. Overall, the paper's central claims stand or fall on the reproducibility of its hardware measurements and benchmark results, not on circular reasoning.
Assumptions & free parameters
free parameters (6)
- distillation loss weight gamma =
0.1
- flow refiner channel widths =
128, 64, 32, 2
- flow refiner depth =
4 layers (from 7)
- backbone choice =
MobileNetV3 large (about 3.12M params)
- training iterations and learning schedule =
1.2M, 750k, 1.2M steps; OneCycle; weight decay 0 or 1e-5
- augmentation parameters =
not included in the arXiv submission
assumptions (5)
- domain assumption PWC-Net is a suitable base model for mobile real-time optical flow.
- domain assumption Conversion to TFLite is faithful and fair across all compared models.
- domain assumption 25 frames per second (about 40ms) is the right real-time threshold at 512x512 input.
- domain assumption Published benchmark scores of competitor methods are comparable to the authors' own runs.
- domain assumption The teacher model is a good target for distillation.
Cite this review
Pith. "Pith review of CompactFlowNet: Efficient Real-time Optical Flow Estimation on Mobile Devices." pith.science (2026). https://pith.science/paper/7AC4WZMS
@misc{pith2026241213273,
author = {Pith},
title = {Pith review of: CompactFlowNet: Efficient Real-time Optical Flow Estimation on Mobile Devices},
year = {2026},
howpublished = {\url{https://pith.science/paper/7AC4WZMS}},
note = {Machine review of arXiv:2412.13273}
}
read the original abstract
We present CompactFlowNet, the first real-time mobile neural network for optical flow prediction, which involves determining the displacement of each pixel in an initial frame relative to the corresponding pixel in a subsequent frame. Optical flow serves as a fundamental building block for various video-related tasks, such as video restoration, motion estimation, video stabilization, object tracking, action recognition, and video generation. While current state-of-the-art methods prioritize accuracy, they often overlook constraints regarding speed and memory usage. Existing light models typically focus on reducing size but still exhibit high latency, compromise significantly on quality, or are optimized for high-performance GPUs, resulting in sub-optimal performance on mobile devices. This study aims to develop a mobile-optimized optical flow model by proposing a novel mobile device-compatible architecture, as well as enhancements to the training pipeline, which optimize the model for reduced weight, low memory utilization, and increased speed while maintaining minimal error. Our approach demonstrates superior or comparable performance to the state-of-the-art lightweight models on the challenging KITTI and Sintel benchmarks. Furthermore, it attains a significantly accelerated inference speed, thereby yielding real-time operational efficiency on the iPhone 8, while surpassing real-time performance levels on more advanced mobile devices.
Figures
Reference graph
Works this paper leans on
-
[1]
Google llc. tensorflow lite. https://www.tensorflow. org/lite. Accessed: 2019-04-08. 4
work page 2019
-
[2]
Martin Abadi, Paul Barham, Jianmin Chen, Zhifeng Chen, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghe- mawat, Geoffrey Irving, Michael Isard, Manjunath Kudlur, Josh Levenberg, Rajat Monga, Sherry Moore, Derek G. Mur- ray, Benoit Steiner, Paul Tucker, Vijay Vasudevan, Pete War- den, Martin Wicke, Yuan Yu, and Xiaoqiang Zheng. Ten- sorflow: A system f...
work page 2016
-
[3]
Video stabilization us- ing raft-based optical flow
Rana Ashar, Burhan Sadiq, Hira Mohiuddin, Saniya Ashraf, Muhammad Imran, and Anayat Ullah. Video stabilization us- ing raft-based optical flow. In 2023 International Conference on Robotics and Automation in Industry (ICRAI), pages 1–5,
work page 2023
-
[4]
Large displacement optical flow: descriptor matching in variational motion estimation
Thomas Brox and Jitendra Malik. Large displacement optical flow: descriptor matching in variational motion estimation. IEEE transactions on pattern analysis and machine intelli- gence, 33(3):500–513, 2010. 1
work page 2010
-
[5]
Mpi-sintel optical flow benchmark: Supplemental material
D Butler, Jonas Wulff, G Stanley, and M Black. Mpi-sintel optical flow benchmark: Supplemental material. In MPI-IS- TR-006, MPI for Intelligent Systems (2012 . Citeseer, 2012. 5
work page 2012
-
[6]
Basicvsr: The search for essential compo- nents in video super-resolution and beyond
Kelvin CK Chan, Xintao Wang, Ke Yu, Chao Dong, and Chen Change Loy. Basicvsr: The search for essential compo- nents in video super-resolution and beyond. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 4947–4956, 2021. 2, 3
work page 2021
-
[7]
Improved optical flow for gesture-based human-robot interaction
Jen-Yen Chang, Antonio Tejero-de Pablos, and Tatsuya Harada. Improved optical flow for gesture-based human-robot interaction. In 2019 International Conference on Robotics and Automation (ICRA), pages 7983–7989. IEEE, 2019. 1, 2
work page 2019
-
[8]
Learning on-road visual control for self-driving vehicles with auxiliary tasks
Yilun Chen, Palanisamy Praveen, Mudalige Priyantha, Kathe- rina Muelling, and John Dolan. Learning on-road visual control for self-driving vehicles with auxiliary tasks. In 2019 IEEE Winter Conference on Applications of Computer Vision (WACV), pages 331–338. IEEE, 2019. 1
work page 2019
Show all 72 references
-
[9]
Mfcflow: A motion feature compensated multi-frame recurrent network for optical flow estimation
Yonghu Chen, Dongchen Zhu, Wenjun Shi, Guanghui Zhang, Tianyu Zhang, Xiaolin Zhang, and Jiamao Li. Mfcflow: A motion feature compensated multi-frame recurrent network for optical flow estimation. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Visi...
2023
-
[10]
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In 2009 IEEE conference on computer vision and pattern recognition, pages 248–255. Ieee, 2009. 5
2009
-
[11]
Flownet: Learning optical flow with convolutional networks
Alexey Dosovitskiy, Philipp Fischer, Eddy Ilg, Philip Hausser, Caner Hazirbas, Vladimir Golkov, Patrick Van Der Smagt, Daniel Cremers, and Thomas Brox. Flownet: Learning optical flow with convolutional networks. In Proceedings of the IEEE international conference on computer v...
2015
-
[12]
End-to-end learning of motion representation for video understanding
Lijie Fan, Wenbing Huang, Chuang Gan, Stefano Ermon, Boqing Gong, and Junzhou Huang. End-to-end learning of motion representation for video understanding. In Proceed- ings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 6016–6025, 2018. 2
2018
-
[13]
Starflow: A spatiotemporal recurrent cell for lightweight multi-frame optical flow estimation
Pierre Godet, Alexandre Boulch, Aur ´elien Plyer, and Guy Le Besnerais. Starflow: A spatiotemporal recurrent cell for lightweight multi-frame optical flow estimation. In 2020 25th International Conference on Pattern Recognition (ICPR), pages 2462–2469. IEEE, 2021. 2
2020
-
[14]
Rethinking channel dimensions for efficient model design
Dongyoon Han, Sangdoo Yun, Byeongho Heo, and YoungJoon Yoo. Rethinking channel dimensions for efficient model design. In Proceedings of the IEEE/CVF conference on Computer Vision and Pattern Recognition, pages 732–741,
-
[15]
Improving optical flow on a pyramid level
Markus Hofinger, Samuel Rota Bul `o, Lorenzo Porzi, Arno Knapitsch, Thomas Pock, and Peter Kontschieder. Improving optical flow on a pyramid level. In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23– 28, 2020, Proceedings, Part XXVIII, pages 770–786...
2020
-
[16]
Determining optical flow
Berthold KP Horn and Brian G Schunck. Determining optical flow. Artificial intelligence, 17(1-3):185–203, 1981. 1, 2
1981
-
[17]
Searching for mo- bilenetv3
Andrew Howard, Mark Sandler, Grace Chu, Liang-Chieh Chen, Bo Chen, Mingxing Tan, Weijun Wang, Yukun Zhu, Ruoming Pang, Vijay Vasudevan, et al. Searching for mo- bilenetv3. In Proceedings of the IEEE/CVF international conference on computer vision, pages 1314–1324, 2019. 4, 5
2019
-
[18]
Mobilenets: Efficient convolu- tional neural networks for mobile vision applications
Andrew G Howard, Menglong Zhu, Bo Chen, Dmitry Kalenichenko, Weijun Wang, Tobias Weyand, Marco An- dreetto, and Hartwig Adam. Mobilenets: Efficient convolu- tional neural networks for mobile vision applications. arXiv preprint arXiv:1704.04861, 2017. 2, 4
2017 arXiv
-
[19]
Flowformer: A transformer architecture for optical flow
Zhaoyang Huang, Xiaoyu Shi, Chao Zhang, Qiang Wang, Ka Chun Cheung, Hongwei Qin, Jifeng Dai, and Hongsheng Li. Flowformer: A transformer architecture for optical flow. In Computer Vision–ECCV 2022: 17th European Conference, Tel Aviv, Israel, October 23–27, 2022, Proceedings, P...
2022
-
[20]
Liteflownet3: Resolving correspondence ambiguity for more accurate optical flow estimation
Tak-Wai Hui and Chen Change Loy. Liteflownet3: Resolving correspondence ambiguity for more accurate optical flow estimation. In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XX 16, pages 169–184. Springer, 2020. 1
2020
-
[21]
Lite- FlowNet: A Lightweight Convolutional Neural Network for Optical Flow Estimation
Tak-Wai Hui, Xiaoou Tang, and Chen Change Loy. Lite- FlowNet: A Lightweight Convolutional Neural Network for Optical Flow Estimation. In Proceedings of IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 8981–8989, 2018. 1, 3
2018
-
[22]
Iterative residual refinement for joint optical flow and occlusion estimation
Junhwa Hur and Stefan Roth. Iterative residual refinement for joint optical flow and occlusion estimation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 5754–5763, 2019. 1, 2
2019
-
[23]
Ai benchmark: Running deep neural networks on android smartphones
Andrey Ignatov, Radu Timofte, William Chou, Ke Wang, Max Wu, Tim Hartley, and Luc Van Gool. Ai benchmark: Running deep neural networks on android smartphones. In Proceedings of the European Conference on Computer Vision (ECCV) Workshops, pages 0–0, 2018. 2
2018
-
[24]
Flownet 2.0: Evolution of optical flow estimation with deep networks
Eddy Ilg, Nikolaus Mayer, Tonmoy Saikia, Margret Keu- per, Alexey Dosovitskiy, and Thomas Brox. Flownet 2.0: Evolution of optical flow estimation with deep networks. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 2462–2470, 2017. 1, 2
2017
-
[25]
Perceiver io: A general architecture for structured inputs & outputs
Andrew Jaegle, Sebastian Borgeaud, Jean-Baptiste Alayrac, Carl Doersch, Catalin Ionescu, David Ding, Skanda Kop- pula, Daniel Zoran, Andrew Brock, Evan Shelhamer, et al. Perceiver io: A general architecture for structured inputs & outputs. arXiv preprint arXiv:2107.14795, 2021. 1, 3
2021 arXiv
-
[26]
Unsupervised learning of multi-frame optical flow with occlusions
Joel Janai, Fatma Guney, Anurag Ranjan, Michael Black, and Andreas Geiger. Unsupervised learning of multi-frame optical flow with occlusions. In Proceedings of the European conference on computer vision (ECCV), pages 690–706, 2018. 2
2018
-
[27]
Learning to estimate hidden motions with global motion aggregation
Shihao Jiang, Dylan Campbell, Yao Lu, Hongdong Li, and Richard Hartley. Learning to estimate hidden motions with global motion aggregation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 9772– 9781, 2021. 1, 3
2021
-
[28]
The hci benchmark suite: Stereo and flow ground truth with uncertainties for urban autonomous driving
Daniel Kondermann, Rahul Nair, Katrin Honauer, Karsten Krispin, Jonas Andrulis, Alexander Brock, Burkhard Gusse- feld, Mohsen Rahimimoghaddam, Sabine Hofmann, Claus Brenner, et al. The hci benchmark suite: Stereo and flow ground truth with uncertainties for urban autonomous dr...
2016
-
[29]
Fdflownet: Fast optical flow estimation using a deep lightweight network
Lingtong Kong and Jie Yang. Fdflownet: Fast optical flow estimation using a deep lightweight network. In 2020 IEEE International Conference on Image Processing (ICIP), pages 1501–1505. IEEE, 2020. 1, 3
2020
-
[30]
Fastflownet: A lightweight network for fast optical flow estimation
Lingtong Kong, Chunhua Shen, and Jie Yang. Fastflownet: A lightweight network for fast optical flow estimation. In 2021 IEEE International Conference on Robotics and Automation (ICRA), pages 10310–10316. IEEE, 2021. 1, 2, 3
2021
-
[31]
Fast optical flow using dense inverse search
Till Kroeger, Radu Timofte, Dengxin Dai, and Luc Van Gool. Fast optical flow using dense inverse search. In Computer Vision–ECCV 2016: 14th European Conference, Amsterdam, The Netherlands, October 11–14, 2016, Proceedings, Part IV 14, pages 471–488. Springer, 2016. 1
2016
-
[32]
On-device neural net inference with mobile gpus
Juhyun Lee, Nikolay Chirkov, Ekaterina Ignasheva, Yury Pisarchyk, Mogan Shieh, Fabio Riccardi, Raman Sarokin, Andrei Kulik, and Matthias Grundmann. On-device neural net inference with mobile gpus. arXiv preprint arXiv:1907.01989,
1907 arXiv
-
[33]
Centermask: Real-time anchor-free instance segmentation
Youngwan Lee and Jongyoul Park. Centermask: Real-time anchor-free instance segmentation. In Proceedings of the IEEE/CVF conference on computer vision and pattern recog- nition, pages 13906–13915, 2020. 4
2020
-
[34]
An energy and gpu-computation efficient backbone network for real-time object detection
Youngwan Lee, Joong-won Hwang, Sangrok Lee, Yuseok Bae, and Jongyoul Park. An energy and gpu-computation efficient backbone network for real-time object detection. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition workshops, pages 0–0, 2019. 4
2019
-
[35]
Efficient- former: Vision transformers at mobilenet speed
Yanyu Li, Geng Yuan, Yang Wen, Ju Hu, Georgios Evange- lidis, Sergey Tulyakov, Yanzhi Wang, and Jian Ren. Efficient- former: Vision transformers at mobilenet speed. Advances in Neural Information Processing Systems, 35:12934–12949,
-
[36]
Gaflow: Incorporating gaussian attention into optical flow
Ao Luo, Fan Yang, Xin Li, Lang Nie, Chunyu Lin, Haoqiang Fan, and Shuaicheng Liu. Gaflow: Incorporating gaussian attention into optical flow. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 9642– 9651, 2023. 3
2023
-
[37]
A large dataset to train convolutional networks for disparity, optical flow, and scene flow estimation
Nikolaus Mayer, Eddy Ilg, Philip Hausser, Philipp Fischer, Daniel Cremers, Alexey Dosovitskiy, and Thomas Brox. A large dataset to train convolutional networks for disparity, optical flow, and scene flow estimation. In Proceedings of the IEEE conference on computer vision and ...
2016
-
[38]
Mobilevit: light- weight, general-purpose, and mobile-friendly vision trans- former
Sachin Mehta and Mohammad Rastegari. Mobilevit: light- weight, general-purpose, and mobile-friendly vision trans- former. arXiv preprint arXiv:2110.02178, 2021. 5
2021 arXiv
-
[39]
Object scene flow for autonomous vehicles
Moritz Menze and Andreas Geiger. Object scene flow for autonomous vehicles. In Conference on Computer Vision and Pattern Recognition (CVPR), 2015. 5
2015
-
[40]
Hardcore-nas: Hard constrained differentiable neural architec- ture search
Niv Nayman, Yonathan Aflalo, Asaf Noy, and Lihi Zelnik. Hardcore-nas: Hard constrained differentiable neural architec- ture search. In International Conference on Machine Learn- ing, pages 7979–7990. PMLR, 2021. 5
2021
-
[41]
Context-aware synthesis for video frame interpolation
Simon Niklaus and Feng Liu. Context-aware synthesis for video frame interpolation. In Proceedings of the IEEE con- ference on computer vision and pattern recognition , pages 1701–1710, 2018. 2, 3
2018
-
[42]
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Kopf, Edward Yang, Zachary DeVito, Martin Rai- son, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, L...
2019
-
[43]
Optical flow estimation using a spatial pyramid network
Anurag Ranjan and Michael J Black. Optical flow estimation using a spatial pyramid network. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 4161–4170, 2017. 1, 3
2017
-
[44]
Sudderth, and Jan Kautz
Zhile Ren, Orazio Gallo, Deqing Sun, Ming-Hsuan Yang, Erik B. Sudderth, and Jan Kautz. A fusion approach for multi- frame optical flow estimation. In 2019 IEEE Winter Con- ference on Applications of Computer Vision (WACV), pages 2077–2086, 2019. 2
2019
-
[45]
Richter, Zeeshan Hayder, and Vladlen Koltun
Stephan R. Richter, Zeeshan Hayder, and Vladlen Koltun. Playing for benchmarks. In IEEE International Conference on Computer Vision, ICCV 2017, Venice, Italy, October 22-29, 2017, pages 2232–2241, 2017. 5
2017
-
[46]
Mobilenetv2: Inverted residuals and linear bottlenecks
Mark Sandler, Andrew Howard, Menglong Zhu, Andrey Zh- moginov, and Liang-Chieh Chen. Mobilenetv2: Inverted residuals and linear bottlenecks. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 4510–4520, 2018. 4
2018
-
[47]
Videoflow: Exploiting temporal cues for multi-frame optical flow estimation
Xiaoyu Shi, Zhaoyang Huang, Weikang Bian, Dasong Li, Manyuan Zhang, Ka Chun Cheung, Simon See, Hongwei Qin, Jifeng Dai, and Hongsheng Li. Videoflow: Exploiting temporal cues for multi-frame optical flow estimation. arXiv preprint arXiv:2303.08340, 2023. 2, 6
2023 arXiv
-
[48]
Flowformer++: Masked cost volume autoencoding for pretraining optical flow estimation
Xiaoyu Shi, Zhaoyang Huang, Dasong Li, Manyuan Zhang, Ka Chun Cheung, Simon See, Hongwei Qin, Jifeng Dai, and Hongsheng Li. Flowformer++: Masked cost volume autoencoding for pretraining optical flow estimation. arXiv preprint arXiv:2303.01237, 2023. 1, 3, 6
2023 arXiv
-
[49]
Fine-grained motion representation for template-free visual tracking
Kai Shuang, Yuheng Huang, Yue Sun, Zhun Cai, and Hao Guo. Fine-grained motion representation for template-free visual tracking. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pages 671– 680, 2020. 1
2020
-
[50]
Two-stream con- volutional networks for action recognition in videos
Karen Simonyan and Andrew Zisserman. Two-stream con- volutional networks for action recognition in videos. In Pro- ceedings of the 27th International Conference on Neural In- formation Processing Systems - Volume 1 , page 568–576, Cambridge, MA, USA, 2014. MIT Press. 1
2014
-
[51]
Pwc-net: Cnns for optical flow using pyramid, warping, and cost volume
Deqing Sun, Xiaodong Yang, Ming-Yu Liu, and Jan Kautz. Pwc-net: Cnns for optical flow using pyramid, warping, and cost volume. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 8934–8943,
-
[52]
Models matter, so does training: An empirical study of cnns for optical flow estimation
Deqing Sun, Xiaodong Yang, Ming-Yu Liu, and Jan Kautz. Models matter, so does training: An empirical study of cnns for optical flow estimation. IEEE transactions on pattern analysis and machine intelligence, 42(6):1408–1423, 2019. 5
2019
-
[53]
Autoflow: Learning a better training set for optical flow
Deqing Sun, Daniel Vlasic, Charles Herrmann, Varun Jampani, Michael Krainin, Huiwen Chang, Ramin Zabih, William T Freeman, and Ce Liu. Autoflow: Learning a better training set for optical flow. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognitio...
2021
-
[54]
Disentan- gling architecture and training for optical flow
Deqing Sun, Charles Herrmann, Fitsum Reda, Michael Ru- binstein, David J Fleet, and William T Freeman. Disentan- gling architecture and training for optical flow. In Computer Vision–ECCV 2022: 17th European Conference, Tel Aviv, Is- rael, October 23–27, 2022, Proceedings, Part...
2022
-
[55]
Efficientnet: Rethinking model scaling for convolutional neural networks
Mingxing Tan and Quoc Le. Efficientnet: Rethinking model scaling for convolutional neural networks. In International conference on machine learning, pages 6105–6114. PMLR,
-
[56]
Efficientnetv2: Smaller models and faster training
Mingxing Tan and Quoc Le. Efficientnetv2: Smaller models and faster training. In International conference on machine learning, pages 10096–10106. PMLR, 2021. 4
2021
-
[57]
Mixconv: Mixed depth- wise convolutional kernels
Mingxing Tan and Quoc V Le. Mixconv: Mixed depth- wise convolutional kernels. arXiv preprint arXiv:1907.09595,
1907 arXiv
-
[58]
Mnasnet: Platform-aware neural architecture search for mobile
Mingxing Tan, Bo Chen, Ruoming Pang, Vijay Vasudevan, Mark Sandler, Andrew Howard, and Quoc V Le. Mnasnet: Platform-aware neural architecture search for mobile. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 2820–2828, 2019. 4
2019
-
[59]
Raft: Recurrent all-pairs field transforms for optical flow
Zachary Teed and Jia Deng. Raft: Recurrent all-pairs field transforms for optical flow. In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part II 16, pages 402–419. Springer, 2020. 1, 3, 6
2020
-
[60]
Wadekar and Abhishek Chaurasia
Shakti N. Wadekar and Abhishek Chaurasia. Mobilevitv3: Mobile-friendly vision transformer with simple and effective fusion of local, global and input features. 2022. 5
2022
-
[61]
Correlation flow: robust optical flow using kernel cross- correlators
Chen Wang, Tete Ji, Thien-Minh Nguyen, and Lihua Xie. Correlation flow: robust optical flow using kernel cross- correlators. In 2018 IEEE International Conference on Robotics and Automation (ICRA) , pages 836–841. IEEE,
2018
-
[62]
Fbnet: Hardware-aware efficient con- vnet design via differentiable neural architecture search
Bichen Wu, Xiaoliang Dai, Peizhao Zhang, Yanghan Wang, Fei Sun, Yiming Wu, Yuandong Tian, Peter Vajda, Yangqing Jia, and Kurt Keutzer. Fbnet: Hardware-aware efficient con- vnet design via differentiable neural architecture search. In Proceedings of the IEEE/CVF Conference on C...
2019
-
[63]
High-resolution optical flow from 1d attention and correlation
Haofei Xu, Jiaolong Yang, Jianfei Cai, Juyong Zhang, and Xin Tong. High-resolution optical flow from 1d attention and correlation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 10498–10507, 2021. 5
2021
-
[64]
Gmflow: Learning optical flow via global matching
Haofei Xu, Jing Zhang, Jianfei Cai, Hamid Rezatofighi, and Dacheng Tao. Gmflow: Learning optical flow via global matching. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 8121–8130,
-
[65]
Accurate optical flow via direct cost volume processing
Jia Xu, Ren´e Ranftl, and Vladlen Koltun. Accurate optical flow via direct cost volume processing. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recogni- tion, pages 1289–1297, 2017. 1
2017
-
[66]
Video enhancement with task-oriented flow
Tianfan Xue, Baian Chen, Jiajun Wu, Donglai Wei, and William T Freeman. Video enhancement with task-oriented flow. International Journal of Computer Vision, 127:1106– 1125, 2019. 3
2019
-
[67]
V olumetric correspon- dence networks for optical flow
Gengshan Yang and Deva Ramanan. V olumetric correspon- dence networks for optical flow. Advances in neural informa- tion processing systems, 32, 2019. 1
2019
-
[68]
Unsupervised motion representation enhanced network for action recogni- tion
Xiaohang Yang, Lingtong Kong, and Jie Yang. Unsupervised motion representation enhanced network for action recogni- tion. In ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 2445–2449. IEEE, 2021. 1
2021
-
[69]
Learning video stabiliza- tion using optical flow
Jiyang Yu and Ravi Ramamoorthi. Learning video stabiliza- tion using optical flow. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 8159–8167, 2020. 1, 2
2020
-
[70]
A du- ality based approach for realtime tv-l 1 optical flow
Christopher Zach, Thomas Pock, and Horst Bischof. A du- ality based approach for realtime tv-l 1 optical flow. In Pattern Recognition: 29th DAGM Symposium, Heidelberg, Germany, September 12-14, 2007. Proceedings 29 , pages 214–223. Springer, 2007. 1
2007
-
[71]
Separable flow: Learning motion cost volumes for optical flow estimation
Feihu Zhang, Oliver J Woodford, Victor Adrian Prisacariu, and Philip HS Torr. Separable flow: Learning motion cost volumes for optical flow estimation. In Proceedings of the IEEE/CVF International Conference on Computer Vision , pages 10807–10817, 2021. 5
2021
-
[72]
Global matching with overlapping at- tention for optical flow estimation
Shiyu Zhao, Long Zhao, Zhixing Zhang, Enyu Zhou, and Dimitris Metaxas. Global matching with overlapping at- tention for optical flow estimation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 17592–17601, 2022. 1, 3
2022
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.