REVIEW 4 major objections 3 minor 68 references
CoDAT: Collaborative Dual-Attention Transformer with Low-Cost Temporal Modeling for Efficient Edge Action Recognition
T0 review · 4 major / 3 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read A dual-attention transformer brings near-state-of-the-art action recognition to edge hardware at a fraction of the energy cost.
desk verdict Competent edge action-recognition paper with strong ablations; the 6x speedup is plausible but the ONNX inference protocol is under-documented. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the Collaborative Dual-Attention (CoDA) module and its placement relative to the temporal shift. Strided Single-Head Attention (SSHA) projects the input through pointwise convolutions with a stride $r_s$, reducing the attention from $N$ tokens to $M = (H/r_s) \times (W/r_s)$ tokens with value channels $C_v = C/4$, so the attention cost drops from $O(N^2 C)$ to $O(M^2 C_v)$. Spatial Convolutional Attention (SCA) computes a single-channel saliency mask from max and average pooling through a $3\times3$ convolution and sigmoid, costing $O(N C)$, and multiplies it into the features. A learnable projection fuses the two branches. Around each CoDA block, a zero-parameter TShift layer permutes a small fraction of channels across adjacent frames and is placed after CoDA but before the final ConvFFN, so the attention branches always see temporally clean single-frame features while the feed-forward network integrates the shifted temporal context.
What would settle it
Re-measuring the baselines from Tables 4-7 under the exact same inference settings as CoDAT (same single-clip, single-crop input, same exported model format, same hardware and power metering, same warm-up and repetition protocol) would settle the claim. In particular, if TokShift and LAPS on their public checkpoints run at less than 6x the latency and less than 13x the FLOPs of CoDAT-S384 on UCF-101, the headline efficiency claim fails.
Extended reading notes
Core claim
CoDAT's central claim is that the accuracy–efficiency frontier for edge action recognition can be moved by jointly attacking the two dominant costs of video transformers: token count and channel redundancy. The paper shows that a single attention head over spatially strided queries, keys, and values, combined with a lightweight convolutional attention branch, recovers the global and local cues that full multi-head attention provides, while a parameter-free temporal shift captures motion between frames. Empirically, CoDAT-S384 matches TokShift and LAPS on UCF-101 at roughly 6-times lower latency and up to 13-times fewer FLOPs, CoDAT-M runs about 2-times faster than EfficientViT-384 and FastViT-S12 on ImageNet-1K at comparable accuracy, and on Kinetics-400 CoDAT is up to 2.9-times faster than Video Swin Transformer and about 2-times faster than temporal-shift ViT variants while staying within one percentage point of their Top-1 accuracy. The efficiency gains persist on the fine-grained MA-52 benchmark, where CoDAT-M384 reaches the highest fine-grained Top-1 accuracy among the compared models.
Load-bearing premise
The load-bearing premise is that the published baseline numbers were produced under conditions close enough to CoDAT's own training and single-clip inference settings to be directly comparable; if the baselines used different training recipes, resolutions, or multi-view protocols, the claimed accuracy and speed advantages could shrink or vanish.
Editorial extensions
If this is right
- Real-time action recognition becomes practical on low-power hardware: on Jetson AGX Orin, CoDAT-S runs UCF-101 inference at 0.89 ms/frame with 13.79 mJ/frame energy.
- The efficiency gains transfer from image to video: CoDAT-L matches ViT-S/DeiT-S on ImageNet-1K with about 3x fewer parameters and 2.8x higher throughput.
- On Kinetics-400, CoDAT-M384 outperforms or matches CNN baselines such as SlowFast and transformer baselines such as TokShift while requiring over 2x fewer GFLOPs and running 2.75x to 3.9x faster.
- On the fine-grained MA-52 benchmark, CoDAT-M reaches the best fine-grained Top-1 accuracy among the compared models while running 2.7x faster than SlowFast and 5x faster than UniFormer-B.
- The approach also works on compressed video: on MPEG-4-compressed UCF-101, CoDAT-S384 surpasses MTRFN by 0.8 percentage points at 8.8x fewer FLOPs.
Reading between the lines
- If the energy and latency ratios hold under fully standardized benchmarking, the practical meaning is that always-on camera analytics on battery-powered nodes becomes plausible; the implied operating point is several frames per second to real time on boards that previously could only run lightweight CNNs.
- The fixed stride ratio $r_s$ per stage is a knob the paper leaves untouched; a content-adaptive stride that increases compression on easy frames could push the efficiency gains further than the static design the paper evaluates.
- The placement principle found here—keep attention branches temporally clean and shift channels only into the feed-forward network—suggests a general recipe for shift-based video transformers that could be tested on other lightweight backbones.
- A natural testable extension is streaming deployment: the paper measures per-frame cost on fixed clips, but batched frame shifting with a buffer would show whether the $2L+1$ temporal receptive field behaves as claimed under continuous input.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This manuscript introduces CoDAT, a three-stage vision transformer for edge action recognition. The central module, CoDA, combines a strided single-head attention (SSHA) branch that reduces spatial tokens and value channels with a spatial convolutional attention (SCA) branch, and a parameter-free temporal shift (TShift) is inserted before the final ConvFFN of each block. The paper evaluates CoDAT on ImageNet-1K, Kinetics-400, MA-52, and UCF-101, and reports latency, throughput, and energy on Jetson AGX Orin and Raspberry Pi 5. The authors claim that CoDAT matches or approaches the accuracy of much heavier CNN and ViT baselines while reducing FLOPs, latency, and energy substantially, with the strongest claims made for UCF-101, where CoDAT-S384 is reported to match TokShift and LAPS at 95.4% top-1 accuracy while running about 6x faster.
Significance. The architecture is internally coherent and the empirical scope is broader than in many edge action-recognition papers: four datasets, two hardware platforms, an explicit energy-measurement procedure, and systematic ablations of stride ratio, channel ratio, shift placement, and stage design. The SSHA complexity bookkeeping and the TShift receptive-field argument are plausible, and the internal ablations in Tables 8 and 9 support the claim that CoDAT's efficiency comes from the strided and channel-compressed attention design rather than from simple model shrinkage. If the efficiency-accuracy claims survive a fair cross-model benchmark protocol, the paper would be a useful contribution to edge IoT perception. However, the abstract and conclusion currently contain quantitative claims that do not match the tables, and the latency/energy comparisons are not yet protocol-fair as written, so the central claim cannot be accepted without substantial revision.
major comments (4)
- [Section 4.1, Table 5] The latency and energy comparisons are load-bearing but the runtime configuration is incomplete. Section 4.1 states only that inference uses ONNX Runtime with CUDA on Jetson AGX Orin at a 50 W power cap; it does not report numerical precision (FP32/FP16/TF32/INT8), ONNX optimization level, or whether a TensorRT execution provider was used. This matters concretely: in Table 5, the TokShift row reports 135 GFLOPs at 10.15 ms, which implies about 13.3 TFLOPS sustained, well above the practical FP32 throughput of a 50 W AGX Orin; the baselines therefore appear to have been run on a faster non-default path while the paper does not state whether CoDAT used the same one. Unless the same ONNX export options, execution provider, and precision are documented for every model, the '6x faster', '2.9x faster', and energy ratios in the abstract, conclusion, and Section 4.3 are not established. Please provide a per-model runtime configuration table or restrict the speed claims to models measured under identical settings.
- [Abstract, Conclusion, Tables 4-6] Several quantitative claims in the abstract and conclusion do not match the tables. On Kinetics-400, the conclusion says CoDAT is 2.9x faster than VSwin-T and nearly 2x faster than LAPS/TokShift; Table 5 gives CoDAT-L384 at 4.74 ms/F versus VSwin-T at 6.96 ms/F (1.47x), and CoDAT-M384 at 2.56 ms/F versus LAPS at 9.34 ms/F and TokShift at 10.15 ms/F (about 3.6x and 4.0x). On MA-52, the conclusion reports 63.6% fine-grained accuracy and latency reductions of more than 5x and 9x relative to VSwin-T and UniFormer-B, while Table 6 gives 63.19% and ratios of 2.72x and 5x. In the image-classification conclusion, CoDAT-L is said to have 2.5x lower latency than ViT-S, while Table 4 gives 1.71 ms / 1.33 ms = 1.29x. Please correct all abstract and conclusion numbers to agree with the tables, or rerun the measurements and update the tables.
- [Section 4.1, Tables 4-7] The cross-model benchmark parity is not sufficiently documented. Tables 4-7 mix self-reported and published baseline numbers, and Section 4.1 says multi-view configurations are applied 'exclusively for top-1 accuracy evaluation based on each model default configuration,' but it is unclear which view configuration corresponds to each baseline row and whether the baselines were run by the authors or taken from the literature. Several rows, such as S2AFormer-mini in Table 4 and I3D, STM, and MViT in Table 5, have no latency or energy entries, so the efficiency comparisons cover different model subsets. To make the central speed and energy claims benchmark-fair, please specify for every baseline the source of the accuracy number (authors' run vs published), the exact view/crop protocol, and the runtime configuration used for latency and energy.
- [Tables 4 and 7] The accuracy results appear to be single-run point estimates, which is problematic for claims that hinge on small differences. The central UCF-101 claim rests on CoDAT-S384 exactly matching TokShift and LAPS at 95.4% (Table 7), and several Table 4 conclusions rely on 0.1-0.4% accuracy differences (e.g., CoDAT-S versus SHViT-S3, or CoDAT-L versus Swin-T). Without multiple seeds, confidence intervals, or at least a stated fixed-seed protocol, these differences are within typical run-to-run noise, especially for a small dataset like UCF-101. Please report variance or multiple seeds for the main comparisons, or soften claims that depend on sub-1% differences.
minor comments (3)
- [Table 9] The row labeled 'TShift1/4C' in Table 9 reports a latency of 7.2 ms while all neighboring rows report 0.89 ms and the FLOPs are unchanged; this is almost certainly a typo and should be corrected.
- [Section 3.1.2] The text says SSHA gives 'a reduction of r_s^4-times over the standard full-resolution O(N^2 C)', but since the value channels are C_v = C/4, the actual reduction is 4 r_s^4. Please correct the wording or state explicitly that the factor of 4 is absorbed into the channel reduction.
- [Throughout] There are several typos and inconsistent labels, including 'VSwim-T' for VSwin-T in Section 4.3.5, the stray 'Trans.' row in Table 5, and the phrase 'the official train/test splits 1.' Please copyedit the final version.
Circularity Check
No significant circularity: CoDAT's claims are empirical measurements and standard complexity bookkeeping, with no fitted parameter renamed as a prediction and no load-bearing self-citation chain.
full rationale
The paper's central claims are empirical: Top-1 accuracies, latencies, energies, and FLOPs are measured or arithmetically computed from the proposed architecture and compared against external baselines. The analytical statements in the paper are complexity bounds for SSHA (O(M^2 C_v + N C) with M = HW/r_s^2) and the temporal receptive field formula 2L+1 frames, which is imported from the TSM literature and used only as a design motivation, not as a derivation of the benchmark results. The design choices such as C_v = 1/4, stride ratios, and TShift placement are selected via ablations on the same benchmarks, which is standard engineering practice and does not constitute fitting a parameter and then calling the result a prediction. The self-citations in the paper (e.g., MicroViT-S3 as a baseline and prior works by the same authors) are background comparisons and are not load-bearing for the proposed architecture's validity. No uniqueness theorem, ansatz, or renamed known result is invoked to force the conclusions. The reported speedups and energy savings assume protocol parity across baselines, which is a correctness and benchmark-comparability risk rather than a circularity issue. Therefore, the derivation chain is empirically self-contained and no circular step can be exhibited from the paper's own equations or citations.
Assumptions & free parameters
free parameters (6)
- Value channel ratio Cv = 1/4 =
1/4 C
- Per-stage stride ratios rs =
[2,2,1]
- FFN expansion ratio =
2
- Channel dims per stage =
e.g., CoDAT-S [192,384,448]
- Block depth per stage =
e.g., CoDAT-S [1,2,2]
- TShift channel fraction =
1/8 C
assumptions (4)
- domain assumption ONNX Runtime faithfully implements all compared models and the power readings from INA3221/MXL7704 represent deployment energy.
- domain assumption Published baseline accuracies are accurate and obtained under comparable training/evaluation protocols.
- standard math Temporal shift yields TRF = 2L+1 as claimed in [40].
- domain assumption Standard dataset splits and labels are used.
Cite this review
Pith. "Pith review of CoDAT: Collaborative Dual-Attention Transformer with Low-Cost Temporal Modeling for Efficient Edge Action Recognition." pith.science (2026). https://pith.science/paper/DGLLOD2Q
@misc{pith2026260806691,
author = {Pith},
title = {Pith review of: CoDAT: Collaborative Dual-Attention Transformer with Low-Cost Temporal Modeling for Efficient Edge Action Recognition},
year = {2026},
howpublished = {\url{https://pith.science/paper/DGLLOD2Q}},
note = {Machine review of arXiv:2608.06691}
}
read the original abstract
Real-time human action recognition on Internet-of-Things (IoT) edge devices requires models that capture rich spatio-temporal cues within strict latency, memory, and power envelopes. Current 3D CNNs, video transformers, and shift-based ViT deliver high accuracy but come at computational costs that preclude edge IoT deployment. This paper proposes CoDAT, a Collaborative Dual-Attention Transformer that replaces conventional multi-head attention with a lightweight dual-branch module: Spatial Convolutional Attention (SCA) for local aggregation and Strided Single-Head Attention (SSHA) for global context. SSHA jointly compresses the spatial resolution and channel dimensions of the query, key, and value tensors via stride-based sparse projection, then fuses the resulting global and local features at a markedly reduced cost. To enable temporal communication across frames, a parameter-free TShift module is embedded in each block. Extensive experiments on Jetson AGX Orin and Raspberry Pi 5 demonstrate that CoDAT achieves an energy-accuracy balance in both image and action recognition. On ImageNet-1K, CoDAT-M runs 2x faster than EfficientViT384 and FastViT-S12 at comparable accuracy, and CoDAT-L matches ViT-S with 3x fewer parameters at 2x higher throughput. On Kinetics-400 and MA-52, CoDAT achieves competitive Top-1 accuracy against state-of-the-art CNN, transformer, and hybrid baselines while running up to 2.9x faster than VSwin-T, 2x faster than ViT-Temporal-Shift variants, and 5x faster than UniFormer-B. On UCF-101, CoDAT-S384 matches TokShift and LAPS while being 6x faster and requiring up to 13x fewer FLOPs, establishing an efficiency-accuracy balance for real-time action recognition in edge IoT perception systems. Code is available at https://github.com/novendrastywn/CoDAT .
Figures
Figures from the paper (4 more)
Reference graph
Works this paper leans on
-
[1]
H. Sun, W. Shi, X. Liang, and Y. Yu, “Vu: Edge computing-enabled video usefulness detection and its application in large-scale video surveillance systems,”IEEE Internet of Things J., vol. 7, no. 2, pp. 800–817, 2019
work page 2019
-
[2]
Strack: Robust tracking of small objects in low-light conditions,
S. B. J. Khan, Z. Peng, M. M. Kamal, H. G. Mohamed, Q. M. Kharma, M. Sheraz, and T. C. Chuah, “Strack: Robust tracking of small objects in low-light conditions,”IEEE Access, 2025
work page 2025
-
[3]
K. Muhammad, H. Ullah, M. S. Obaidat, A. Ullah, A. Munir, M. Sajjad, and V . H. C. De Albuquerque, “Ai-driven salient soc- cer events recognition framework for next-generation iot-enabled environments,”IEEE Internet of Things J., vol. 10, no. 3, pp. 2202– 2214, 2021
work page 2021
-
[4]
Contactless patient care using hospital iot: Cctv-camera-based physiological monitoring in icu,
H. Wang, J. Huang, G. Wang, H. Lu, and W. Wang, “Contactless patient care using hospital iot: Cctv-camera-based physiological monitoring in icu,”IEEE Internet of Things J., vol. 11, no. 4, pp. 5781–5797, 2023
work page 2023
-
[5]
Optimization for short video propagation based on user interaction analysis in edge networks,
Z. Li and Q. Yang, “Optimization for short video propagation based on user interaction analysis in edge networks,”IEEE Internet of Things J., 2025
work page 2025
-
[6]
Panacea+: Panoramic and con- trollable video generation for autonomous driving,
Y. Wen, Y. Zhao, Y. Liu, B. Huang, F. Jia, Y. Wang, C. Zhang, T. Wang, X. Sun, and X. Zhang, “Panacea+: Panoramic and con- trollable video generation for autonomous driving,”IEEE Trans. on Circuits and Syst. for Video Technol., pp. 1–1, 2025
work page 2025
-
[7]
S. B. J. Khan, P . Zhang, M. M. Kamal, A. Alharbi, A. Tolba, M. Sheraz, and T. C. Chuah, “Tracenet: A novel modular frame- work for robust multi-object tracking in crowded and dynamic environments,”Alexandria Engineering Journal, vol. 137, pp. 401– 413, 2026
work page 2026
-
[8]
Hamot: A hierarchical adaptive framework for robust multi-object tracking in complex environments,
J. K. S. Baz, P . Zhang, M. M. Kamal, H. G. Mohamed, M. Sheraz, and T. C. Chuah, “Hamot: A hierarchical adaptive framework for robust multi-object tracking in complex environments,”CMES- Computer Modeling in Engineering and Sciences, vol. 145, no. 1, pp. 947–969, 2025
work page 2025
Show all 68 references
-
[9]
Imagenet classifi- cation with deep convolutional neural networks,
A. Krizhevsky, I. Sutskever, and G. E. Hinton, “Imagenet classifi- cation with deep convolutional neural networks,”Adv. Neural Inf. Process. Syst., vol. 25, 2012
2012
-
[10]
Facelivt: Face recognition using linear vision transformer with structural reparameterization for mobile device,
N. Setyawan, C.-C. Sun, M.-H. Hsu, W.-K. Kuo, and J.-W. Hsieh, “Facelivt: Face recognition using linear vision transformer with structural reparameterization for mobile device,” inProc. IEEE Int. Conf. Image Process. (ICIP), pp. 1720–1725, 2025
2025
-
[11]
Facelivtv2: An improved hybrid architecture for efficient mobile face recognition,
N. Setyawan, C.-C. Sun, M.-H. Hsu, W.-K. Kuo, and J.-W. Hsieh, “Facelivtv2: An improved hybrid architecture for efficient mobile face recognition,”IEEE Trans. Biom. Behav. Identity Sci., 2026
2026
-
[12]
Video analytics for detecting motorcy- clist helmet rule violations,
C.-M. Tsai, J.-W. Hsieh, M.-C. Chang, G.-L. He, P .-Y. Chen, W.-T. Chang, and Y.-K. Hsieh, “Video analytics for detecting motorcy- clist helmet rule violations,” in2023 IEEE/CVF Conf. Comput. Vis. and Pattern Recognit. Workshops (CVPRW), pp. 5366–5374, 2023
2023
-
[13]
Smiletrack: Similarity learning for occlusion-aware multiple object tracking,
Y.-H. Wang, J.-W. Hsieh, P .-Y. Chen, M.-C. Chang, H.-H. So, and X. Li, “Smiletrack: Similarity learning for occlusion-aware multiple object tracking,” inProc. AAAI Conf. Artif. Intell., vol. 38, pp. 5740– 5748, 2024
2024
-
[14]
Lighttrack-reid: A lightweight and occlusion-robust framework for multi-object tracking,
S. B. J. Khan, P . Zhang, M. M. Kamal, and A. K. J. Saudagar, “Lighttrack-reid: A lightweight and occlusion-robust framework for multi-object tracking,”Plos one, vol. 21, no. 3, p. e0342246, 2026
2026
-
[15]
Ssp-sam: Sam with semantic- spatial prompt for referring expression segmentation,
W. Tang, X. Liu, Y. Sun, and Z. Li, “Ssp-sam: Sam with semantic- spatial prompt for referring expression segmentation,”IEEE Trans. Circuits Syst. Video Technol., pp. 1–1, 2025
2025
-
[16]
Ucf101: A dataset of 101 human actions classes from videos in the wild,
K. Soomro, A. R. Zamir, and M. Shah, “Ucf101: A dataset of 101 human actions classes from videos in the wild,”arXiv preprint arXiv:1212.0402, 2012
2012 arXiv
-
[17]
The kinet- ics human action video dataset,
W. Kay, J. Carreira, K. Simonyan, B. Zhang, C. Hillier, S. Vijaya- narasimhan, F. Viola, T. Green, T. Back, P . Natsev,et al., “The kinet- ics human action video dataset,”arXiv preprint arXiv:1705.06950, 2017
2017 arXiv
-
[18]
Benchmarking micro-action recognition: Dataset, methods, and applications,
D. Guo, K. Li, B. Hu, Y. Zhang, and M. Wang, “Benchmarking micro-action recognition: Dataset, methods, and applications,” IEEE Trans. Circuits Syst. Video Technol., vol. 34, no. 7, pp. 6238– 6252, 2024
2024
-
[19]
Quo vadis, action recognition? a new model and the kinetics dataset,
J. Carreira and A. Zisserman, “Quo vadis, action recognition? a new model and the kinetics dataset,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., pp. 6299–6308, 2017
2017
-
[20]
Slowfast networks for video recognition,
C. Feichtenhofer, H. Fan, J. Malik, and K. He, “Slowfast networks for video recognition,” inProc. IEEE/CVF Int. Conf. on Comput. Vis., pp. 6202–6211, 2019
2019
-
[21]
Spatio- temporal adaptive network with bidirectional temporal difference for action recognition,
Z. Li, J. Li, Y. Ma, R. Wang, Z. Shi, Y. Ding, and X. Liu, “Spatio- temporal adaptive network with bidirectional temporal difference for action recognition,”IEEE Trans. Circuits Syst. Video Technol., vol. 33, no. 9, pp. 5174–5185, 2023
2023
-
[22]
Agpn: Action granu- larity pyramid network for video action recognition,
Y. Chen, H. Ge, Y. Liu, X. Cai, and L. Sun, “Agpn: Action granu- larity pyramid network for video action recognition,”IEEE Trans. Circuits Syst. Video Technol., vol. 33, no. 8, pp. 3912–3923, 2023
2023
-
[23]
An image is worth 16x16 words: Transformers for image recognition at scale,
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, et al., “An image is worth 16x16 words: Transformers for image recognition at scale,”arXiv preprint arXiv:2010.11929, 2020
2010 arXiv
-
[24]
Is space-time attention all you need for video understanding?,
G. Bertasius, H. Wang, and L. Torresani, “Is space-time attention all you need for video understanding?,” inIcml, vol. 2, p. 4, 2021
2021
-
[25]
Video trans- former network,
D. Neimark, O. Bar, M. Zohar, and D. Asselmann, “Video trans- former network,” inProc. IEEE/CVF Int. Conf. on Comput. Vis., pp. 3163–3172, 2021
2021
-
[26]
Video swin transformer,
Z. Liu, J. Ning, Y. Cao, Y. Wei, Z. Zhang, S. Lin, and H. Hu, “Video swin transformer,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., pp. 3202–3211, 2022
2022
-
[27]
Uniformer: Unifying convolution and self-attention for visual recognition,
K. Li, Y. Wang, J. Zhang, P . Gao, G. Song, Y. Liu, H. Li, and Y. Qiao, “Uniformer: Unifying convolution and self-attention for visual recognition,”IEEE Trans. Pattern Anal. Mach. Intell., vol. 45, no. 10, pp. 12581–12600, 2023
2023
-
[28]
Token shift transformer for video classification,
H. Zhang, Y. Hao, and C.-W. Ngo, “Token shift transformer for video classification,” inProc. ACM Int. Conf. Multimed., pp. 917– 925, 2021
2021
-
[29]
Long-term leap attention, short-term periodic shift for video classification,
H. Zhang, L. Cheng, Y. Hao, and C.-w. Ngo, “Long-term leap attention, short-term periodic shift for video classification,” in Proceedings of the 30th acm international conference on multimedia, pp. 5773–5782, 2022
2022
-
[30]
Temporal shift module-based vision transformer network for action recognition,
K. Zhang, M. Lyu, X. Guo, L. Zhang, and C. Liu, “Temporal shift module-based vision transformer network for action recognition,” IEEE Access, vol. 12, pp. 47246–47257, 2024
2024
-
[31]
Tsm: Temporal shift module for efficient and scalable video understanding on edge devices,
J. Lin, C. Gan, K. Wang, and S. Han, “Tsm: Temporal shift module for efficient and scalable video understanding on edge devices,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 44, no. 05, pp. 2760– 2774, 2022
2022
-
[32]
Deep residual learning for image recognition,
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., pp. 770–778, 2016
2016
-
[33]
Swin transformer: Hierarchical vision transformer using shifted windows,
Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo, “Swin transformer: Hierarchical vision transformer using shifted windows,” inProc. IEEE/CVF Int. Conf. on Comput. Vis., pp. 10012– 10022, 2021
2021
-
[34]
Movinets: Mobile video networks for efficient video ACCEPTED ON IEEE INTERNET OF THINGS JOURNAL 15 recognition,
D. Kondratyuk, L. Yuan, Y. Li, L. Zhang, M. Tan, M. Brown, and B. Gong, “Movinets: Mobile video networks for efficient video ACCEPTED ON IEEE INTERNET OF THINGS JOURNAL 15 recognition,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., pp. 16020–16030, 2021
2021
-
[35]
Deepsensemoe: Harnessing power of time series foundation models for few-shot human activity recognition,
Z. Fu, D. Cheng, L. Zhang, W. Huang, Z. Chen, and H. Wu, “Deepsensemoe: Harnessing power of time series foundation models for few-shot human activity recognition,” inProc. AAAI Conf. Artif. Intell., vol. 40, pp. 292–299, 2026
2026
-
[36]
Sensor- prompt tuning: Aligning time series foundational models with motion sensors for few-shot activity recognition,
X. Liu, D. Cheng, Z. Fu, L. Zhang, H. Wu, and A. Song, “Sensor- prompt tuning: Aligning time series foundational models with motion sensors for few-shot activity recognition,”IEEE Trans. Mob. Comput., 2026
2026
-
[37]
Deep convolutional state space model as human activity recognizer,
L. Wang, C. Bu, M. Yao, D. Xiong, S. Wang, D. Cheng, L. Zhang, H. Wu, and A. Song, “Deep convolutional state space model as human activity recognizer,”Information Fusion, p. 103982, 2025
2025
-
[38]
Learn- ing spatiotemporal features with 3d convolutional networks,
D. Tran, L. Bourdev, R. Fergus, L. Torresani, and M. Paluri, “Learn- ing spatiotemporal features with 3d convolutional networks,” in Proc. IEEE Int. Conf. Comput. Vis., pp. 4489–4497, 2015
2015
-
[39]
A closer look at spatiotemporal convolutions for action recognition,
D. Tran, H. Wang, L. Torresani, J. Ray, Y. LeCun, and M. Paluri, “A closer look at spatiotemporal convolutions for action recognition,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., pp. 6450– 6459, 2018
2018
-
[40]
Tsm: Temporal shift module for efficient video understanding,
J. Lin, C. Gan, and S. Han, “Tsm: Temporal shift module for efficient video understanding,” inProc. IEEE/CVF Int. Conf. on Comput. Vis., pp. 7083–7093, 2019
2019
-
[41]
X3d: Expanding architectures for efficient video recognition,
C. Feichtenhofer, “X3d: Expanding architectures for efficient video recognition,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., pp. 203–213, 2020
2020
-
[42]
Mtrfn: Multiscale temporal receptive field network for compressed video action recognition at edge servers,
L. He, M. Zhang, S. Zhang, L. Wang, and F. Li, “Mtrfn: Multiscale temporal receptive field network for compressed video action recognition at edge servers,”IEEE Internet of Things J., vol. 9, no. 15, pp. 13965–13977, 2022
2022
-
[43]
Temporal transformer networks with self-supervision for action recognition,
Y. Zhang, J. Li, N. Jiang, G. Wu, H. Zhang, Z. Shi, Z. Liu, Z. Wu, and X. Liu, “Temporal transformer networks with self-supervision for action recognition,”IEEE Internet of Things J., vol. 10, no. 14, pp. 12999–13011, 2023
2023
-
[44]
Pyramid vision transformer: A versatile backbone for dense prediction without convolutions,
W. Wang, E. Xie, X. Li, D.-P . Fan, K. Song, D. Liang, T. Lu, P . Luo, and L. Shao, “Pyramid vision transformer: A versatile backbone for dense prediction without convolutions,” inProc. IEEE/CVF Int. Conf. on Comput. Vis., pp. 568–578, 2021
2021
-
[45]
Pvt v2: Improved baselines with pyramid vision transformer,
W. Wang, E. Xie, X. Li, D.-P . Fan, K. Song, D. Liang, T. Lu, P . Luo, and L. Shao, “Pvt v2: Improved baselines with pyramid vision transformer,”Comput. Vis. Medi, vol. 8, no. 3, pp. 415–424, 2022
2022
-
[46]
Edgenext: efficiently amalgamated cnn-transformer architecture for mobile vision applications,
M. Maaz, A. Shaker, H. Cholakkal, S. Khan, S. W. Zamir, R. M. Anwer, and F. Shahbaz Khan, “Edgenext: efficiently amalgamated cnn-transformer architecture for mobile vision applications,” in European conference on computer vision, pp. 3–20, Springer, 2022
2022
-
[47]
Fastvit: A fast hybrid vision transformer using structural reparameteriza- tion,
P . K. A. Vasu, J. Gabriel, J. Zhu, O. Tuzel, and A. Ranjan, “Fastvit: A fast hybrid vision transformer using structural reparameteriza- tion,” inProc. IEEE/CVF Int. Conf. on Comput. Vis., pp. 5785–5795, 2023
2023
-
[48]
Effi- cientvit: Memory efficient vision transformer with cascaded group attention,
X. Liu, H. Peng, N. Zheng, Y. Yang, H. Hu, and Y. Yuan, “Effi- cientvit: Memory efficient vision transformer with cascaded group attention,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., pp. 14420–14430, 2023
2023
-
[49]
Rethinking vision transformers for mobilenet size and speed,
Y. Li, J. Hu, Y. Wen, G. Evangelidis, K. Salahi, Y. Wang, S. Tulyakov, and J. Ren, “Rethinking vision transformers for mobilenet size and speed,” inProc. IEEE/CVF Int. Conf. on Comput. Vis., pp. 16889– 16900, 2023
2023
-
[50]
Shvit: Single-head vision transformer with memory efficient macro design,
S. Yun and Y. Ro, “Shvit: Single-head vision transformer with memory efficient macro design,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., pp. 5756–5767, 2024
2024
-
[51]
S2aformer: Strip self-attention for efficient vision transformer,
G. Xu, W. Huang, W. Jia, J. Li, G. Gao, and G.-J. Qi, “S2aformer: Strip self-attention for efficient vision transformer,”IEEE Trans. Image Process., pp. 1–1, 2025
2025
-
[52]
Training data-efficient image transformers & distilla- tion through attention,
H. Touvron, M. Cord, M. Douze, F. Massa, A. Sablayrolles, and H. J ´egou, “Training data-efficient image transformers & distilla- tion through attention,” inProc. Int. Conf. Mach. Learn., pp. 10347– 10357, PMLR, 2021
2021
-
[53]
Tlee: Temporal-wise and layer-wise early exiting network for efficient video recognition on edge devices,
Q. Wang, W. Fang, and N. N. Xiong, “Tlee: Temporal-wise and layer-wise early exiting network for efficient video recognition on edge devices,”IEEE Internet of Things J., vol. 11, no. 2, pp. 2842– 2854, 2023
2023
-
[54]
Conditional positional encodings for vision transformers,
X. Chu, Z. Tian, B. Zhang, X. Wang, and C. Shen, “Conditional positional encodings for vision transformers,”Int. Conf. Learn. Represent. (ICLR) 2021, 2021
2021
-
[55]
Imagenet large scale visual recognition challenge,
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein,et al., “Imagenet large scale visual recognition challenge,”Int. J. Comput. Vis., vol. 115, pp. 211–252, 2015
2015
-
[56]
Iformer: Integrating convnet and transformer for mo- bile application,
C. Zheng, “Iformer: Integrating convnet and transformer for mo- bile application,”Int. Conf. Learn. Represent. (ICLR), 2025
2025
-
[57]
Microvit: a vision transformer with low complexity self attention for edge device,
N. Setyawan, C.-C. Sun, M.-H. Hsu, W.-K. Kuo, and J.-W. Hsieh, “Microvit: a vision transformer with low complexity self attention for edge device,” inProc. IEEE Int. Symp. Circuits Syst. (ISCAS), pp. 1–5, IEEE, 2025
2025
-
[58]
Large batch optimization for deep learning: Training bert in 76 minutes,
Y. You, J. Li, S. Reddi, J. Hseu, S. Kumar, S. Bhojanapalli, X. Song, J. Demmel, K. Keutzer, and C.-J. Hsieh, “Large batch optimization for deep learning: Training bert in 76 minutes,”arXiv preprint arXiv:1904.00962, 2019
1904 arXiv
-
[59]
Group contextualization for video recognition,
Y. Hao, H. Zhang, C.-W. Ngo, and X. He, “Group contextualization for video recognition,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., pp. 928–938, 2022
2022
-
[60]
Learning spatiotem- poral and motion features in a unified 2d network for action recognition,
M. Wang, J. Xing, J. Su, J. Chen, and Y. Liu, “Learning spatiotem- poral and motion features in a unified 2d network for action recognition,”IEEE Trans. Pattern Anal. Mach. Intell., vol. 45, no. 3, pp. 3347–3362, 2022
2022
-
[61]
Multiscale vision transformers,
H. Fan, B. Xiong, K. Mangalam, Y. Li, Z. Yan, J. Malik, and C. Fe- ichtenhofer, “Multiscale vision transformers,” inProc. IEEE/CVF Int. Conf. on Comput. Vis., pp. 6824–6835, 2021
2021
-
[62]
Dualactnet: Exploiting slowfast architecture for micro-action recognition,
C. Yu, Y. Ru, Z. Xu, H. Wu, H. Yang, and Z. He, “Dualactnet: Exploiting slowfast architecture for micro-action recognition,” in Chinese Conference on Biometric Recognition, pp. 59–68, Springer, 2024
2024
-
[63]
Tdn: Temporal difference networks for efficient action recognition,
L. Wang, Z. Tong, B. Ji, and G. Wu, “Tdn: Temporal difference networks for efficient action recognition,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., pp. 1895–1904, 2021
1904
-
[64]
Mobile video action recognition,
Y. Huo, X. Xu, Y. Lu, Y. Niu, Z. Lu, and J.-R. Wen, “Mobile video action recognition,”arXiv preprint arXiv:1908.10155, 2019
1908 arXiv
-
[65]
Compressed video action recognition,
C.-Y. Wu, M. Zaheer, H. Hu, R. Manmatha, A. J. Smola, and P . Kr¨ahenb ¨uhl, “Compressed video action recognition,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., pp. 6026–6035, 2018
2018
-
[66]
Afd- former: A hybrid transformer with asymmetric flow division for synthesized view quality enhancement,
X. Zhang, N. Cai, H. Zhang, Y. Zhang, J. Di, and W. Lin, “Afd- former: A hybrid transformer with asymmetric flow division for synthesized view quality enhancement,”IEEE Trans. Circuits Syst. Video Technol., vol. 33, no. 8, pp. 3786–3798, 2023
2023
-
[67]
Token fusion: Bridging the gap between token pruning and token merging,
M. Kim, S. Gao, Y.-C. Hsu, Y. Shen, and H. Jin, “Token fusion: Bridging the gap between token pruning and token merging,” in Proc. IEEE/CVF Winter Conf. on Appl. Comput. Vis., pp. 1383–1392, 2024. Novendra Setyawan (Student Member, IEEE) received the B.Eng. degree in Electrica...
2024
-
[2011]
He is currently a Full Professor with the Department of Electronic and Computer Engineering, National Taiwan University of Science and Technology
He was a Principal Engineer with TSMC, specializing in EDA design. He is currently a Full Professor with the Department of Electronic and Computer Engineering, National Taiwan University of Science and Technology. His research interests include image processing, system integra...
2005
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.