Pith. sign in

REVIEW 4 major objections 3 minor 68 references

CoDAT: Collaborative Dual-Attention Transformer with Low-Cost Temporal Modeling for Efficient Edge Action Recognition

T0 review · 4 major / 3 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read A dual-attention transformer brings near-state-of-the-art action recognition to edge hardware at a fraction of the energy cost.

desk verdict Competent edge action-recognition paper with strong ablations; the 6x speedup is plausible but the ONNX inference protocol is under-documented. read the letter →

arxiv 2608.06691 v1 pith:DGLLOD2Q submitted 2026-08-07 cs.CV

classification cs.CV
keywords actionrecognitionvisiontransformeredgecomputingtemporalshiftefficientattentionIoTsingle-headenergyefficiency
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper claims that accurate video action recognition can run in real time on low-power IoT hardware if the standard multi-head self-attention is replaced by a collaborative dual-branch design: a strided single-head attention that compresses both spatial and channel dimensions for global context, paired with a convolutional spatial attention for local detail. A zero-parameter temporal shift placed after the attention block supplies inter-frame motion information at negligible cost. On ImageNet-1K, Kinetics-400, MA-52, and UCF-101, the resulting CoDAT family matches or approaches state-of-the-art accuracy while running several times faster and consuming far less energy per frame than comparable CNN, transformer, and hybrid baselines on Jetson AGX Orin and Raspberry Pi 5. If these results hold, they shift the practical deployment envelope for action recognition toward battery- and power-limited edge devices in surveillance, healthcare, and industrial monitoring.

What carries the argument

The load-bearing object is the Collaborative Dual-Attention (CoDA) module and its placement relative to the temporal shift. Strided Single-Head Attention (SSHA) projects the input through pointwise convolutions with a stride $r_s$, reducing the attention from $N$ tokens to $M = (H/r_s) \times (W/r_s)$ tokens with value channels $C_v = C/4$, so the attention cost drops from $O(N^2 C)$ to $O(M^2 C_v)$. Spatial Convolutional Attention (SCA) computes a single-channel saliency mask from max and average pooling through a $3\times3$ convolution and sigmoid, costing $O(N C)$, and multiplies it into the features. A learnable projection fuses the two branches. Around each CoDA block, a zero-parameter TShift layer permutes a small fraction of channels across adjacent frames and is placed after CoDA but before the final ConvFFN, so the attention branches always see temporally clean single-frame features while the feed-forward network integrates the shifted temporal context.

What would settle it

Re-measuring the baselines from Tables 4-7 under the exact same inference settings as CoDAT (same single-clip, single-crop input, same exported model format, same hardware and power metering, same warm-up and repetition protocol) would settle the claim. In particular, if TokShift and LAPS on their public checkpoints run at less than 6x the latency and less than 13x the FLOPs of CoDAT-S384 on UCF-101, the headline efficiency claim fails.

Watch

Extended reading notes

Core claim

CoDAT's central claim is that the accuracy–efficiency frontier for edge action recognition can be moved by jointly attacking the two dominant costs of video transformers: token count and channel redundancy. The paper shows that a single attention head over spatially strided queries, keys, and values, combined with a lightweight convolutional attention branch, recovers the global and local cues that full multi-head attention provides, while a parameter-free temporal shift captures motion between frames. Empirically, CoDAT-S384 matches TokShift and LAPS on UCF-101 at roughly 6-times lower latency and up to 13-times fewer FLOPs, CoDAT-M runs about 2-times faster than EfficientViT-384 and FastViT-S12 on ImageNet-1K at comparable accuracy, and on Kinetics-400 CoDAT is up to 2.9-times faster than Video Swin Transformer and about 2-times faster than temporal-shift ViT variants while staying within one percentage point of their Top-1 accuracy. The efficiency gains persist on the fine-grained MA-52 benchmark, where CoDAT-M384 reaches the highest fine-grained Top-1 accuracy among the compared models.

Load-bearing premise

The load-bearing premise is that the published baseline numbers were produced under conditions close enough to CoDAT's own training and single-clip inference settings to be directly comparable; if the baselines used different training recipes, resolutions, or multi-view protocols, the claimed accuracy and speed advantages could shrink or vanish.

Editorial extensions

If this is right

  • Real-time action recognition becomes practical on low-power hardware: on Jetson AGX Orin, CoDAT-S runs UCF-101 inference at 0.89 ms/frame with 13.79 mJ/frame energy.
  • The efficiency gains transfer from image to video: CoDAT-L matches ViT-S/DeiT-S on ImageNet-1K with about 3x fewer parameters and 2.8x higher throughput.
  • On Kinetics-400, CoDAT-M384 outperforms or matches CNN baselines such as SlowFast and transformer baselines such as TokShift while requiring over 2x fewer GFLOPs and running 2.75x to 3.9x faster.
  • On the fine-grained MA-52 benchmark, CoDAT-M reaches the best fine-grained Top-1 accuracy among the compared models while running 2.7x faster than SlowFast and 5x faster than UniFormer-B.
  • The approach also works on compressed video: on MPEG-4-compressed UCF-101, CoDAT-S384 surpasses MTRFN by 0.8 percentage points at 8.8x fewer FLOPs.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the energy and latency ratios hold under fully standardized benchmarking, the practical meaning is that always-on camera analytics on battery-powered nodes becomes plausible; the implied operating point is several frames per second to real time on boards that previously could only run lightweight CNNs.
  • The fixed stride ratio $r_s$ per stage is a knob the paper leaves untouched; a content-adaptive stride that increases compression on easy frames could push the efficiency gains further than the static design the paper evaluates.
  • The placement principle found here—keep attention branches temporally clean and shift channels only into the feed-forward network—suggests a general recipe for shift-based video transformers that could be tested on other lightweight backbones.
  • A natural testable extension is streaming deployment: the paper measures per-frame cost on fixed clips, but batched frame shifting with a buffer would show whether the $2L+1$ temporal receptive field behaves as claimed under continuous input.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 3 minor

Summary. This manuscript introduces CoDAT, a three-stage vision transformer for edge action recognition. The central module, CoDA, combines a strided single-head attention (SSHA) branch that reduces spatial tokens and value channels with a spatial convolutional attention (SCA) branch, and a parameter-free temporal shift (TShift) is inserted before the final ConvFFN of each block. The paper evaluates CoDAT on ImageNet-1K, Kinetics-400, MA-52, and UCF-101, and reports latency, throughput, and energy on Jetson AGX Orin and Raspberry Pi 5. The authors claim that CoDAT matches or approaches the accuracy of much heavier CNN and ViT baselines while reducing FLOPs, latency, and energy substantially, with the strongest claims made for UCF-101, where CoDAT-S384 is reported to match TokShift and LAPS at 95.4% top-1 accuracy while running about 6x faster.

Significance. The architecture is internally coherent and the empirical scope is broader than in many edge action-recognition papers: four datasets, two hardware platforms, an explicit energy-measurement procedure, and systematic ablations of stride ratio, channel ratio, shift placement, and stage design. The SSHA complexity bookkeeping and the TShift receptive-field argument are plausible, and the internal ablations in Tables 8 and 9 support the claim that CoDAT's efficiency comes from the strided and channel-compressed attention design rather than from simple model shrinkage. If the efficiency-accuracy claims survive a fair cross-model benchmark protocol, the paper would be a useful contribution to edge IoT perception. However, the abstract and conclusion currently contain quantitative claims that do not match the tables, and the latency/energy comparisons are not yet protocol-fair as written, so the central claim cannot be accepted without substantial revision.

major comments (4)
  1. [Section 4.1, Table 5] The latency and energy comparisons are load-bearing but the runtime configuration is incomplete. Section 4.1 states only that inference uses ONNX Runtime with CUDA on Jetson AGX Orin at a 50 W power cap; it does not report numerical precision (FP32/FP16/TF32/INT8), ONNX optimization level, or whether a TensorRT execution provider was used. This matters concretely: in Table 5, the TokShift row reports 135 GFLOPs at 10.15 ms, which implies about 13.3 TFLOPS sustained, well above the practical FP32 throughput of a 50 W AGX Orin; the baselines therefore appear to have been run on a faster non-default path while the paper does not state whether CoDAT used the same one. Unless the same ONNX export options, execution provider, and precision are documented for every model, the '6x faster', '2.9x faster', and energy ratios in the abstract, conclusion, and Section 4.3 are not established. Please provide a per-model runtime configuration table or restrict the speed claims to models measured under identical settings.
  2. [Abstract, Conclusion, Tables 4-6] Several quantitative claims in the abstract and conclusion do not match the tables. On Kinetics-400, the conclusion says CoDAT is 2.9x faster than VSwin-T and nearly 2x faster than LAPS/TokShift; Table 5 gives CoDAT-L384 at 4.74 ms/F versus VSwin-T at 6.96 ms/F (1.47x), and CoDAT-M384 at 2.56 ms/F versus LAPS at 9.34 ms/F and TokShift at 10.15 ms/F (about 3.6x and 4.0x). On MA-52, the conclusion reports 63.6% fine-grained accuracy and latency reductions of more than 5x and 9x relative to VSwin-T and UniFormer-B, while Table 6 gives 63.19% and ratios of 2.72x and 5x. In the image-classification conclusion, CoDAT-L is said to have 2.5x lower latency than ViT-S, while Table 4 gives 1.71 ms / 1.33 ms = 1.29x. Please correct all abstract and conclusion numbers to agree with the tables, or rerun the measurements and update the tables.
  3. [Section 4.1, Tables 4-7] The cross-model benchmark parity is not sufficiently documented. Tables 4-7 mix self-reported and published baseline numbers, and Section 4.1 says multi-view configurations are applied 'exclusively for top-1 accuracy evaluation based on each model default configuration,' but it is unclear which view configuration corresponds to each baseline row and whether the baselines were run by the authors or taken from the literature. Several rows, such as S2AFormer-mini in Table 4 and I3D, STM, and MViT in Table 5, have no latency or energy entries, so the efficiency comparisons cover different model subsets. To make the central speed and energy claims benchmark-fair, please specify for every baseline the source of the accuracy number (authors' run vs published), the exact view/crop protocol, and the runtime configuration used for latency and energy.
  4. [Tables 4 and 7] The accuracy results appear to be single-run point estimates, which is problematic for claims that hinge on small differences. The central UCF-101 claim rests on CoDAT-S384 exactly matching TokShift and LAPS at 95.4% (Table 7), and several Table 4 conclusions rely on 0.1-0.4% accuracy differences (e.g., CoDAT-S versus SHViT-S3, or CoDAT-L versus Swin-T). Without multiple seeds, confidence intervals, or at least a stated fixed-seed protocol, these differences are within typical run-to-run noise, especially for a small dataset like UCF-101. Please report variance or multiple seeds for the main comparisons, or soften claims that depend on sub-1% differences.
minor comments (3)
  1. [Table 9] The row labeled 'TShift1/4C' in Table 9 reports a latency of 7.2 ms while all neighboring rows report 0.89 ms and the FLOPs are unchanged; this is almost certainly a typo and should be corrected.
  2. [Section 3.1.2] The text says SSHA gives 'a reduction of r_s^4-times over the standard full-resolution O(N^2 C)', but since the value channels are C_v = C/4, the actual reduction is 4 r_s^4. Please correct the wording or state explicitly that the factor of 4 is absorbed into the channel reduction.
  3. [Throughout] There are several typos and inconsistent labels, including 'VSwim-T' for VSwin-T in Section 4.3.5, the stray 'Trans.' row in Table 5, and the phrase 'the official train/test splits 1.' Please copyedit the final version.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: CoDAT's claims are empirical measurements and standard complexity bookkeeping, with no fitted parameter renamed as a prediction and no load-bearing self-citation chain.

full rationale

The paper's central claims are empirical: Top-1 accuracies, latencies, energies, and FLOPs are measured or arithmetically computed from the proposed architecture and compared against external baselines. The analytical statements in the paper are complexity bounds for SSHA (O(M^2 C_v + N C) with M = HW/r_s^2) and the temporal receptive field formula 2L+1 frames, which is imported from the TSM literature and used only as a design motivation, not as a derivation of the benchmark results. The design choices such as C_v = 1/4, stride ratios, and TShift placement are selected via ablations on the same benchmarks, which is standard engineering practice and does not constitute fitting a parameter and then calling the result a prediction. The self-citations in the paper (e.g., MicroViT-S3 as a baseline and prior works by the same authors) are background comparisons and are not load-bearing for the proposed architecture's validity. No uniqueness theorem, ansatz, or renamed known result is invoked to force the conclusions. The reported speedups and energy savings assume protocol parity across baselines, which is a correctness and benchmark-comparability risk rather than a circularity issue. Therefore, the derivation chain is empirically self-contained and no circular step can be exhibited from the paper's own equations or citations.

Assumptions & free parameters 6 free parameters · 4 assumptions · 0 invented entities

The paper's central claims depend on a set of hand-chosen hyperparameters (Cv, stride ratios, depth, channel widths) that are tuned through ablations, and on domain assumptions about benchmark comparability and hardware measurement fidelity. No parameters are fitted to the final accuracy numbers, so the empirical claims are not circular. The largest unstated risk is that published baseline numbers come from heterogeneous training/evaluation pipelines.

free parameters (6)
  • Value channel ratio Cv = 1/4 = 1/4 C
    Chosen by ablation in Table 8; used in SSHA to compress value channels. The central complexity and accuracy claims depend on this choice.
  • Per-stage stride ratios rs = [2,2,1]
    Hand-coded in Table 3 for the three pyramid stages; sets the token reduction in SSHA.
  • FFN expansion ratio = 2
    From Table 3, all variants use FFN expansion 2, chosen without sensitivity analysis.
  • Channel dims per stage = e.g., CoDAT-S [192,384,448]
    Capacity allocation across stages, hand-tuned; affects the parameter counts in the efficiency comparisons.
  • Block depth per stage = e.g., CoDAT-S [1,2,2]
    Depth schedule from Table 3; the TShift TRF analysis depends on total depth L=5.
  • TShift channel fraction = 1/8 C
    Ablation in Table 9 picks 1/8 C for the final design, giving a small accuracy gain over 1/16 C.
assumptions (4)
  • domain assumption ONNX Runtime faithfully implements all compared models and the power readings from INA3221/MXL7704 represent deployment energy.
    Section 4.1 describes the setup but does not include the ONNX export flags or sensor calibration for each model.
  • domain assumption Published baseline accuracies are accurate and obtained under comparable training/evaluation protocols.
    Tables 4-7 rely on numbers from the original papers, which use different epochs, augmentations, and multi-view testing.
  • standard math Temporal shift yields TRF = 2L+1 as claimed in [40].
    Section 3.3 uses this to state CoDAT-S (L=5) has a TRF of 11 frames; the formula is taken from TSM [40].
  • domain assumption Standard dataset splits and labels are used.
    The paper follows standard ImageNet, Kinetics-400, MA-52, and UCF-101 protocols without re-annotation.

how reviews work

0 comments
Cite this review

Pith. "Pith review of CoDAT: Collaborative Dual-Attention Transformer with Low-Cost Temporal Modeling for Efficient Edge Action Recognition." pith.science (2026). https://pith.science/paper/DGLLOD2Q

@misc{pith2026260806691,
  author       = {Pith},
  title        = {Pith review of: CoDAT: Collaborative Dual-Attention Transformer with Low-Cost Temporal Modeling for Efficient Edge Action Recognition},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/DGLLOD2Q}},
  note         = {Machine review of arXiv:2608.06691}
}
read the original abstract

Real-time human action recognition on Internet-of-Things (IoT) edge devices requires models that capture rich spatio-temporal cues within strict latency, memory, and power envelopes. Current 3D CNNs, video transformers, and shift-based ViT deliver high accuracy but come at computational costs that preclude edge IoT deployment. This paper proposes CoDAT, a Collaborative Dual-Attention Transformer that replaces conventional multi-head attention with a lightweight dual-branch module: Spatial Convolutional Attention (SCA) for local aggregation and Strided Single-Head Attention (SSHA) for global context. SSHA jointly compresses the spatial resolution and channel dimensions of the query, key, and value tensors via stride-based sparse projection, then fuses the resulting global and local features at a markedly reduced cost. To enable temporal communication across frames, a parameter-free TShift module is embedded in each block. Extensive experiments on Jetson AGX Orin and Raspberry Pi 5 demonstrate that CoDAT achieves an energy-accuracy balance in both image and action recognition. On ImageNet-1K, CoDAT-M runs 2x faster than EfficientViT384 and FastViT-S12 at comparable accuracy, and CoDAT-L matches ViT-S with 3x fewer parameters at 2x higher throughput. On Kinetics-400 and MA-52, CoDAT achieves competitive Top-1 accuracy against state-of-the-art CNN, transformer, and hybrid baselines while running up to 2.9x faster than VSwin-T, 2x faster than ViT-Temporal-Shift variants, and 5x faster than UniFormer-B. On UCF-101, CoDAT-S384 matches TokShift and LAPS while being 6x faster and requiring up to 13x fewer FLOPs, establishing an efficiency-accuracy balance for real-time action recognition in edge IoT perception systems. Code is available at https://github.com/novendrastywn/CoDAT .

Figures

Figures reproduced from arXiv: 2608.06691 by the authors.

Figure 1
Figure 1. Comparison of our proposed CoDAT model with SOTA methods. [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 1
Figure 1. This paper is organized as follows. Section II reviews re￾lated work on CNNs for video understanding, transformer￾based video recognition, and lightweight backbone ap￾proaches. Section III presents the proposed CoDAT archi￾tectures in detail. Section IV describes the experimental setup, benchmarks, evaluation results, and ablation studies. Finally, Section V concludes the paper and outlines future research direction… view at source ↗
Figure 2
Figure 2. Cloud edge collaborative intelligent activity monitoring system [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figures from the paper (4 more)
Figure 3
Figure 3. Figure 3: The overall CoDAT Architecture, CoDAT used pyramid architecture with 3 stages where each stage consist CoDA module, which operates two lightweight branches: Strided Single-Head Attention (SSHA) for efficient global spatial dependency modeling and Spatial Convolutional …
Figure 4
Figure 4. Figure 4: Comparison of different self-attention mechanisms: (a) Multi-Head Self-Attention (MHSA) [52], (b) Cascaded Group Attention [48], (c) Strip [PITH_FULL_IMAGE:figures/full_fig_p005_4.png]
Figure 5
Figure 5. Figure 5: CoDAT Temporal Block variants. (a) Full-shift A. (b) Full-shift B. (c) Attention-shift. and (d) Final ConvFFN-shift. subset of feature channels across adjacent frames without introducing any parameters. Given video features X ∈ R B·T ×C×H×W , TShift will reshaped it in…
Figure 6
Figure 6. Figure 6: The attention map visualizations on two examples from the Kinetics-400 validation set of CoDAT in comparison with UniFormer. (a): Example [PITH_FULL_IMAGE:figures/full_fig_p013_6.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

68 extracted references · 64 canonical work pages

  1. [1]

    Vu: Edge computing-enabled video usefulness detection and its application in large-scale video surveillance systems,

    H. Sun, W. Shi, X. Liang, and Y. Yu, “Vu: Edge computing-enabled video usefulness detection and its application in large-scale video surveillance systems,”IEEE Internet of Things J., vol. 7, no. 2, pp. 800–817, 2019

  2. [2]

    Strack: Robust tracking of small objects in low-light conditions,

    S. B. J. Khan, Z. Peng, M. M. Kamal, H. G. Mohamed, Q. M. Kharma, M. Sheraz, and T. C. Chuah, “Strack: Robust tracking of small objects in low-light conditions,”IEEE Access, 2025

  3. [3]

    Ai-driven salient soc- cer events recognition framework for next-generation iot-enabled environments,

    K. Muhammad, H. Ullah, M. S. Obaidat, A. Ullah, A. Munir, M. Sajjad, and V . H. C. De Albuquerque, “Ai-driven salient soc- cer events recognition framework for next-generation iot-enabled environments,”IEEE Internet of Things J., vol. 10, no. 3, pp. 2202– 2214, 2021

  4. [4]

    Contactless patient care using hospital iot: Cctv-camera-based physiological monitoring in icu,

    H. Wang, J. Huang, G. Wang, H. Lu, and W. Wang, “Contactless patient care using hospital iot: Cctv-camera-based physiological monitoring in icu,”IEEE Internet of Things J., vol. 11, no. 4, pp. 5781–5797, 2023

  5. [5]

    Optimization for short video propagation based on user interaction analysis in edge networks,

    Z. Li and Q. Yang, “Optimization for short video propagation based on user interaction analysis in edge networks,”IEEE Internet of Things J., 2025

  6. [6]

    Panacea+: Panoramic and con- trollable video generation for autonomous driving,

    Y. Wen, Y. Zhao, Y. Liu, B. Huang, F. Jia, Y. Wang, C. Zhang, T. Wang, X. Sun, and X. Zhang, “Panacea+: Panoramic and con- trollable video generation for autonomous driving,”IEEE Trans. on Circuits and Syst. for Video Technol., pp. 1–1, 2025

  7. [7]

    Tracenet: A novel modular frame- work for robust multi-object tracking in crowded and dynamic environments,

    S. B. J. Khan, P . Zhang, M. M. Kamal, A. Alharbi, A. Tolba, M. Sheraz, and T. C. Chuah, “Tracenet: A novel modular frame- work for robust multi-object tracking in crowded and dynamic environments,”Alexandria Engineering Journal, vol. 137, pp. 401– 413, 2026

  8. [8]

    Hamot: A hierarchical adaptive framework for robust multi-object tracking in complex environments,

    J. K. S. Baz, P . Zhang, M. M. Kamal, H. G. Mohamed, M. Sheraz, and T. C. Chuah, “Hamot: A hierarchical adaptive framework for robust multi-object tracking in complex environments,”CMES- Computer Modeling in Engineering and Sciences, vol. 145, no. 1, pp. 947–969, 2025

Show all 68 references
  1. [9]

    Imagenet classifi- cation with deep convolutional neural networks,

    A. Krizhevsky, I. Sutskever, and G. E. Hinton, “Imagenet classifi- cation with deep convolutional neural networks,”Adv. Neural Inf. Process. Syst., vol. 25, 2012

  2. [10]

    Facelivt: Face recognition using linear vision transformer with structural reparameterization for mobile device,

    N. Setyawan, C.-C. Sun, M.-H. Hsu, W.-K. Kuo, and J.-W. Hsieh, “Facelivt: Face recognition using linear vision transformer with structural reparameterization for mobile device,” inProc. IEEE Int. Conf. Image Process. (ICIP), pp. 1720–1725, 2025

  3. [11]

    Facelivtv2: An improved hybrid architecture for efficient mobile face recognition,

    N. Setyawan, C.-C. Sun, M.-H. Hsu, W.-K. Kuo, and J.-W. Hsieh, “Facelivtv2: An improved hybrid architecture for efficient mobile face recognition,”IEEE Trans. Biom. Behav. Identity Sci., 2026

  4. [12]

    Video analytics for detecting motorcy- clist helmet rule violations,

    C.-M. Tsai, J.-W. Hsieh, M.-C. Chang, G.-L. He, P .-Y. Chen, W.-T. Chang, and Y.-K. Hsieh, “Video analytics for detecting motorcy- clist helmet rule violations,” in2023 IEEE/CVF Conf. Comput. Vis. and Pattern Recognit. Workshops (CVPRW), pp. 5366–5374, 2023

  5. [13]

    Smiletrack: Similarity learning for occlusion-aware multiple object tracking,

    Y.-H. Wang, J.-W. Hsieh, P .-Y. Chen, M.-C. Chang, H.-H. So, and X. Li, “Smiletrack: Similarity learning for occlusion-aware multiple object tracking,” inProc. AAAI Conf. Artif. Intell., vol. 38, pp. 5740– 5748, 2024

  6. [14]

    Lighttrack-reid: A lightweight and occlusion-robust framework for multi-object tracking,

    S. B. J. Khan, P . Zhang, M. M. Kamal, and A. K. J. Saudagar, “Lighttrack-reid: A lightweight and occlusion-robust framework for multi-object tracking,”Plos one, vol. 21, no. 3, p. e0342246, 2026

  7. [15]

    Ssp-sam: Sam with semantic- spatial prompt for referring expression segmentation,

    W. Tang, X. Liu, Y. Sun, and Z. Li, “Ssp-sam: Sam with semantic- spatial prompt for referring expression segmentation,”IEEE Trans. Circuits Syst. Video Technol., pp. 1–1, 2025

  8. [16]

    Ucf101: A dataset of 101 human actions classes from videos in the wild,

    K. Soomro, A. R. Zamir, and M. Shah, “Ucf101: A dataset of 101 human actions classes from videos in the wild,”arXiv preprint arXiv:1212.0402, 2012

  9. [17]

    The kinet- ics human action video dataset,

    W. Kay, J. Carreira, K. Simonyan, B. Zhang, C. Hillier, S. Vijaya- narasimhan, F. Viola, T. Green, T. Back, P . Natsev,et al., “The kinet- ics human action video dataset,”arXiv preprint arXiv:1705.06950, 2017

  10. [18]

    Benchmarking micro-action recognition: Dataset, methods, and applications,

    D. Guo, K. Li, B. Hu, Y. Zhang, and M. Wang, “Benchmarking micro-action recognition: Dataset, methods, and applications,” IEEE Trans. Circuits Syst. Video Technol., vol. 34, no. 7, pp. 6238– 6252, 2024

  11. [19]

    Quo vadis, action recognition? a new model and the kinetics dataset,

    J. Carreira and A. Zisserman, “Quo vadis, action recognition? a new model and the kinetics dataset,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., pp. 6299–6308, 2017

  12. [20]

    Slowfast networks for video recognition,

    C. Feichtenhofer, H. Fan, J. Malik, and K. He, “Slowfast networks for video recognition,” inProc. IEEE/CVF Int. Conf. on Comput. Vis., pp. 6202–6211, 2019

  13. [21]

    Spatio- temporal adaptive network with bidirectional temporal difference for action recognition,

    Z. Li, J. Li, Y. Ma, R. Wang, Z. Shi, Y. Ding, and X. Liu, “Spatio- temporal adaptive network with bidirectional temporal difference for action recognition,”IEEE Trans. Circuits Syst. Video Technol., vol. 33, no. 9, pp. 5174–5185, 2023

  14. [22]

    Agpn: Action granu- larity pyramid network for video action recognition,

    Y. Chen, H. Ge, Y. Liu, X. Cai, and L. Sun, “Agpn: Action granu- larity pyramid network for video action recognition,”IEEE Trans. Circuits Syst. Video Technol., vol. 33, no. 8, pp. 3912–3923, 2023

  15. [23]

    An image is worth 16x16 words: Transformers for image recognition at scale,

    A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, et al., “An image is worth 16x16 words: Transformers for image recognition at scale,”arXiv preprint arXiv:2010.11929, 2020

  16. [24]

    Is space-time attention all you need for video understanding?,

    G. Bertasius, H. Wang, and L. Torresani, “Is space-time attention all you need for video understanding?,” inIcml, vol. 2, p. 4, 2021

  17. [25]

    Video trans- former network,

    D. Neimark, O. Bar, M. Zohar, and D. Asselmann, “Video trans- former network,” inProc. IEEE/CVF Int. Conf. on Comput. Vis., pp. 3163–3172, 2021

  18. [26]

    Video swin transformer,

    Z. Liu, J. Ning, Y. Cao, Y. Wei, Z. Zhang, S. Lin, and H. Hu, “Video swin transformer,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., pp. 3202–3211, 2022

  19. [27]

    Uniformer: Unifying convolution and self-attention for visual recognition,

    K. Li, Y. Wang, J. Zhang, P . Gao, G. Song, Y. Liu, H. Li, and Y. Qiao, “Uniformer: Unifying convolution and self-attention for visual recognition,”IEEE Trans. Pattern Anal. Mach. Intell., vol. 45, no. 10, pp. 12581–12600, 2023

  20. [28]

    Token shift transformer for video classification,

    H. Zhang, Y. Hao, and C.-W. Ngo, “Token shift transformer for video classification,” inProc. ACM Int. Conf. Multimed., pp. 917– 925, 2021

  21. [29]

    Long-term leap attention, short-term periodic shift for video classification,

    H. Zhang, L. Cheng, Y. Hao, and C.-w. Ngo, “Long-term leap attention, short-term periodic shift for video classification,” in Proceedings of the 30th acm international conference on multimedia, pp. 5773–5782, 2022

  22. [30]

    Temporal shift module-based vision transformer network for action recognition,

    K. Zhang, M. Lyu, X. Guo, L. Zhang, and C. Liu, “Temporal shift module-based vision transformer network for action recognition,” IEEE Access, vol. 12, pp. 47246–47257, 2024

  23. [31]

    Tsm: Temporal shift module for efficient and scalable video understanding on edge devices,

    J. Lin, C. Gan, K. Wang, and S. Han, “Tsm: Temporal shift module for efficient and scalable video understanding on edge devices,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 44, no. 05, pp. 2760– 2774, 2022

  24. [32]

    Deep residual learning for image recognition,

    K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., pp. 770–778, 2016

  25. [33]

    Swin transformer: Hierarchical vision transformer using shifted windows,

    Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo, “Swin transformer: Hierarchical vision transformer using shifted windows,” inProc. IEEE/CVF Int. Conf. on Comput. Vis., pp. 10012– 10022, 2021

  26. [34]

    Movinets: Mobile video networks for efficient video ACCEPTED ON IEEE INTERNET OF THINGS JOURNAL 15 recognition,

    D. Kondratyuk, L. Yuan, Y. Li, L. Zhang, M. Tan, M. Brown, and B. Gong, “Movinets: Mobile video networks for efficient video ACCEPTED ON IEEE INTERNET OF THINGS JOURNAL 15 recognition,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., pp. 16020–16030, 2021

  27. [35]

    Deepsensemoe: Harnessing power of time series foundation models for few-shot human activity recognition,

    Z. Fu, D. Cheng, L. Zhang, W. Huang, Z. Chen, and H. Wu, “Deepsensemoe: Harnessing power of time series foundation models for few-shot human activity recognition,” inProc. AAAI Conf. Artif. Intell., vol. 40, pp. 292–299, 2026

  28. [36]

    Sensor- prompt tuning: Aligning time series foundational models with motion sensors for few-shot activity recognition,

    X. Liu, D. Cheng, Z. Fu, L. Zhang, H. Wu, and A. Song, “Sensor- prompt tuning: Aligning time series foundational models with motion sensors for few-shot activity recognition,”IEEE Trans. Mob. Comput., 2026

  29. [37]

    Deep convolutional state space model as human activity recognizer,

    L. Wang, C. Bu, M. Yao, D. Xiong, S. Wang, D. Cheng, L. Zhang, H. Wu, and A. Song, “Deep convolutional state space model as human activity recognizer,”Information Fusion, p. 103982, 2025

  30. [38]

    Learn- ing spatiotemporal features with 3d convolutional networks,

    D. Tran, L. Bourdev, R. Fergus, L. Torresani, and M. Paluri, “Learn- ing spatiotemporal features with 3d convolutional networks,” in Proc. IEEE Int. Conf. Comput. Vis., pp. 4489–4497, 2015

  31. [39]

    A closer look at spatiotemporal convolutions for action recognition,

    D. Tran, H. Wang, L. Torresani, J. Ray, Y. LeCun, and M. Paluri, “A closer look at spatiotemporal convolutions for action recognition,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., pp. 6450– 6459, 2018

  32. [40]

    Tsm: Temporal shift module for efficient video understanding,

    J. Lin, C. Gan, and S. Han, “Tsm: Temporal shift module for efficient video understanding,” inProc. IEEE/CVF Int. Conf. on Comput. Vis., pp. 7083–7093, 2019

  33. [41]

    X3d: Expanding architectures for efficient video recognition,

    C. Feichtenhofer, “X3d: Expanding architectures for efficient video recognition,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., pp. 203–213, 2020

  34. [42]

    Mtrfn: Multiscale temporal receptive field network for compressed video action recognition at edge servers,

    L. He, M. Zhang, S. Zhang, L. Wang, and F. Li, “Mtrfn: Multiscale temporal receptive field network for compressed video action recognition at edge servers,”IEEE Internet of Things J., vol. 9, no. 15, pp. 13965–13977, 2022

  35. [43]

    Temporal transformer networks with self-supervision for action recognition,

    Y. Zhang, J. Li, N. Jiang, G. Wu, H. Zhang, Z. Shi, Z. Liu, Z. Wu, and X. Liu, “Temporal transformer networks with self-supervision for action recognition,”IEEE Internet of Things J., vol. 10, no. 14, pp. 12999–13011, 2023

  36. [44]

    Pyramid vision transformer: A versatile backbone for dense prediction without convolutions,

    W. Wang, E. Xie, X. Li, D.-P . Fan, K. Song, D. Liang, T. Lu, P . Luo, and L. Shao, “Pyramid vision transformer: A versatile backbone for dense prediction without convolutions,” inProc. IEEE/CVF Int. Conf. on Comput. Vis., pp. 568–578, 2021

  37. [45]

    Pvt v2: Improved baselines with pyramid vision transformer,

    W. Wang, E. Xie, X. Li, D.-P . Fan, K. Song, D. Liang, T. Lu, P . Luo, and L. Shao, “Pvt v2: Improved baselines with pyramid vision transformer,”Comput. Vis. Medi, vol. 8, no. 3, pp. 415–424, 2022

  38. [46]

    Edgenext: efficiently amalgamated cnn-transformer architecture for mobile vision applications,

    M. Maaz, A. Shaker, H. Cholakkal, S. Khan, S. W. Zamir, R. M. Anwer, and F. Shahbaz Khan, “Edgenext: efficiently amalgamated cnn-transformer architecture for mobile vision applications,” in European conference on computer vision, pp. 3–20, Springer, 2022

  39. [47]

    Fastvit: A fast hybrid vision transformer using structural reparameteriza- tion,

    P . K. A. Vasu, J. Gabriel, J. Zhu, O. Tuzel, and A. Ranjan, “Fastvit: A fast hybrid vision transformer using structural reparameteriza- tion,” inProc. IEEE/CVF Int. Conf. on Comput. Vis., pp. 5785–5795, 2023

  40. [48]

    Effi- cientvit: Memory efficient vision transformer with cascaded group attention,

    X. Liu, H. Peng, N. Zheng, Y. Yang, H. Hu, and Y. Yuan, “Effi- cientvit: Memory efficient vision transformer with cascaded group attention,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., pp. 14420–14430, 2023

  41. [49]

    Rethinking vision transformers for mobilenet size and speed,

    Y. Li, J. Hu, Y. Wen, G. Evangelidis, K. Salahi, Y. Wang, S. Tulyakov, and J. Ren, “Rethinking vision transformers for mobilenet size and speed,” inProc. IEEE/CVF Int. Conf. on Comput. Vis., pp. 16889– 16900, 2023

  42. [50]

    Shvit: Single-head vision transformer with memory efficient macro design,

    S. Yun and Y. Ro, “Shvit: Single-head vision transformer with memory efficient macro design,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., pp. 5756–5767, 2024

  43. [51]

    S2aformer: Strip self-attention for efficient vision transformer,

    G. Xu, W. Huang, W. Jia, J. Li, G. Gao, and G.-J. Qi, “S2aformer: Strip self-attention for efficient vision transformer,”IEEE Trans. Image Process., pp. 1–1, 2025

  44. [52]

    Training data-efficient image transformers & distilla- tion through attention,

    H. Touvron, M. Cord, M. Douze, F. Massa, A. Sablayrolles, and H. J ´egou, “Training data-efficient image transformers & distilla- tion through attention,” inProc. Int. Conf. Mach. Learn., pp. 10347– 10357, PMLR, 2021

  45. [53]

    Tlee: Temporal-wise and layer-wise early exiting network for efficient video recognition on edge devices,

    Q. Wang, W. Fang, and N. N. Xiong, “Tlee: Temporal-wise and layer-wise early exiting network for efficient video recognition on edge devices,”IEEE Internet of Things J., vol. 11, no. 2, pp. 2842– 2854, 2023

  46. [54]

    Conditional positional encodings for vision transformers,

    X. Chu, Z. Tian, B. Zhang, X. Wang, and C. Shen, “Conditional positional encodings for vision transformers,”Int. Conf. Learn. Represent. (ICLR) 2021, 2021

  47. [55]

    Imagenet large scale visual recognition challenge,

    O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein,et al., “Imagenet large scale visual recognition challenge,”Int. J. Comput. Vis., vol. 115, pp. 211–252, 2015

  48. [56]

    Iformer: Integrating convnet and transformer for mo- bile application,

    C. Zheng, “Iformer: Integrating convnet and transformer for mo- bile application,”Int. Conf. Learn. Represent. (ICLR), 2025

  49. [57]

    Microvit: a vision transformer with low complexity self attention for edge device,

    N. Setyawan, C.-C. Sun, M.-H. Hsu, W.-K. Kuo, and J.-W. Hsieh, “Microvit: a vision transformer with low complexity self attention for edge device,” inProc. IEEE Int. Symp. Circuits Syst. (ISCAS), pp. 1–5, IEEE, 2025

  50. [58]

    Large batch optimization for deep learning: Training bert in 76 minutes,

    Y. You, J. Li, S. Reddi, J. Hseu, S. Kumar, S. Bhojanapalli, X. Song, J. Demmel, K. Keutzer, and C.-J. Hsieh, “Large batch optimization for deep learning: Training bert in 76 minutes,”arXiv preprint arXiv:1904.00962, 2019

  51. [59]

    Group contextualization for video recognition,

    Y. Hao, H. Zhang, C.-W. Ngo, and X. He, “Group contextualization for video recognition,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., pp. 928–938, 2022

  52. [60]

    Learning spatiotem- poral and motion features in a unified 2d network for action recognition,

    M. Wang, J. Xing, J. Su, J. Chen, and Y. Liu, “Learning spatiotem- poral and motion features in a unified 2d network for action recognition,”IEEE Trans. Pattern Anal. Mach. Intell., vol. 45, no. 3, pp. 3347–3362, 2022

  53. [61]

    Multiscale vision transformers,

    H. Fan, B. Xiong, K. Mangalam, Y. Li, Z. Yan, J. Malik, and C. Fe- ichtenhofer, “Multiscale vision transformers,” inProc. IEEE/CVF Int. Conf. on Comput. Vis., pp. 6824–6835, 2021

  54. [62]

    Dualactnet: Exploiting slowfast architecture for micro-action recognition,

    C. Yu, Y. Ru, Z. Xu, H. Wu, H. Yang, and Z. He, “Dualactnet: Exploiting slowfast architecture for micro-action recognition,” in Chinese Conference on Biometric Recognition, pp. 59–68, Springer, 2024

  55. [63]

    Tdn: Temporal difference networks for efficient action recognition,

    L. Wang, Z. Tong, B. Ji, and G. Wu, “Tdn: Temporal difference networks for efficient action recognition,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., pp. 1895–1904, 2021

  56. [64]

    Mobile video action recognition,

    Y. Huo, X. Xu, Y. Lu, Y. Niu, Z. Lu, and J.-R. Wen, “Mobile video action recognition,”arXiv preprint arXiv:1908.10155, 2019

  57. [65]

    Compressed video action recognition,

    C.-Y. Wu, M. Zaheer, H. Hu, R. Manmatha, A. J. Smola, and P . Kr¨ahenb ¨uhl, “Compressed video action recognition,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit., pp. 6026–6035, 2018

  58. [66]

    Afd- former: A hybrid transformer with asymmetric flow division for synthesized view quality enhancement,

    X. Zhang, N. Cai, H. Zhang, Y. Zhang, J. Di, and W. Lin, “Afd- former: A hybrid transformer with asymmetric flow division for synthesized view quality enhancement,”IEEE Trans. Circuits Syst. Video Technol., vol. 33, no. 8, pp. 3786–3798, 2023

  59. [67]

    Token fusion: Bridging the gap between token pruning and token merging,

    M. Kim, S. Gao, Y.-C. Hsu, Y. Shen, and H. Jin, “Token fusion: Bridging the gap between token pruning and token merging,” in Proc. IEEE/CVF Winter Conf. on Appl. Comput. Vis., pp. 1383–1392, 2024. Novendra Setyawan (Student Member, IEEE) received the B.Eng. degree in Electrica...

  60. [2011]

    He is currently a Full Professor with the Department of Electronic and Computer Engineering, National Taiwan University of Science and Technology

    He was a Principal Engineer with TSMC, specializing in EDA design. He is currently a Full Professor with the Department of Electronic and Computer Engineering, National Taiwan University of Science and Technology. His research interests include image processing, system integra...

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.