REVIEW 3 major objections 5 minor 86 references
ProTAL: A Drag-and-Link Video Programming Framework for Temporal Action Localization
T0 review · 3 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read A drag-and-link rule interface can generate enough labels from unlabeled video to train temporal action localization models nearly as well as full manual annotation.
desk verdict A genuinely useful video programming framework with a solid usability study, but Table 1's accuracy comparison conflates the new labels with a modified training objective, so the central claim needs an ablation. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The key machinery is the key event: a labeled atomic substructure of an action written as $state_1 \rightarrow state_2 \rightarrow \dots \rightarrow state_n$, where each state is a graph whose nodes are visual elements and whose edges carry constraints of direction, relative distance, contact, and person association. A frame matches a state when the state graph is a subgraph of the frame's extracted element-relation graph, so label generation reduces to subgraph matching with pruning. Training then extends the base model by adding a state-order loss that exploits the natural temporal order of states within a key event on unlabeled videos.
What would settle it
Run the ProTAL pipeline on a video set with dense overlapping people or strong viewpoint changes, the failure cases the paper itself lists, and compare the resulting TAL model against one trained on ground-truth labels for the same set; if detector errors propagate so that the model's mAP collapses toward the single-frame baseline while full supervision still succeeds, the central claim is falsified.
Extended reading notes
Core claim
The central claim is that an action can be decomposed into key events: sequences of two or more states, each defined by relations among extracted visual elements such as body parts and objects, and that these state definitions act as labeling functions that generate frame-wise weak labels for whole video datasets. The authors show that matching a state definition against a frame is subgraph matching, since every frame is represented as a graph of visual elements with relational edges. They further show that sparse key-event labels suffice to train a TAL model when the base single-frame-supervision model is modified to predict key-event state order on unlabeled frames. On 470 unlabeled table tennis clips, this pipeline produced a serve-localization model with average mAP 0.825, essentially matching full supervision at 0.833 while using over 30 times less annotation time.
Load-bearing premise
The pipeline assumes that automatic per-frame pose estimation and object detection are reliable enough that the subgraph matcher retrieves the correct key-event frames; when occlusion, dense scenes, or viewpoint changes break those detectors, the generated labels degrade and the trained model cannot recover.
Editorial extensions
If this is right
- Users can build TAL training sets from scratch with a few minutes of rule definition rather than hours of per-video annotation.
- Because the same key-event definitions are applied automatically to every video, annotation effort stays roughly constant as dataset size grows.
- Key-event labels sit between full supervision and single-frame supervision in label density, and the state-order loss lets unlabeled frames contribute to training.
- The drag-and-link interaction reduced task completion time from 622.1 to 518.8 seconds and iterations from 2.8 to 1.7 in the user study.
- The framework generalizes across single-human actions, human-human interactions, and human-object interactions, with examples given for high jump, golf swing, arm wrestling, and tumbling.
Reading between the lines
- If pose estimation and object detection keep improving, the same state-graph formulation could extend to 3D or 4D annotation, where key events are defined by relations in space rather than in a 2D frame.
- The framework likely transfers to other video tasks with identifiable key events, such as action quality scoring or spatial action segmentation, because the constraint types are not TAL-specific.
- A natural next step is to let multimodal models translate natural-language action descriptions into these relation graphs, which would reduce the user's burden of choosing elements and constraints.
- The reported 30x annotation speedup compares against one expert annotating one dataset; the real gain depends on detector reliability and action complexity, so the scalability claim deserves testing across more domains.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents ProTAL, an interactive data-programming framework for temporal action localization. ProTAL decomposes actions into user-defined 'key events,' each represented as an ordered sequence of states; states are defined by dragging visual-element nodes and linking them with relational constraints (direction, relative distance, contact, association). Generated frame labels are used to train a semi-supervised TAL model based on SF-Net with a modified classification target and a state-order loss. The authors evaluate ProTAL through a single expert usage scenario on unlabeled table-tennis videos (comparing against SF-Net with full and single-frame supervision, Table 1) and through a 12-participant user study comparing drag-and-link interaction with a form-based baseline.
Significance. The core idea—extending data programming to TAL through decomposable key events and an interactive graph-based constraint definition—is novel and timely, and the drag-and-link interaction design is well motivated by prior work. The user study is carefully controlled: it uses a counterbalanced within-subject design, objective completion-time and iteration metrics, and a predefined label-quality threshold. The authors also candidly state limitations in Section 8.3.1. However, the central quantitative claim that ProTAL effectively constructs TAL models is currently underdetermined: the headline comparison in Table 1 confounds the label-generation method with the modified training objective, and the evaluation rests on a single expert, one dataset, and one run per model. The paper's scientific contribution would be substantially strengthened by an ablation that isolates these factors.
major comments (3)
- [Section 5.3.2, Table 1] The ProTAL model is trained with a modified SF-Net objective (state-level classification target and state-order loss on unlabeled videos), while the two SF-Net baselines in Table 1 use the original objective. Therefore the reported avg-mAP 0.825 vs. 0.833 and 0.650 conflates the contribution of ProTAL's generated key-event labels with the contribution of the modified semi-supervised training procedure. Section 5.3.2 also describes the state-order loss only verbally, with no equation or implementation detail, so it cannot be reproduced or controlled. Please add an ablation that trains (a) the original SF-Net objective on ProTAL-generated labels and (b) the modified objective on single-frame labels, or otherwise hold the training objective fixed across comparisons.
- [Section 6.3] The comparative evaluation in Section 6.3 relies on one expert, one table-tennis dataset, and a single run per model; Table 1 reports no variance estimates or significance tests, yet the text says Alex's model 'significantly outperformed' the single-frame baseline. Given that Table 1 is the primary quantitative support for the central effectiveness claim, the absence of any error bars, repeated runs, or statistical testing makes the comparison difficult to interpret. At minimum, report multiple runs with seeds and the corresponding distribution of mAP.
- [Section 7.2, Section 7.4.4] In the user study, a task is considered complete when the generated labels satisfy accuracy≥0.8 and recall≥0.2. The recall threshold of 0.2 is very permissive: a key-event definition that retrieves only one fifth of the relevant frames passes the quality gate. The claim in Section 7.4.4 that drag-and-link 'helps define key events accurately' is therefore based mainly on iteration count and completion time rather than on the quality of the resulting labels or downstream model performance. Please either tighten the quality criteria or reframe the claim as an efficiency benefit, not an accuracy benefit.
minor comments (5)
- [Section 5.2.2] The subgraph matching algorithm Φ is described at a high level; please specify the pruning strategy or give an algorithmic sketch, since the correctness and speed of key-event frame retrieval depend on it.
- [Equation (4)] The notation 'state_k t≤thr' is hard to parse; please define the threshold semantics explicitly, for example as an upper bound on the elapsed time between a frame matching state k and a frame matching state k+1.
- [Section 7.3] The paired t-test is described as comparing tasks completed by 'the same group,' but the reported test appears to pool across groups; please clarify the exact pairing and whether N=12 or N=6 per comparison.
- [Section 6.2.3] In the usage scenario, the expert handles the two viewpoints by defining two separate key events (K1 and K2); a sentence connecting this practice to the viewpoint limitation acknowledged in Section 8.3.1 would help readers understand the proposed mitigation.
- [Figure 6] The additional usage examples (high jump, golf swing, arm wrestling, standing ab twist) are illustrative screenshots without quantitative evaluation; please label them as demonstrations rather than validated results.
Circularity Check
No significant circularity: label generation is user-rule-based and validated against external ground truth.
full rationale
ProTAL's central contribution is an interactive data programming pipeline: users define key events as sequences of visual element states with constraints (Equations 3-9), the system retrieves matching frames by subgraph matching (Equation 10), and the resulting sparse frame labels are used to train a TAL model. This chain does not derive a target quantity from a first-principles model; the labels are generated by user-specified rules, and the framework does not claim that these labels are mathematically entailed by anything other than the user's own definitions. The headline quantitative evidence (Table 1) compares the ProTAL-trained model against SF-Net trained with full supervision and single-frame supervision, with mAP computed on manually annotated test videos; therefore the reported performance is not a by-construction restatement of the generated labels. The only self-citations (EventAnchor, VideoModerator, Smartboard) appear in related work and are not load-bearing for the paper's claims. The acknowledged limitations in Section 8.3.1 concern upstream perception robustness (occlusion, dense scenes, viewpoint changes), which is an external failure mode rather than circular reasoning. One legitimate evaluation concern is that Section 5.3.2 extends the SF-Net objective with state-level classification and a state-order loss on unlabeled videos, so Table 1 does not fully isolate the label-generation contribution from the modified training objective; however, this is an experimental confound rather than a circularity, because the training objective is not derived from the labels by construction and the evaluation still uses external ground truth. No fitted parameter is renamed as a prediction, and no load-bearing claim reduces to a self-citation or to the paper's own definitions. Therefore no circular step is identified.
Assumptions & free parameters
free parameters (4)
- time interval threshold between key event states =
0.3 seconds in the usage scenario (user-set)
- direction angle range for relational constraints =
70-degree interval in state 1 example
- contact IoU threshold =
not specified in text
- distance constraint order =
unspecified
assumptions (3)
- domain assumption Actions can be decomposed into key events, which are atomic spatiotemporal units defined by changes in relations between visual elements.
- domain assumption A video frame can be represented as a graph of visual elements and their computable relations, and matching a state definition is a subgraph matching problem.
- domain assumption A semi-supervised SF-Net extension with a state-order loss can train an effective TAL model from sparse frame labels generated by ProTAL.
invented entities (1)
-
key event
Cite this review
Pith. "Pith review of ProTAL: A Drag-and-Link Video Programming Framework for Temporal Action Localization." pith.science (2026). https://pith.science/paper/LEOXLPGL
@misc{pith2026250517555,
author = {Pith},
title = {Pith review of: ProTAL: A Drag-and-Link Video Programming Framework for Temporal Action Localization},
year = {2026},
howpublished = {\url{https://pith.science/paper/LEOXLPGL}},
note = {Machine review of arXiv:2505.17555}
}
read the original abstract
Temporal Action Localization (TAL) aims to detect the start and end timestamps of actions in a video. However, the training of TAL models requires a substantial amount of manually annotated data. Data programming is an efficient method to create training labels with a series of human-defined labeling functions. However, its application in TAL faces difficulties of defining complex actions in the context of temporal video frames. In this paper, we propose ProTAL, a drag-and-link video programming framework for TAL. ProTAL enables users to define \textbf{key events} by dragging nodes representing body parts and objects and linking them to constrain the relations (direction, distance, etc.). These definitions are used to generate action labels for large-scale unlabelled videos. A semi-supervised method is then employed to train TAL models with such labels. We demonstrate the effectiveness of ProTAL through a usage scenario and a user study, providing insights into designing video programming framework.
Figures
Figures from the paper (6 more)
Reference graph
Works this paper leans on
-
[1]
Maya Antoun and Daniel Asmar. 2023. Human object interaction detection: Design and survey. Image and Vision Computing 130, C (2023), 104617. https: //doi.org/10.1016/J.IMAVIS.2022.104617
arXiv 2023
-
[2]
Bach, Bryan Dawei He, Alexander Ratner, and Christopher Ré
Stephen H. Bach, Bryan Dawei He, Alexander Ratner, and Christopher Ré. 2017. Learning the Structure of Generative Models without Labeled Data. InProceedings of the 34th International Conference on Machine Learning . 273–282
2017
-
[3]
Djamila Romaissa Beddiar, Brahim Nini, Mohammad Sabokrou, and Abdenour Hadid. 2020. Vision-based human activity recognition: a survey.Multimedia Tools and Applications 79, 41-42 (2020), 30509–30555. https://doi.org/10.1007/S11042- 020-09004-3
doi:10.1007/s11042- 2020
-
[4]
João Carreira and Andrew Zisserman. 2017. Quo Vadis, Action Recognition? A New Model and the Kinetics Dataset. In 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) . 4724–4733. https://doi.org/10.1109/CVPR.2017. 502
-
[5]
Ross, Jia Deng, and Rahul Sukthankar
Yu-Wei Chao, Sudheendra Vijayanarasimhan, Bryan Seybold, David A. Ross, Jia Deng, and Rahul Sukthankar. 2018. Rethinking the Faster R-CNN Architecture for Temporal Action Localization. In Proceedings of 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition . 1130–1139. https://doi.org/10.1109/ CVPR.2018.00124
-
[6]
Changjian Chen, Jiashu Chen, Weikai Yang, Haoze Wang, Johannes Knittel, Xibin Zhao, Steffen Koch, Thomas Ertl, and Shixia Liu. 2024. Enhancing Single- Frame Supervision for Better Temporal Action Localization. IEEE Transactions on Visualization and Computer Graphics 30, 6 (2024), 2903–2915. https://doi.org/ 10.1109/TVCG.2024.3388521
arXiv 2024
-
[7]
Lu Chen, Sida Peng, and Xiaowei Zhou. 2021. Towards efficient and photorealistic 3D human reconstruction: A brief survey. Visual Informatics 5, 4 (2021), 11–19. https://doi.org/10.1016/j.visinf.2021.10.003
- [8]
Show all 86 references
-
[9]
Minh Dang, Kyungbok Min, Hanxiang Wang, Md
L. Minh Dang, Kyungbok Min, Hanxiang Wang, Md. Jalil Piran, Cheol Hee Lee, and Hyeonjoon Moon. 2020. Sensor-based and vision-based human activity recognition: A comprehensive survey. Pattern Recognition 108 (2020), 107561. https://doi.org/10.1016/J.PATCOG.2020.107561
2020
-
[10]
Dazhen Deng, Jiang Wu, Jiachen Wang, Yihong Wu, Xiao Xie, Zheng Zhou, Hui Zhang, Xiaolong (Luke) Zhang, and Yingcai Wu. 2021. EventAnchor: Reducing Human Interactions in Event Annotation of Racket Sports Videos. In Proceedings of the 2021 CHI Conference on Human Factors in Com...
2021
-
[11]
Haodong Duan, Yue Zhao, Kai Chen, Dahua Lin, and Bo Dai. 2022. Revisiting Skeleton-based Action Recognition. In Proceedings of 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . 2959–2968. https://doi.org/ 10.1109/CVPR52688.2022.00298
2022
-
[12]
Victor Escorcia, Fabian Caba Heilbron, Juan Carlos Niebles, and Bernard Ghanem
-
[13]
Sara Evensen, Chang Ge, and Çagatay Demiralp. 2020. Ruler: Data Programming by Demonstration for Document Labeling. In Findings of the Association for Computational Linguistics: EMNLP 2020 . 1996–2005. https://doi.org/10.18653/V1/ 2020.FINDINGS-EMNLP.181
2020 doi
-
[14]
Bernard Ghanem Fabian Caba Heilbron, Victor Escorcia and Juan Carlos Niebles
-
[15]
Gueter Josmy Faure, Min-Hung Chen, and Shang-Hong Lai. 2023. Holistic Interaction Transformer Network for Action Detection. In 2023 IEEE/CVF Winter Conference on Applications of Computer Vision (W ACV). 3329–3339. https://doi. org/10.1109/WACV56688.2023.00334
2023
-
[16]
Yutong Feng, Jianwen Jiang, Ziyuan Huang, Zhiwu Qing, Xiang Wang, Shiwei Zhang, Mingqian Tang, and Yue Gao. 2021. Relation Modeling in Spatio-Temporal Action Localization
2021
-
[17]
Jiyang Gao, Zhenheng Yang, and Ram Nevatia. 2017. Cascaded Boundary Re- gression for Temporal Action Detection. In Proceedings of British Machine Vision Conference 2017
2017
-
[18]
Girshick, Piotr Dollár, and Kaiming He
Georgia Gkioxari, Ross B. Girshick, Piotr Dollár, and Kaiming He. 2018. Detecting and Recognizing Human-Object Interactions. InProceedings of the IEEE conference on computer vision and pattern recognition . 8359–8367. https://doi.org/10.1109/ CVPR.2018.00872
2018
-
[19]
C. Gu, C. Sun, D. A. Ross, C. Vondrick, C. Pantofaru, Y. Li, S. Vijayanarasimhan, G. Toderici, S. Ricco, R. Sukthankar, C. Schmid, and J. Malik. 2018. AVA: A Video Dataset of Spatio-Temporally Localized Atomic Visual Actions. In 2018 IEEE/CVF Conference on Computer Vision and ...
2018
-
[20]
Dongming Han, Jiacheng Pan, Xiaodong Zhao, and Wei Chen. 2021. NetV.js: A web-based library for high-efficiency visualization of large-scale graphs and networks. Visual Informatics 5, 1 (2021), 61–66. https://doi.org/10.1016/j.visinf. 2021.01.002
2021 doi
-
[21]
Jianben He, Xingbo Wang, Kam Kwai Wong, Xijie Huang, Changjian Chen, Zixin Chen, Fengjie Wang, Min Zhu, and Huamin Qu. 2024. VideoPro: A Visual Analytics Approach for Interactive Video Programming. 30, 1 (2024), 87–97. https://doi.org/10.1109/TVCG.2023.3326586
2024
-
[22]
Md Naimul Hoque, Wenbin He, Arvind Kumar Shekar, Liang Gou, and Liu Ren
-
[23]
Edwin L Hutchins, James D Hollan, and Donald A Norman. 1985. Direct manipu- lation interfaces. Human–computer interaction 1, 4 (1985), 311–338
1985
- [24]
-
[25]
THUMOS Challenge: Action Recognition with a Large Number of Classes
Y.-G. Jiang, J. Liu, A. Roshan Zamir, G. Toderici, I. Laptev, M. Shah, and R. Suk- thankar. 2014. "THUMOS Challenge: Action Recognition with a Large Number of Classes". http://crcv.ucf.edu/THUMOS14/
2014
-
[26]
Pushpajit Khaire and Praveen Kumar. 2022. Deep learning and RGB-D based hu- man action, human-human and human-object interaction recognition: A survey. Journal of Visual Communication and Image Representation 86, C (2022), 103531. https://doi.org/10.1016/J.JVCIR.2022.103531
2022
-
[27]
Bumsoo Kim, Junhyun Lee, Jaewoo Kang, Eun-Sol Kim, and Hyunwoo J. Kim
-
[28]
Kuno Kurzhals, Marcel Hlawatsch, Christof Seeger, and Daniel Weiskopf. 2017. Visual Analytics for Mobile Eye Tracking. IEEE Transactions on Visualization and Computer Graphics 23, 1 (2017), 301–310. https://doi.org/10.1109/TVCG.2016. 2598695
2017 doi
-
[30]
Zhengyang Li, Jie Li, and Xinying Ma. 2025. Representing multi-dimensional data as graph to visualize and analyze subset communities. Journal of Visualization (2025). https://doi.org/10.1007/s12650-025-01045-w
2025 doi
-
[31]
Qinying Liu, Zilei Wang, and Shenghai Rong. 2023. Improve Temporal Action Proposals using Hierarchical Context. Pattern Recognition 140 (2023), 109560. https://doi.org/10.1016/j.patcog.2023.109560
2023
-
[32]
Shilong Liu, Zhaoyang Zeng, Tianhe Ren, Feng Li, Hao Zhang, Jie Yang, Chunyuan Li, Jianwei Yang, Hang Su, Jun Zhu, et al . 2023. Grounding dino: Marrying dino with grounded pre-training for open-set object detection. arXiv preprint arXiv:2303.05499 (2023)
2023 arXiv
-
[33]
Shuming Liu, Chen-Lin Zhang, Chen Zhao, and Bernard Ghanem. 2023. End-to- End Temporal Action Detection with 1B Parameters Across 1000 Frames. arXiv preprint arXiv: 2311.17241 (2023), 18591–18601
2023 arXiv
-
[35]
Ziao Liu, Xiao Xie, Moqi He, Wenshuo Zhao, Yihong Wu, Liqi Cheng, Hui Zhang, and Yingcai Wu. 2024. Smartboard: Visual Exploration of Team Tactics with LLM Agent. IEEE Transactions on Visualization and Computer Graphics (2024), 1–11. https://doi.org/10.1109/TVCG.2024.3456200
2024
-
[36]
Júlio Castro Lopes and Rui Pedro Lopes. 2024. Computer Vision in Augmented, Virtual, Mixed and Extended Reality environments—A bibliometric review.Visual Informatics 8, 4 (2024), 13–22. https://doi.org/10.1016/j.visinf.2024.11.002
2024 doi
-
[37]
Fan Ma, Linchao Zhu, Yi Yang, Shengxin Zha, Gourab Kundu, Matt Feiszli, and Zheng Shou. 2020. SF-Net: Single-Frame Supervision for Temporal Action Local- ization. In Proceedings of Computer Vision - ECCV 2020 - 16th European Conference . 420–437. https://doi.org/10.1007/978-3-...
2020 doi
-
[38]
Ayana Murakami and Takayuki Itoh. 2025. Flexible optimization of hierarchical graph layout by genetic algorithm with various conditions. Journal of Visualiza- tion 28, 1 (2025), 181–204. https://doi.org/10.1007/s12650-024-01018-5
2025 doi
-
[39]
Sauradip Nag, Xiatian Zhu, Yi-Zhe Song, and Tao Xiang. 2022. Semi-supervised Temporal Action Detection with Proposal-Free Masking. In Proceedings of Com- puter Vision – ECCV 2022. 663–680. https://doi.org/10.1007/978-3-031-20062-5_38
2022 doi
-
[40]
Phuc Nguyen, Ting Liu, Gautam Prasad, and Bohyung Han. 2018. Weakly Super- vised Action Localization by Sparse Temporal Pooling Network. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) . 6752–
2018
-
[41]
Phuc Nguyen, Deva Ramanan, and Charless Fowlkes. 2019. Weakly-Supervised Action Localization With Background Modeling. In Proceedings of 2019 IEEE/CVF International Conference on Computer Vision (ICCV) . 5501–5510. https://doi.org/ 10.1109/ICCV.2019.00560
2019
-
[42]
Dietrich, and Cláu- dio T
Jorge Piazentin Ono, Arvi Gjoka, Justin Salamon, Carlos A. Dietrich, and Cláu- dio T. Silva. 2019. HistoryTracker: Minimizing Human Interactions in Baseball Game Annotation. In Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems. 63. https://doi.org/10...
2019
-
[43]
Junting Pan, Siyu Chen, Mike Zheng Shou, Yu Liu, Jing Shao, and Hongsheng Li. 2021. Actor-Context-Actor Relation Network for Spatio-Temporal Action Localization. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2021, virtual, June 19-25, 2021 . 464–474. http...
2021
-
[44]
Paoletti, J
G. Paoletti, J. Cavazza, C. Beyan, and A. Del Bue. 2021. Subspace Clustering for Action Recognition with Covariance Representations and Temporal Pruning. In 2020 25th International Conference on Pattern Recognition (ICPR) . 6035–6042. https://doi.org/10.1109/ICPR48806.2021.9412060
2021
-
[46]
Bach, Henry Ehrenberg, Jason Fries, Sen Wu, and Christopher Ré
Alexander Ratner, Stephen H. Bach, Henry Ehrenberg, Jason Fries, Sen Wu, and Christopher Ré. 2017. Snorkel: Rapid Training Data Creation with Weak Supervision. Proceedings of the VLDB Endowment 11, 3 (2017), 269–282. https: //doi.org/10.14778/3157794.3157797
2017
-
[47]
Alexander Ratner, Braden Hancock, Jared Dunnmon, Frederic Sala, Shreyash Pandey, and Christopher Ré. 2019. Training Complex Models with Multi-Task Weak Supervision. InProceedings of the Thirty-Third AAAI Conference on Artificial Intelligence. 4763–4771. https://doi.org/10.1609...
2019 doi
-
[48]
Alexander J Ratner, Christopher M De Sa, Sen Wu, Daniel Selsam, and Christopher Ré. 2016. Data Programming: Creating Large Training Sets, Quickly. In Advances in Neural Information Processing Systems , Vol. 29
2016
-
[49]
Tianhe Ren, Shilong Liu, Ailing Zeng, Jing Lin, Kunchang Li, He Cao, Jiayu Chen, Xinyu Huang, Yukang Chen, Feng Yan, Zhaoyang Zeng, Hao Zhang, Feng Li, Jie Yang, Hongyang Li, Qing Jiang, and Lei Zhang. 2024. Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks. ...
2024 arXiv
-
[50]
Benjamin Renoust, Haolin Ren, and Guy Melançon. 2019. Animated Drag and Drop Interaction for Dynamic Multidimensional Graphs. arXiv preprint arXiv: 1902.01564 (2019)
2019 arXiv
-
[51]
Dian Shao, Yue Zhao, Bo Dai, and Dahua Lin. 2020. FineGym: A Hierarchi- cal Video Dataset for Fine-Grained Action Understanding. In 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . 2613–2622. https://doi.org/10.1109/CVPR42600.2020.00269
2020
-
[52]
Dingfeng Shi, Yujie Zhong, Qiong Cao, Lin Ma, Jia Li, and Dacheng Tao. 2023. Temporal Action Localization with Enhanced Instant Discriminability. arXiv preprint arXiv: 2309.05590 (2023)
2023 arXiv
-
[53]
Dingfeng Shi, Yujie Zhong, Qiong Cao, Lin Ma, Jia Lit, and Dacheng Tao. 2023. TriDet: Temporal Action Detection with Relative Boundary Modeling. In Pro- ceedings of 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 18857–18866. https://doi.org/10.1109...
2023
-
[54]
Zheng Shou, Dongang Wang, and Shih-Fu Chang. 2016. Temporal Action Lo- calization in Untrimmed Videos via Multi-stage CNNs. In Proceedings of 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) . 1049–1058. https://doi.org/10.1109/CVPR.2016.119
2016 doi
-
[55]
Alexandros Stergiou and Ronald Poppe. 2019. Analyzing human-human interac- tions: A survey. Computer Vision and Image Understanding 188, C (2019), 102799. https://doi.org/10.1016/J.CVIU.2019.102799
2019
-
[56]
Tan Tang, Yanhong Wu, Yingcai Wu, Lingyun Yu, and Yuhong Li. 2022. Video- Moderator: A Risk-aware Framework for Multimodal Video Moderation in E- Commerce. IEEE Transactions on Visualization and Computer Graphics 28, 1 (2022), 846–856. https://doi.org/10.1109/TVCG.2021.3114781
2022
-
[57]
Zhan Tong, Yibing Song, Jue Wang, and Limin Wang. 2022. VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training. In Advances in Neural Information Processing Systems . 10078–10093
2022
-
[58]
Du Tran, Lubomir Bourdev, Rob Fergus, Lorenzo Torresani, and Manohar Paluri
-
[59]
Paroma Varma, Frederic Sala, Ann He, Alexander Ratner, and Christopher Ré
-
[60]
Binglu Wang, Yongqiang Zhao, Le Yang, Teng Long, and Xuelong Li. 2024. Tem- poral Action Localization in the Deep Learning Era: A Survey. IEEE Trans- actions on Pattern Analysis and Machine Intelligence 46, 4 (2024), 2171–2190. https://doi.org/10.1109/TPAMI.2023.3330794
2024
-
[61]
Limin Wang, Yuanjun Xiong, Dahua Lin, and Luc Van Gool. 2017. UntrimmedNets for Weakly Supervised Action Recognition and Detection. In Proceedings of 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) . 6402–6411. https://doi.org/10.1109/CVPR.2017.678
2017 doi
-
[62]
Limin Wang, Yuanjun Xiong, Zhe Wang, Yu Qiao, Dahua Lin, Xiaoou Tang, and Luc Van Gool. 2019. Temporal Segment Networks for Action Recognition in Videos. IEEE Transactions on Pattern Analysis and Machine Intelligence 41, 11 (2019), 2740–2755. https://doi.org/10.1109/TPAMI.2018.2868668
2019
-
[63]
Tiancai Wang, Tong Yang, Martin Danelljan, Fahad Shahbaz Khan, Xiangyu Zhang, and Jian Sun. 2020. Learning Human-Object Interaction Detection Using Interaction Points. In Proceedings of 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . 4115–4124. htt...
2020
-
[64]
Yanyan Wang, Zhanning Bai, Zhifeng Lin, Xiaoqing Dong, Yingchaojie Feng, Jiacheng Pan, and Wei Chen. 2021. G6: A web-based library for graph visualization. Visual Informatics 5, 4 (2021), 49–55. https://doi.org/10.1016/j.visinf.2021.12.003
2021 doi
-
[65]
In Proceedings of 2015 IEEE International Conference on Computer Vision (ICCV)
Learning Spatiotemporal Features with 3D Convolutional Networks. In Proceedings of 2015 IEEE International Conference on Computer Vision (ICCV) . 4489–4497. https://doi.org/10.1109/ICCV.2015.510
2015 doi
-
[66]
Yufei Xu, Jing Zhang, Qiming Zhang, and Dacheng Tao. 2024. ViTPose: Simple Vision Transformer Baselines for Human Pose Estimation. In Proceedings of the 36th International Conference on Neural Information Processing Systems . 38571– 38584
2024
-
[67]
Katsu Yamane and Yoshihiko Nakamura. 2003. Natural Motion Animation through Constraining and Deconstraining at Will. IEEE Trans. Vis. Comput. Graph. 9, 3 (2003), 352–360. https://doi.org/10.1109/TVCG.2003.1207443
2003 arXiv
-
[68]
Sijie Yan, Yuanjun Xiong, and Dahua Lin. 2018. Spatial Temporal Graph Convo- lutional Networks for Skeleton-Based Action Recognition. Proceedings of AAAI Conference on Artificial Intelligence, 7444–7452. https://doi.org/10.1609/aaai.v32i1. 12328
2018 doi
-
[69]
Le Yang, Junwei Han, Tao Zhao, Tianwei Lin, Dingwen Zhang, and Jianxin Chen
-
[70]
Bangpeng Yao and Li Fei-Fei. 2010. Modeling mutual context of object and human pose in human-object interaction activities. In2010 IEEE Computer Society Conference on Computer Vision and Pattern Recognition . 17–24. https://doi.org/ 10.1109/CVPR.2010.5540235
2010
-
[71]
Runhao Zeng, Wenbing Huang, Chuang Gan, Mingkui Tan, Yu Rong, Peilin Zhao, and Junzhou Huang. 2019. Graph Convolutional Networks for Temporal Action Localization. In Proceedings of 2019 IEEE/CVF International Conference on Computer Vision (ICCV). 7093–7102. https://doi.org/10....
2019
-
[72]
Zhao, Junzhou Huang, and Chuang Gan
Runhao Zeng, Wenbing Huang, Mingkui Tan, Yu Rong, P. Zhao, Junzhou Huang, and Chuang Gan. 2022. Graph Convolutional Module for Temporal Action Local- ization in Videos. IEEE Transactions on Pattern Analysis and Machine Intelligence 44, 10 (2022), 6209–6223. https://doi.org/10....
2022
-
[73]
Kankanhalli
Bingjie Xu, Yongkang Wong, Junnan Li, Qi Zhao, and Mohan S. Kankanhalli. 2019. Learning to Detect Human-Object Interactions With Knowledge. InProceedings of CHI ’25, April 26-May 1, 2025, Yokohama, Japan He et al. the IEEE/CVF Conference on Computer Vision and Pattern Recognit...
2019
-
[74]
Chen-Lin Zhang, Jianxin Wu, and Yin Li. 2022. ActionFormer: Localizing Mo- ments of Actions with Transformers. In Proceedings of Computer Vision - ECCV 2022 - 17th European Conference . 492–510. https://doi.org/10.1007/978-3-031- 19772-7_29
2022 doi
-
[75]
Yue Zhang, Zhenyuan Wang, Jinhui Zhang, Guihua Shan, and Dong Tian. 2023. A survey of immersive visualization: Focus on perception and interaction. Visual Informatics 7, 4 (2023), 22–35. https://doi.org/10.1016/j.visinf.2023.10.003
2023 doi
-
[76]
Chen Zhao, Ali Thabet, and Bernard Ghanem. 2021. Video Self-Stitching Graph Network for Temporal Action Localization. In2021 IEEE/CVF International Confer- ence on Computer Vision (ICCV). 13638–13647. https://doi.org/10.1109/ICCV48922. 2021.01340
2021
-
[77]
Thabet, and Bernard Ghanem
Chen Zhao, Ali K. Thabet, and Bernard Ghanem. 2021. Video Self-Stitching Graph Network for Temporal Action Localization. InProceedings of 2021 IEEE/CVF International Conference on Computer Vision (ICCV) . 13638–13647. https://doi. org/10.1109/ICCV48922.2021.01340
2021
-
[78]
Yue Zhao, Yuanjun Xiong, Limin Wang, Zhirong Wu, Xiaoou Tang, and Dahua Lin. 2017. Temporal Action Detection with Structured Segment Networks. In Proceedings of 2017 IEEE International Conference on Computer Vision (ICCV) . 2933–2942. https://doi.org/10.1109/ICCV.2017.317
2017 doi
-
[79]
Yue Zhao, Yuanjun Xiong, Limin Wang, Zhirong Wu, Xiaoou Tang, and Dahua Lin. 2020. Temporal Action Detection with Structured Segment Networks. Inter- national Journal of Computer Vision 128, 1 (2020), 74–95. https://doi.org/10.1007/ S11263-019-01211-2
2020
-
[80]
Qian Zhou, David Ledo, George Fitzmaurice, and Fraser Anderson. 2024. TimeTun- nel: Integrating Spatial and Temporal Motion Editing for Character Animation in Virtual Reality. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems. 101:1–101:17. https:...
2024
-
[81]
Cheng Zou, Bohan Wang, Yue Hu, Junqi Liu, Qian Wu, Yu Zhao, Boxun Li, Chenguang Zhang, Chi Zhang, Yichen Wei, and Jian Sun. 2021. End-to-End Human Object Interaction Detection With HOI Transformer. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recogn...
2021
-
[82]
Daochen Zha, Zaid Pervaiz Bhat, Kwei-Herng Lai, Fan Yang, Zhimeng Jiang, Shaochen Zhong, and Xia Hu. 2023. Data-centric Artificial Intelligence: A Survey. arXiv preprint arXiv: 2303.10158 (2023)
2023 arXiv
-
[2015]
In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition
ActivityNet: A Large-Scale Video Benchmark for Human Activity Under- standing. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 961–970
-
[2016]
In Proceedings of Computer Vision – ECCV 2016
DAPs: Deep Action Proposals for Action Understanding. In Proceedings of Computer Vision – ECCV 2016 . 768–784. https://doi.org/10.1007/978-3-319- 46487-9_47
2016 doi
-
[2019]
In Proceed- ings of the 36th International Conference on Machine Learning
Learning Dependency Structures for Weak Supervision Models. In Proceed- ings of the 36th International Conference on Machine Learning . 6418–6427
-
[2022]
IEEE Transactions on Pattern Analysis and Machine Intelligence 44, 12 (2022), 9814–9829
Background-Click Supervision for Temporal Action Localization. IEEE Transactions on Pattern Analysis and Machine Intelligence 44, 12 (2022), 9814–9829. https://doi.org/10.1109/TPAMI.2021.3132058
2022
-
[2023]
Visual concept programming: A visual analytics approach to injecting human intelligence at scale. IEEE Transactions on Visualization and Computer ProTAL: A Drag-and-Link Video Programming Framework for Temporal Action Localization CHI ’25, April 26-May 1, 2025, Yokohama, Japan...
2023
-
[6056]
https://doi.org/10.1109/CVPR.2018.00633
2018
-
[6761]
https://doi.org/10.1109/CVPR.2018.00706
2018
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.