REVIEW 3 major objections 7 minor 56 references
Proposal-free Temporal Moment Localization of a Natural-Language Query in Video using Guided Attention
T0 review · 3 major / 7 minor · reviewed 2026-08-14 · deepseek-v4-flash
Pith's one-line read A proposal-free attention model overtakes previous methods for temporal moment localization in video.
desk verdict Clear architecture and two reusable loss ideas, but the SOTA claim is overbroad and the comparison isn't feature-controlled. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the guided-attention dynamic filter. The filter is a single-layer network $\theta(\bar{h}) = \tanh(W_\theta \bar{h} + b_\theta)$ that turns the mean-pooled sentence representation into a query-specific vector; that vector is multiplied against the video feature matrix to produce the attention map $A = \operatorname{softmax}(G^{\top}\theta(\bar{h})/\sqrt{n})$, and the attended sequence $\bar{G} = A \odot G$ is passed to the localization layer. The localization layer uses a two-layer bidirectional GRU followed by two linear heads, one for start and one for end. Two auxiliary losses carry much of the effect: the attention loss $\mathcal{L}_{att}$ suppresses attention mass outside the annotated interval, and the soft-label loss $\mathcal{L}_{KL}$ minimizes KL divergence between the predicted start/end distributions and quantized Gaussian targets centered at the annotated boundaries.
What would settle it
Re-run every baseline in Tables 2, 3, and 4 with the same pre-extracted video features and feature sampling rate used for the proposed model, or re-run the proposed model with the features used by the baselines. If the intersection-over-union margins against ExCL, MAN, CTRL, and ABLR vanish or invert, the central claim is an artifact of feature choice rather than of the proposed architecture.
Extended reading notes
Core claim
The paper's central claim is that moment localization can be done as direct boundary prediction rather than retrieve-and-rank. From the pooled sentence encoding, a filter function produces a query-specific vector; taking a softmax over its inner products with the video features gives a temporal attention map, the attended features are fed through a two-layer bidirectional GRU, and two linear heads predict categorical distributions for the start and end positions. The argument is that two losses make this succeed: an attention loss that penalizes attention mass outside the ground-truth segment, enforcing the hypothesis that the most generalizable features for a query lie inside the target boundaries, and a KL-divergence loss that matches the predicted start and end distributions to quantized Gaussians centered at the annotated boundaries, absorbing annotation uncertainty. With these components, the paper reports best published results on Charades-STA at all reported thresholds and on TACoS at the strictest threshold, plus the best mean intersection-over-union on ActivityNet-Captions.
Load-bearing premise
The central claim assumes all compared methods were given comparable video features; the paper does not report the features used by each baseline, so the state-of-the-art margins could in principle come from the proposed model's stronger pre-extracted video features rather than from the dynamic filter, attention loss, or soft labels.
Editorial extensions
If this is right
- The proposal-free design removes the proposal-generation and ranking stages, so the architecture can be applied to videos of very different lengths without retuning a proposal grid or sliding-window schedule.
- The ablation shows the soft-label KL loss improves accuracy over a hard-index likelihood loss, so treating boundary annotations as uncertain targets rather than exact points is a directly useful recipe for this task.
- The attention loss improves performance with and without soft labels, supporting the paper's hypothesis that query-relevant, generalizable video features concentrate inside the annotated segment.
- The authors state that the dynamic-filter guided-attention mechanism is modular and expect it to transfer to other vision-and-language tasks beyond moment localization.
- On ActivityNet-Captions the method achieves higher mean intersection-over-union than all compared methods, indicating that even where it does not win at every threshold, its predicted boundaries are on average closer to the annotated intervals.
Reading between the lines
- If the gains are due to the proposed losses rather than the feature encoder, the same attention loss could act as a light regularizer for any query-conditioned video encoder, including weakly supervised settings where exact boundary labels are unavailable.
- Because the method outputs full start and end distributions, one could read off calibration or ambiguity directly: a wide predicted distribution would flag queries where multiple moments could match, potentially supporting ranked or confidence-annotated localizations.
- A network with the same soft-label machinery could pool multiple human annotations per moment by fitting a mixture of Gaussians whose variance reflects measured inter-annotator disagreement, a natural next test on datasets with known low agreement.
- The comparison tables leave the visual features of each baseline unspecified; a controlled re-implementation with matched features would reveal how much of the reported margin is architectural.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a proposal-free model for temporal moment localization from natural language queries. The model encodes video with I3D and sentences with a BiGRU over GloVe embeddings, uses a dynamic filter to produce a query-conditioned temporal attention over video features, and feeds the attended features to a localization layer that predicts start and end distributions. Two auxiliary design choices are introduced: a KL loss against Gaussian soft labels to model annotation uncertainty, and an attention loss that encourages attention mass to fall inside the ground-truth moment. The method is evaluated on Charades-STA, TACoS, and ActivityNet-Captions, with an ablation study on Charades-STA. The authors claim state-of-the-art performance on all three datasets.
Significance. The architecture is simple, end-to-end trainable, and proposal-free, and the two proposed losses are well motivated by annotation subjectivity and generalization of attended features. The paper releases code and features, and the ablation in Table 1 cleanly demonstrates the internal value of the KL soft-label loss and the attention loss on Charades-STA. If the state-of-the-art claim were supported, this would be a useful contribution to the moment-localization literature. However, as detailed in the major comments, the headline SOTA claim is not supported by the paper's own reported numbers, and the comparisons are confounded by visual-feature differences, so the significance is currently limited to the architectural and loss-design insights rather than the claimed empirical dominance.
major comments (3)
- [Abstract; Section 4.5, Tables 3 and 4] The blanket claim that the method 'outperforms state-of-the-art methods on these datasets' is contradicted by the paper's own results. On TACoS, ExCL is dramatically better at alpha=0.3 (44.20 vs 24.54) and alpha=0.5 (28.00 vs 21.65), while the proposed method leads only at alpha=0.7. On ANet-Cap, ABLR is better at alpha=0.3 (55.67 vs 51.28) and alpha=0.5 (36.79 vs 33.04). The mIoU margin over ABLR is only 0.79 (37.78 vs 36.99), with no error bars or significance test reported. The abstract and conclusion should be reworded to state where the method leads and where it trails, rather than claiming overall state-of-the-art performance.
- [Section 4.2; Section 4.5] The state-of-the-art comparison is not feature-controlled. The proposed model uses I3D features (for ANet-Cap, fine-tuned on the target dataset), but the visual features used by each baseline are not reported. Several baselines (e.g., CTRL, MCN, ACRN) are known in the literature to use weaker pre-extracted features such as C3D. The large margins on Charades-STA could therefore be partly attributable to feature quality rather than to the proposed guided attention and soft labels. To support the SOTA claim, the paper should either report the features used by every baseline or rerun at least the proposal-free baselines (ABLR, ExCL) with the same I3D features as the proposed model.
- [Section 4.5, Table 5 and text] The claim of state-of-the-art on ANet-Cap rests on a very small mIoU advantage (37.78 vs 36.99 for ABLR), and the paper reports no variance, confidence intervals, or significance tests across multiple runs. Given that the differences at alpha=0.3 and alpha=0.5 favor ABLR, the conclusion that the proposed method 'consistently outperforms' others on this dataset is not justified by the presented evidence. The authors should either report multiple-seed statistics or temper the claim to reflect that the method is competitive with, but not clearly superior to, ABLR on ANet-Cap.
minor comments (7)
- [Introduction, Section 1] The sentence claiming 'to the best of our knowledge, our approach is the first to do so' (i.e., the first proposal-free method) is contradicted by the later description of ABLR and ExCL as proposal-free methods in Section 4.5; this novelty claim should be removed or qualified.
- [Section 3.3, Eq. (2)-(3)] Equation (3) calls A ⊙ G a Hadamard product, but A is an n-dimensional vector and G is an n×d matrix, so the operation is column-wise scaling rather than a true elementwise Hadamard product; the notation should be clarified.
- [Section 3.4, Eq. (5) and nearby text] The notation τs ∼ N(τs, 1) reuses τs for both the target distribution and its mean; this is confusing and should be replaced with distinct symbols, e.g., p_s for the soft-label distribution.
- [Section 3.1] The mapping τ = (t·n·fps)/l appears dimensionally inconsistent if t is a time in seconds; the paper should state explicitly whether t is a frame index and how fps and l relate to the feature sampling rate.
- [Section 4.5, Tables 2-4] Several citation numbers in the tables appear inconsistent with the reference list: Table 2 lists 'ABLR [8]' but ABLR is reference [53], and Table 3 lists 'ACRN [53]' but ACRN is reference [31]; these should be corrected.
- [Section 4.1, TACoS description] The TACoS dataset description contains an unresolved reference placeholder '[?]' for the MPII Compositive dataset; this should be fixed before publication.
- [Conclusion] The conclusion contains a typo: 'archives state-of-the-art performance' should read 'achieves state-of-the-art performance'.
Circularity Check
No circularity: the derivation is a standard supervised-learning pipeline with no self-referential reduction.
full rationale
The paper proposes an end-to-end trainable model for temporal moment localization. The model is trained on labeled start/end annotations and evaluated on held-out test data, which is a standard supervised-learning setup. The attention loss (Eq. 4) and soft-label targets (quantized Gaussians centered at ground-truth boundaries, Eq. 5) are constructed from the training labels and used as supervision; they are not predictions derived from the model. At inference, the model outputs a distribution and takes an argmax (Eq. 7), which is not equivalent to any input by construction. The dynamic filter and guided attention are architectural choices grounded in prior work, not smuggled via self-citation, and the authors do not invoke any uniqueness theorem or rename a known result as a new contribution. The only substantive criticism of the paper is that the claimed state-of-the-art advantage is not fully established because the comparisons do not control for visual-feature quality across baselines (e.g., the proposed method uses I3D features, while some baselines may use weaker features), and on TACoS and ANet-Cap the proposed method loses to ExCL/ABLR at several thresholds. That is a correctness or evaluation-fairness concern, not a circularity. No step in the paper reduces a prediction to a fitted parameter or to a self-citation, and the core derivation is self-contained against external benchmarks. Therefore the circularity score is 0.
Assumptions & free parameters
free parameters (6)
- soft label Gaussian std (sigma) =
1
- attention loss weight =
1 (implicit)
- hidden size =
256
- learning rate =
1e-4
- weight decay =
1e-3
- dropout =
0.5
assumptions (4)
- domain assumption Pre-extracted I3D features are sufficiently discriminative for temporal localization.
- domain assumption The soft-label Gaussian with fixed sigma=1 adequately models annotation uncertainty.
- domain assumption Suppressing attention outside the ground-truth segment during training improves generalization.
- domain assumption Baselines in the comparison tables use comparable visual features and evaluation protocols.
Cite this review
Pith. "Pith review of Proposal-free Temporal Moment Localization of a Natural-Language Query in Video using Guided Attention." pith.science (2026). https://pith.science/paper/CHJTQQQ6
@misc{pith2026190807236,
author = {Pith},
title = {Pith review of: Proposal-free Temporal Moment Localization of a Natural-Language Query in Video using Guided Attention},
year = {2026},
howpublished = {\url{https://pith.science/paper/CHJTQQQ6}},
note = {Machine review of arXiv:1908.07236}
}
read the original abstract
This paper studies the problem of temporal moment localization in a long untrimmed video using natural language as the query. Given an untrimmed video and a sentence as the query, the goal is to determine the starting, and the ending, of the relevant visual moment in the video, that corresponds to the query sentence. While previous works have tackled this task by a propose-and-rank approach, we introduce a more efficient, end-to-end trainable, and {\em proposal-free approach} that relies on three key components: a dynamic filter to transfer language information to the visual domain, a new loss function to guide our model to attend the most relevant parts of the video, and soft labels to model annotation uncertainty. We evaluate our method on two benchmark datasets, Charades-STA and ActivityNet-Captions. Experimental results show that our approach outperforms state-of-the-art methods on both datasets.
Figures
Reference graph
Works this paper leans on
-
[1]
H. Alwassel, F. Caba Heilbron, V . Escorcia, and B. Ghanem. Diagnosing error in temporal action detectors. In The Eu- ropean Conference on Computer Vision (ECCV), September
-
[2]
H. Alwassel, F. Caba Heilbron, V . Escorcia, and B. Ghanem. Diagnosing Error in Temporal Action Detectors. In V . Fer- rari, M. Hebert, C. Sminchisescu, and Y . Weiss, editors, Computer Vision ECCV 2018 , volume 11207, pages 264–
work page 2018
-
[3]
F. Caba Heilbron, V . Escorcia, B. Ghanem, and J. Car- los Niebles. Activitynet: A large-scale video benchmark for human activity understanding. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , pages 961–970, 2015. 5, 6
work page 2015
-
[4]
J. Carreira and A. Zisserman. Quo vadis, action recognition? a new model and the kinetics dataset. In CVPR, 2017. 4, 6
work page 2017
-
[5]
Y . Chao, S. Vijayanarasimhan, B. Seybold, D. A. Ross, J. Deng, and R. Sukthankar. Rethinking the faster R-CNN architecture for temporal action localization. CVPR, 2018. 1
work page 2018
-
[6]
D. Chen, A. Fisch, J. Weston, and A. Bordes. Reading Wikipedia to Answer Open-Domain Questions. pages 1870– 1879, July 2017. 2
work page 2017
-
[7]
J. Chen, X. Chen, L. Ma, Z. Jie, and T.-S. Chua. Tempo- rally grounding natural sentence in video. In Proceedings of the 2018 Conference on Empirical Methods in Natural Lan- guage Processing, pages 162–171, Brussels, Belgium, 2018. Association for Computational Linguistics. 3, 7
work page 2018
-
[8]
S. Chen and Y .-G. Jiang. Semantic proposal for activity lo- calizaiton in videos via sentence query. AAAI, 2019. 3, 7
work page 2019
Show all 56 references
-
[9]
Chung, C
J. Chung, C. Gulcehre, K. Cho, and Y . Bengio. Empirical evaluation of gated recurrent neural networks on sequence modeling. arXiv preprint arXiv:1412.3555, 2014. 4, 5
2014 arXiv
-
[10]
Escorcia, F
V . Escorcia, F. Caba Heilbron, J. C. Niebles, and B. Ghanem. DAPs: Deep Action Proposals for Action Understanding. In B. Leibe, J. Matas, N. Sebe, and M. Welling, editors, Computer Vision ECCV 2016 , Lecture Notes in Computer Science, pages 768–784. Springer International Publishing,
2016
-
[11]
Escorcia, F
V . Escorcia, F. C. Heilbron, J. C. Niebles, and B. Ghanem. DAPs: Deep Action Proposals for Action Understanding. ECCV, 2016. 1
2016
-
[12]
J. Gao, C. Sun, Z. Yang, and R. Nevatia. Tall: Temporal activity localization via language query. In ICCV, 2017. 1, 2, 3, 5, 6, 7
2017
-
[13]
J. Gao, Z. Yang, C. Sun, K. Chen, and R. Nevatia. TURN TAP: temporal unit regression network for temporal action proposals. ICCV, 2017. 1
2017
-
[14]
Gavrilyuk, A
K. Gavrilyuk, A. Ghodrati, Z. Li, and C. G. M. Snoek. Ac- tor and action video segmentation from a sentence. In The IEEE Conference on Computer Vision and Pattern Recogni- tion (CVPR), June 2018. 4
2018
-
[15]
R. Ge, J. Gao, K. Chen, and R. Nevatia. Mac: Mining ac- tivity concepts for language-based temporal localization. In WACV, 2019. 3
2019
-
[16]
Ghosh, A
S. Ghosh, A. Agarwal, Z. Parekh, and A. Hauptmann. Excl: Extractive clip localization using natural language descrip- tions. arXiv preprint arXiv:1904.02755, 2019. 2, 3, 6, 7
1904 arXiv
-
[17]
M. Hahn, A. Kadav, J. M. Rehg, and H. P. Graf. Tripping through time: Efficient localization of activities in videos. arXiv preprint arXiv:1904.09936, 2019. 3, 7
1904 arXiv
-
[18]
K. He, X. Zhang, S. Ren, and J. Sun. Deep residual learning for image recognition. InThe IEEE Conference on Computer Vision and Pattern Recognition (CVPR), June 2016. 4
2016
-
[19]
L. A. Hendricks, O. Wang, E. Shechtman, J. Sivic, T. Darrell, and B. Russell. Localizing moments in video with natural language. In ICCV, 2017. 1, 2, 3, 7
2017
-
[20]
L. A. Hendricks, O. Wang, E. Shechtman, J. Sivic, T. Darrell, and B. Russell. Localizing moments in video with temporal language. In EMNLP, 2018. 3
2018
-
[21]
Idrees, A
H. Idrees, A. R. Zamir, Y .-G. Jiang, A. Gorban, I. Laptev, R. Sukthankar, and M. Shah. The THUMOS challenge on action recognition for videos in the wild. Computer Vision and Image Understanding, 155:1–23, Feb. 2017. 2
2017
-
[22]
Ioffe and C
S. Ioffe and C. Szegedy. Batch normalization: Accelerating deep network training by reducing internal covariate shift. arXiv preprint arXiv:1502.03167, 2015. 4
2015 arXiv
-
[23]
X. Jia, B. De Brabandere, T. Tuytelaars, and L. V . Gool. Dy- namic filter networks. In Advances in Neural Information Processing Systems, pages 667–675, 2016. 4
2016
-
[24]
W. Kay, J. Carreira, K. Simonyan, B. Zhang, C. Hillier, S. Vijayanarasimhan, F. Viola, T. Green, T. Back, P. Natsev, M. Suleyman, and A. Zisserman. The kinetics human action video dataset. CoRR, 2017. 6
2017
-
[25]
D. P. Kingma and J. Ba. Adam: A method for stochastic optimization. CoRR, 2014. 6
2014
-
[26]
Krishna, K
R. Krishna, K. Hata, F. Ren, L. Fei-Fei, and J. C. Niebles. Dense-captioning events in videos. In ICCV, 2017. 2, 5
2017
-
[27]
Krishna, Y
R. Krishna, Y . Zhu, O. Groth, J. Johnson, K. Hata, J. Kravitz, S. Chen, Y . Kalantidis, L.-J. Li, D. A. Shamma, M. Bern- stein, and L. Fei-Fei. Visual genome: Connecting language and vision using crowdsourced dense image annotations
-
[28]
Z. Li, R. Tao, E. Gavves, C. G. Snoek, and A. W. Smeulders. Tracking by natural language specification. In CVPR, pages 6495–6503, 2017. 4
2017
-
[29]
T. Lin, X. Zhao, and Z. Shou. Single Shot Temporal Action Detection. ACMMM, 2017. 1
2017
-
[30]
T. Lin, X. Zhao, and Z. Shou. Single Shot Temporal Action Detection. In Proceedings of the 25th ACM International Conference on Multimedia , MM ’17, pages 988–996, New York, NY , USA, 2017. ACM. event-place: Mountain View, California, USA. 2
2017
-
[31]
M. Liu, X. Wang, L. Nie, X. He, B. Chen, and T.-S. Chua. Attentive moment retrieval in videos. In The 41st Interna- tional ACM SIGIR Conference on Research & Development in Information Retrieval, pages 15–24. ACM, 2018. 3, 7
2018
-
[32]
Y . Liu, A. Gupta, P. Abbeel, and S. Levine. Imitation from observation: Learning to imitate behaviors from raw video via context translation. 2019. 1
2019
-
[33]
S. Ma, L. Sigal, and S. Sclaroff. Learning Activity Progres- sion in LSTMs for Activity Detection and Early Detection. pages 1942–1950, 2016. 2
1942
-
[34]
Pennington, R
J. Pennington, R. Socher, and C. Manning. Glove: Global vectors for word representation. In Proceedings of the 2014 conference on empirical methods in natural language pro- cessing (EMNLP), pages 1532–1543, 2014. 4
2014
-
[35]
S. Ren, K. He, R. Girshick, and J. Sun. Faster R-CNN: To- wards real-time object detection with region proposal net- works. In NIPS, 2015. 3
2015
-
[36]
Richard, H
A. Richard, H. Kuehne, A. Iqbal, and J. Gall. Neuralnetwork-viterbi: A framework for weakly super- vised video learning. In IEEE Conf. on Computer Vision and Pattern Recognition, volume 2, 2018. 1
2018
-
[37]
Rohrbach, M
A. Rohrbach, M. Rohrbach, W. Qiu, A. Friedrich, M. Pinkal, and B. Schiele. Coherent multi-sentence video description with variable level of detail. In X. Jiang, J. Hornegger, and R. Koch, editors, Pattern Recognition, 2014. 2, 5, 6
2014
-
[38]
Rohrbach, M
A. Rohrbach, M. Rohrbach, N. Tandon, and B. Schiele. A Dataset for Movie Description. pages 3202–3212, 2015. 2
2015
-
[39]
Salimans, I
T. Salimans, I. Goodfellow, W. Zaremba, V . Cheung, A. Rad- ford, and X. Chen. Improved techniques for training gans. In Advances in neural information processing systems , pages 2234–2242, 2016. 5
2016
-
[40]
Z. Shou, D. Wang, and S.-F. Chang. Temporal action lo- calization in untrimmed videos via multi-stage cnns. In The IEEE Conference on Computer Vision and Pattern Recogni- tion (CVPR), June 2016. 2
2016
-
[41]
G. A. Sigurdsson, O. Russakovsky, and A. Gupta. What ac- tions are needed for understanding human actions in videos? In ICCV, 2017. 2
2017
-
[42]
G. A. Sigurdsson, O. Russakovsky, and A. Gupta. What ac- tions are needed for understanding human actions in videos? In Proceedings of the IEEE International Conference on Computer Vision, pages 2137–2146, 2017. 5
2017
-
[43]
G. A. Sigurdsson, G. Varol, X. Wang, A. Farhadi, I. Laptev, and A. Gupta. Hollywood in homes: Crowdsourcing data collection for activity understanding. In European Confer- ence on Computer Vision, 2016. 5
2016
-
[44]
Simonyan and A
K. Simonyan and A. Zisserman. Two-stream convolutional networks for action recognition in videos. In Advances in neural information processing systems , pages 568–576,
-
[45]
Singh, T
B. Singh, T. K. Marks, M. Jones, O. Tuzel, and M. Shao. A Multi-Stream Bi-Directional Recurrent Neural Network for Fine-Grained Action Detection. pages 1961–1970, 2016. 2
1961
-
[46]
Szegedy, W
C. Szegedy, W. Liu, Y . Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V . Vanhoucke, and A. Rabinovich. Going deeper with convolutions. In Proceedings of the IEEE conference on computer vision and pattern recogni- tion, pages 1–9, 2015. 4
2015
-
[47]
Szegedy, V
C. Szegedy, V . Vanhoucke, S. Ioffe, J. Shlens, and Z. Wojna. Rethinking the Inception Architecture for Computer Vision. pages 2818–2826, 2016. 5
2016
-
[48]
D. Tran, L. Bourdev, R. Fergus, L. Torresani, and M. Paluri. Learning spatiotemporal features with 3d convolutional net- works. In Proceedings of the IEEE international conference on computer vision, pages 4489–4497, 2015. 2, 4
2015
-
[49]
Vaswani, N
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin. Attention is all you need. In Advances in neural information processing sys- tems, pages 5998–6008, 2017. 5
2017
-
[50]
W. Wang, Y . Huang, and L. Wang. Language-driven tempo- ral activity localization: A semantic matching reinforcement learning model. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , pages 334–343,
-
[51]
H. Xu, A. Das, and K. Saenko. R-c3d: Region convolutional 3d network for temporal activity detection. ICCV, 2017. 1
2017
-
[52]
H. Xu, K. He, L. Sigal, S. Sclaroff, and K. Saenko. Mul- tilevel language and vision integration for text-to-clip re- trieval. In AAAI, 2019. 3, 7
2019
-
[53]
Y . Yuan, T. Mei, and W. Zhu. To find where you talk: Tem- poral sentence localization in video with attention based lo- cation regression. AAAI, 2019. 3, 7
2019
-
[54]
Zhang, X
D. Zhang, X. Dai, X. Wang, Y .-F. Wang, and L. S. Davis. Man: Moment alignment network for natural language mo- ment retrieval via iterative graph adjustment. CVPR, 2019. 3, 4, 7
2019
-
[55]
Y . Zhao, Y . Xiong, L. Wang, Z. W. . . . . V . (ICCV), and U. 2017. Temporal action detection with structured segment networks. ICCV, 2017. 1
2017
-
[280]
Springer International Publishing, Cham, 2018. 2
2018
Reviewed August 14, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.