REVIEW 4 major objections 7 minor 1 cited by
Dual-task Mutual Reinforcing Embedded Joint Video Paragraph Retrieval and Grounding
T0 review · 4 major / 7 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read The paper defines VPRG, a task that jointly retrieves a video for a paragraph query and localizes each sentence's event, and claims DMR-JRG solves it without temporal labels.
desk verdict VPRG is a sensible new task, but the reported gains are confounded by query format and a suspect baseline row, and the temporal-synchronization module has a training-loop problem as written. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the mutual-reinforcement loop between a coarse and a fine cross-modal feature space. The retrieval branch (Video-Paragraph Cross Modal Retrieval) constructs the coarse space with class tokens and InfoNCE plus triplet losses. The grounding branch constructs the fine space with three modules: Visual-Textual Consistency on Local Dimension aligns sentence features to candidate moments via 2D-TAN-style proposals and masked-language reconstruction; Visual-Textual Feature Alignment on Global Dimension fuses top candidate-moment features and aligns a fused class token to the paragraph token; and Bidirectional Temporal Synchronization of Events on Temporal Dimension runs forward and reverse branches whose predictions become soft pseudo-labels supervising the main score map. The Grounding Reinforcement Retrieval Module closes the loop by using grounding scores as pseudo-labels for retrieval scores with an MSE loss.
What would settle it
Randomly shuffle the sentence order of test paragraphs in ActivityNet Captions before feeding them to a trained DMR-JRG model and compare grounding accuracy at IoU thresholds 0.3, 0.5, and 0.7 with the unshuffled results; a sharp drop, particularly relative to an ablation without the bidirectional synchronization module, would show that the chronological-order assumption carries the reported gains.
Extended reading notes
Core claim
The paper's central claim is that VPRG can be solved under weak supervision by treating retrieval and grounding as mutually reinforcing tasks. This is, by the authors' account, the first attempt at the combined task. The retrieval branch uses class tokens and contrastive losses to build a coarse paragraph-video feature space; the grounding branch builds a fine-grained space through local sentence-video consistency, global class-token alignment, and bidirectional temporal synchronization, with the forward and reverse synchronization branches producing pseudo-labels that train the main score maps. A grounding-reinforcement-retrieval module regresses retrieval scores toward grounding scores, completing the mutual-reinforcement loop. The paper reports that DMR-JRG outperforms existing sentence-based retrieval-and-grounding methods on ActivityNet Captions, Charades-STA, and TaCoS, and interprets the gains as evidence that paragraph context plus multi-dimensional consistency helps both tasks.
Load-bearing premise
The load-bearing premise is that the sentences of a paragraph describe events in the same chronological order in which they occur in the video; if a paragraph is not ordered that way, the forward and reverse synchronization branches generate wrong pseudo-labels and the grounding score maps are trained with incorrect supervision.
Editorial extensions
If this is right
- A single weakly supervised model can retrieve a video from a paragraph query and localize every sentence's event, using only paragraph-video correspondence during training.
- Paragraph-level queries can beat sentence-level retrieval-and-grounding systems on the same datasets, since the extra context improves both retrieval and grounding.
- Grounding quality transfers to retrieval: the pseudo-label feedback from grounding to retrieval (GRRM) improves retrieval scores in the ablations.
- Bidirectional chronological synchronization contributes more than either direction alone, indicating that exploiting sentence order is useful for grounding.
- The approach remains effective on a visually homogeneous dataset (TaCoS), where distinguishing events requires language as much as vision.
Reading between the lines
- A natural next test is to compare DMR-JRG against fully supervised VPG methods under a shared protocol; the paper only compares with VSRG baselines, so the gap between weak and full supervision is not yet measured.
- If the chronological-order assumption transfers, the bidirectional synchronization idea could be applied to other ordered-alignment problems, such as aligning instruction text to procedural videos or narrated slides to recordings.
- The method's reliance on paragraph-video pairs during training still assumes one paragraph maps to one video in the batch; extending it to partially relevant or multi-paragraph queries would require reformulating the contrastive negatives.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper defines a new task, Video Paragraph Retrieval and Grounding (VPRG), and proposes DMR-JRG, a weakly supervised dual-branch method. The retrieval branch aligns global paragraph and video features via InfoNCE and triplet losses; the grounding branch uses local, global, and temporal consistency modules (VTC-LD, VTFA-GD, BTSE-TD), with a GRRM module distilling grounding scores back into retrieval scores. Training requires only paragraph-video correspondence, not temporal boundary labels. The method is evaluated on ActivityNet Captions, Charades-STA, and TaCoS against VSRG baselines, with ablations and hyperparameter studies. The code is released.
Significance. If the central empirical claim were cleanly supported, this would be a useful contribution: it defines a sensible new task that combines paragraph-level retrieval with sentence-level grounding, it operates under a weak-supervision setting that reduces annotation cost, and it openly provides code. The ablation study (Table IV) and hyperparameter analyses (Tables V, Fig. 7-8) are thorough and internally consistent. However, the headline comparisons in Tables II-III are confounded by an unfair query-setting mismatch, and the TaCoS baseline row appears to be a verbatim copy of the Charades-STA row. The self-training nature of BTSE-TD (Eqs. 18-21) and GRRM (Eq. 17) also deserves closer scrutiny. Because the empirical superiority claim is the paper's third stated contribution, these issues are load-bearing for the current form of the paper.
major comments (4)
- [§IV.C, Tables II-III] The comparison is confounded: DMR-JRG uses full paragraph queries (§III.A) while every baseline (MCN, CAL, XML, HMAN, ReLoCLNet, MS-SL, JSG) is a VSRG method that takes a single-sentence query. The paper acknowledges this mismatch in §IV.C and asserts 'relative fairness,' but paragraph queries contain strictly more information: sentences in a paragraph can mutually disambiguate events and constrain the temporal layout, which helps both retrieval and grounding. The reported gains therefore conflate the query-format advantage with the proposed architecture. The claim that DMR-JRG 'significantly outperforms several latest VSRG methods' (Sec. I, contribution 3) is not established without paragraph-query baselines (e.g., JSG or SCN with sentence-level scores pooled over the paragraph) or a sentence-query variant of DMR-JRG.
- [Table III] The row labeled 'JSG* [31]' on TaCoS lists values 7.23, 28.71, 5.67, 22.50, 3.28, 12.34, which are exactly the values reported for JSG on Charades-STA in Table II. Since TaCoS and Charades-STA differ substantially in video count, paragraph length, and domain, identical results are implausible and indicate a replication/copy error. As a result, all TaCoS improvement percentages in §IV.C (e.g., 11.51%, 16.74%) are computed against an invalid baseline, and the TaCoS superiority claim is unsupported until JSG is rerun on TaCoS and the table is corrected.
- [§III.G, Eqs. (18)-(21)] BTSE-TD generates soft labels gt^f and gt^r by taking argmax over P^f_m and P^r_m, which are themselves predictions of the same network (Eqs. 18-19), and then uses those labels to supervise the main score map P_m via binary cross-entropy (Eqs. 20-21). This is a self-training loop with no ground-truth anchor; its correctness hinges on the assumption stated in §III.G that 'sentences describing these events in text paragraphs follow a certain logical sequence.' The assumption is not verified on the experimental datasets, and ActivityNet Captions in particular does not guarantee chronological caption order. If a paragraph is not temporally ordered, the pseudo-labels will be systematically wrong and the BCE losses will train P_m with incorrect supervision. The paper should either provide evidence that paragraphs in the three datasets are consistently ordered, or include a diagnostic showing that BTSE-TD improves over a version trained without this self-supervision.
- [§III.F, Eq. (17)] The GRRM loss forces the retrieval score Sr(·) to match the grounding score Sg(·), which is itself a product of the same network. The paper states that grounding scores are used as 'pseudo-labels' to reinforce retrieval, but it does not show that Sg is more reliable than Sr before distillation, nor does it evaluate the retrieval branch in the setting where the correspondence between videos and paragraphs is truly unknown at training time. The claim in Sec. I that the method works 'when the correspondence between videos and paragraphs is unknown' is only demonstrated in the standard scenario where correspondence is known during training and removed at inference. The paper should clarify the intended meaning of 'correspondence unknown' and, if the stronger claim is intended, provide an experiment that trains without paragraph-video pairing, or at least justify why the current training setting suffices.
minor comments (7)
- [Sec. I, contribution list] The method name is typed inconsistently: the first contribution bullet reads 'Dual-task Mutual Reinforcing Embedded Joint Video Paragraph Retrieval and Grounding (DMP-JRG)', while the rest of the paper uses DMR-JRG.
- [Index Terms] The index term 'Video Paragraph Retrival' contains a misspelling; it should read 'Retrieval'.
- [§III.B, Eq. (1)] The text says 'E1 and E1 are two Transformer encoders' but the two encoders should be E1 and E2, based on Eq. (1).
- [§III.H.2] The inference section refers to 'TSVT-TD', which appears to be a typo for the BTSE-TD module.
- [§IV.E] The hyperparameter section repeatedly writes 'ActiveNet Captions'; the correct dataset name is 'ActivityNet Captions'.
- [Fig. 3 caption] The caption reads 'It comprises there core parts'; 'there' should be 'three'.
- [§IV.C (ActivityNet discussion)] The sentence 'our approach shows significant improvements in the VSRG task' should refer to the VPRG task, since that is the task being evaluated.
Circularity Check
No circular derivation: the claimed VPRG performance is an empirical result evaluated on external ground truth, and the internal pseudo-label and distillation loops are training mechanisms rather than inferential reductions.
full rationale
The paper's central claim is that DMR-JRG solves the new VPRG task and outperforms existing VSRG methods. This is an empirical claim measured by R@K IoU against held-out temporal annotations and retrieval ground truth, not a closed-form derivation from the model's own outputs. The BTSE-TD module (Eqs. 18-21) generates soft labels from its own forward/reverse predictions and uses them to supervise the main score maps, and GRRM (Eq. 17) distills grounding similarities into retrieval similarities. These are self-referential training objectives, but they do not make the test-time prediction equivalent to the training input: evaluation still depends on external labels, and the reported numbers are not forced by construction. There is no load-bearing self-citation or imported uniqueness theorem; citations such as SCN [49] and 2D-TAN [32] are used for standard component choices (masking, proposal construction, reward heuristics). The most serious issues are external validity concerns, not circularity. The paper itself admits in Sec. IV-C that 'due to the lack of existing research on VPRG task, to validate the effectiveness of our approach, we choose to compare it with existing VSRG methods,' and notes that VSRG methods use single-sentence queries while DMR-JRG uses paragraph queries. That asymmetry confounds the claimed improvements but does not constitute a circular derivation. Additionally, the TaCoS baseline row 'JSG* [31]' in Table III is numerically identical to JSG's Charades-STA row in Table II (7.23, 28.71, 5.67, 22.50, 3.28, 12.34), which appears to be a data-replication error rather than a circular step. Under the stated standard requiring an exhibited equation-level reduction or fitted-parameter-renamed-as-prediction, the paper exhibits no significant circularity.
Assumptions & free parameters
free parameters (8)
- Q =
3
- w_i =
{0.4, 0.3, 0.3}
- beta_1 =
0.04
- beta_2 =
unspecified
- Delta =
0.2
- lambda =
learnable
- mask_ratio =
1/3
- reward_scores =
1 with step 1/(Q-1)
assumptions (8)
- domain assumption C3D features pre-trained on UCF101 are sufficient video representations.
- domain assumption Word2Vec plus Bi-LSTM captures sentence semantics for paragraphs.
- domain assumption The 2D-TAN candidate segment set covers the true moments.
- domain assumption Sentences in a paragraph occur in the same chronological order as the corresponding events in the video.
- ad hoc to paper The top Q candidate moments selected from the predicted score map contain the correct moments.
- ad hoc to paper Pseudo-labels produced by the forward and reverse branches are reliable enough to supervise the main score map.
- domain assumption Known paragraph-video correspondence is available during training.
- domain assumption The VSRG evaluation metric transfers meaningfully to paragraph queries.
invented entities (3)
-
Learnable class tokens (f_cls,v, f_cls,t, f'_cls,v, f_cls,f)
-
Grounding Reinforcement Retrieval Module (GRRM)
-
Bidirectional Temporal Synchronization module (BTSE-TD)
Cite this review
Pith. "Pith review of Dual-task Mutual Reinforcing Embedded Joint Video Paragraph Retrieval and Grounding." pith.science (2026). https://pith.science/paper/E5TGUWQH
@misc{pith2026241117481,
author = {Pith},
title = {Pith review of: Dual-task Mutual Reinforcing Embedded Joint Video Paragraph Retrieval and Grounding},
year = {2026},
howpublished = {\url{https://pith.science/paper/E5TGUWQH}},
note = {Machine review of arXiv:2411.17481}
}
read the original abstract
Video Paragraph Grounding (VPG) aims to precisely locate the most appropriate moments within a video that are relevant to a given textual paragraph query. However, existing methods typically rely on large-scale annotated temporal labels and assume that the correspondence between videos and paragraphs is known. This is impractical in real-world applications, as constructing temporal labels requires significant labor costs, and the correspondence is often unknown. To address this issue, we propose a Dual-task Mutual Reinforcing Embedded Joint Video Paragraph Retrieval and Grounding method (DMR-JRG). In this method, retrieval and grounding tasks are mutually reinforced rather than being treated as separate issues. DMR-JRG mainly consists of two branches: a retrieval branch and a grounding branch. The retrieval branch uses inter-video contrastive learning to roughly align the global features of paragraphs and videos, reducing modality differences and constructing a coarse-grained feature space to break free from the need for correspondence between paragraphs and videos. Additionally, this coarse-grained feature space further facilitates the grounding branch in extracting fine-grained contextual representations. In the grounding branch, we achieve precise cross-modal matching and grounding by exploring the consistency between local, global, and temporal dimensions of video segments and textual paragraphs. By synergizing these dimensions, we construct a fine-grained feature space for video and textual features, greatly reducing the need for large-scale annotated temporal labels.
Figures
Figures from the paper (7 more)
Forward citations
Cited by 1 Pith paper
-
Query-centric Audio-Visual Cognition Network for Moment Retrieval, Segmentation and Step-Captioning
QUAG, a query-centric audio-visual cognition network, reports state-of-the-art moment retrieval and segmentation on HIREST and competitive video summarization on TVSum.
Reference graph
Works this paper leans on
-
[31]
Joint searching and grounding: Multi-granularity video content retrieval,
Z. Chen, X. Jiang, X. Xu, Z. Cao, Y . Mo, and H. T. Shen, “Joint searching and grounding: Multi-granularity video content retrieval,” in Proceedings of the 31st ACM International Conference on Multimedia , 2023, pp. 975–983
work page 2023
-
[1]
Real-world anomaly detection in surveillance videos,
W. Sultani, C. Chen, and M. Shah, “Real-world anomaly detection in surveillance videos,” in 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2018, pp. 6479–6488
work page 2018
-
[2]
Toward video anomaly retrieval from video anomaly detection: New benchmarks and model,
P. Wu, J. Liu, X. He, Y . Peng, P. Wang, and Y . Zhang, “Toward video anomaly retrieval from video anomaly detection: New benchmarks and model,” IEEE Transactions on Image Processing , vol. 33, pp. 2213– 2225, 2024
2024
-
[3]
Robust multi-drone multi-target tracking to resolve target occlusion: A benchmark,
Z. Liu, Y . Shang, T. Li, G. Chen, Y . Wang, Q. Hu, and P. Zhu, “Robust multi-drone multi-target tracking to resolve target occlusion: A benchmark,” IEEE Transactions on Multimedia , vol. 25, pp. 1462– 1476, 2023
work page 2023
-
[4]
Yolov3-mt: A yolov3 using multi-target tracking for vehicle visual detection,
K. Wang and M. Liu, “Yolov3-mt: A yolov3 using multi-target tracking for vehicle visual detection,” Applied Intelligence , vol. 52, no. 2, pp. 2070–2091, 2022
work page 2022
-
[5]
Deepmtt: A deep learning maneuvering target-tracking algorithm based on bidirectional lstm network,
J. Liu, Z. Wang, and M. Xu, “Deepmtt: A deep learning maneuvering target-tracking algorithm based on bidirectional lstm network,” Informa- tion Fusion, vol. 53, pp. 289–304, 2020
work page 2020
-
[6]
Robust obstacle detection and recognition for driver assistance systems,
J. Leng, Y . Liu, D. Du, T. Zhang, and P. Quan, “Robust obstacle detection and recognition for driver assistance systems,” IEEE Transactions on Intelligent Transportation Systems, vol. 21, no. 4, pp. 1560–1571, 2020
work page 2020
-
[7]
End-to-end object detection with transformers,
N. Carion, F. Massa, G. Synnaeve, N. Usunier, A. Kirillov, and S. Zagoruyko, “End-to-end object detection with transformers,” in European conference on computer vision , 2020, pp. 213–229
2020
Show all 68 references
-
[8]
Pareto refocusing for drone-view object detection,
J. Leng, M. Mo, Y . Zhou, C. Gao, W. Li, and X. Gao, “Pareto refocusing for drone-view object detection,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 33, no. 3, pp. 1320–1334, 2023
2023
-
[9]
Sparse r-cnn: End-to-end object detection with learnable proposals,
P. Sun, R. Zhang, Y . Jiang, T. Kong, C. Xu, W. Zhan, M. Tomizuka, L. Li, Z. Yuan, C. Wang et al. , “Sparse r-cnn: End-to-end object detection with learnable proposals,” in 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2021, pp. 14 454– 14 463
2021
-
[10]
Recent advances for aerial object detection: A survey,
J. Leng, Y . Ye, M. Mo, C. Gao, J. Gan, B. Xiao, and X. Gao, “Recent advances for aerial object detection: A survey,”ACM Computing Surveys, vol. 56, no. 12, pp. 1–36, 2024
2024
-
[11]
Cascade r-cnn: Delving into high quality object detection,
Z. Cai and N. Vasconcelos, “Cascade r-cnn: Delving into high quality object detection,” in 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2018, pp. 6154–6162
2018
-
[12]
Crnet: Context-guided reasoning network for detecting hard objects,
J. Leng, Y . Liu, X. Gao, and Z. Wang, “Crnet: Context-guided reasoning network for detecting hard objects,” IEEE Transactions on Multimedia , vol. 26, pp. 3765–3777, 2024
2024
-
[13]
Triple adversarial learning and multi-view imaginative reasoning for unsupervised domain adaptation person re-identification,
H. Li, N. Dong, Z. Yu, D. Tao, and G. Qi, “Triple adversarial learning and multi-view imaginative reasoning for unsupervised domain adaptation person re-identification,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 32, no. 5, pp. 2814–2830, 2021
2021
-
[14]
Logical relation inference and multiview information interaction for domain adaptation person re-identification,
S. Li, F. Li, J. Li, H. Li, B. Zhang, D. Tao, and X. Gao, “Logical relation inference and multiview information interaction for domain adaptation person re-identification,” IEEE Transactions on Neural Networks and Learning Systems, vol. 35, no. 10, pp. 14 770–14 782, 2024
2024
-
[15]
Attribute-aligned domain- invariant feature learning for unsupervised domain adaptation person re-identification,
H. Li, Y . Chen, D. Tao, Z. Yu, and G. Qi, “Attribute-aligned domain- invariant feature learning for unsupervised domain adaptation person re-identification,” IEEE Transactions on Information Forensics and Security, vol. 16, pp. 1480–1494, 2021
2021
-
[16]
Intermediary-guided bidi- rectional spatial–temporal aggregation network for video-based visible- infrared person re-identification,
H. Li, M. Liu, Z. Hu, F. Nie, and Z. Yu, “Intermediary-guided bidi- rectional spatial–temporal aggregation network for video-based visible- infrared person re-identification,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 33, no. 9, pp. 4962–4972, 2023
2023
-
[17]
Video moment retrieval from text queries via single frame annotation,
R. Cui, T. Qian, P. Peng, E. Daskalaki, J. Chen, X. Guo, H. Sun, and Y .-G. Jiang, “Video moment retrieval from text queries via single frame annotation,” in Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval , 2022,...
2022
-
[18]
Text-based local- ization of moments in a video corpus,
S. Paul, N. C. Mithun, and A. K. Roy-Chowdhury, “Text-based local- ization of moments in a video corpus,” IEEE Transactions on Image Processing, vol. 30, pp. 8886–8899, 2021
2021
-
[19]
Language-guided multi-granularity con- text aggregation for temporal sentence grounding,
G. Gong, L. Zhu, and Y . Mu, “Language-guided multi-granularity con- text aggregation for temporal sentence grounding,” IEEE Transactions on Multimedia, vol. 25, pp. 7402–7414, 2023
2023
-
[20]
Conditional video diffusion network for fine-grained temporal sentence grounding,
D. Liu, J. Zhu, X. Fang, Z. Xiong, H. Wang, R. Li, and P. Zhou, “Conditional video diffusion network for fine-grained temporal sentence grounding,” IEEE Transactions on Multimedia , vol. 26, pp. 5461–5476, 2024
2024
-
[21]
Relational net- work via cascade crf for video language grounding,
T. Zhang, X. Lu, H. Zhang, X. Nie, Y . Yin, and J. Shen, “Relational net- work via cascade crf for video language grounding,” IEEE Transactions on Multimedia, vol. 26, pp. 8297–8311, 2024. JOURNAL OF LATEX CLASS FILES 15
2024
-
[22]
Self-supervised learn- ing for semi-supervised temporal language grounding,
F. Luo, S. Chen, J. Chen, Z. Wu, and Y .-G. Jiang, “Self-supervised learn- ing for semi-supervised temporal language grounding,” IEEE Transac- tions on Multimedia , vol. 25, pp. 7747–7757, 2023
2023
-
[23]
Zero-shot video moment retrieval with angular reconstructive text embeddings,
X. Jiang, X. Xu, Z. Zhou, Y . Yang, F. Shen, and H. T. Shen, “Zero-shot video moment retrieval with angular reconstructive text embeddings,” IEEE Transactions on Multimedia , pp. 1–14, 2024
2024
-
[24]
Point-supervised video temporal grounding,
Z. Xu, K. Wei, X. Yang, and C. Deng, “Point-supervised video temporal grounding,” IEEE Transactions on Multimedia , vol. 25, pp. 6121–6131, 2023
2023
-
[25]
Siamese learning with joint alignment and regression for weakly-supervised video paragraph grounding,
C. Tan, J. Lai, W.-S. Zheng, and J.-F. Hu, “Siamese learning with joint alignment and regression for weakly-supervised video paragraph grounding,” in 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2024, pp. 13 569–13 580
2024
-
[26]
Dense events grounding in video,
P. Bao, Q. Zheng, and Y . Mu, “Dense events grounding in video,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 35, no. 2, 2021, pp. 920–928
2021
-
[27]
Semi- supervised video paragraph grounding with contrastive encoder,
X. Jiang, X. Xu, J. Zhang, F. Shen, Z. Cao, and H. T. Shen, “Semi- supervised video paragraph grounding with contrastive encoder,” in2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022, pp. 2466–2475
2022
-
[28]
Gtlr: Graph- based transformer with language reconstruction for video paragraph grounding,
X. Jiang, X. Xu, J. Zhang, F. Shen, Z. Cao, and X. Cai, “Gtlr: Graph- based transformer with language reconstruction for video paragraph grounding,” in 2022 IEEE International Conference on Multimedia and Expo (ICME), 2022, pp. 1–6
2022
-
[29]
Hierarchical semantic correspondence networks for video paragraph grounding,
C. Tan, Z. Lin, J.-F. Hu, W.-S. Zheng, and J. Lai, “Hierarchical semantic correspondence networks for video paragraph grounding,” in 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023, pp. 18 973–18 982
2023
-
[30]
End-to-end dense video grounding via parallel regression,
F. Shi, W. Huang, and L. Wang, “End-to-end dense video grounding via parallel regression,” Computer Vision and Image Understanding , vol. 242, p. 103980, 2024
2024
-
[32]
Learning 2d temporal adjacent networks for moment localization with natural language,
S. Zhang, H. Peng, J. Fu, and J. Luo, “Learning 2d temporal adjacent networks for moment localization with natural language,” inProceedings of the AAAI Conference on Artificial Intelligence , vol. 34, no. 07, 2020, pp. 12 870–12 877
2020
-
[33]
Multi- stage aggregated transformer network for temporal language localization in videos,
M. Zhang, Y . Yang, X. Chen, Y . Ji, X. Xu, J. Li, and H. T. Shen, “Multi- stage aggregated transformer network for temporal language localization in videos,” in 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2021, pp. 12 669–12 678
2021
-
[34]
Structured multi- level interaction network for video moment localization via language query,
H. Wang, Z.-J. Zha, L. Li, D. Liu, and J. Luo, “Structured multi- level interaction network for video moment localization via language query,” in 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2021, pp. 7026–7035
2021
-
[35]
Progressive localization networks for language-based moment localization,
Q. Zheng, J. Dong, X. Qu, X. Yang, Y . Wang, P. Zhou, B. Liu, and X. Wang, “Progressive localization networks for language-based moment localization,” ACM Transactions on Multimedia Computing, Communications and Applications , vol. 19, no. 2, pp. 1–21, 2023
2023
-
[36]
Fast video moment retrieval,
J. Gao and C. Xu, “Fast video moment retrieval,” in 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2021, pp. 1523–1532
2021
-
[37]
Exploring optical-flow-guided motion and detection-based appearance for temporal sentence ground- ing,
D. Liu, X. Fang, W. Hu, and P. Zhou, “Exploring optical-flow-guided motion and detection-based appearance for temporal sentence ground- ing,” IEEE Transactions on Multimedia , vol. 25, pp. 8539–8553, 2023
2023
-
[38]
Temporally language grounding with multi-modal multi-prompt tuning,
Y . Zeng, N. Han, K. Pan, and Q. Jin, “Temporally language grounding with multi-modal multi-prompt tuning,” IEEE Transactions on Multime- dia, vol. 26, pp. 3366–3377, 2024
2024
-
[39]
Dynamic pathway for query- aware feature learning in language-driven action localization,
S. Yang, X. Wu, Z. Shang, and J. Luo, “Dynamic pathway for query- aware feature learning in language-driven action localization,” IEEE Transactions on Multimedia , vol. 26, pp. 7451–7461, 2024
2024
-
[40]
Hierarchical local-global transformer for temporal sentence grounding,
X. Fang, D. Liu, P. Zhou, Z. Xu, and R. Li, “Hierarchical local-global transformer for temporal sentence grounding,” IEEE Transactions on Multimedia, vol. 26, pp. 3263–3277, 2024
2024
-
[41]
Local-global video-text interactions for temporal grounding,
J. Mun, M. Cho, and B. Han, “Local-global video-text interactions for temporal grounding,” in 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2020, pp. 10 810–10 819
2020
-
[42]
Proposal-free video grounding with contextual pyramid network,
K. Li, D. Guo, and M. Wang, “Proposal-free video grounding with contextual pyramid network,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 35, no. 3, 2021, pp. 1902–1910
2021
-
[43]
Hisa: Hierarchically semantic associating for video temporal grounding,
Z. Xu, D. Chen, K. Wei, C. Deng, and H. Xue, “Hisa: Hierarchically semantic associating for video temporal grounding,” IEEE Transactions on Image Processing , vol. 31, pp. 5178–5188, 2022
2022
-
[44]
Siamese alignment network for weakly supervised video moment retrieval,
Y . Wang, M. Liu, Y . Wei, Z. Cheng, Y . Wang, and L. Nie, “Siamese alignment network for weakly supervised video moment retrieval,” IEEE Transactions on Multimedia , vol. 25, pp. 3921–3933, 2022
2022
-
[45]
Weakly supervised temporal adjacent network for language grounding,
Y . Wang, J. Deng, W. Zhou, and H. Li, “Weakly supervised temporal adjacent network for language grounding,” IEEE Transactions on Mul- timedia, vol. 24, pp. 3276–3286, 2021
2021
-
[46]
Asynce: Disentangling false-positives for weakly-supervised video grounding,
C. Da, Y . Zhang, Y . Zheng, P. Pan, Y . Xu, and C. Pan, “Asynce: Disentangling false-positives for weakly-supervised video grounding,” in Proceedings of the 29th ACM International Conference on Multimedia , 2021, pp. 1129–1137
2021
-
[47]
Weakly supervised video moment retrieval from text queries,
N. C. Mithun, S. Paul, and A. K. Roy-Chowdhury, “Weakly supervised video moment retrieval from text queries,” in 2019 IEEE/CVF Confer- ence on Computer Vision and Pattern Recognition (CVPR) , 2019, pp. 11 592–11 601
2019
-
[48]
Dual masked modeling for weakly-supervised temporal boundary discovery,
Y . Ma, Y . Liu, L. Wang, W. Kang, Y . Qiao, and Y . Wang, “Dual masked modeling for weakly-supervised temporal boundary discovery,” IEEE Transactions on Multimedia , vol. 26, pp. 5694–5704, 2024
2024
-
[49]
Weakly-supervised video moment retrieval via semantic completion network,
Z. Lin, Z. Zhao, Z. Zhang, Q. Wang, and H. Liu, “Weakly-supervised video moment retrieval via semantic completion network,” in Proceed- ings of the AAAI Conference on Artificial Intelligence , vol. 34, no. 07, 2020, pp. 11 539–11 546
2020
-
[50]
Counterfactual cross-modality reasoning for weakly supervised video moment localization,
Z. Lv, B. Su, and J.-R. Wen, “Counterfactual cross-modality reasoning for weakly supervised video moment localization,” in Proceedings of the 31st ACM International Conference on Multimedia, 2023, pp. 6539– 6547
2023
-
[51]
Weakly supervised video moment localization with contrastive negative sample mining,
M. Zheng, Y . Huang, Q. Chen, and Y . Liu, “Weakly supervised video moment localization with contrastive negative sample mining,” in Pro- ceedings of the AAAI Conference on Artificial Intelligence, vol. 36, no. 3, 2022, pp. 3517–3525
2022
-
[52]
Weakly supervised temporal sentence grounding with gaussian-based contrastive proposal learning,
M. Zheng, Y . Huang, Q. Chen, Y . Peng, and Y . Liu, “Weakly supervised temporal sentence grounding with gaussian-based contrastive proposal learning,” in 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022, pp. 15 555–15 564
2022
-
[53]
Long short-term memory,
S. Hochreiter and J. Schmidhuber, “Long short-term memory,” Neural computation, vol. 9, no. 8, pp. 1735–1780, 1997
1997
-
[54]
Distributed representations of words and phrases and their composi- tionality,
T. Mikolov, I. Sutskever, K. Chen, G. S. Corrado, and J. Dean, “Distributed representations of words and phrases and their composi- tionality,” Advances in neural information processing systems , vol. 26, 2013
2013
-
[55]
Learning spatiotemporal features with 3d convolutional networks,
D. Tran, L. Bourdev, R. Fergus, L. Torresani, and M. Paluri, “Learning spatiotemporal features with 3d convolutional networks,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2015, pp. 4489–4497
2015
-
[56]
Attention is all you need,
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” Advances in neural information processing systems , vol. 30, 2017
2017
-
[57]
Momentum contrast for unsupervised visual representation learning,
K. He, H. Fan, Y . Wu, S. Xie, and R. Girshick, “Momentum contrast for unsupervised visual representation learning,” in 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2020, pp. 9729–9738
2020
-
[58]
Facenet: A unified embed- ding for face recognition and clustering,
F. Schroff, D. Kalenichenko, and J. Philbin, “Facenet: A unified embed- ding for face recognition and clustering,” in 2015 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2015, pp. 815– 823
2015
-
[59]
Localizing moments in video with natural language,
L. A. Hendricks, O. Wang, E. Shechtman, J. Sivic, T. Darrell, and B. Russell, “Localizing moments in video with natural language,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2017, pp. 5803–5812
2017
-
[60]
Temporal localization of moments in video collections with natural language,
E. Victor, S. Mattia, S. Josef, G. Bernard, and R. Bryan, “Temporal localization of moments in video collections with natural language,” arXiv preprint arXiv:1907.12763 , 2019
1907 arXiv
-
[61]
Tvr: A large-scale dataset for video-subtitle moment retrieval,
J. Lei, L. Yu, T. L. Berg, and M. Bansal, “Tvr: A large-scale dataset for video-subtitle moment retrieval,” in European Conference on Computer Vision. Springer, 2020, pp. 447–463
2020
-
[62]
Video corpus moment retrieval with contrastive learning,
H. Zhang, A. Sun, W. Jing, G. Nan, L. Zhen, J. T. Zhou, and R. S. M. Goh, “Video corpus moment retrieval with contrastive learning,” in Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval , 2021, pp. 685– 695
2021
-
[63]
Partially relevant video retrieval,
J. Dong, X. Chen, M. Zhang, X. Yang, S. Chen, X. Li, and X. Wang, “Partially relevant video retrieval,” in Proceedings of the 30th ACM International Conference on Multimedia , 2022, pp. 246–257
2022
-
[64]
Activitynet: A large-scale video benchmark for human activity un- derstanding,
F. Caba Heilbron, V . Escorcia, B. Ghanem, and J. Carlos Niebles, “Activitynet: A large-scale video benchmark for human activity un- derstanding,” in 2015 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2015. JOURNAL OF LATEX CLASS FILES 16
2015
-
[65]
Script data for attribute-based recognition of composite activities,
M. Rohrbach, M. Regneri, M. Andriluka, S. Amin, M. Pinkal, and B. Schiele, “Script data for attribute-based recognition of composite activities,” in European Conference on Computer Vision, 2012, pp. 144– 157
2012
-
[66]
Grounding action descriptions in videos,
M. Regneri, M. Rohrbach, D. Wetzel, S. Thater, B. Schiele, and M. Pinkal, “Grounding action descriptions in videos,” Transactions of the Association for Computational Linguistics , vol. 1, pp. 25–36, 2013
2013
-
[67]
Large-scale video classification with convolutional neural networks,
A. Karpathy, G. Toderici, S. Shetty, T. Leung, R. Sukthankar, and L. Fei-Fei, “Large-scale video classification with convolutional neural networks,” in 2014 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2014, pp. 1725–1732
2014
-
[68]
Bert: Pre-training of deep bidirectional transformers for language understanding,
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “Bert: Pre-training of deep bidirectional transformers for language understanding,” in Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologi...
2019
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.