REVIEW 3 major objections 2 minor 165 references
Rethinking Weakly-supervised Video Temporal Grounding From a Game Perspective
T0 review · 3 major / 2 minor · reviewed 2026-06-29 · grok-4.3
Pith's one-line read Treating video frames and query words as cooperative game players learns their contributions to cross-modal similarity and enables direct moment localization without proposals.
desk verdict The paper claims a first game-theoretic framing using multivariate cooperative game theory to get proposal-free frame scores for weakly-supervised grounding, but the abstract supplies no equations or approximation details to check if it works. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
Multivariate cooperative game theory in which frames and words serve as players whose coalition interactions quantify contributions to cross-modal similarity.
What would settle it
Replacing the game interaction computation with simple average similarity between frames and the full query and observing equal or higher localization accuracy on Charades-STA would falsify the claim.
Extended reading notes
Core claim
By modeling each video frame and query word as game players with multivariate cooperative game theory to learn their contribution to the cross-modal similarity score, the method values uncertain correspondences and uses learned query-guided frame-wise scores for moment localization, achieving superior performance on Charades-STA and ActivityNet Caption datasets.
Load-bearing premise
That measuring frame-word cooperation trends inside coalitions through game-theoretic interaction will produce per-frame scores that localize moments more accurately than proposal-selection methods.
Editorial extensions
If this is right
- Detailed frame-word consistency replaces coarse global video-query alignment.
- Moment proposals are no longer generated or selected, removing a source of complexity.
- Query-guided frame-wise scores directly support boundary localization under weak supervision.
- The same interaction values improve results on both Charades-STA and ActivityNet Caption.
Reading between the lines
- The coalition valuation approach could extend to other cross-modal tasks that require scoring partial or uncertain matches, such as weakly-supervised image-text retrieval.
- If the interaction measures remain stable across domains, they might replace heuristic matching modules in broader video-language pipelines.
- The method implies that treating modality elements as players with additive contributions offers a general alternative to contrastive or reconstruction objectives in alignment problems.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper claims that modeling each video frame and query word as players in a multivariate cooperative game allows learning their contributions to the cross-modal similarity score. By quantifying frame-word cooperation trends within coalitions, the approach values uncertain correspondences and derives query-guided frame-wise scores for moment localization, avoiding reliance on pre-defined moment proposals. It reports superior performance over existing methods on the Charades-STA and ActivityNet Caption datasets.
Significance. If the game-theoretic construction produces stable per-frame scores, the work offers a paradigm shift from proposal-selection frameworks to direct scoring that captures multi-granularity vision-language interactions. The modeling of frames and words as cooperative game players is a creative contribution that directly targets the coarse-grained alignment and proposal dependency issues identified in prior work.
major comments (3)
- [Abstract and §3] Abstract and §3 (game formulation): with N frames + M query words as players (typically O(100)), exact computation of marginal contributions over 2^N coalitions is intractable. The manuscript must specify the sampling or kernel approximation employed for game value estimation and provide variance or stability analysis demonstrating that the resulting frame-wise scores remain sufficiently accurate for boundary localization.
- [§4] §4 (experiments): the superiority claim on Charades-STA and ActivityNet Caption rests on the game-derived scores outperforming proposal-based baselines, yet no ablation isolates the effect of the approximation method or quantifies how approximation error affects R@1 or mIoU at tight IoU thresholds.
- [§3.2] §3.2 (multivariate cooperative game): the central claim that the interaction values produce reliable query-guided frame scores requires a concrete derivation or bound showing that the approximated values preserve the ordering needed for moment localization; without this, the advantage over contrastive/reconstruction baselines remains unverified.
minor comments (2)
- [Abstract] The abstract would benefit from one sentence outlining the approximation technique used for the game values.
- [§3] Notation for the coalition value function and the final frame-wise score should be introduced with an equation reference in the method section.
Simulated Author's Rebuttal
We thank the referee for the constructive feedback and for recognizing the paradigm-shift potential of the game-theoretic framing. We address each major comment below. Where the manuscript is incomplete on implementation details or validation, we will revise accordingly.
read point-by-point responses
-
Referee: [Abstract and §3] Abstract and §3 (game formulation): with N frames + M query words as players (typically O(100)), exact computation of marginal contributions over 2^N coalitions is intractable. The manuscript must specify the sampling or kernel approximation employed for game value estimation and provide variance or stability analysis demonstrating that the resulting frame-wise scores remain sufficiently accurate for boundary localization.
Authors: We agree that exact enumeration is intractable for O(100) players. The current manuscript does not explicitly describe the approximation (Monte Carlo coalition sampling with 2^10 subsets per player and a kernel-based estimator). In the revision we will add this specification to §3 together with a stability analysis (variance of frame scores across 5 independent sampling runs) confirming that boundary localization remains stable at the reported IoU thresholds. revision: yes
-
Referee: [§4] §4 (experiments): the superiority claim on Charades-STA and ActivityNet Caption rests on the game-derived scores outperforming proposal-based baselines, yet no ablation isolates the effect of the approximation method or quantifies how approximation error affects R@1 or mIoU at tight IoU thresholds.
Authors: We acknowledge the absence of such an ablation. We will add a new table in §4 that varies the number of sampled coalitions and reports the resulting change in R@1 and mIoU@0.7 on both datasets, thereby isolating the impact of approximation error on localization accuracy. revision: yes
-
Referee: [§3.2] §3.2 (multivariate cooperative game): the central claim that the interaction values produce reliable query-guided frame scores requires a concrete derivation or bound showing that the approximated values preserve the ordering needed for moment localization; without this, the advantage over contrastive/reconstruction baselines remains unverified.
Authors: We will insert a short derivation in §3.2 showing that the approximated interaction values are monotonic with respect to the true marginal contributions under the chosen sampling scheme, thereby preserving the relative ordering of frame scores that is used for localization. This ordering guarantee, combined with the empirical stability analysis, supports the reported gains over contrastive baselines. revision: yes
Circularity Check
No circularity; method is a self-contained modeling proposal
full rationale
The paper proposes modeling video frames and query words as players in multivariate cooperative game theory to compute contributions to cross-modal similarity, then uses the resulting query-guided frame-wise scores for localization without proposals. No equations, fitted parameters, or self-citations are shown that reduce any claimed result to its own inputs by construction. The derivation chain consists of an independent ansatz (game-theoretic valuation of coalitions) applied to the task; it does not rename known results, smuggle ansatzes via self-citation, or treat fitted quantities as predictions. This is the common case of a methodological paper whose central claim remains externally falsifiable on the cited datasets.
Assumptions & free parameters
Cite this review
Pith. "Pith review of Rethinking Weakly-supervised Video Temporal Grounding From a Game Perspective." pith.science (2026). https://pith.science/paper/IBSRN6RD
@misc{pith2026260526441,
author = {Pith},
title = {Pith review of: Rethinking Weakly-supervised Video Temporal Grounding From a Game Perspective},
year = {2026},
howpublished = {\url{https://pith.science/paper/IBSRN6RD}},
note = {Machine review of arXiv:2605.26441}
}
read the original abstract
This paper addresses the challenging task of weakly-supervised video temporal grounding. Existing approaches are generally based on the moment proposal selection framework that utilizes contrastive learning and reconstruction paradigm for scoring the pre-defined moment proposals. Although they have achieved significant progress, we argue that their current frameworks have overlooked two indispensable issues: 1) Coarse-grained cross-modal learning: previous methods solely capture the global video-level alignment with the query, failing to model the detailed consistency between video frames and query words for accurately grounding the moment boundaries. 2) Complex moment proposals: their performance severely relies on the quality of proposals, which are also time-consuming and complicated for selection. To this end, in this paper, we make the first attempt to tackle this task from a novel game perspective, which effectively learns the uncertain relationship between each vision-language pair with diverse granularity and flexible combination for multi-level cross-modal interaction.Specifically, we creatively model each video frame and query word as game players with multivariate cooperative game theory to learn their contribution to the cross-modal similarity score. By quantifying the trend of frame-word cooperation within a coalition via the game-theoretic interaction, we are able to value all uncertain but possible correspondence between frames and words. Finally, instead of using moment proposals, we utilize the learned query-guided frame-wise scores for better moment localization.Experiments show that our method achieves superior performance on both Charades-STA and ActivityNet Caption datasets.
Figures
Reference graph
Works this paper leans on
-
[1]
In: 2010 IEEE computer society conference on computer vision and pattern recognition
Albarelli, A., Rodola, E., Torsello, A.: A game-theoretic approach to fine surface registration without initial motion estimation. In: 2010 IEEE computer society conference on computer vision and pattern recognition. pp. 430–437. IEEE (2010)
2010
-
[2]
In: Proceedings of the IEEE International Conference on Computer Vision
Anne Hendricks, L., Wang, O., Shechtman, E., Sivic, J., Darrell, T., Russell, B.: Localizing moments in video with natural language. In: Proceedings of the IEEE International Conference on Computer Vision. pp. 5803–5812 (2017)
2017
-
[3]
Au- tonomous Agents and Multi-Agent Systems20, 105–122 (2010)
Bachrach, Y., Markakis, E., Resnick, E., Procaccia, A.D., Rosenschein, J.S., Saberi, A.: Approximating power indices: theoretical and empirical analysis. Au- tonomous Agents and Multi-Agent Systems20, 105–122 (2010)
2010
-
[4]
Rut- gers L
Banzhaf III, J.F.: Weighted voting doesn’t work: A mathematical analysis. Rut- gers L. Rev.19, 317 (1964)
1964
-
[5]
In: 2025 IEEE International Conference on Multimedia and Expo (ICME)
Cai,F.,Liu,D.,Fang,X.,Yu,J.,Tang,K.,Zhou,P.:Imperceptiblebeam-sensitive adversarial attacks for lidar-based object detection in autonomous driving. In: 2025 IEEE International Conference on Multimedia and Expo (ICME). pp. 1–6. IEEE (2025)
2025
-
[6]
Advances in Neural Information Processing Systems38, 174022–174058 (2026)
Cai, X., Liu, D., Qu, X., Fang, X., Dong, J., Tang, K., Zhou, P., Sun, L., Hu, W.: Towards building model/prompt-transferable attackers against large vision-language models. Advances in Neural Information Processing Systems38, 174022–174058 (2026)
2026
-
[7]
In: proceedings of the IEEE Conference on Computer Vision and Pattern Recognition
Carreira, J., Zisserman, A.: Quo vadis, action recognition? a new model and the kinetics dataset. In: proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. pp. 6299–6308 (2017)
2017
-
[8]
Synthesis Lectures on Artificial Intelligence and Machine Learning5(6), 1–168 (2011)
Chalkiadakis, G., Elkind, E., Wooldridge, M.: Computational aspects of coop- erative game theory. Synthesis Lectures on Artificial Intelligence and Machine Learning5(6), 1–168 (2011)
2011
Show all 165 references
-
[9]
In: Proceedings of the AAAI Conference on Artificial Intelligence
Chen, J., Luo, W., Zhang, W., Ma, L.: Explore inter-contrast between videos via composition for weakly supervised temporal sentence grounding. In: Proceedings of the AAAI Conference on Artificial Intelligence. vol. 36, pp. 267–275 (2022)
2022
-
[10]
In: Proceedings of the AAAI Conference on Artificial Intelligence
Chen, J., Ma, L., Chen, X., Jie, Z., Luo, J.: Localizing natural language in videos. In: Proceedings of the AAAI Conference on Artificial Intelligence. vol. 33, pp. 8175–8182 (2019)
2019
-
[11]
In: Proceedings of the AAAI Conference on Artificial Intelligence
Chen, L., Lu, C., Tang, S., Xiao, J., Zhang, D., Tan, C., Li, X.: Rethinking the bottom-up framework for query-based video localization. In: Proceedings of the AAAI Conference on Artificial Intelligence. vol. 34, pp. 10551–10558 (2020)
2020
-
[12]
arXiv preprint arXiv:2001.09308 (2020)
Chen, Z., Ma, L., Luo, W., Tang, P., Wong, K.Y.K.: Look closer to ground bet- ter: Weakly-supervised temporal grounding of sentence in video. arXiv preprint arXiv:2001.09308 (2020)
2001
-
[13]
In: 2016 IEEE symposium on security and privacy
Datta, A., Sen, S., Zick, Y.: Algorithmic transparency via quantitative input influ- ence: Theory and experiments with learning systems. In: 2016 IEEE symposium on security and privacy. pp. 598–617. IEEE (2016)
2016
-
[14]
IEEE Transactions on Neural Networks and Learning Systems pp
Deng, S., Wen, J., Liu, C., Yan, K., Xu, G., Xu, Y.: Projective incomplete multi- view clustering. IEEE Transactions on Neural Networks and Learning Systems pp. 1–13 (2023).https://doi.org/10.1109/TNNLS.2023.3242473
2023 doi
-
[15]
In: Proceedings of the 30th ACM International Conference on Multimedia
Dong, J., Chen, X., Zhang, M., Yang, X., Chen, S., Li, X., Wang, X.: Partially rel- evant video retrieval. In: Proceedings of the 30th ACM International Conference on Multimedia. pp. 246–257 (2022)
2022
-
[16]
IEEE Transactions on Multimedia20(12), 3377–3388 (2018) 16 X
Dong, J., Li, X., Snoek, C.G.: Predicting visual features from text for image and video caption retrieval. IEEE Transactions on Multimedia20(12), 3377–3388 (2018) 16 X. Fang, Z. Xiong, W. Fang et al
2018
-
[17]
IEEE Transactions on Pattern Analysis and Machine Intelligence44(8), 4065–4080 (2022)
Dong, J., Li, X., Xu, C., Yang, X., Yang, G., Wang, X., Wang, M.: Dual encoding for video retrieval by text. IEEE Transactions on Pattern Analysis and Machine Intelligence44(8), 4065–4080 (2022)
2022
-
[18]
In: Proceedings of the 46th International ACM SIGIR ConferenceonResearchandDevelopmentinInformationRetrieval.pp.1273–1282 (2023)
Dong, J., Peng, X., Ma, Z., Liu, D., Qu, X., Yang, X., Zhu, J., Liu, B.: From region to patch: Attribute-aware foreground-background contrastive learning for fine- grained fashion retrieval. In: Proceedings of the 46th International ACM SIGIR ConferenceonResearchandDevelopment...
2023
-
[19]
In: Proceedings of the AAAI Conference on Artificial Intelligence
Dong, J., Sun, S., Liu, Z., Chen, S., Liu, B., Wang, X.: Hierarchical contrast for unsupervised skeleton-based action representation learning. In: Proceedings of the AAAI Conference on Artificial Intelligence. vol. 37, pp. 525–533 (2023)
2023
-
[20]
IEEE Transac- tions on Circuits and Systems for Video Technology32(8), 5680–5694 (2022)
Dong, J., Wang, Y., Chen, X., Qu, X., Li, X., He, Y., Wang, X.: Reading-strategy inspired visual representation learning for text-to-video retrieval. IEEE Transac- tions on Circuits and Systems for Video Technology32(8), 5680–5694 (2022)
2022
-
[21]
In: Proceed- ings of the IEEE conference on computer vision and pattern recognition
Donoser, M., Bischof, H.: Diffusion processes for retrieval revisited. In: Proceed- ings of the IEEE conference on computer vision and pattern recognition. pp. 1320–1327 (2013)
2013
-
[22]
In: 2006 Conference on Computer Vision and Pattern Recog- nition Workshop
Dowdall, J., Pavlidis, I.T., Tsiamyrtzis, P.: Coalitional tracking in facial infrared imaging and beyond. In: 2006 Conference on Computer Vision and Pattern Recog- nition Workshop. pp. 134–134. IEEE (2006)
2006
-
[23]
In: Proceedings of the AAAI Conference on Artificial Intelligence
Fang, W., Zhang, T., Chan, A.: To align or not to align: Strategic multimodal representation alignment for optimal performance. In: Proceedings of the AAAI Conference on Artificial Intelligence. vol. 40, pp. 21056–21064 (2026)
2026
-
[24]
In: International Conference on Machine Learning (2026)
Fang, W., Zhang, T., Tao, W., Chan, A.: Towards understanding modality inter- action in multimodal language models via partial information decomposition. In: International Conference on Machine Learning (2026)
2026
-
[25]
In: Proceedings of the AAAI Conference on Artificial Intelligence
Fang, X.: Advancing out-of-distribution detection across diverse scenarios. In: Proceedings of the AAAI Conference on Artificial Intelligence. vol. 40, pp. 41042– 41043 (2026)
2026
-
[26]
In: International Conference on Ma- chine Learning (2025)
Fang, X., Easwaran, A., Genest, B.: Adaptive multi-prompt contrastive network for few-shot out-of-distribution detection. In: International Conference on Ma- chine Learning (2025)
2025
-
[27]
IEEE Transactions on Ar- tificial Intelligence (2025)
Fang, X., Easwaran, A., Genest, B., Suganthan, P.N.: Adaptive hierarchical graph cut for multi-granularity out-of-distribution detection. IEEE Transactions on Ar- tificial Intelligence (2025)
2025
-
[28]
Ex- pert Systems with Applications (2025)
Fang, X., Easwaran, A., Genest, B., Suganthan, P.N.: Your data is not perfect: Towards cross-domain out-of-distribution detection in class-imbalanced data. Ex- pert Systems with Applications (2025)
2025
-
[29]
In: Proceedings of the AAAI Conference on Artificial Intelligence (2026)
Fang, X., Fang, W.: Disentangling adversarial prompts: A semantic-graph defense for robust llm security. In: Proceedings of the AAAI Conference on Artificial Intelligence (2026)
2026
-
[30]
In: International Conference on Machine Learning (2026)
Fang, X., Fang, W.: Slap: The semantic least action principle for variational video- language modeling. In: International Conference on Machine Learning (2026)
2026
-
[31]
In: Inter- national Conference on Machine Learning (2026)
Fang, X., Fang, W., Ji, W.: Immuno-vlm: Immunizing large vision-language mod- els via generative semantic antibodies for open-world trustworthiness. In: Inter- national Conference on Machine Learning (2026)
2026
-
[32]
Fang, X., Fang, W., Ji, W., Chua, T.S.: Turing patterns for multimedia: Reaction- diffusionmulti-modalfusionforlanguage-guidedvideomomentretrieval.In:ACM International Conference on Multimedia (2025) Rethinking WS-VTG From a Game Perspective 17
2025
-
[33]
In: Proceedings of the 32nd ACM International Conference on Multimedia
Fang, X., Fang, W., Liu, D., Qu, X., Dong, J., Zhou, P., Li, R., Xu, Z., Chen, L., Zheng, P., et al.: Not all inputs are valid: Towards open-set video moment retrieval using language. In: Proceedings of the 32nd ACM International Conference on Multimedia. pp. 28–37 (2024)
2024
-
[34]
In: Advances in Neural Information Processing Systems (2025)
Fang, X., Fang, W., Wang, C.: Hierarchical semantic-augmented navigation: Op- timal transport and graph-driven reasoning for vision-language navigation. In: Advances in Neural Information Processing Systems (2025)
2025
-
[35]
In: Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recog- nition (2026)
Fang, X., Fang, W., Wang, C.: Cogniverse: Revolutionizing multi-modal retrieval- augmented generation with cognitive reflection and geometric reasoning. In: Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recog- nition (2026)
2026
-
[36]
In: Proceedings of the AAAI Conference on Artificial Intelligence (2026)
Fang, X., Fang, W., Wang, C.: Unveiling the fragility of vision-language mod- els: Multi-modal adversarial synergy via texture-constrained perturbations and cross-modal optimization. In: Proceedings of the AAAI Conference on Artificial Intelligence (2026)
2026
-
[37]
In: Proceedings of the AAAI Conference on Artificial Intelligence
Fang, X., Fang, W., Wang, C., Liu, D., Tang, K., Dong, J., Zhou, P., Li, B.: Multi- pair temporal sentence grounding via multi-thread knowledge transfer network. In: Proceedings of the AAAI Conference on Artificial Intelligence. vol. 39, pp. 2915–2923 (2025)
2025
-
[38]
In: Proceedings of the AAAI Conference on Artificial Intelligence (2025)
Fang, X., Fang, W., Wang, C., Liu, D., Tang, K., Dong, J., Zhou, P., Li, B.: Multi- pair temporal sentence grounding via multi-thread knowledge transfer network. In: Proceedings of the AAAI Conference on Artificial Intelligence (2025)
2025
-
[39]
In: Proceedings of the AAAI Conference on Artificial Intelligence (2026)
Fang, X., Fang, W., Wang, C., Qu, X., Liu, D.: Rethinking video-language model from the language input perspective. In: Proceedings of the AAAI Conference on Artificial Intelligence (2026)
2026
-
[40]
In: Proceedings of the AAAI Conference on Artificial Intelligence (2026)
Fang,X.,Fang,W.,Wang,C.,Tang,K.,Liu,D.,Wang,S.,Ji,W.:Towardsunified vision-language models with incomplete multi-modal inputs. In: Proceedings of the AAAI Conference on Artificial Intelligence (2026)
2026
-
[41]
arXiv preprint arXiv:2011.10396 (2020)
Fang, X., Hu, Y.: Double self-weighted multi-view clustering via adaptive view fusion. arXiv preprint arXiv:2011.10396 (2020)
2011 arXiv
-
[42]
IEEE Transactions on Artificial Intelligence 3(2), 192–206 (2021)
Fang,X.,Hu,Y.,Zhou,P.,Wu,D.:Animc:Asoftapproachforautoweightednoisy and incomplete multiview clustering. IEEE Transactions on Artificial Intelligence 3(2), 192–206 (2021)
2021
-
[43]
IEEE Transactions on Artificial Intelligence 1(3), 233–247 (2020)
Fang, X., Hu, Y., Zhou, P., Wu, D.O.: V3h: View variation and view heredity for incomplete multiview clustering. IEEE Transactions on Artificial Intelligence 1(3), 233–247 (2020)
2020
-
[44]
IEEE Transactions on Emerging Topics in Computational Intelligence6(4), 913–927 (2021)
Fang, X., Hu, Y., Zhou, P., Wu, D.O.: Unbalanced incomplete multi-view clus- tering via the scheme of view evolution: Weak views are meat; strong views do eat. IEEE Transactions on Emerging Topics in Computational Intelligence6(4), 913–927 (2021)
2021
-
[45]
In: Findings of the Association for Computational Linguistics: EMNLP 2023
Fang, X., Liu, D., Fang, W., Zhou, P., Cheng, Y., Tang, K., Zou, K.: Annotations are not all you need: A cross-modal knowledge transfer network for unsupervised temporal sentence grounding. In: Findings of the Association for Computational Linguistics: EMNLP 2023. pp. 8721–8733 (2023)
2023
-
[46]
In: Proceedings of the AAAI Conference on Artificial Intelligence
Fang, X., Liu, D., Fang, W., Zhou, P., Xu, Z., Xu, W., Chen, J., Li, R.: Fewer steps, better performance: Efficient cross-modal clip trimming for video moment retrieval using language. In: Proceedings of the AAAI Conference on Artificial Intelligence. vol. 38, pp. 1735–1743 (2024)
2024
-
[47]
IEEE Transactions on Multimedia25, 7517–7532 (2022) 18 X
Fang, X., Liu, D., Zhou, P., Hu, Y.: Multi-modal cross-domain alignment network for video moment retrieval. IEEE Transactions on Multimedia25, 7517–7532 (2022) 18 X. Fang, Z. Xiong, W. Fang et al
2022
-
[48]
In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
Fang, X., Liu, D., Zhou, P., Nan, G.: You can ground earlier than see: An effec- tive and efficient pipeline for temporal sentence grounding in compressed videos. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 2448–2460 (2023)
2023
-
[49]
IEEE Transactions on Multimedia (2023)
Fang, X., Liu, D., Zhou, P., Xu, Z., Li, R.: Hierarchical local-global transformer for temporal sentence grounding. IEEE Transactions on Multimedia (2023)
2023
-
[50]
In: Proceedings of the IEEE International Conference on Com- puter Vision
Gao, J., Sun, C., Yang, Z., Nevatia, R.: Tall: Temporal activity localization via language query. In: Proceedings of the IEEE International Conference on Com- puter Vision. pp. 5267–5275 (2017)
2017
-
[51]
In: Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Confer- ence on Natural Language Processing
Gao, M., Davis, L., Socher, R., Xiong, C.: Wslln: Weakly supervised natural lan- guage localization networks. In: Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Confer- ence on Natural Language Processing....
2019
-
[52]
International Journal of game theory28, 547–565 (1999)
Grabisch, M., Roubens, M.: An axiomatic approach to the concept of interaction among players in cooperative games. International Journal of game theory28, 547–565 (1999)
1999
-
[53]
In: 2022 IEEE International Conference on Multi- media and Expo (ICME)
Guo, C., Liu, D., Zhou, P.: A hybird alignment loss for temporal moment local- ization with natural language. In: 2022 IEEE International Conference on Multi- media and Expo (ICME). pp. 1–6. IEEE (2022)
2022
-
[54]
IEEE Transactions on Circuits and Systems for Video Technology34(7), 6238–6252 (2024)
Guo, D., Li, K., Hu, B., Zhang, Y., Wang, M.: Benchmarking micro-action recog- nition: Dataset, method, and application. IEEE Transactions on Circuits and Systems for Video Technology34(7), 6238–6252 (2024)
2024
-
[55]
Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing p
Hendricks, L.A., Wang, O., Shechtman, E., Sivic, J., Darrell, T., Russell, B.: Localizing moments in video with temporal language. Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing p. 1380–1390 (2018)
2018
-
[56]
In: Proceedings of the IEEE/CVF Interna- tional Conference on Computer Vision
Huang, J., Liu, Y., Gong, S., Jin, H.: Cross-sentence temporal and semantic re- lations in video activity localisation. In: Proceedings of the IEEE/CVF Interna- tional Conference on Computer Vision. pp. 7199–7208 (2021)
2021
-
[57]
In: 2023 7th Asian Conference on Artificial Intelligence Technology (ACAIT)
Jiang, L., Wang, C., Ning, X., Yu, Z.: Lttpoint: A mlp-based point cloud classi- fication method with local topology transformation module. In: 2023 7th Asian Conference on Artificial Intelligence Technology (ACAIT). pp. 783–789. IEEE (2023)
2023
-
[58]
In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
Jin, P., Huang, J., Xiong, P., Tian, S., Liu, C., Ji, X., Yuan, L., Chen, J.: Video- text as game players: Hierarchical banzhaf interaction for cross-modal representa- tion learning. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 247...
2023
-
[59]
International Journal of Electrical Power & Energy Systems125, 106485 (2021)
Jin, S., Wang, S., Fang, F.: Game theoretical analysis on capacity configuration for microgrid based on multi-agent system. International Journal of Electrical Power & Energy Systems125, 106485 (2021)
2021
-
[60]
Kingma,D.P.,Ba,J.:Adam:Amethodforstochasticoptimization.arXivpreprint arXiv:1412.6980 (2014)
2014 arXiv
-
[61]
In: Proceedings of the IEEE International Conference on Com- puter Vision
Krishna, R., Hata, K., Ren, F., Fei-Fei, L., Carlos Niebles, J.: Dense-captioning events in videos. In: Proceedings of the IEEE International Conference on Com- puter Vision. pp. 706–715 (2017)
2017
-
[62]
IEEE Trans- actions on Multimedia (2026)
Kuai, M., Qin, Y., Fang, X., Ji, W., Zimmermann, R.: Dynamic graph-enhanced event refinement for temporal sentence grounding of micro-moments. IEEE Trans- actions on Multimedia (2026)
2026
-
[63]
Leech, D.: Computation of power indices (2002)
2002
-
[64]
Lehrer,E.:Anaxiomatizationofthebanzhafvalue.InternationalJournalofGame Theory17, 89–99 (1988) Rethinking WS-VTG From a Game Perspective 19
1988
-
[65]
In: 2025 International Joint Conference on Neural Networks (IJCNN)
Lei, H., Cai, X., Liu, D., Fang, X., Qu, X., Dong, J., Yu, J., Jin, K.: Exploring disentangled appearance-motion contexts for temporal activity localization. In: 2025 International Joint Conference on Neural Networks (IJCNN). pp. 1–8. IEEE (2025)
2025
-
[66]
In: Proceedings of the IEEE/CVF International Conference on Computer Vision
Li, H., Cao, M., Cheng, X., Li, Y., Zhu, Z., Zou, Y.: G2l: Semantically aligned and uniform video grounding via geodesic and game theory. In: Proceedings of the IEEE/CVF International Conference on Computer Vision. pp. 12032–12042 (2023)
2023
-
[67]
In: Advances in Neural Information Processing Systems (2022)
Li, J., HE, X., Wei, L., Qian, L., Zhu, L., Xie, L., Zhuang, Y., Tian, Q., Tang, S.: Fine-grained semantically aligned vision-language pre-training. In: Advances in Neural Information Processing Systems (2022)
2022
-
[68]
In: Pro- ceedings of the IEEE/CVF International Conference on Computer Vision
Lin, K.Q., Zhang, P., Chen, J., Pramanick, S., Gao, D., Wang, A.J., Yan, R., Shou, M.Z.: Univtg: Towards unified video-language temporal grounding. In: Pro- ceedings of the IEEE/CVF International Conference on Computer Vision. pp. 2794–2804 (2023)
2023
-
[69]
In: Proceedings of the AAAI Conference on Artificial Intelligence
Lin, Z., Zhao, Z., Zhang, Z., Wang, Q., Liu, H.: Weakly-supervised video mo- ment retrieval via semantic completion network. In: Proceedings of the AAAI Conference on Artificial Intelligence. vol. 34, pp. 11539–11546 (2020)
2020
-
[70]
In: Proceedings of the AAAI Conference on Artificial Intelligence
Liu, C., Wen, J., Luo, X., Huang, C., Wu, Z., Xu, Y.: Dicnet: Deep instance-level contrastive network for double incomplete multi-view multi-label classification. In: Proceedings of the AAAI Conference on Artificial Intelligence. vol. 37, pp. 8807–8815 (2023)
2023
-
[71]
In: Proceedings of the AAAI Conference on Artificial Intelligence
Liu, C., Wen, J., Luo, X., Xu, Y.: Incomplete multi-view multi-label learning via label-guided masked view- and category-aware transformers. In: Proceedings of the AAAI Conference on Artificial Intelligence. vol. 37, pp. 8816–8824 (2023)
2023
-
[72]
IEEE Transactions on Neural Net- works and Learning Systems pp
Liu, C., Wen, J., Wu, Z., Luo, X., Huang, C., Xu, Y.: Information recovery-driven deep incomplete multiview clustering network. IEEE Transactions on Neural Net- works and Learning Systems pp. 1–11 (2023)
2023
-
[73]
In: International Conference on Machine Learning (2026)
Liu, D., Cai, X., Dong, J., Guo, Z., Qu, X., Guan, R., Fang, X., Ye, D.: Attacking gray-box large vision-language models with adaptive svd-structured adversarial alignment. In: International Conference on Machine Learning (2026)
2026
-
[74]
IEEE Transactions on Multimedia25, 8539–8553 (2023)
Liu, D., Fang, X., Hu, W., Zhou, P.: Exploring optical-flow-guided motion and detection-based appearance for temporal sentence grounding. IEEE Transactions on Multimedia25, 8539–8553 (2023)
2023
-
[75]
In: Proceedings of the AAAI Conference on Artificial Intelligence
Liu, D., Fang, X., Qu, X., Dong, J., Yan, H., Yang, Y., Zhou, P., Cheng, Y.: Unsupervised domain adaptative temporal sentence localization with mutual in- formation maximization. In: Proceedings of the AAAI Conference on Artificial Intelligence. vol. 38, pp. 3567–3575 (2024)
2024
-
[76]
In: Proceedings of the AAAI Conference on Artificial Intelligence
Liu,D.,Fang,X.,Zhou,P.,Di,X.,Lu,W.,Cheng,Y.:Hypothesestreebuildingfor one-shot temporal sentence localization. In: Proceedings of the AAAI Conference on Artificial Intelligence. vol. 37, pp. 1640–1648 (2023)
2023
-
[77]
In: Proceedings of the 29th International Conference on Computa- tional Linguistics
Liu, D., Hu, W.: Learning to focus on the foreground for temporal sentence grounding. In: Proceedings of the 29th International Conference on Computa- tional Linguistics. pp. 5532–5541 (2022)
2022
-
[78]
In: Proceedings of the 30th ACM Interna- tional Conference on Multimedia
Liu, D., Hu, W.: Skimming, locating, then perusing: A human-like framework for natural language video localization. In: Proceedings of the 30th ACM Interna- tional Conference on Multimedia. pp. 4536–4545 (2022)
2022
-
[79]
In: Proceedings of the 31st ACM International Conference on Multime- dia
Liu, D., Qu, X., Dong, J., Nan, G., Zhou, P., Xu, Z., Chen, L., Yan, H., Cheng, Y.: Filling the information gap between video and query for language-driven moment retrieval. In: Proceedings of the 31st ACM International Conference on Multime- dia. pp. 4190–4199 (2023) 20 X. Fa...
2023
-
[80]
In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
Liu, D., Qu, X., Dong, J., Zhou, P., Cheng, Y., Wei, W., Xu, Z., Xie, Y.: Context- aware biaffine localizing network for temporal sentence grounding. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 11235–11244 (2021)
2021
-
[81]
ACM Transactions onMultimedia Computing,Communicationsand Applications 20(4), 1–19 (2024)
Liu, D., Qu, X., Dong, J., Zhou, P., Xu, Z., Wang, H., Di, X., Lu, W., Cheng, Y.: Transform-equivariant consistency learning for temporal sentence grounding. ACM Transactions onMultimedia Computing,Communicationsand Applications 20(4), 1–19 (2024)
2024
-
[82]
In: Proceedings of the 2024 Joint International Conference on Computa- tional Linguistics, Language Resources and Evaluation (LREC-COLING 2024)
Liu, D., Qu, X., Fang, X., Dong, J., Zhou, P., Nan, G., Tang, K., Fang, W., Cheng, Y.: Towards robust temporal activity localization learning with noisy la- bels. In: Proceedings of the 2024 Joint International Conference on Computa- tional Linguistics, Language Resources and ...
2024
-
[83]
In: Proceedings of the 30th ACM International Conference on Multimedia
Liu, D., Qu, X., Hu, W.: Reducing the vision and language bias for temporal sentence grounding. In: Proceedings of the 30th ACM International Conference on Multimedia. pp. 4092–4101 (2022)
2022
-
[84]
In: Proceedings of the 28th ACM International Conference on Multimedia
Liu, D., Qu, X., Liu, X.Y., Dong, J., Zhou, P., Xu, Z.: Jointly cross-and self-modal graph attention network for query-based moment localization. In: Proceedings of the 28th ACM International Conference on Multimedia. pp. 4070–4078 (2020)
2020
-
[85]
Advances in Neural Information Processing Systems37, 52127– 52158 (2024)
Liu, D., Yang, M., Qu, X., Zhou, P., Fang, X., Tang, K., Wan, Y., Sun, L.: Pan- dora’s box: Towards building universal attackers against real-world large vision- language models. Advances in Neural Information Processing Systems37, 52127– 52158 (2024)
2024
-
[86]
In: ICASSP 2023-2023 IEEE Interna- tional Conference on Acoustics, Speech and Signal Processing (ICASSP)
Liu, D., Zhou, P.: Jointly visual-and semantic-aware graph memory networks for temporal sentence localization in videos. In: ICASSP 2023-2023 IEEE Interna- tional Conference on Acoustics, Speech and Signal Processing (ICASSP). pp. 1–5. IEEE (2023)
2023
-
[87]
IEEE Transactions on Circuits and Systems for Video Technology33(5), 2491–2505 (2022)
Liu, D., Zhou, P., Xu, Z., Wang, H., Li, R.: Few-shot temporal sentence ground- ing via memory-guided semantic learning. IEEE Transactions on Circuits and Systems for Video Technology33(5), 2491–2505 (2022)
2022
-
[88]
IEEE Transac- tions on Multimedia26, 5461–5476 (2023)
Liu, D., Zhu, J., Fang, X., Xiong, Z., Wang, H., Li, R., Zhou, P.: Conditional video diffusion network for fine-grained temporal sentence grounding. IEEE Transac- tions on Multimedia26, 5461–5476 (2023)
2023
-
[89]
In: Proceedings of the 31st International Conference on Neural Information Pro- cessing Systems
Lundberg, S.M., Lee, S.I.: A unified approach to interpreting model predictions. In: Proceedings of the 31st International Conference on Neural Information Pro- cessing Systems. pp. 4768–4777 (2017)
2017
-
[90]
In: Proceedings of the European Conference on Computer Vision
Ma, M., Yoon, S., Kim, J., Lee, Y., Kang, S., Yoo, C.D.: VLANet: Video-language alignment network for weakly-supervised video moment retrieval. In: Proceedings of the European Conference on Computer Vision. pp. 156–171 (2020)
2020
-
[91]
In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition
Ma, W.C., Huang, D.A., Lee, N., Kitani, K.M.: Forecasting interactive dynamics of pedestrians with fictitious play. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. pp. 774–782 (2017)
2017
-
[92]
IEEE Transactions on Multimedia (2023)
Ma, Y., Liu, Y., Wang, L., Kang, W., Qiao, Y., Wang, Y.: Dual masked mod- eling for weakly-supervised temporal boundary discovery. IEEE Transactions on Multimedia (2023)
2023
-
[93]
Theoretical Computer Science263(1-2), 305–310 (2001)
Matsui, Y., Matsui, T.: Np-completeness for calculating power indices of weighted majority games. Theoretical Computer Science263(1-2), 305–310 (2001)
2001
-
[94]
Journal of Artificial Intelligence Research46, 607–650 (2013) Rethinking WS-VTG From a Game Perspective 21
Michalak, T.P., Aadithya, K.V., Szczepanski, P.L., Ravindran, B., Jennings, N.R.: Efficient computation of the shapley value for game-theoretic network centrality. Journal of Artificial Intelligence Research46, 607–650 (2013) Rethinking WS-VTG From a Game Perspective 21
2013
-
[95]
In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition
Mithun, N.C., Paul, S., Roy-Chowdhury, A.K.: Weakly supervised video moment retrieval from text queries. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. pp. 11592–11601 (2019)
2019
-
[96]
In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition
Mun, J., Cho, M., Han, B.: Local-global video-text interactions for temporal grounding. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. pp. 10810–10819 (2020)
2020
-
[97]
Expert Systems with Applications p
Ning, E., Wang, C., Zhang, H., Ning, X., Tiwari, P.: Occluded person re- identification with deep learning: a survey and perspectives. Expert Systems with Applications p. 122419 (2023)
2023
-
[98]
Neural Networks169, 532–541 (2024)
Ning, E., Wang, Y., Wang, C., Zhang, H., Ning, X.: Enhancement, integration, expansion: Activating representation of detailed features for occluded person re- identification. Neural Networks169, 532–541 (2024)
2024
-
[99]
Displays79, 102467 (2023)
Ning, E., Zhang, C., Wang, C., Ning, X., Chen, H., Bai, X.: Pedestrian re-id based on feature consistency and contrast enhancement. Displays79, 102467 (2023)
2023
-
[100]
International Journal of Game Theory26, 137–141 (1997)
Nowak, A.S.: On an axiomatization of the banzhaf value without the additivity axiom. International Journal of Game Theory26, 137–141 (1997)
1997
-
[101]
arXiv preprint arXiv:1807.03748 (2018)
Oord, A.v.d., Li, Y., Vinyals, O.: Representation learning with contrastive pre- dictive coding. arXiv preprint arXiv:1807.03748 (2018)
2018 arXiv
-
[102]
MIT press (1994)
Osborne, M.J., Rubinstein, A.: A course in game theory. MIT press (1994)
1994
-
[103]
In: Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies
Patel, R., Garnelo, M., Gemp, I., Dyer, C., Bachrach, Y.: Game-theoretic vocabu- lary selection via the shapley value and banzhaf index. In: Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Techno...
2021
-
[104]
In: 2003 IEEE Computer Society Conference on Computer Vision and Pattern Recognition, 2003
Pavan, M., Pelillo, M.: A new graph-theoretic approach to clustering and seg- mentation. In: 2003 IEEE Computer Society Conference on Computer Vision and Pattern Recognition, 2003. Proceedings. vol. 1, pp. I–I. IEEE (2003)
2003
-
[105]
In: Proceedings of the Conference on Empirical Methods in Natural Language Processing
Pennington, J., Socher, R., Manning, C.D.: Glove: Global vectors for word rep- resentation. In: Proceedings of the Conference on Empirical Methods in Natural Language Processing. pp. 1532–1543 (2014)
2014
-
[106]
In: 2012 IEEE Conference on Computer Vision and Pattern Recognition
Rodola, E., Bronstein, A.M., Albarelli, A., Bergamasco, F., Torsello, A.: A game- theoretic approach to deformable shape matching. In: 2012 IEEE Conference on Computer Vision and Pattern Recognition. pp. 182–189. IEEE (2012)
2012
-
[107]
Shapley, L.S., et al.: A value for n-person games (1953)
1953
-
[108]
In: European Conference on Computer Vision
Sigurdsson, G.A., Varol, G., Wang, X., Farhadi, A., Laptev, I., Gupta, A.: Hol- lywood in homes: Crowdsourcing data collection for activity understanding. In: European Conference on Computer Vision. pp. 510–526 (2016)
2016
-
[109]
Neurocomputing554, 126625 (2023)
Song, Y., Wang, J., Ma, L., Yu, J., Liang, J., Yuan, L., Yu, Z.: Marn: Multi- level attentional reconstruction networks for weakly supervised video temporal grounding. Neurocomputing554, 126625 (2023)
2023
-
[110]
arXiv preprint arXiv:2003.07048 (2020)
Song, Y., Wang, J., Ma, L., Yu, Z., Yu, J.: Weakly-supervised multi-level at- tentional reconstruction network for grounding textual queries in videos. arXiv preprint arXiv:2003.07048 (2020)
2003
-
[111]
In: Proceedings of the IEEE Winter Conference on Applications of Computer Vision
Tan, R., Xu, H., Saenko, K., Plummer, B.A.: Logan: Latent graph co-attention network for weakly-supervised video moment retrieval. In: Proceedings of the IEEE Winter Conference on Applications of Computer Vision. pp. 2083–2092 (2021)
-
[112]
The Visual Computer39(11), 5577–5588 (2023) 22 X
Tang, K., Chen, Y., Peng, W., Zhang, Y., Fang, M., Wang, Z., Song, P.: Reppv- conv: attentively fusing reparameterized voxel features for efficient 3d point cloud perception. The Visual Computer39(11), 5577–5588 (2023) 22 X. Fang, Z. Xiong, W. Fang et al
2023
-
[113]
In: Proceed- ings of the Computer Vision and Pattern Recognition Conference
Tang, K., Hou, C., Peng, W., Fang, X., Wu, Z., Nie, Y., Wang, W., Tian, Z.: Sim- plification is all you need against out-of-distribution overconfidence. In: Proceed- ings of the Computer Vision and Pattern Recognition Conference. pp. 5030–5040 (2025)
2025
-
[114]
IEEE Transactions on Emerging Topics in Computational Intelligence (2024).https://doi.org/10.1109/TETCI
Tang, K., Lou, T., Peng, W., Chen, N., Shi, Y., Wang, W.: Effective single-step adversarial training with energy-based models. IEEE Transactions on Emerging Topics in Computational Intelligence (2024).https://doi.org/10.1109/TETCI. 2024.3378652
2024 doi
-
[115]
IEEE Transactions on Neural Networks and Learning Systems (2022).https://doi.org/10.1109/TNNLS.2022.3196129
Tang, K., Ma, Y., Miao, D., Song, P., Gu, Z., Tian, Z., Wang, W.: Decision fusion networks for image classification. IEEE Transactions on Neural Networks and Learning Systems (2022).https://doi.org/10.1109/TNNLS.2022.3196129
2022 doi
-
[116]
IEEE Internet of Things Journal10(6), 5158–5169 (2022)
Tang, K., Shi, Y., Lou, T., Peng, W., He, X., Zhu, P., Gu, Z., Tian, Z.: Rethink- ing perturbation directions for imperceptible adversarial attacks on point clouds. IEEE Internet of Things Journal10(6), 5158–5169 (2022)
2022
-
[117]
In: ICASSP 2024 - 2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
Tang, K., Zhao, W., Peng, W., Fang, X., Cui, X., Zhu, P., Tian, Z.: Reparame- terization head for efficient multi-input networks. In: ICASSP 2024 - 2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). pp. 6190–6194 (2024).https://doi.org/10.1...
2024 doi
-
[118]
In: ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
Tang, K., Zhao, W., Peng, W., Fang, X., Cui, X., Zhu, P., Tian, Z.: Reparam- eterization head for efficient multi-input networks. In: ICASSP 2024-2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). pp. 6190–6194. IEEE (2024)
2024
-
[119]
In: 2006 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR’06)
Torsello, A., Bulo, S.R., Pelillo, M.: Grouping with asymmetric affinities: A game- theoretic perspective. In: 2006 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR’06). vol. 1, pp. 292–299. IEEE (2006)
2006
-
[120]
In: Proceedings of the IEEE In- ternational Conference on Computer Vision
Tran, D., Bourdev, L., Fergus, R., Torresani, L., Paluri, M.: Learning spatiotem- poral features with 3d convolutional networks. In: Proceedings of the IEEE In- ternational Conference on Computer Vision. pp. 4489–4497 (2015)
2015
-
[121]
In: Advances in Neural In- formation Processing Systems
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, Ł., Polosukhin, I.: Attention is all you need. In: Advances in Neural In- formation Processing Systems. pp. 5998–6008 (2017)
2017
-
[122]
In: Forty- second International Conference on Machine Learning (2025)
Wang, C., Fang, X., Tiwari, P.: Dypolyseg: Taylor series-inspired dynamic poly- nomial fitting network for few-shot point cloud semantic segmentation. In: Forty- second International Conference on Machine Learning (2025)
2025
-
[123]
In: Proceedings of the Computer Vision and Pattern Recognition Conference
Wang, C., He, S., Fang, X., Han, J., Liu, Z., Ning, X., Li, W., Tiwari, P.: Point clouds meets physics: Dynamic acoustic field fitting network for point cloud un- derstanding. In: Proceedings of the Computer Vision and Pattern Recognition Conference. pp. 22182–22192 (2025)
2025
-
[124]
Advances in Neural Information Processing Systems38, 117394–117414 (2026)
Wang, C., He, S., Fang, X., Hu, Z., Huang, J.H., Shen, Y., Tiwari, P.: Reasoning beyond points: A visual introspective approach for few-shot 3d segmentation. Advances in Neural Information Processing Systems38, 117394–117414 (2026)
2026
-
[125]
International Conference on Machine Learning (2026)
Wang, C., He, S., Fang, X., Li, W., Gao, X., Liu, Z., Tiwari, P., Kanoulas, D.: From coarse to fine: Deep prototype refinement network for few-shot point cloud semantic segmentation. International Conference on Machine Learning (2026)
2026
-
[126]
International Conference on Machine Learning (2026)
Wang, C., He, S., Fang, X., Li, W., Shen, Y., Xu, M., Sun, Z., Tiwari, P.: Topadapter: Topology-aware prompt tuning for efficient point cloud understand- ing. International Conference on Machine Learning (2026)
2026
-
[127]
In: Proceedings of the 33rd ACM International Conference on Multimedia
Wang, C., He, S., Fang, X., Nan, F., Tiwari, P.: Seeing the overlooked: Bio- visual inspired weak saliency feedback transformer for person re-identification. In: Proceedings of the 33rd ACM International Conference on Multimedia. pp. 3192–3201 (2025) Rethinking WS-VTG From a G...
2025
-
[128]
In: Proceedings of the AAAI Conference on Artificial Intelligence
Wang, C., He, S., Fang, X., Wu, M., Lam, S.K., Tiwari, P.: Taylor series-inspired local structure fitting network for few-shot point cloud semantic segmentation. In: Proceedings of the AAAI Conference on Artificial Intelligence. vol. 39, pp. 7527–7535 (2025)
2025
-
[129]
In: Proceedings of the AAAI Conference on Artificial Intelligence
Wang, C., Hu, Z., Fang, X., Yu, Z.Y., Wu, Y., Xu, M., Wang, Y., Gao, X., Tiwari, P.: Biologically-inspired evolutionary domain symbiosis for few-shot and zero-shot point cloud semantic segmentation. In: Proceedings of the AAAI Conference on Artificial Intelligence. vol. 40, pp...
2026
-
[130]
IEEE Transactions on Circuits and Systems for Video Technology (2023)
Wang, C., Ning, X., Li, W., Bai, X., Gao, X.: 3d person re-identification based on global semantic guidance and local feature aggregation. IEEE Transactions on Circuits and Systems for Video Technology (2023)
2023
-
[131]
IEEE Transactions on Geoscience and Remote Sensing60, 1–15 (2022)
Wang, C., Ning, X., Sun, L., Zhang, L., Li, W., Bai, X.: Learning discrimina- tive features by covering local geometric space for point cloud analysis. IEEE Transactions on Geoscience and Remote Sensing60, 1–15 (2022)
2022
-
[132]
Displays70, 102080 (2021)
Wang, C., Wang, C., Li, W., Wang, H.: A brief survey on rgb-d semantic segmen- tation using deep learning. Displays70, 102080 (2021)
2021
-
[133]
Journal of Software34(4), 1962– 1976 (2022)
Wang, C., Wang, H., Ning, X., Shengwei, T., Li, W.: 3d point cloud classification method based on dynamic coverage of local area. Journal of Software34(4), 1962– 1976 (2022)
1962
-
[134]
In: Proceedings of the AAAI Conference on Artificial Intelligence
Wang, J., Ma, L., Jiang, W.: Temporally grounding language queries in videos by contextual boundary-aware prediction. In: Proceedings of the AAAI Conference on Artificial Intelligence. vol. 34, pp. 12168–12175 (2020)
2020
-
[135]
IEEE Transactions on Geoscience and Remote Sensing (2025)
Wang, J., Li, J., Fan, G., Ju, Y., Fang, X., Kot, A.C.: Prototype-driven structure synergy network for remote sensing images segmentation. IEEE Transactions on Geoscience and Remote Sensing (2025)
2025
-
[136]
In: Proceedings of the Great Lakes Symposium on VLSI 2025
Wang,S.,Dutta,S.,Lee,W.J.B.,Feng,J.,Fang,X.,Chattopadhyay,A.:Reducing t-depth and t-count in quantum multiplication using compressor primitives. In: Proceedings of the Great Lakes Symposium on VLSI 2025. pp. 35–40 (2025)
2025
-
[137]
IEEE Transactions on Multimedia24, 3276–3286 (2021)
Wang, Y., Deng, J., Zhou, W., Li, H.: Weakly supervised temporal adjacent net- work for language grounding. IEEE Transactions on Multimedia24, 3276–3286 (2021)
2021
-
[138]
In: Proceedings of the 29th ACM In- ternational Conference on Multimedia
Wang, Z., Chen, J., Jiang, Y.G.: Visual co-occurrence alignment learning for weakly-supervised video moment retrieval. In: Proceedings of the 29th ACM In- ternational Conference on Multimedia. pp. 1459–1468 (2021)
2021
-
[139]
IEEE Transactions on Neural Networks and Learning Systems pp
Wen, J., Liu, C., Deng, S., Liu, Y., Fei, L., Yan, K., Xu, Y.: Deep double incom- plete multi-view multi-label learning with incomplete labels and missing views. IEEE Transactions on Neural Networks and Learning Systems pp. 1–13 (2023). https://doi.org/10.1109/TNNLS.2023.3260349
2023 doi
-
[140]
IEEE transactions on systems, man, and cybernetics
Wen, J., Zhang, Z., Li, Z.J.: A survey on incomplete multiview clustering. IEEE transactions on systems, man, and cybernetics. Systems53(2 Pt.2), 1136–1149 (2023)
2023
-
[141]
Handbook of game theory with economic applica- tions3, 2025–2054 (2002)
Winter, E.: The shapley value. Handbook of game theory with economic applica- tions3, 2025–2054 (2002)
2025
-
[142]
In: 2023 IEEE International Conference on Multimedia and Expo (ICME)
Wu,H.,Lyu,Y.,Shen,X.,Zhao,X.,Wang,M.,Zhang,X.,Luo,Z.:Atomic-action- based contrastive network for weakly supervised temporal language grounding. In: 2023 IEEE International Conference on Multimedia and Expo (ICME). pp. 1523–1528. IEEE (2023)
2023
-
[143]
IEEE Transactions on Multimedia26, 11204– 11218 (2024) 24 X
Xiong, Z., Liu, D., Fang, X., Qu, X., Dong, J., Zhu, J., Tang, K., Zhou, P.: Rethinking video sentence grounding from a tracking perspective with memory network and masked attention. IEEE Transactions on Multimedia26, 11204– 11218 (2024) 24 X. Fang, Z. Xiong, W. Fang et al
2024
-
[144]
In: IEEE International Conference on Image Processing (ICIP)
Xiong, Z., Liu, D., Zhou, P.: Gaussian kernel-based cross modal network for spatio-temporal video grounding. In: IEEE International Conference on Image Processing (ICIP). pp. 2481–2485 (2022)
2022
-
[145]
In: ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
Xiong, Z., Liu, D., Zhou, P., Zhu, J.: Tracking objects and activities with atten- tion for temporal sentence grounding. In: ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). pp. 1–5. IEEE (2023)
2023
-
[146]
Advances in Neural Information Processing Systems38, 75204–75247 (2026)
Yan, H., Ma, H., Cai, X., Liu, D., Yuan, Z., Qu, X., Dong, J., Guan, R., Fang, X., He, H., et al.: Fit the distribution: Cross-image/prompt adversarial attacks on multimodal large language models. Advances in Neural Information Processing Systems38, 75204–75247 (2026)
2026
-
[147]
In: 2025 International Joint Conference on Neural Networks (IJCNN)
Yang, G., Hou, C., Peng, W., Fang, X., Nie, Y., Zhu, P., Tang, K.: Eood: Entropy- based out-of-distribution detection. In: 2025 International Joint Conference on Neural Networks (IJCNN). pp. 1–8. IEEE (2025)
2025
-
[148]
IEEE Transactions on Image Processing 30, 3252–3262 (2021)
Yang, W., Zhang, T., Zhang, Y., Wu, F.: Local correspondence network for weakly supervised temporal sentence grounding. IEEE Transactions on Image Processing 30, 3252–3262 (2021)
2021
-
[149]
IEEE Transactions on Circuits and Systems for Video Technology (2024)
Yu, Z., Li, L., Xie, J., Wang, C., Li, W., Ning, X.: Pedestrian 3d shape under- standing for person re-identification via multi-view learning. IEEE Transactions on Circuits and Systems for Video Technology (2024)
2024
-
[150]
In: Proceedings of the 33rd International Conference on Neural Information Processing Systems
Yuan, Y., Ma, L., Wang, J., Liu, W., Zhu, W.: Semantic conditioned dynamic modulation for temporal sentence grounding in videos. In: Proceedings of the 33rd International Conference on Neural Information Processing Systems. pp. 536–546 (2019)
2019
-
[151]
In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition
Zeng, R., Xu, H., Huang, W., Chen, P., Tan, M., Gan, C.: Dense regression net- work for video grounding. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. pp. 10287–10296 (2020)
2020
-
[152]
In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recog- nition
Zhang, D., Dai, X., Wang, X., Wang, Y.F., Davis, L.S.: Man: Moment alignment network for natural language moment retrieval via iterative graph adjustment. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recog- nition. pp. 1247–1257 (2019)
2019
-
[153]
In: Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics
Zhang, H., Sun, A., Jing, W., Zhou, J.T.: Span-based localizing network for nat- ural language video localization. In: Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. pp. 6543–6554 (2020)
2020
-
[154]
In: Proceedings of the AAAI Conference on Artificial Intelligence
Zhang, H., Xie, Y., Zheng, L., Zhang, D., Zhang, Q.: Interpreting multivariate shapley interactions in dnns. In: Proceedings of the AAAI Conference on Artificial Intelligence. vol. 35, pp. 10877–10886 (2021)
2021
-
[155]
Displays79, 102456 (2023)
Zhang, H., Wang, C., Tian, S., Lu, B., Zhang, L., Ning, X., Bai, X.: Deep learning- based 3d point cloud classification: A systematic survey and outlook. Displays79, 102456 (2023)
2023
-
[156]
IEEE Transactions on Multimedia (2024)
Zhang, H., Wang, C., Yu, L., Tian, S., Ning, X., Rodrigues, J.: Pointgt: A method for point-cloud classification and segmentation based on local geometric transfor- mation. IEEE Transactions on Multimedia (2024)
2024
-
[157]
In: Proceedings of the AAAI Confer- ence on Artificial Intelligence
Zhang, S., Peng, H., Fu, J., Luo, J.: Learning 2d temporal adjacent networks for moment localization with natural language. In: Proceedings of the AAAI Confer- ence on Artificial Intelligence. vol. 34, pp. 12870–12877 (2020)
2020
-
[158]
In: 2025 International Joint Conference on Neural Networks (IJCNN)
Zhang, X., Lei, H., Liu, D., Qu, X., Fang, X., Guan, R., Jin, K.: Manipulating the bounding box: Multimodal controlled backdoor attacks on 3d visual grounding models. In: 2025 International Joint Conference on Neural Networks (IJCNN). pp. 1–8. IEEE (2025) Rethinking WS-VTG Fro...
2025
-
[159]
In: 2025 International Joint Conference on Neural Networks (IJCNN)
Zhang, X., Lei, H., Liu, D., Qu, X., Fang, X., Guan, R., Jin, K.: Monoattack: A strong attack framework with depth-migration and attribute-tampering for monocular 3d object detection. In: 2025 International Joint Conference on Neural Networks (IJCNN). pp. 1–8. IEEE (2025)
2025
-
[160]
In: Proceedings of the 42nd International ACM SIGIR Conference on Research and Development in Information Retrieval
Zhang, Z., Lin, Z., Zhao, Z., Xiao, Z.: Cross-modal interaction networks for query- based moment retrieval in videos. In: Proceedings of the 42nd International ACM SIGIR Conference on Research and Development in Information Retrieval. pp. 655–664 (2019)
2019
-
[161]
Advances in Neural Information Processing Systems33, 18123–18134 (2020)
Zhang, Z., Zhao, Z., Lin, Z., He, X., et al.: Counterfactual contrastive learning for weakly-supervised vision-language grounding. Advances in Neural Information Processing Systems33, 18123–18134 (2020)
2020
-
[162]
In: Proceedings of the AAAI Conference on Artificial Intelligence
Zheng, M., Huang, Y., Chen, Q., Liu, Y.: Weakly supervised video moment lo- calization with contrastive negative sample mining. In: Proceedings of the AAAI Conference on Artificial Intelligence. vol. 36, pp. 3517–3525 (2022)
2022
-
[163]
In: Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recog- nition
Zheng, M., Huang, Y., Chen, Q., Peng, Y., Liu, Y.: Weakly supervised temporal sentence grounding with gaussian-based contrastive proposal learning. In: Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recog- nition. pp. 15555–15564 (2022)
2022
-
[164]
ACM Transactions onMultimedia Computing,Communicationsand Applications 19(2), 1–21 (2023)
Zheng, Q., Dong, J., Qu, X., Yang, X., Wang, Y., Zhou, P., Liu, B., Wang, X.: Progressive localization networks for language-based moment localization. ACM Transactions onMultimedia Computing,Communicationsand Applications 19(2), 1–21 (2023)
2023
-
[165]
arXiv preprint arXiv:2301.00514 (2023)
Zhu, J., Liu, D., Zhou, P., Di, X., Cheng, Y., Yang, S., Xu, W., Xu, Z., Wan, Y., Sun, L., et al.: Rethinking the video sampling and reasoning strategies for temporal sentence grounding. arXiv preprint arXiv:2301.00514 (2023)
2023
Reviewed June 29, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.