REVIEW 3 major objections 2 minor 54 references
Registering the 4D Millimeter Wave Radar Point Clouds Via Generalized Method of Moments
T0 review · 3 major / 2 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read The paper proposes a correspondence-free point cloud registration framework for 4D millimeter wave radar built on the Generalized Method of Moments, claims consistency for the estimator, and reports accuracy comparable to LiDAR-based…
desk verdict The full text under this arXiv ID is a different paper, so the only reviewable content is an abstract; the idea is plausible but there is no manuscript. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The Generalized Method of Moments (GMM) estimator: instead of matching points, the method matches empirical moments of the source cloud's positions and radial velocities against those of the target cloud under a candidate rigid transformation. The moment conditions are the central object; they are what make correspondences unnecessary and what the consistency guarantee applies to. The radial-velocity component is distinctive to 4D radar and is what the method uses to extract additional geometric information beyond raw positions.
What would settle it
Find a pair of distinct rigid transformations that produce exactly the same empirical moments on a real 4D radar cloud pair — for example, a symmetric cloud with radial velocities that are invariant under a small rotation — and show the estimator cannot choose between them. More concretely, run the method on a sequence with known ground-truth motion and check whether the recovered rotation and translation match the ground truth to within the sensor noise floor; a systematic bias would indicate misspecified moments.
Extended reading notes
Core claim
The central claim is that a Generalized Method of Moments estimator, using moment conditions built from the 3D positions and radial velocities of 4D radar points, can register two radar point clouds consistently and accurately without correspondences. The method avoids the fragile nearest-neighbor or feature-matching steps that typically fail on sparse data. The authors show consistency of the proposed estimator and support the claim with experiments on synthetic and real-world datasets, where the approach outperforms benchmark radar registration methods and reaches accuracy comparable to LiDAR-based frameworks.
Load-bearing premise
The moment equations formed from point positions and radial velocities uniquely pin down the true rigid transformation between the two clouds under realistic radar noise, so that the Generalized Method of Moments is consistent for the correct alignment rather than for a wrong one.
Editorial extensions
If this is right
- Radar-only SLAM and odometry become feasible without a separate correspondence or feature module, simplifying the perception pipeline.
- Registration no longer degrades sharply when clouds are extremely sparse, because the moment equations average over all points rather than relying on individual matches.
- The same framework may transfer to other sensors that provide velocity-like measurements, such as Doppler lidar or automotive radar with Doppler.
- Adopting 4D radar over LiDAR becomes more attractive for all-weather robot perception, since the registration front end no longer sacrifices accuracy.
Reading between the lines
- The consistency of GMM is standard asymptotic theory; the scientific risk is that the chosen moments may be weakly identifying for small, symmetric clouds. A reader should check whether the moment equations uniquely determine the pose for the sparse, noisy case.
- If the radial-velocity moments are the main source of information, then a degenerate scene with few stationary scatterers or with velocities that are invariant under rotation could make the objective flat; testing on such degenerate scenes would clarify the method's limits.
- The claim that accuracy is comparable to LiDAR-based frameworks is benchmark-dependent; a fair test would compare against LiDAR registration on the same trajectories and error metrics, not just against published numbers.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The abstract claims a correspondence-free 4D millimeter-wave radar point cloud registration framework based on the Generalized Method of Moments, including a consistency result and experiments on synthetic and real-world data that show higher accuracy than benchmarks and accuracy comparable to LiDAR-based frameworks. The body of the submission, however, is an entirely unrelated manuscript on weakly supervised multimodal temporal forgery localization (Xu, Lu, Luo). The radar-specific moment conditions, radial-velocity measurement model, estimator derivation, consistency proof, and all radar registration experiments are absent from the submitted text, so the claims in the abstract cannot be checked against any supporting material.
Significance. If the claimed result holds, a consistent, correspondence-free GMM estimator that exploits 4D radar position and radial-velocity moments would be a genuinely useful contribution to radar-based SLAM, since sparse and noisy radar clouds defeat standard correspondence-based registration. The manuscript, however, offers no way to evaluate that contribution: the central methodological content and every radar-specific experiment are missing. The idea is plausible in principle, but plausibility is not a substitute for the derivations, identification conditions, and benchmark comparisons that the abstract promises.
major comments (3)
- [Full text (entirety)] The body of arXiv:2508.02187 is not the paper described in the abstract; it is "Weakly Supervised Multimodal Temporal Forgery Localization via Multitask Learning" by Wenbo Xu, Wei Lu, and Xiangyang Luo, with its own abstract, its own introduction on Deepfake detection, and its own experimental tables. None of the radar registration content — the moment functions, the radial-velocity model, the estimator, the consistency theorem, or the radar experiments — appears anywhere in the submitted text. This makes the central claim unverifiable, and the problem is not a localized issue that a revision could repair.
- [Abstract vs. body] The statement "we show the consistency of the proposed method" is unsupported by any theorem, proof, or even a specification of the moment conditions and identification assumptions. Standard GMM consistency requires a unique population moment zero at the true parameter and an appropriate rank condition; the manuscript provides none of these, and the body contains no equations from which they could be inferred. The consistency claim is therefore not a contribution that can be assessed on the submitted material.
- [Experiments] The abstract's claims of higher accuracy and robustness than benchmarks and of LiDAR-comparable performance are not backed by any radar experiment in the body. The only experimental tables (Tables I–IV) report temporal forgery localization metrics (mAP@IoU and AR@Proposals) on LAV-DF and AV-Deepfake1M; there is no synthetic radar dataset, no real-world radar dataset, no registration baselines, no error bars, and no registration evaluation protocol. The experimental claims cannot be evaluated or reproduced from the submitted manuscript.
minor comments (2)
- [Full text (running header)] The running header of the body cites "arXiv:2508.02179v1 [cs.CV] 4 Aug 2025", which does not match the identifier arXiv:2508.02187 under review; this mismatch should be corrected in any future submission.
- [Title and authorship] The title, author list, and subject area of the body differ completely from those implied by the abstract, so a reader cannot determine from the submission which paper is actually intended for review.
Circularity Check
No circularity identified in the available text; the supplied full text is an unrelated paper, so no specific reduction of the claimed GMM result to its inputs can be exhibited.
full rationale
The abstract for arXiv:2508.02187 claims a correspondence-free GMM-based 4D radar registration framework with consistency and accuracy comparable to LiDAR. The body text supplied under this arXiv ID is actually "Weakly Supervised Multimodal Temporal Forgery Localization via Multitask Learning" (Xu, Lu, Luo), which contains no definitions of the radar moment conditions, no radial-velocity measurement model, no identification conditions, and no consistency proof for the GMM estimator. Without those equations or any self-contained derivation, there is no quotable step in which a fitted parameter, a moment condition, or a cited previous result is shown to be equivalent by construction to the claimed prediction. Standard GMM asymptotics are not circular merely because they are invoked; circularity would require the specific moment functions to be defined in terms of the target transformation or fitted to the same data used for the accuracy comparison. Since no such reduction can be quoted from the available manuscript, the appropriate finding is no significant circularity. This is an absence-of-information result about the supplied text, not a demonstration that the radar registration derivation is circular.
Assumptions & free parameters
assumptions (3)
- domain assumption The chosen moment conditions, built from point positions and radial velocities, identify the true rigid transformation between the source and target radar clouds.
- standard math Standard GMM regularity conditions hold for the radar point cloud setting, including existence, identification, and a uniform law of large numbers over the compact parameter space.
- domain assumption Both radar scans observe the same underlying scene and the radar measurement noise model is correctly specified.
Cite this review
Pith. "Pith review of Registering the 4D Millimeter Wave Radar Point Clouds Via Generalized Method of Moments." pith.science (2026). https://pith.science/paper/SZTCDQXK
@misc{pith2026250802187,
author = {Pith},
title = {Pith review of: Registering the 4D Millimeter Wave Radar Point Clouds Via Generalized Method of Moments},
year = {2026},
howpublished = {\url{https://pith.science/paper/SZTCDQXK}},
note = {Machine review of arXiv:2508.02187}
}
read the original abstract
4D millimeter wave radars (4D radars) are new emerging sensors that provide point clouds of objects with both position and radial velocity measurements. Compared to LiDARs, they are more affordable and reliable sensors for robots' perception under extreme weather conditions. On the other hand, point cloud registration is an essential perception module that provides robot's pose feedback information in applications such as Simultaneous Localization and Mapping (SLAM). Nevertheless, the 4D radar point clouds are sparse and noisy compared to those of LiDAR, and hence we shall confront great challenges in registering the radar point clouds. To address this issue, we propose a point cloud registration framework for 4D radars based on Generalized Method of Moments. The method does not require explicit point-to-point correspondences between the source and target point clouds, which is difficult to compute for sparse 4D radar point clouds. Moreover, we show the consistency of the proposed method. Experiments on both synthetic and real-world datasets show that our approach achieves higher accuracy and robustness than benchmarks, and the accuracy is even comparable to LiDAR-based frameworks.
Reference graph
Works this paper leans on
-
[1]
Dynamic difference learning with spatio–temporal correlation for deepfake video detection,
Q. Yin, W. Lu, B. Li, and J. Huang, “Dynamic difference learning with spatio–temporal correlation for deepfake video detection,” IEEE Transactions on Information Forensics and Security, vol. 18, pp. 4046–4058, 2023
work page 2023
-
[2]
Istvt: Interpretable spatial-temporal video transformer for deepfake detection,
C. Zhao, C. Wang, G. Hu, H. Chen, C. Liu, and J. Tang, “Istvt: Interpretable spatial-temporal video transformer for deepfake detection,” IEEE Transactions on Information Forensics and Security, vol. 18, pp. 1335–1348, 2023
work page 2023
-
[3]
J. Wu, W. Xu, W. Lu, X. Luo, R. Yang, and S. Guo, “Weakly-supervised audio temporal forgery localization via pro- gressive audio-language co-learning network,” arXiv preprint arXiv:2505.01880, 2025
work page Pith review arXiv 2025
-
[4]
Fine- grained multimodal deepfake classification via heterogeneous graphs,
Q. Yin, W. Lu, X. Cao, X. Luo, Y . Zhou, and J. Huang, “Fine- grained multimodal deepfake classification via heterogeneous graphs,” International Journal of Computer Vision , pp. 1–15, 2024
work page 2024
-
[5]
Improving gener- alization of deepfake detectors by imposing gradient regulariza- tion,
W. Guan, W. Wang, J. Dong, and B. Peng, “Improving gener- alization of deepfake detectors by imposing gradient regulariza- tion,” IEEE Transactions on Information Forensics and Security , vol. 19, pp. 5345–5356, 2024
work page 2024
-
[6]
Deepfake detection and localization using multi-view inconsistency measurement,
B. Zhang, Q. Yin, W. Lu, and X. Luo, “Deepfake detection and localization using multi-view inconsistency measurement,” IEEE Transactions on Dependable and Secure Computing , vol. 22, no. 2, pp. 1796–1809, 2025
work page 2025
-
[7]
Not made for each other-audio-visual dissonance-based deepfake detection and localization,
K. Chugh, P. Gupta, A. Dhall, and R. Subramanian, “Not made for each other-audio-visual dissonance-based deepfake detection and localization,” in Proceedings of the 28th ACM international conference on multimedia , 2020, pp. 439–447
work page 2020
-
[8]
Z. Cai, S. Ghosh, A. Dhall, T. Gedeon, K. Stefanov, and M. Hayat, “Glitch in the matrix: A large scale benchmark for content driven audio–visual forgery detection and localization,” Computer Vision and Image Understanding, vol. 236, p. 103818, 2023
work page 2023
Show all 54 references
-
[9]
Audio-visual tempo- ral forgery detection using embedding-level fusion and multi- dimensional contrastive loss,
M. Liu, J. Wang, X. Qian, and H. Li, “Audio-visual tempo- ral forgery detection using embedding-level fusion and multi- dimensional contrastive loss,” IEEE Transactions on Circuits and Systems for Video Technology, 2023
2023
-
[10]
Um- maformer: A universal multimodal-adaptive transformer frame- work for temporal forgery localization,
R. Zhang, H. Wang, M. Du, H. Liu, Y . Zhou, and Q. Zeng, “Um- maformer: A universal multimodal-adaptive transformer frame- work for temporal forgery localization,” in Proceedings of the 31st ACM International Conference on Multimedia , 2023, pp. 8749–8759
2023
-
[11]
Bi-stream coteaching network for weakly-supervised deepfake localization in videos,
Z. Li, Z. Teng, B. Zhang, and J. Fan, “Bi-stream coteaching network for weakly-supervised deepfake localization in videos,” Trans. Info. For. Sec., vol. 20, p. 1724–1738, Jan. 2025
2025
-
[12]
Cpl: Curriculum pseudo labeling for weakly supervised temporal forgery localization,
D. Zhang, M. Fang, Z. Lu, and H. Xie, “Cpl: Curriculum pseudo labeling for weakly supervised temporal forgery localization,” in ICASSP 2025 - 2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2025, pp. 1–5
2025
-
[13]
A multimodal deviation perceiving framework for weakly-supervised temporal forgery localization,
W. Xu, J. Wu, W. Lu, X. Luo, and Q. Wang, “A multimodal deviation perceiving framework for weakly-supervised temporal forgery localization,” arXiv preprint arXiv:2507.16596 , 2025
2025 arXiv
-
[14]
Can chatgpt detect deepfakes? a study of using multimodal large language models for media forensics,
S. Jia, R. Lyu, K. Zhao, Y . Chen, Z. Yan, Y . Ju, C. Hu, X. Li, B. Wu, and S. Lyu, “Can chatgpt detect deepfakes? a study of using multimodal large language models for media forensics,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024...
2024
-
[15]
Mmnet: multi-collaboration and multi-supervision network for sequential deepfake detection,
R. Xia, D. Liu, J. Li, L. Yuan, N. Wang, and X. Gao, “Mmnet: multi-collaboration and multi-supervision network for sequential deepfake detection,” IEEE Transactions on Information Forensics and Security, 2024
2024
-
[16]
Jointly defending deepfake manip- ulation and adversarial attack using decoy mechanism,
G.-L. Chen and C.-C. Hsu, “Jointly defending deepfake manip- ulation and adversarial attack using decoy mechanism,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 45, no. 8, pp. 9922–9931, 2023
2023
-
[17]
Deepfake de- tection based on discrepancies between faces and their context,
Y . Nirkin, L. Wolf, Y . Keller, and T. Hassner, “Deepfake de- tection based on discrepancies between faces and their context,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 44, no. 10, pp. 6111–6121, 2021
2021
-
[19]
Hearing lips and seeing voices,
H. McGurk and J. MacDonald, “Hearing lips and seeing voices,” Nature, vol. 264, no. 5588, pp. 746–748, 1976
1976
-
[20]
Joint audio-visual deepfake detection,
Y . Zhou and S.-N. Lim, “Joint audio-visual deepfake detection,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2021, pp. 14 800–14 809. 13
2021
-
[21]
Do you really mean that? content driven audio-visual deepfake dataset and multimodal method for temporal forgery localization,
Z. Cai, K. Stefanov, A. Dhall, and M. Hayat, “Do you really mean that? content driven audio-visual deepfake dataset and multimodal method for temporal forgery localization,” in 2022 International Conference on Digital Image Computing: Tech- niques and Applications (DICTA) . IE...
2022
-
[22]
Frade: Forgery- aware audio-distilled multimodal learning for deepfake detec- tion,
F. Nie, J. Ni, J. Zhang, B. Zhang, and W. Zhang, “Frade: Forgery- aware audio-distilled multimodal learning for deepfake detec- tion,” in Proceedings of the 32nd ACM International Conference on Multimedia, 2024, pp. 6297–6306
2024
-
[23]
A brief introduction to weakly supervised learning,
Z.-H. Zhou, “A brief introduction to weakly supervised learning,” National science review , vol. 5, no. 1, pp. 44–53, 2018
2018
-
[24]
Cf- deformable detr: an end-to-end alignment-free model for weakly aligned visible-infrared object detection,
H. Fu, J. Yuan, G. Zhong, X. He, J. Lin, and Z. Li, “Cf- deformable detr: an end-to-end alignment-free model for weakly aligned visible-infrared object detection,” in Proceedings of the Thirty-Third International Joint Conference on Artificial Intelli- gence, 2024, pp. 758–766
2024
-
[25]
A consistency and integration model with adaptive thresholds for weakly supervised object localization,
H. Su and M. Yang, “A consistency and integration model with adaptive thresholds for weakly supervised object localization,” in Proceedings of the Thirty-Third International Joint Conference on Artificial Intelligence , 2024, pp. 1281–1289
2024
-
[26]
Complete instances mining for weakly supervised instance segmentation,
Z. Li, Z. Zeng, Y . Liang, and J.-G. Yu, “Complete instances mining for weakly supervised instance segmentation,” arXiv preprint arXiv:2402.07633, 2024
2024 arXiv
-
[27]
Weakly super- vised object localization and detection: A survey,
D. Zhang, J. Han, G. Cheng, and M.-H. Yang, “Weakly super- vised object localization and detection: A survey,” IEEE transac- tions on pattern analysis and machine intelligence, vol. 44, no. 9, pp. 5866–5885, 2021
2021
-
[28]
Adaptive zone learning for weakly supervised object localization,
Z. Chen, S. Wang, L. Cao, Y . Shen, and R. Ji, “Adaptive zone learning for weakly supervised object localization,” IEEE Transactions on Neural Networks and Learning Systems , 2024
2024
-
[29]
Temporal action localization in the deep learning era: A survey,
B. Wang, Y . Zhao, L. Yang, T. Long, and X. Li, “Temporal action localization in the deep learning era: A survey,” IEEE Transactions on Pattern Analysis and Machine Intelligence, 2023
2023
-
[30]
Multitask learning,
R. Caruana, “Multitask learning,” Machine learning , vol. 28, no. 1, pp. 41–75, 1997
1997
-
[31]
A survey on multi-task learning,
Y . Zhang and Q. Yang, “A survey on multi-task learning,” IEEE transactions on knowledge and data engineering , vol. 34, no. 12, pp. 5586–5609, 2021
2021
-
[32]
Multi-task learning in natural language processing: An overview,
S. Chen, Y . Zhang, and Q. Yang, “Multi-task learning in natural language processing: An overview,” ACM Comput. Surv., vol. 56, no. 12, Jul. 2024
2024
-
[33]
Multi-task and multi-lingual joint learning of neural lexical utterance classification based on partially-shared model- ing,
R. Masumura, T. Tanaka, R. Higashinaka, H. Masataki, and Y . Aono, “Multi-task and multi-lingual joint learning of neural lexical utterance classification based on partially-shared model- ing,” in Proceedings of the 27th International Conference on Computational Linguistics , ...
2018
-
[34]
Person- alized microblog sentiment classification via adversarial cross- lingual multi-task learning,
W. Wang, S. Feng, W. Gao, D. Wang, and Y . Zhang, “Person- alized microblog sentiment classification via adversarial cross- lingual multi-task learning,” in Proceedings of the 2018 Con- ference on Empirical Methods in Natural Language Processing , E. Riloff, D. Chiang, J. Hock...
2018
-
[35]
Worse wer, but better bleu? leveraging word embedding as intermedi- ate in multitask end-to-end speech translation,
S.-P. Chuang, T.-W. Sung, A. H. Liu, and H.-y. Lee, “Worse wer, but better bleu? leveraging word embedding as intermedi- ate in multitask end-to-end speech translation,” arXiv preprint arXiv:2005.10678, 2020
2005 arXiv
-
[36]
Learning modality- specific representations with self-supervised multi-task learning for multimodal sentiment analysis,
W. Yu, H. Xu, Z. Yuan, and J. Wu, “Learning modality- specific representations with self-supervised multi-task learning for multimodal sentiment analysis,” in Proceedings of the AAAI conference on artificial intelligence , vol. 35, no. 12, 2021, pp. 10 790–10 797
2021
-
[37]
Mod- eling task relationships in multi-task learning with multi-gate mixture-of-experts,
J. Ma, Z. Zhao, X. Yi, J. Chen, L. Hong, and E. H. Chi, “Mod- eling task relationships in multi-task learning with multi-gate mixture-of-experts,” in Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery & Data Min- ing, ser. KDD ’18. New York, NY ...
2018
-
[38]
Decouple and resolve: Transformer-based models for online anomaly detection from weakly labeled videos,
T. Liu, C. Zhang, K.-M. Lam, and J. Kong, “Decouple and resolve: Transformer-based models for online anomaly detection from weakly labeled videos,” IEEE Transactions on Information Forensics and Security, vol. 18, pp. 15–28, 2023
2023
-
[39]
Bi-stream coteach- ing network for weakly-supervised deepfake localization in videos,
Z. Li, Z. Teng, B. Zhang, and J. Fan, “Bi-stream coteach- ing network for weakly-supervised deepfake localization in videos,” IEEE Transactions on Information Forensics and Se- curity, vol. 20, pp. 1724–1738, 2025
2025
-
[40]
Multi-task gans for view-specific feature learning in gait recognition,
Y . He, J. Zhang, H. Shan, and L. Wang, “Multi-task gans for view-specific feature learning in gait recognition,” IEEE Transactions on Information Forensics and Security , vol. 14, no. 1, pp. 102–113, 2019
2019
-
[41]
Avoid-df: Audio-visual joint learning for detecting deepfake,
W. Yang, X. Zhou, Z. Chen, B. Guo, Z. Ba, Z. Xia, X. Cao, and K. Ren, “Avoid-df: Audio-visual joint learning for detecting deepfake,” IEEE Transactions on Information Forensics and Security, vol. 18, pp. 2015–2029, 2023
2015
-
[42]
Multi-modal deep- fake detection via multi-task audio-visual prompt learning,
H. Miao, Y . Guo, Z. Liu, and Y . Wang, “Multi-modal deep- fake detection via multi-task audio-visual prompt learning,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 39, no. 1, 2025, pp. 612–621
2025
-
[43]
Feature selection: A data perspective,
J. Li, K. Cheng, S. Wang, F. Morstatter, R. P. Trevino, J. Tang, and H. Liu, “Feature selection: A data perspective,” ACM com- puting surveys (CSUR) , vol. 50, no. 6, pp. 1–45, 2017
2017
-
[44]
Feature selection techniques for machine learning: a survey of more than two decades of research,
D. Theng and K. K. Bhoyar, “Feature selection techniques for machine learning: a survey of more than two decades of research,” Knowledge and Information Systems , vol. 66, no. 3, pp. 1575–1637, 2024
2024
-
[45]
Av-deepfake1m: A large-scale llm-driven audio-visual deepfake dataset,
Z. Cai, S. Ghosh, A. P. Adatia, M. Hayat, A. Dhall, T. Gedeon, and K. Stefanov, “Av-deepfake1m: A large-scale llm-driven audio-visual deepfake dataset,” in Proceedings of the 32nd ACM International Conference on Multimedia , 2024, pp. 7414–7423
2024
-
[46]
Deepfake detection via inter-frame inconsistency recomposition and en- hancement,
C. Zhu, B. Zhang, Q. Yin, C. Yin, and W. Lu, “Deepfake detection via inter-frame inconsistency recomposition and en- hancement,” Pattern Recognition, vol. 147, p. 110077, 2024
2024
-
[47]
Detection of deepfake videos using long-distance attention,
W. Lu, L. Liu, B. Zhang, J. Luo, X. Zhao, Y . Zhou, and J. Huang, “Detection of deepfake videos using long-distance attention,” IEEE Transactions on Neural Networks and Learning Systems , vol. 35, no. 7, pp. 9366–9379, 2024
2024
-
[48]
Temporal segment networks for action recognition in videos,
L. Wang, Y . Xiong, Z. Wang, Y . Qiao, D. Lin, X. Tang, and L. Van Gool, “Temporal segment networks for action recognition in videos,” IEEE transactions on pattern analysis and machine intelligence, vol. 41, no. 11, pp. 2740–2755, 2018
2018
-
[49]
wav2vec 2.0: A framework for self-supervised learning of speech repre- sentations,
A. Baevski, Y . Zhou, A. Mohamed, and M. Auli, “wav2vec 2.0: A framework for self-supervised learning of speech repre- sentations,” Advances in neural information processing systems , vol. 33, pp. 12 449–12 460, 2020
2020
-
[50]
Actionformer: Localizing mo- ments of actions with transformers,
C.-L. Zhang, J. Wu, and Y . Li, “Actionformer: Localizing mo- ments of actions with transformers,” in European Conference on Computer Vision. Springer, 2022, pp. 492–510
2022
-
[51]
Tridet: Temporal action detection with relative boundary modeling,
D. Shi, Y . Zhong, Q. Cao, L. Ma, J. Li, and D. Tao, “Tridet: Temporal action detection with relative boundary modeling,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 18 857–18 866
2023
-
[52]
Mfms: Learning modality-fused and modality-specific features for deepfake detection and localization tasks,
Y . Zhang, C. Miao, M. Luo, J. Li, W. Deng, W. Yao, Z. Li, B. Hu, W. Feng, T. Gong et al. , “Mfms: Learning modality-fused and modality-specific features for deepfake detection and localization tasks,” in Proceedings of the 32nd ACM International Conference on Multimedia, 2024...
2024
-
[53]
Cola: Weakly- supervised temporal action localization with snippet contrastive learning,
C. Zhang, M. Cao, D. Yang, J. Chen, and Y . Zou, “Cola: Weakly- supervised temporal action localization with snippet contrastive learning,” in Proceedings of the IEEE/CVF conference on com- puter vision and pattern recognition , 2021, pp. 16 010–16 019
2021
-
[54]
Full-stage pseudo label quality enhancement for weakly-supervised temporal action localization,
Q. Feng, W. Li, T. Lin, and X. Chen, “Full-stage pseudo label quality enhancement for weakly-supervised temporal action localization,” arXiv preprint arXiv:2407.08971 , 2024
2024 arXiv
-
[55]
Multilevel semantic and adap- tive actionness learning for weakly supervised temporal action localization,
Z. Li, Z. Wang, and C. Dong, “Multilevel semantic and adap- tive actionness learning for weakly supervised temporal action localization,” Neural Networks, vol. 182, p. 106905, 2025
2025
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.