Pith. sign in

REVIEW 3 major objections 2 minor 54 references

Registering the 4D Millimeter Wave Radar Point Clouds Via Generalized Method of Moments

T0 review · 3 major / 2 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read The paper proposes a correspondence-free point cloud registration framework for 4D millimeter wave radar built on the Generalized Method of Moments, claims consistency for the estimator, and reports accuracy comparable to LiDAR-based…

desk verdict The full text under this arXiv ID is a different paper, so the only reviewable content is an abstract; the idea is plausible but there is no manuscript. read the letter →

arxiv 2508.02187 v3 pith:SZTCDQXK submitted 2025-08-04 cs.RO cs.CV

classification cs.ROcs.CV
keywords 4DmillimeterwaveradarpointcloudregistrationGeneralizedMethodofMomentscorrespondence-freeradialvelocityconsistencySLAMsparseclouds
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper proposes a point cloud registration framework for 4D millimeter wave radar, built on the Generalized Method of Moments, that aligns a source cloud to a target cloud without ever computing explicit point-to-point correspondences. The authors argue that correspondences are the bottleneck for sparse, noisy radar clouds, and that moments formed from point positions and radial velocities carry enough information to recover the rigid transformation. They claim consistency of the estimator and report that on synthetic and real-world benchmarks it is more accurate and robust than existing radar registration methods, with accuracy close to LiDAR-based registration. If correct, this would make radar a viable perception sensor for pose estimation and SLAM in bad weather and other conditions where LiDAR is unreliable.

What carries the argument

The Generalized Method of Moments (GMM) estimator: instead of matching points, the method matches empirical moments of the source cloud's positions and radial velocities against those of the target cloud under a candidate rigid transformation. The moment conditions are the central object; they are what make correspondences unnecessary and what the consistency guarantee applies to. The radial-velocity component is distinctive to 4D radar and is what the method uses to extract additional geometric information beyond raw positions.

What would settle it

Find a pair of distinct rigid transformations that produce exactly the same empirical moments on a real 4D radar cloud pair — for example, a symmetric cloud with radial velocities that are invariant under a small rotation — and show the estimator cannot choose between them. More concretely, run the method on a sequence with known ground-truth motion and check whether the recovered rotation and translation match the ground truth to within the sensor noise floor; a systematic bias would indicate misspecified moments.

Watch

Extended reading notes

Core claim

The central claim is that a Generalized Method of Moments estimator, using moment conditions built from the 3D positions and radial velocities of 4D radar points, can register two radar point clouds consistently and accurately without correspondences. The method avoids the fragile nearest-neighbor or feature-matching steps that typically fail on sparse data. The authors show consistency of the proposed estimator and support the claim with experiments on synthetic and real-world datasets, where the approach outperforms benchmark radar registration methods and reaches accuracy comparable to LiDAR-based frameworks.

Load-bearing premise

The moment equations formed from point positions and radial velocities uniquely pin down the true rigid transformation between the two clouds under realistic radar noise, so that the Generalized Method of Moments is consistent for the correct alignment rather than for a wrong one.

Editorial extensions

If this is right

  • Radar-only SLAM and odometry become feasible without a separate correspondence or feature module, simplifying the perception pipeline.
  • Registration no longer degrades sharply when clouds are extremely sparse, because the moment equations average over all points rather than relying on individual matches.
  • The same framework may transfer to other sensors that provide velocity-like measurements, such as Doppler lidar or automotive radar with Doppler.
  • Adopting 4D radar over LiDAR becomes more attractive for all-weather robot perception, since the registration front end no longer sacrifices accuracy.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The consistency of GMM is standard asymptotic theory; the scientific risk is that the chosen moments may be weakly identifying for small, symmetric clouds. A reader should check whether the moment equations uniquely determine the pose for the sparse, noisy case.
  • If the radial-velocity moments are the main source of information, then a degenerate scene with few stationary scatterers or with velocities that are invariant under rotation could make the objective flat; testing on such degenerate scenes would clarify the method's limits.
  • The claim that accuracy is comparable to LiDAR-based frameworks is benchmark-dependent; a fair test would compare against LiDAR registration on the same trajectories and error metrics, not just against published numbers.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 2 minor

Summary. The abstract claims a correspondence-free 4D millimeter-wave radar point cloud registration framework based on the Generalized Method of Moments, including a consistency result and experiments on synthetic and real-world data that show higher accuracy than benchmarks and accuracy comparable to LiDAR-based frameworks. The body of the submission, however, is an entirely unrelated manuscript on weakly supervised multimodal temporal forgery localization (Xu, Lu, Luo). The radar-specific moment conditions, radial-velocity measurement model, estimator derivation, consistency proof, and all radar registration experiments are absent from the submitted text, so the claims in the abstract cannot be checked against any supporting material.

Significance. If the claimed result holds, a consistent, correspondence-free GMM estimator that exploits 4D radar position and radial-velocity moments would be a genuinely useful contribution to radar-based SLAM, since sparse and noisy radar clouds defeat standard correspondence-based registration. The manuscript, however, offers no way to evaluate that contribution: the central methodological content and every radar-specific experiment are missing. The idea is plausible in principle, but plausibility is not a substitute for the derivations, identification conditions, and benchmark comparisons that the abstract promises.

major comments (3)
  1. [Full text (entirety)] The body of arXiv:2508.02187 is not the paper described in the abstract; it is "Weakly Supervised Multimodal Temporal Forgery Localization via Multitask Learning" by Wenbo Xu, Wei Lu, and Xiangyang Luo, with its own abstract, its own introduction on Deepfake detection, and its own experimental tables. None of the radar registration content — the moment functions, the radial-velocity model, the estimator, the consistency theorem, or the radar experiments — appears anywhere in the submitted text. This makes the central claim unverifiable, and the problem is not a localized issue that a revision could repair.
  2. [Abstract vs. body] The statement "we show the consistency of the proposed method" is unsupported by any theorem, proof, or even a specification of the moment conditions and identification assumptions. Standard GMM consistency requires a unique population moment zero at the true parameter and an appropriate rank condition; the manuscript provides none of these, and the body contains no equations from which they could be inferred. The consistency claim is therefore not a contribution that can be assessed on the submitted material.
  3. [Experiments] The abstract's claims of higher accuracy and robustness than benchmarks and of LiDAR-comparable performance are not backed by any radar experiment in the body. The only experimental tables (Tables I–IV) report temporal forgery localization metrics (mAP@IoU and AR@Proposals) on LAV-DF and AV-Deepfake1M; there is no synthetic radar dataset, no real-world radar dataset, no registration baselines, no error bars, and no registration evaluation protocol. The experimental claims cannot be evaluated or reproduced from the submitted manuscript.
minor comments (2)
  1. [Full text (running header)] The running header of the body cites "arXiv:2508.02179v1 [cs.CV] 4 Aug 2025", which does not match the identifier arXiv:2508.02187 under review; this mismatch should be corrected in any future submission.
  2. [Title and authorship] The title, author list, and subject area of the body differ completely from those implied by the abstract, so a reader cannot determine from the submission which paper is actually intended for review.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity identified in the available text; the supplied full text is an unrelated paper, so no specific reduction of the claimed GMM result to its inputs can be exhibited.

full rationale

The abstract for arXiv:2508.02187 claims a correspondence-free GMM-based 4D radar registration framework with consistency and accuracy comparable to LiDAR. The body text supplied under this arXiv ID is actually "Weakly Supervised Multimodal Temporal Forgery Localization via Multitask Learning" (Xu, Lu, Luo), which contains no definitions of the radar moment conditions, no radial-velocity measurement model, no identification conditions, and no consistency proof for the GMM estimator. Without those equations or any self-contained derivation, there is no quotable step in which a fitted parameter, a moment condition, or a cited previous result is shown to be equivalent by construction to the claimed prediction. Standard GMM asymptotics are not circular merely because they are invoked; circularity would require the specific moment functions to be defined in terms of the target transformation or fitted to the same data used for the accuracy comparison. Since no such reduction can be quoted from the available manuscript, the appropriate finding is no significant circularity. This is an absence-of-information result about the supplied text, not a demonstration that the radar registration derivation is circular.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The central claim rests almost entirely on content that is not visible in this submission. The abstract commits to GMM-based consistency and empirical superiority, which requires: an identifying set of moment conditions, standard GMM regularity conditions, and an overlapping scene with a correctly specified radar noise model. No free parameters can be enumerated because the manuscript body was not provided. No new physical entities are introduced. If the actual paper exists, this ledger would need to be rebuilt from its estimator section and experimental protocol.

assumptions (3)
  • domain assumption The chosen moment conditions, built from point positions and radial velocities, identify the true rigid transformation between the source and target radar clouds.
    GMM consistency only tells you that the estimator converges to the point where the moments hold; for the registration to be correct, those moments must be informative about the pose. The abstract's phrase 'we show the consistency of the proposed method' does not state or prove this identification condition.
  • standard math Standard GMM regularity conditions hold for the radar point cloud setting, including existence, identification, and a uniform law of large numbers over the compact parameter space.
    A GMM consistency proof in this setting would be an application of textbook asymptotic theory; the abstract does not list the conditions, and the full text that would contain them is absent.
  • domain assumption Both radar scans observe the same underlying scene and the radar measurement noise model is correctly specified.
    Moment-based alignment without correspondences requires scene overlap and an accurate noise model. A misspecified model can make the estimator consistent for the wrong transformation. Neither condition is visible in the abstract.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Registering the 4D Millimeter Wave Radar Point Clouds Via Generalized Method of Moments." pith.science (2026). https://pith.science/paper/SZTCDQXK

@misc{pith2026250802187,
  author       = {Pith},
  title        = {Pith review of: Registering the 4D Millimeter Wave Radar Point Clouds Via Generalized Method of Moments},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/SZTCDQXK}},
  note         = {Machine review of arXiv:2508.02187}
}
read the original abstract

4D millimeter wave radars (4D radars) are new emerging sensors that provide point clouds of objects with both position and radial velocity measurements. Compared to LiDARs, they are more affordable and reliable sensors for robots' perception under extreme weather conditions. On the other hand, point cloud registration is an essential perception module that provides robot's pose feedback information in applications such as Simultaneous Localization and Mapping (SLAM). Nevertheless, the 4D radar point clouds are sparse and noisy compared to those of LiDAR, and hence we shall confront great challenges in registering the radar point clouds. To address this issue, we propose a point cloud registration framework for 4D radars based on Generalized Method of Moments. The method does not require explicit point-to-point correspondences between the source and target point clouds, which is difficult to compute for sparse 4D radar point clouds. Moreover, we show the consistency of the proposed method. Experiments on both synthetic and real-world datasets show that our approach achieves higher accuracy and robustness than benchmarks, and the accuracy is even comparable to LiDAR-based frameworks.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

54 extracted references · 51 canonical work pages

  1. [1]

    Dynamic difference learning with spatio–temporal correlation for deepfake video detection,

    Q. Yin, W. Lu, B. Li, and J. Huang, “Dynamic difference learning with spatio–temporal correlation for deepfake video detection,” IEEE Transactions on Information Forensics and Security, vol. 18, pp. 4046–4058, 2023

  2. [2]

    Istvt: Interpretable spatial-temporal video transformer for deepfake detection,

    C. Zhao, C. Wang, G. Hu, H. Chen, C. Liu, and J. Tang, “Istvt: Interpretable spatial-temporal video transformer for deepfake detection,” IEEE Transactions on Information Forensics and Security, vol. 18, pp. 1335–1348, 2023

  3. [3]

    Weakly-supervised Audio Temporal Forgery Localization via Progressive Audio-language Co-learning Network

    J. Wu, W. Xu, W. Lu, X. Luo, R. Yang, and S. Guo, “Weakly-supervised audio temporal forgery localization via pro- gressive audio-language co-learning network,” arXiv preprint arXiv:2505.01880, 2025

  4. [4]

    Fine- grained multimodal deepfake classification via heterogeneous graphs,

    Q. Yin, W. Lu, X. Cao, X. Luo, Y . Zhou, and J. Huang, “Fine- grained multimodal deepfake classification via heterogeneous graphs,” International Journal of Computer Vision , pp. 1–15, 2024

  5. [5]

    Improving gener- alization of deepfake detectors by imposing gradient regulariza- tion,

    W. Guan, W. Wang, J. Dong, and B. Peng, “Improving gener- alization of deepfake detectors by imposing gradient regulariza- tion,” IEEE Transactions on Information Forensics and Security , vol. 19, pp. 5345–5356, 2024

  6. [6]

    Deepfake detection and localization using multi-view inconsistency measurement,

    B. Zhang, Q. Yin, W. Lu, and X. Luo, “Deepfake detection and localization using multi-view inconsistency measurement,” IEEE Transactions on Dependable and Secure Computing , vol. 22, no. 2, pp. 1796–1809, 2025

  7. [7]

    Not made for each other-audio-visual dissonance-based deepfake detection and localization,

    K. Chugh, P. Gupta, A. Dhall, and R. Subramanian, “Not made for each other-audio-visual dissonance-based deepfake detection and localization,” in Proceedings of the 28th ACM international conference on multimedia , 2020, pp. 439–447

  8. [8]

    Glitch in the matrix: A large scale benchmark for content driven audio–visual forgery detection and localization,

    Z. Cai, S. Ghosh, A. Dhall, T. Gedeon, K. Stefanov, and M. Hayat, “Glitch in the matrix: A large scale benchmark for content driven audio–visual forgery detection and localization,” Computer Vision and Image Understanding, vol. 236, p. 103818, 2023

Show all 54 references
  1. [9]

    Audio-visual tempo- ral forgery detection using embedding-level fusion and multi- dimensional contrastive loss,

    M. Liu, J. Wang, X. Qian, and H. Li, “Audio-visual tempo- ral forgery detection using embedding-level fusion and multi- dimensional contrastive loss,” IEEE Transactions on Circuits and Systems for Video Technology, 2023

  2. [10]

    Um- maformer: A universal multimodal-adaptive transformer frame- work for temporal forgery localization,

    R. Zhang, H. Wang, M. Du, H. Liu, Y . Zhou, and Q. Zeng, “Um- maformer: A universal multimodal-adaptive transformer frame- work for temporal forgery localization,” in Proceedings of the 31st ACM International Conference on Multimedia , 2023, pp. 8749–8759

  3. [11]

    Bi-stream coteaching network for weakly-supervised deepfake localization in videos,

    Z. Li, Z. Teng, B. Zhang, and J. Fan, “Bi-stream coteaching network for weakly-supervised deepfake localization in videos,” Trans. Info. For. Sec., vol. 20, p. 1724–1738, Jan. 2025

  4. [12]

    Cpl: Curriculum pseudo labeling for weakly supervised temporal forgery localization,

    D. Zhang, M. Fang, Z. Lu, and H. Xie, “Cpl: Curriculum pseudo labeling for weakly supervised temporal forgery localization,” in ICASSP 2025 - 2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , 2025, pp. 1–5

  5. [13]

    A multimodal deviation perceiving framework for weakly-supervised temporal forgery localization,

    W. Xu, J. Wu, W. Lu, X. Luo, and Q. Wang, “A multimodal deviation perceiving framework for weakly-supervised temporal forgery localization,” arXiv preprint arXiv:2507.16596 , 2025

  6. [14]

    Can chatgpt detect deepfakes? a study of using multimodal large language models for media forensics,

    S. Jia, R. Lyu, K. Zhao, Y . Chen, Z. Yan, Y . Ju, C. Hu, X. Li, B. Wu, and S. Lyu, “Can chatgpt detect deepfakes? a study of using multimodal large language models for media forensics,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024...

  7. [15]

    Mmnet: multi-collaboration and multi-supervision network for sequential deepfake detection,

    R. Xia, D. Liu, J. Li, L. Yuan, N. Wang, and X. Gao, “Mmnet: multi-collaboration and multi-supervision network for sequential deepfake detection,” IEEE Transactions on Information Forensics and Security, 2024

  8. [16]

    Jointly defending deepfake manip- ulation and adversarial attack using decoy mechanism,

    G.-L. Chen and C.-C. Hsu, “Jointly defending deepfake manip- ulation and adversarial attack using decoy mechanism,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 45, no. 8, pp. 9922–9931, 2023

  9. [17]

    Deepfake de- tection based on discrepancies between faces and their context,

    Y . Nirkin, L. Wolf, Y . Keller, and T. Hassner, “Deepfake de- tection based on discrepancies between faces and their context,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 44, no. 10, pp. 6111–6121, 2021

  10. [19]

    Hearing lips and seeing voices,

    H. McGurk and J. MacDonald, “Hearing lips and seeing voices,” Nature, vol. 264, no. 5588, pp. 746–748, 1976

  11. [20]

    Joint audio-visual deepfake detection,

    Y . Zhou and S.-N. Lim, “Joint audio-visual deepfake detection,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2021, pp. 14 800–14 809. 13

  12. [21]

    Do you really mean that? content driven audio-visual deepfake dataset and multimodal method for temporal forgery localization,

    Z. Cai, K. Stefanov, A. Dhall, and M. Hayat, “Do you really mean that? content driven audio-visual deepfake dataset and multimodal method for temporal forgery localization,” in 2022 International Conference on Digital Image Computing: Tech- niques and Applications (DICTA) . IE...

  13. [22]

    Frade: Forgery- aware audio-distilled multimodal learning for deepfake detec- tion,

    F. Nie, J. Ni, J. Zhang, B. Zhang, and W. Zhang, “Frade: Forgery- aware audio-distilled multimodal learning for deepfake detec- tion,” in Proceedings of the 32nd ACM International Conference on Multimedia, 2024, pp. 6297–6306

  14. [23]

    A brief introduction to weakly supervised learning,

    Z.-H. Zhou, “A brief introduction to weakly supervised learning,” National science review , vol. 5, no. 1, pp. 44–53, 2018

  15. [24]

    Cf- deformable detr: an end-to-end alignment-free model for weakly aligned visible-infrared object detection,

    H. Fu, J. Yuan, G. Zhong, X. He, J. Lin, and Z. Li, “Cf- deformable detr: an end-to-end alignment-free model for weakly aligned visible-infrared object detection,” in Proceedings of the Thirty-Third International Joint Conference on Artificial Intelli- gence, 2024, pp. 758–766

  16. [25]

    A consistency and integration model with adaptive thresholds for weakly supervised object localization,

    H. Su and M. Yang, “A consistency and integration model with adaptive thresholds for weakly supervised object localization,” in Proceedings of the Thirty-Third International Joint Conference on Artificial Intelligence , 2024, pp. 1281–1289

  17. [26]

    Complete instances mining for weakly supervised instance segmentation,

    Z. Li, Z. Zeng, Y . Liang, and J.-G. Yu, “Complete instances mining for weakly supervised instance segmentation,” arXiv preprint arXiv:2402.07633, 2024

  18. [27]

    Weakly super- vised object localization and detection: A survey,

    D. Zhang, J. Han, G. Cheng, and M.-H. Yang, “Weakly super- vised object localization and detection: A survey,” IEEE transac- tions on pattern analysis and machine intelligence, vol. 44, no. 9, pp. 5866–5885, 2021

  19. [28]

    Adaptive zone learning for weakly supervised object localization,

    Z. Chen, S. Wang, L. Cao, Y . Shen, and R. Ji, “Adaptive zone learning for weakly supervised object localization,” IEEE Transactions on Neural Networks and Learning Systems , 2024

  20. [29]

    Temporal action localization in the deep learning era: A survey,

    B. Wang, Y . Zhao, L. Yang, T. Long, and X. Li, “Temporal action localization in the deep learning era: A survey,” IEEE Transactions on Pattern Analysis and Machine Intelligence, 2023

  21. [30]

    Multitask learning,

    R. Caruana, “Multitask learning,” Machine learning , vol. 28, no. 1, pp. 41–75, 1997

  22. [31]

    A survey on multi-task learning,

    Y . Zhang and Q. Yang, “A survey on multi-task learning,” IEEE transactions on knowledge and data engineering , vol. 34, no. 12, pp. 5586–5609, 2021

  23. [32]

    Multi-task learning in natural language processing: An overview,

    S. Chen, Y . Zhang, and Q. Yang, “Multi-task learning in natural language processing: An overview,” ACM Comput. Surv., vol. 56, no. 12, Jul. 2024

  24. [33]

    Multi-task and multi-lingual joint learning of neural lexical utterance classification based on partially-shared model- ing,

    R. Masumura, T. Tanaka, R. Higashinaka, H. Masataki, and Y . Aono, “Multi-task and multi-lingual joint learning of neural lexical utterance classification based on partially-shared model- ing,” in Proceedings of the 27th International Conference on Computational Linguistics , ...

  25. [34]

    Person- alized microblog sentiment classification via adversarial cross- lingual multi-task learning,

    W. Wang, S. Feng, W. Gao, D. Wang, and Y . Zhang, “Person- alized microblog sentiment classification via adversarial cross- lingual multi-task learning,” in Proceedings of the 2018 Con- ference on Empirical Methods in Natural Language Processing , E. Riloff, D. Chiang, J. Hock...

  26. [35]

    Worse wer, but better bleu? leveraging word embedding as intermedi- ate in multitask end-to-end speech translation,

    S.-P. Chuang, T.-W. Sung, A. H. Liu, and H.-y. Lee, “Worse wer, but better bleu? leveraging word embedding as intermedi- ate in multitask end-to-end speech translation,” arXiv preprint arXiv:2005.10678, 2020

  27. [36]

    Learning modality- specific representations with self-supervised multi-task learning for multimodal sentiment analysis,

    W. Yu, H. Xu, Z. Yuan, and J. Wu, “Learning modality- specific representations with self-supervised multi-task learning for multimodal sentiment analysis,” in Proceedings of the AAAI conference on artificial intelligence , vol. 35, no. 12, 2021, pp. 10 790–10 797

  28. [37]

    Mod- eling task relationships in multi-task learning with multi-gate mixture-of-experts,

    J. Ma, Z. Zhao, X. Yi, J. Chen, L. Hong, and E. H. Chi, “Mod- eling task relationships in multi-task learning with multi-gate mixture-of-experts,” in Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery & Data Min- ing, ser. KDD ’18. New York, NY ...

  29. [38]

    Decouple and resolve: Transformer-based models for online anomaly detection from weakly labeled videos,

    T. Liu, C. Zhang, K.-M. Lam, and J. Kong, “Decouple and resolve: Transformer-based models for online anomaly detection from weakly labeled videos,” IEEE Transactions on Information Forensics and Security, vol. 18, pp. 15–28, 2023

  30. [39]

    Bi-stream coteach- ing network for weakly-supervised deepfake localization in videos,

    Z. Li, Z. Teng, B. Zhang, and J. Fan, “Bi-stream coteach- ing network for weakly-supervised deepfake localization in videos,” IEEE Transactions on Information Forensics and Se- curity, vol. 20, pp. 1724–1738, 2025

  31. [40]

    Multi-task gans for view-specific feature learning in gait recognition,

    Y . He, J. Zhang, H. Shan, and L. Wang, “Multi-task gans for view-specific feature learning in gait recognition,” IEEE Transactions on Information Forensics and Security , vol. 14, no. 1, pp. 102–113, 2019

  32. [41]

    Avoid-df: Audio-visual joint learning for detecting deepfake,

    W. Yang, X. Zhou, Z. Chen, B. Guo, Z. Ba, Z. Xia, X. Cao, and K. Ren, “Avoid-df: Audio-visual joint learning for detecting deepfake,” IEEE Transactions on Information Forensics and Security, vol. 18, pp. 2015–2029, 2023

  33. [42]

    Multi-modal deep- fake detection via multi-task audio-visual prompt learning,

    H. Miao, Y . Guo, Z. Liu, and Y . Wang, “Multi-modal deep- fake detection via multi-task audio-visual prompt learning,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 39, no. 1, 2025, pp. 612–621

  34. [43]

    Feature selection: A data perspective,

    J. Li, K. Cheng, S. Wang, F. Morstatter, R. P. Trevino, J. Tang, and H. Liu, “Feature selection: A data perspective,” ACM com- puting surveys (CSUR) , vol. 50, no. 6, pp. 1–45, 2017

  35. [44]

    Feature selection techniques for machine learning: a survey of more than two decades of research,

    D. Theng and K. K. Bhoyar, “Feature selection techniques for machine learning: a survey of more than two decades of research,” Knowledge and Information Systems , vol. 66, no. 3, pp. 1575–1637, 2024

  36. [45]

    Av-deepfake1m: A large-scale llm-driven audio-visual deepfake dataset,

    Z. Cai, S. Ghosh, A. P. Adatia, M. Hayat, A. Dhall, T. Gedeon, and K. Stefanov, “Av-deepfake1m: A large-scale llm-driven audio-visual deepfake dataset,” in Proceedings of the 32nd ACM International Conference on Multimedia , 2024, pp. 7414–7423

  37. [46]

    Deepfake detection via inter-frame inconsistency recomposition and en- hancement,

    C. Zhu, B. Zhang, Q. Yin, C. Yin, and W. Lu, “Deepfake detection via inter-frame inconsistency recomposition and en- hancement,” Pattern Recognition, vol. 147, p. 110077, 2024

  38. [47]

    Detection of deepfake videos using long-distance attention,

    W. Lu, L. Liu, B. Zhang, J. Luo, X. Zhao, Y . Zhou, and J. Huang, “Detection of deepfake videos using long-distance attention,” IEEE Transactions on Neural Networks and Learning Systems , vol. 35, no. 7, pp. 9366–9379, 2024

  39. [48]

    Temporal segment networks for action recognition in videos,

    L. Wang, Y . Xiong, Z. Wang, Y . Qiao, D. Lin, X. Tang, and L. Van Gool, “Temporal segment networks for action recognition in videos,” IEEE transactions on pattern analysis and machine intelligence, vol. 41, no. 11, pp. 2740–2755, 2018

  40. [49]

    wav2vec 2.0: A framework for self-supervised learning of speech repre- sentations,

    A. Baevski, Y . Zhou, A. Mohamed, and M. Auli, “wav2vec 2.0: A framework for self-supervised learning of speech repre- sentations,” Advances in neural information processing systems , vol. 33, pp. 12 449–12 460, 2020

  41. [50]

    Actionformer: Localizing mo- ments of actions with transformers,

    C.-L. Zhang, J. Wu, and Y . Li, “Actionformer: Localizing mo- ments of actions with transformers,” in European Conference on Computer Vision. Springer, 2022, pp. 492–510

  42. [51]

    Tridet: Temporal action detection with relative boundary modeling,

    D. Shi, Y . Zhong, Q. Cao, L. Ma, J. Li, and D. Tao, “Tridet: Temporal action detection with relative boundary modeling,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 18 857–18 866

  43. [52]

    Mfms: Learning modality-fused and modality-specific features for deepfake detection and localization tasks,

    Y . Zhang, C. Miao, M. Luo, J. Li, W. Deng, W. Yao, Z. Li, B. Hu, W. Feng, T. Gong et al. , “Mfms: Learning modality-fused and modality-specific features for deepfake detection and localization tasks,” in Proceedings of the 32nd ACM International Conference on Multimedia, 2024...

  44. [53]

    Cola: Weakly- supervised temporal action localization with snippet contrastive learning,

    C. Zhang, M. Cao, D. Yang, J. Chen, and Y . Zou, “Cola: Weakly- supervised temporal action localization with snippet contrastive learning,” in Proceedings of the IEEE/CVF conference on com- puter vision and pattern recognition , 2021, pp. 16 010–16 019

  45. [54]

    Full-stage pseudo label quality enhancement for weakly-supervised temporal action localization,

    Q. Feng, W. Li, T. Lin, and X. Chen, “Full-stage pseudo label quality enhancement for weakly-supervised temporal action localization,” arXiv preprint arXiv:2407.08971 , 2024

  46. [55]

    Multilevel semantic and adap- tive actionness learning for weakly supervised temporal action localization,

    Z. Li, Z. Wang, and C. Dong, “Multilevel semantic and adap- tive actionness learning for weakly supervised temporal action localization,” Neural Networks, vol. 182, p. 106905, 2025

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.