REVIEW 4 major objections 4 minor 50 references
Multiscale Adaptive Conflict-Balancing Model For Multimedia Deepfake Detection
T0 review · 4 major / 4 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read The paper proposes MACB-DF, a contrastive, conflict-balancing audio-visual fusion method that reports state-of-the-art deepfake detection across three benchmarks and improved cross-dataset generalization.
desk verdict A plausible incremental architecture with a genuine cross-dataset story, undercut by an internal number discrepancy and missing reproducibility details. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the multi-scale fusion loop built from three interacting components. First, Multi-modal Adaptive Contrast Learning (MACL) aligns video and audio embeddings with an adaptive temperature derived from variance, weighted skewness, and attention entropy, and additionally clusters embeddings using a composite Mahalanobis-cosine distance to produce each modality's fusion weight $w_{m_i}$. Second, the weighted fusion $x^{\text{fused}}_i = w_{v_i}V_{i-1} + w_{a_i}A_{i-1}$ combines deep video and audio features, and the fused representation is multiplied back into both streams at successive layers, so each modality learns from the joint signal. Third, the Orthogonalization-multimodal Pareto module solves $\min_{\alpha_m,\alpha_u}\|\alpha_m g_m+\alpha_u g_u\|^2$ under a simplex constraint, adding an orthogonality penalty whose strength is modulated by cosine similarity between the unimodal and multimodal gradients; this is the component that is supposed to stop one modality's gradient from overriding the other.
What would settle it
A reader can settle the cross-dataset claim by re-running the DFDC-to-DefakeAVMiT and DFDC-to-FakeAVCeleb experiments under the paper's 80/20 split and checking whether ACC reproduces at 91.2% and 89.2%; a second check is to replace the weighted feature sum in Eq. (17) with concatenation plus a linear layer and see whether accuracy survives.
Extended reading notes
Core claim
The central claim is that the main obstacle to accurate audio-visual deepfake detection is not the lack of fusion capacity but the imbalance and conflict between modalities during training. The paper's MACB-DF pipeline extracts spatiotemporal video features and log Mel-spectrogram audio features, then applies multi-modal adaptive contrast learning (MACL) that aligns positive audio-video pairs, separates negatives with a margin, and uses an adaptive temperature to control gradient scale. Clustering in the aligned space produces composite distances, Mahalanobis and cosine, that feed per-modality fusion weights, which are then used in a weighted sum of deep features; the fused representation is re-injected into both streams at multiple scales. A final orthogonalization-multimodal Pareto module treats the joint update from unimodal and multimodal gradients as a constrained minimization and applies orthogonal regularization when the gradients point in conflicting directions. The paper claims this design yields state-of-the-art accuracy on DefakeAVMiT, FakeAVCeleb, and DFDC, and better cross-dataset generalization when trained on DFDC.
Load-bearing premise
The load-bearing premise is that the video and audio feature tensors can be added together after being multiplied by learned weights; the paper never explains what projection or normalization makes that addition meaningful.
Editorial extensions
If this is right
- If the reported 95.5% average accuracy is reproducible, MACB-DF would be the best-performing audio-visual deepfake detector on these three benchmarks, and the cross-dataset results imply detectors trained on DFDC can transfer to unseen forgery datasets with only a modest drop.
- The ablations attribute clear accuracy gains to each component: removing contrastive learning drops ACC by 3.2 points on FakeAVCeleb and 4.7 points on DFDC, removing intra-modal losses costs 1.0 to 2.0 AUC points, and removing adaptive fusion weights costs 0.6 ACC points on DFDC.
- Because the Pareto module's main observed benefit is faster and more stable early convergence, the method implies that addressing gradient conflict is a training-dynamics problem as much as a final-performance problem.
- The multi-scale re-injection of the fused feature into both unimodal streams, Eq. (18), suggests that balanced multimodal learning requires each modality to see the joint representation, not just the final classifier.
Reading between the lines
- If Eq. (17)'s missing projection is filled in, the same fusion and Pareto machinery could be lifted into other audio-visual tasks, such as fake speech detection or speaker-video verification, where modality imbalance is also reported.
- The discrepancy between the abstract's 8.0/7.7 percentage-point cross-dataset gains and Table 3's 6.8/6.4-point values means the exact size of the claimed generalization improvement is not yet stable; a reader should treat the table as the conservative number until reproduced.
- A direct test of the balancing story would be to feed deliberately corrupted audio or video during inference and check whether the fusion weights $w_{v_i}, w_{a_i}$ shift to down-weight the corrupted modality; if they do not, the method balances losses during training but does not adapt at test time.
- The contrastive clustering step could be evaluated as a standalone module on unimodal deepfake benchmarks; if it improves unimodal accuracy too, its benefit is not specific to cross-modal fusion.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes MACB-DF, a multimodal audio-visual deepfake detection framework combining adaptive contrastive learning (MACL), multi-scale feature fusion with learnable fusion weights, a representation fusion module (RFMF), and an orthogonalization-multimodal Pareto module (OM-Pareto). The authors claim consistent state-of-the-art performance on DefakeAVMiT, FakeAVCeleb, and DFDC, with 95.5% average accuracy, and report cross-dataset gains of 8.0% and 7.7% ACC in the abstract when training on DFDC and testing on DefakeAVMiT and FakeAVCeleb. Section 6 presents the same cross-dataset experiment with gains of 6.8% and 6.4%, creating a direct internal contradiction in the paper's headline claim. Extensive ablations are presented for the contrastive losses, the fusion weighting, and the OM-Pareto module.
Significance. If the claimed results are reproducible, MACB-DF would be a meaningful step for audio-visual deepfake detection, particularly its cross-dataset generalization, which is more practically relevant than intra-dataset accuracy. The paper's strengths are the broad ablation coverage (contrastive losses, fusion weighting, OM-Pareto), the evaluation on three benchmarks, and the explicit formulation of the training objectives. However, the central quantitative claim is internally inconsistent, and no code, seeds, or error bars are provided, so the current evidence is not sufficient to confirm the stated state-of-the-art performance.
major comments (4)
- [Abstract / Section 6 (Table 3)] The headline cross-dataset claim is internally inconsistent: the Abstract and Conclusion state absolute ACC improvements of 8.0% and 7.7% over AVoiD-DF when trained on DFDC, but Table 3 and the Section 6 text report margins of 6.8% and 6.4% (91.2−84.4=6.8; 89.2−82.8=6.4). Because the state-of-the-art and generalization claims rest on these margins, the authors must reconcile the two sets of numbers, report the correct values in all locations, and state whether Table 3 or the abstract was based on a different evaluation protocol.
- [Section 3.2.4, Eq. (17)] The weighted fusion x_fused_i = w_v_i * V_{i-1} + w_a_i * A_{i-1} presupposes that video and audio feature maps have identical shape and comparable scales, but no projection, normalization, or dimension-matching operation is specified for V_{i-1} and A_{i-1}. Similarly, the clustering in Eq. (9) requires the cluster count K and per-cluster covariances Sigma_k at inference time, yet the manuscript does not state how K is chosen or how Sigma_k is estimated on a test batch. Without these details, the fusion and contrastive-alignment pipeline is not fully specified.
- [Section 3.5, Eqs. (21)-(24)] The OM-Pareto module is central to the claim of resolving gradient conflicts, but the paper does not provide the actual optimization procedure: Eq. (21) is a constrained optimization problem, yet the text only says 'Execute complete PGD iterations' without giving the projection operator, step size, or termination criterion, and Eq. (22) introduces lambda_orth without specifying how it is incorporated into the final update. The reader cannot reproduce or verify the claimed conflict resolution from the text.
- [Section 5.1 / Tables 1-2] The experimental reporting is not sufficient to support the quantitative point estimates: no error bars, number of seeds, or standard deviations are reported for any table, and Section 5.1 contains an unfinished placeholder 'amplify decision boundaries by a%'. Since the paper's central claims are quantitative, the authors should provide multi-seed results, variance measures, complete text for the ablation analysis, and a clear statement of code or implementation availability.
minor comments (4)
- [Throughout] There are several typographical issues: 'INTRODDUCTION' in Section 1, 'deptly balances' in the Figure 1 caption, and 'FakeA VCeleb' in Section 4.1; these should be corrected.
- [Table 3] The caption of Table 3 is malformed: 'THE TRAINING SETS AND TESTING SETS FOR CROSS-DATASET' appears to be a fragment of a sentence rather than a proper caption, and the table does not explicitly state that the rows are methods trained on DFDC.
- [Section 4.1] The LAV-DF dataset is described in Section 4.1 but never used in any experiment or ablation; the authors should either include results on it or explain why it is omitted.
- [Section 5.1] The ablation discussion states that contrastive learning 'amplifies decision boundaries by a% compared to baseline models'; this is an unfinished placeholder and must be replaced with the actual measured value or a precise description.
Circularity Check
No significant circularity: all proposed modules are heuristic design choices evaluated empirically, and no claimed prediction reduces by construction to a fitted input.
full rationale
I found no load-bearing circular step. The paper's core claims are empirical benchmark accuracies (Tables 1 and 3) obtained by training MACB-DF with the proposed losses and fusion rules; the equations (2), (7), (8), (17), and (21)-(24) are architectural and optimization choices, not reductions of the target accuracy. The cross-dataset numbers are measured outcomes, not fitted parameters renamed as predictions. The cited prior work is external and is not used as a self-citation chain to force the design. Two reporting issues are noted but are not circularity: the Abstract's '8.0% and 7.7%' improvements disagree with Table 3's own 6.8% and 6.4% margins over AVoiD-DF (Sec. 6), and Section 5.1 contains an unfinished placeholder ('amplify decision boundaries by a%'), marking incomplete verification. These should be corrected, but neither makes the derivation equivalent to its inputs. Accordingly, the circularity score is 0.
Assumptions & free parameters
free parameters (5)
- beta_k (k=1..3): adaptive temperature coefficients =
learned, values not reported
- gamma_t: temporal temperature gate =
learned, values not reported
- lambda_0, kappa: OM-Pareto regularization strength and curvature =
not reported
- margin m and negative count K in contrastive losses =
not reported
- beta (Eq. 9) and gamma (Eq. 10) balance parameters =
not reported
assumptions (3)
- domain assumption Audio and video features can be projected into a shared space where weighted summation is a valid fusion operator.
- ad hoc to paper The gradient conflict between modalities is fully characterized by the two vectors g_m and g_u and their cosine similarity, and solving Eq. (21) via PGD improves joint training.
- domain assumption Adaptive temperature based on variance, skewness, and attention entropy stabilizes contrastive learning.
Cite this review
Pith. "Pith review of Multiscale Adaptive Conflict-Balancing Model For Multimedia Deepfake Detection." pith.science (2026). https://pith.science/paper/FNUAGAZ7
@misc{pith2026250512966,
author = {Pith},
title = {Pith review of: Multiscale Adaptive Conflict-Balancing Model For Multimedia Deepfake Detection},
year = {2026},
howpublished = {\url{https://pith.science/paper/FNUAGAZ7}},
note = {Machine review of arXiv:2505.12966}
}
read the original abstract
Advances in computer vision and deep learning have blurred the line between deepfakes and authentic media, undermining multimedia credibility through audio-visual forgery. Current multimodal detection methods remain limited by unbalanced learning between modalities. To tackle this issue, we propose an Audio-Visual Joint Learning Method (MACB-DF) to better mitigate modality conflicts and neglect by leveraging contrastive learning to assist in multi-level and cross-modal fusion, thereby fully balancing and exploiting information from each modality. Additionally, we designed an orthogonalization-multimodal pareto module that preserves unimodal information while addressing gradient conflicts in audio-video encoders caused by differing optimization targets of the loss functions. Extensive experiments and ablation studies conducted on mainstream deepfake datasets demonstrate consistent performance gains of our model across key evaluation metrics, achieving an average accuracy of 95.5% across multiple datasets. Notably, our method exhibits superior cross-dataset generalization capabilities, with absolute improvements of 8.0% and 7.7% in ACC scores over the previous best-performing approach when trained on DFDC and tested on DefakeAVMiT and FakeAVCeleb datasets.
Figures
Figures from the paper (3 more)
Reference graph
Works this paper leans on
-
[1]
Tadas Baltrušaitis, Chaitanya Ahuja, and Louis-Philippe Morency. 2019. Multi- modal Machine Learning: A Survey and Taxonomy. IEEE Transactions on Pattern Analysis and Machine Intelligence 41, 2 (2019), 423–443. doi:10.1109/TPAMI.2018. 2798607
-
[2]
Remi Cadene, Corentin Dancette, Matthieu Cord, Devi Parikh, et al. 2019. Rubi: Reducing unimodal biases for visual question answering. Advances in neural information processing systems 32 (2019)
2019
-
[3]
Zhixi Cai, Kalin Stefanov, Abhinav Dhall, and Munawar Hayat. 2022. Do you really mean that? content driven audio-visual deepfake dataset and multimodal method for temporal forgery localization. In 2022 International Conference on Digital Image Computing: Techniques and Applications (DICTA). IEEE, IEEE, 1–10
work page 2022
-
[4]
Caroline Chan, Shiry Ginosar, Tinghui Zhou, and Alexei A Efros. 2019. Everybody dance now. In Proceedings of the IEEE/CVF international conference on computer vision. 5933–5942
2019
-
[5]
Liang Chen, Yong Zhang, Yibing Song, Lingqiao Liu, and Jue Wang. 2022. Self- supervised Learning of Adversarial Example: Towards Good Generalizations for Deepfake Detection. In 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 18689–18698. doi:10.1109/CVPR52688.2022.01815
arXiv 2022
-
[6]
Harry Cheng, Yangyang Guo, Tianyi Wang, Qi Li, Xiaojun Chang, and Liqiang Nie
-
[7]
Komal Chugh, Parul Gupta, Abhinav Dhall, and Ramanathan Subramanian. 2020. Not made for each other-audio-visual dissonance-based deepfake detection and localization. In Proceedings of the 28th ACM international conference on multimedia. 439–447
work page 2020
-
[8]
Hao Dang, Feng Liu, Joel Stehouwer, Xiaoming Liu, and Anil K. Jain. 2020. On the Detection of Digital Face Manipulation. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
work page 2020
Show all 50 references
-
[9]
Brecht Desplanques, Jenthe Thienpondt, and Kris Demuynck. 2020. Ecapa- tdnn: Emphasized channel attention, propagation and aggregation in tdnn based speaker verification. arXiv preprint arXiv:2005.07143
2020 arXiv
-
[10]
Brian Dolhansky, Joanna Bitton, Ben Pflaum, Jikuo Lu, Russ Howes, Menglin Wang, and Cristian Canton Ferrer. 2020. The deepfake detection challenge (dfdc) dataset. arXiv preprint arXiv:2006.07397 (2020)
2020 arXiv
-
[11]
Tharindu Fernando, Clinton Fookes, Simon Denman, and Sridha Sridharan. 2019. Exploiting human social cognition for the detection of fake and fraudulent faces via memory networks. arXiv preprint arXiv:1911.07844 (2019)
2019 arXiv
-
[12]
Qiqi Gu, Shen Chen, Taiping Yao, Yang Chen, Shouhong Ding, and Ran Yi. 2022. Exploiting fine-grained face forgery clues via progressive enhancement learning. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 36. 735–743
2022
-
[13]
Ruidong Han, Xiaofeng Wang, Ningning Bai, Qin Wang, Zinian Liu, and Jianru Xue. 2023. FCD-Net: Learning to detect multiple types of homologous deepfake face images. IEEE Transactions on Information Forensics and Security 18 (2023), 2653–2666
2023
-
[14]
Peisong He, Haoliang Li, and Hongxia Wang. 2019. Detection of fake images via the ensemble of deep representations from multi color spaces. In 2019 IEEE international conference on image processing (ICIP) . IEEE, 2299–2303
2019
-
[15]
R Huang, MWY Lam, J Wang, D Su, D Yu, Y Ren, and Z Zhao. 2022. FastDiff: A Fast Conditional Diffusion Model for High-Quality Speech Synthesis. In IJCAI Multiscale Adaptive Conflict-Balancing Model For Multimedia Deepfake Detection ICMR ’25, June 30-July 3, 2025, Chicago, IL, U...
2022
-
[16]
Hafsa Ilyas, Ali Javed, and Khalid Mahmood Malik. 2023. AVFakeNet: A uni- fied end-to-end Dense Swin Transformer deep learning model for audio–visual deepfakes detection. Applied Soft Computing 136 (2023), 110124
2023
-
[17]
Jee-weon Jung, Hee-Soo Heo, Hemlata Tak, Hye-jin Shim, Joon Son Chung, Bong- Jin Lee, Ha-Jin Yu, and Nicholas Evans. 2022. Aasist: Audio anti-spoofing using integrated spectro-temporal graph attention networks. In ICASSP 2022-2022 IEEE international conference on acoustics, sp...
2022
-
[18]
Hasam Khalid, Shahroz Tariq, Minha Kim, and Simon S Woo. 2021. FakeAVCeleb: A novel audio-video multimodal deepfake dataset.arXiv preprint arXiv:2108.05080 (2021)
2021 arXiv
-
[19]
Lingzhi Li, Jianmin Bao, Hao Yang, Dong Chen, and Fang Wen. 2020. Advanc- ing high fidelity identity swapping for forgery detection. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition . 5074–5083
2020
-
[20]
Yuanman Li and Jiantao Zhou. 2018. Fast and effective image copy-move forgery detection via hierarchical feature point matching. IEEE Transactions on Informa- tion Forensics and Security 14, 5 (2018), 1307–1322
2018
-
[21]
Paul Pu Liang, Amir Zadeh, and Louis-Philippe Morency. 2024. Foundations & trends in multimodal machine learning: Principles, challenges, and open ques- tions. Comput. Surveys 56, 10 (2024), 1–42
2024
-
[22]
T Lin. 2017. Focal Loss for Dense Object Detection.arXiv preprint arXiv:1708.02002 (2017)
2017 arXiv
-
[23]
Xiaolong Liu, Yang Yu, Xiaolong Li, and Yao Zhao. 2024. MCL: Multimodal Contrastive Learning for Deepfake Detection. IEEE Transactions on Circuits and Systems for Video Technology 34, 4 (2024), 2803–2813. doi:10.1109/TCSVT.2023. 3312738
2024 doi
-
[24]
Dhruv Malik, Ashwin Pananjady, Kush Bhatia, Koulik Khamaru, Peter L Bartlett, and Martin J Wainwright. 2020. Derivative-free methods for policy optimization: Guarantees for linear quadratic systems. Journal of Machine Learning Research 21, 21 (2020), 1–51
2020
-
[25]
Fatemeh Zare Mehrjardi, Ali Mohammad Latif, Mohsen Sardari Zarchi, and Razieh Sheikhpour. 2023. A survey on deep learning-based image forgery detection. Pattern Recognition (2023), 109778
2023
-
[26]
Trisha Mittal, Uttaran Bhattacharya, Rohan Chandra, Aniket Bera, and Dinesh Manocha. 2020. Emotions don’t lie: An audio-visual deepfake detection method using affective cues. In Proceedings of the 28th ACM international conference on multimedia. 2823–2832
2020
-
[27]
Huy H Nguyen, Junichi Yamagishi, and Isao Echizen. 2019. Capsule-forensics: Using capsule networks to detect forged images and videos. In ICASSP 2019-2019 IEEE international conference on acoustics, speech and signal processing (ICASSP) . IEEE, 2307–2311
2019
-
[28]
Fan Nie, Jiangqun Ni, Jian Zhang, Bin Zhang, and Weizhe Zhang. 2024. FRADE: Forgery-aware Audio-distilled Multimodal Learning for Deepfake Detection. In Proceedings of the 32nd ACM International Conference on Multimedia . 6297–6306
2024
-
[29]
Xiaokang Peng, Yake Wei, Andong Deng, Dong Wang, and Di Hu. 2022. Balanced multimodal learning via on-the-fly gradient modulation. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition . 8238–8247
2022
-
[30]
Jinghan Ru, Jun Tian, Chengwei Xiao, Jingjing Li, and Heng Tao Shen. 2024. Imbalanced Open Set Domain Adaptation via Moving-Threshold Estimation and Gradual Alignment. IEEE Transactions on Multimedia 26 (2024), 2504–2514. doi:10.1109/TMM.2023.3297768
2024
-
[31]
Jinghan Ru, Yuxin Xie, Xianwei Zhuang, Yuguo Yin, and Yuexian Zou. 2025. Do we really have to filter out random noise in pre-training data for language models? arXiv:2502.06604 [cs.CL] https://arxiv.org/abs/2502.06604
2025 arXiv
-
[32]
Ekraam Sabir, Jiaxin Cheng, Ayush Jaiswal, Wael AbdAlmageed, Iacopo Masi, and Prem Natarajan. 2019. Recurrent convolutional strategies for face manipulation detection in videos. Interfaces (GUI) 3, 1 (2019), 80–87
2019
-
[33]
Saniat Javid Sohrawardi, Akash Chintha, Bao Thai, Sovantharith Seng, Andrea Hickerson, Raymond Ptucha, and Matthew Wright. 2019. Poster: Towards robust open-world detection of deepfakes. In Proceedings of the 2019 ACM SIGSAC conference on computer and communications security ....
2019
-
[34]
Shahroz Tariq, Sangyup Lee, Hoyoung Kim, Youjin Shin, and Simon S Woo. 2018. Detecting both machine and human created fake face images in the wild. In Proceedings of the 2nd international workshop on multimedia privacy and security . 81–87
2018
-
[35]
Shahroz Tariq, Sangyup Lee, Hoyoung Kim, Youjin Shin, and Simon S Woo
-
[36]
Weiyao Wang, Du Tran, and Matt Feiszli. 2020. What makes training multi- modal classification networks hard?. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition . 12695–12705
2020
-
[37]
Yaohui Wang, Piotr Bilinski, Francois Bremond, and Antitza Dantcheva. 2020. Imaginator: Conditional spatio-temporal gan for video generation. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision. 1160–1169
2020
-
[38]
Nan Wu, Stanislaw Jastrzebski, Kyunghyun Cho, and Krzysztof J Geras. 2022. Characterizing and overcoming the greedy nature of learning in multi-modal deep neural networks. In International Conference on Machine Learning . PMLR, 24043–24055
2022
-
[39]
Yuxin Xie, Zhihong Zhu, Xianwei Zhuang, Liming Liang, Zhichang Wang, and Yuexian Zou. 2024. GPA: Global and Prototype Alignment for Audio-Text Re- trieval. In Interspeech 2024. 5078–5082. doi:10.21437/Interspeech.2024-1642
2024 doi
-
[40]
Wenyuan Yang, Xiaoyu Zhou, Zhikai Chen, Bofei Guo, Zhongjie Ba, Zhihua Xia, Xiaochun Cao, and Kui Ren. 2023. AVoiD-DF: Audio-Visual Joint Learning for Detecting Deepfake. IEEE Transactions on Information Forensics and Security 18 (2023), 2015–2029. doi:10.1109/TIFS.2023.3262148
2023
-
[41]
Yuguo Yin, Yuxin Xie, Wenyuan Yang, Dongchao Yang, Jinghan Ru, Xianwei Zhuang, Liming Liang, and Yuexian Zou. 2025. ATRI: Mitigating Multilin- gual Audio Text Retrieval Inconsistencies by Reducing Data Distribution Errors. arXiv:2502.14627 [cs.SD] https://arxiv.org/abs/2502.14627
2025 arXiv
-
[42]
Yang Yu, Xiaolong Liu, Rongrong Ni, Siyuan Yang, Yao Zhao, and Alex C. Kot
-
[43]
Hanqing Zhao, Wenbo Zhou, Dongdong Chen, Tianyi Wei, Weiming Zhang, and Nenghai Yu. 2021. Multi-attentional Deepfake Detection. arXiv:2103.02406 [cs.CV] https://arxiv.org/abs/2103.02406
2021 arXiv
-
[44]
Tianfei Zhou, Wenguan Wang, Zhiyuan Liang, and Jianbing Shen. 2021. Face forensics in the wild. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 5778–5788
2021
-
[45]
Xianwei Zhuang, Hongxiang Li, Xuxin Cheng, Zhihong Zhu, Yuxin Xie, and Yuexian Zou. 2025. KDProR: A Knowledge-Decoupling Probabilistic Frame- work for Video-Text Retrieval. In Computer Vision – ECCV 2024 , Aleš Leonardis, Elisa Ricci, Stefan Roth, Olga Russakovsky, Torsten Sat...
2025
-
[46]
Xianwei Zhuang, Yuxin Xie, Yufan Deng, Dongchao Yang, Liming Liang, Jinghan Ru, Yuguo Yin, and Yuexian Zou. 2025. VARGPT-v1.1: Improve Visual Autore- gressive Large Unified Model via Iterative Instruction Tuning and Reinforcement Learning. arXiv:2504.02949 [cs.CV] https://arxi...
2025 arXiv
-
[47]
Xianwei Zhuang, Zhihong Zhu, Yuxin Xie, Liming Liang, and Yuexian Zou. 2025. VASparse: Towards Efficient Visual Hallucination Mitigation via Visual-Aware Token Sparsification. arXiv:2501.06553 [cs.CV] https://arxiv.org/abs/2501.06553
2025 arXiv
-
[2019]
In Proceedings of the 34th ACM/SIGAPP Symposium on Applied Computing
Gan is a friend or foe? a framework to detect various fake face images. In Proceedings of the 34th ACM/SIGAPP Symposium on Applied Computing . 1296– 1303
-
[2023]
ACM Transactions on Multimedia Computing, Communications and Applications 20, 3 (2023), 1–22
Voice-face homogeneity tells deepfake. ACM Transactions on Multimedia Computing, Communications and Applications 20, 3 (2023), 1–22
2023
-
[2024]
IEEE Transactions on Circuits and Systems for Video Technology 34, 8 (2024), 6926–6936
PVASS-MDD: Predictive Visual-Audio Alignment Self-Supervision for Multimodal Deepfake Detection. IEEE Transactions on Circuits and Systems for Video Technology 34, 8 (2024), 6926–6936. doi:10.1109/TCSVT.2023.3309899
2024
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.