Pith. sign in

REVIEW 3 major objections 6 minor 35 references

Multimedia Verification Through Multi-Agent Deep Research Multimodal Large Language Models

T0 review · 3 major / 6 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read Multi-agent MLLM system claims verifiable multimedia forensics, from geolocation to source tracking.

desk verdict A coherent challenge-system description with a useful architecture, but its central authenticity claim rests on one unvalidated case and an invalid overlay-consistency heuristic. read the letter →

arxiv 2507.04410 v1 pith:BFTB57P7 submitted 2025-07-06 cs.CV cs.AIcs.IR

classification cs.CVcs.AIcs.IR
keywords multimediaverificationmultimodallargelanguagemodelsmulti-agentsystemfact-checkingreverseimagesearchmetadataanalysisdeepresearchmisinformation
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper describes a submission to the ACM Multimedia 2025 Grand Challenge on Multimedia Verification, aiming to show that a six-stage pipeline combining multimodal large language models with external verification tools can authenticate real-world multimedia content. The authors claim their system, built around a Deep Researcher Agent with four tools, can extract precise geolocation, timing, and source attribution while detecting synthetic manipulation. They demonstrate it on a single case study of a missile-strike video, labeling it 'Verified' based on converging evidence. The relevance is that most fact-checked misinformation involves images or videos, yet prior methods often isolate technical forensics from contextual analysis. If the approach works, it would offer an integrated, evidence-grounded alternative to hallucination-prone MLLM reasoning.

What carries the argument

The central mechanism is the Deep Researcher Agent, an LLM-driven iterative search-and-analysis engine equipped with four tools: reverse image search, metadata analysis, fact-checking databases, and a verified-news processor that systematically extracts spatial (Where?), temporal (When?), attribution (Who?), and motivational (Why?) context from trusted news sources. This agent operates inside a six-stage pipeline—raw data processing, planning, information extraction, deep research, evidence collection, and report generation—that separates video and image inputs, uses a frame extractor to capture key moments, and coordinates verification through a Planner Agent. The system claims to reduce MLLM hallucination by anchoring reasoning to externally retrieved evidence and by documenting the full provenance chain.

What would settle it

Run the system on a deliberately fabricated video that includes consistent, realistic embedded overlays—such as a forged timestamp and location text—and check whether it is still classified as 'Verified.' Alternatively, compare the system's verdicts against ground-truth labels across the full 50-sample challenge set; if any manipulated clip with consistent overlays is labeled Verified, the paper's authenticity criterion fails.

Watch

Extended reading notes

Core claim

The paper claims that combining an MLLM orchestrator with tool-assisted verification—reverse image search, metadata analysis, fact-checking databases, and a verified-news processor that extracts where, when, who, and why—can verify multimedia authenticity, geolocation, timing, and source provenance in a challenging real-world case. The system processed a video of a missile strike on a bridge in Dnipro, Ukraine, and produced a structured report that located the event at approximately 48.4647° N, 35.0462° E, dated it to 04/05/2022 at 19:58:37 local time, traced its original posting to a Twitter account, and concluded the footage was not synthetically manipulated. The authors frame this as evidence that a multi-agent deep-research pipeline can handle complex geopolitical misinformation scenarios while maintaining transparency through provenance tracking.

Load-bearing premise

The verdict that footage is authentic depends on assuming that consistent embedded overlays (time, location, MAC address) and matching rendering cues prove a video was not synthetically manipulated, but the paper does not test whether a manipulated video carrying believable, consistent overlays would fool the system.

Editorial extensions

If this is right

  • If the approach is correct, automated verification of geolocation, timing, and source attribution can be performed end-to-end from raw video and images, producing structured reports suitable for fact-checking workflows.
  • The four-tool Deep Researcher Agent suggests a template for evidence-grounded MLLM use that reduces reliance on parametric memory and lowers the risk of fabricated justifications.
  • The verified-news processor's extraction of where, when, who, and why aligns verification with the standard journalistic questions, making outputs more interpretable to human fact-checkers.
  • Successful handling of a geopolitical event case implies the system could generalize to other high-stakes verification scenarios, such as electoral claims or disaster footage, where multiple corroborating sources exist.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A testable extension would be running the system across the full 50-sample challenge dataset and comparing its verdicts against ground-truth labels; the paper's single-case demonstration leaves the overall accuracy unmeasured.
  • The forensic trust in consistent embedded overlays implies a vulnerability: a deliberately fabricated video with believable, consistent timestamps and location text could be misclassified as authentic, since the paper's evidence for 'Verified' relies precisely on such overlays.
  • The modular architecture could be adapted to incorporate independent forensic detectors or human-in-the-loop appeal mechanisms, connecting to the contestable-AI direction the authors sketch for future work.
  • The system's dependence on online sources means verification quality will vary with source availability; for novel or obscure events, the 'Who' and 'Why' dimensions may remain unresolved despite strong visual evidence.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The paper describes a submission to the ACMMM25 Grand Challenge on Multimedia Verification. The authors propose a six-stage multi-agent pipeline that combines multimodal large language models with external verification tools (reverse image search, metadata analysis, fact-checking databases, and a verified-news processor). The system is demonstrated on a single challenge case, ID43-3, for which the generated report classifies the content as 'Verified', extracts geolocation and timestamps, and attributes the source to a Twitter account. The paper claims that the system successfully verifies content authenticity and effectively addresses real-world multimedia verification scenarios.

Significance. If substantiated, the contribution would be a useful integration of MLLM-based reasoning with tool-assisted verification for multimedia fact-checking. The architecture is described coherently, and the case study illustrates how contextual evidence can be organized into a structured verification report. However, the paper provides no quantitative evaluation, no comparison with ground-truth labels, no baseline, and no independent validation of the forensic conclusions. The central effectiveness claim therefore rests on a single self-generated report. The paper does not ship code, checkable proofs, or a benchmark evaluation; its value is currently limited to a system description and an illustrative example.

major comments (3)
  1. [Abstract and Section 5] The central claim that the system 'successfully verified content authenticity' and 'effectively addressing real-world multimedia verification scenarios' is supported only by a single case study (ID43-3). The dataset is described as containing 50 samples (Section 3), and the conclusion refers to evaluation on 'challenge dataset cases' in the plural, yet no aggregate statistics, no ground-truth comparison, and no error analysis are reported. Without quantitative evidence or at least a comparison of the system's output against the official challenge labels, the effectiveness claim is not established.
  2. [Section 5, Forensic Analysis] The authenticity verdict is based on the consistency of embedded overlays (time, location, MAC-address information) across video segments, together with lighting and shadow cues. This is not a valid forensic argument: a manipulated or fully synthetic video can carry consistent overlays, and overlay consistency does not rule out pixel-level manipulation. The report also asserts that 'state-of-the-art detection tools' found no synthetic or deepfake anomalies, but no such tool is named, and Section 4.4 lists no manipulation detector among the four verification tools. The 'Verified' classification is therefore an unsupported MLLM judgment rather than a measured forensic result.
  3. [Section 5, Verified Evidence] The report presents approximate coordinates (48.4647° N, 35.0462° E) and an exact timestamp (04/05/2022, 19:58:37) as 'Verified', but the surrounding text indicates these values come from embedded overlays and video metadata. No independent confirmation is shown, such as matching the timestamp to known news reports, comparing the location with satellite imagery or map data, or cross-validating across multiple independent sources. The claim that geolocation and timing information were 'extracted' is credible, but the stronger claim that they were independently verified is not supported by the evidence presented.
minor comments (6)
  1. [Abstract] There is a typo in the phrase 'theACMMM25'; a space is missing between 'the' and 'ACMMM25'.
  2. [Section 5 heading] The heading 'Verificiation Report' contains a typo; it should be 'Verification Report'.
  3. [Section 4.4] The bullet list states that the verified-news processing tool extracts 'four critical source details' but then lists five items: Source detail, Where?, When?, Who?, and Why?. The count should be corrected to five, or the items should be regrouped.
  4. [Section 5, Evidence Images] The evidence image links reference 'ID43-2.mp4' and 'ID43-1.mp4' while the case under analysis is ID43-3; the relationship between these files and the submitted case should be clarified.
  5. [References] Reference [8] contains placeholder page numbers 'xxxx–yyyy' and should be completed before publication.
  6. [Section 4] The paper does not specify the MLLM versions, prompt templates, or tool configuration details beyond naming Gemini 2.0 Flash and Yandex Image Search API; additional implementation details would improve reproducibility.

Circularity Check

2 steps flagged · score 4.0 of 10

Authenticity verdict is partly self-confirmatory: the report uses the system's own YAML outputs as supporting evidence and measures success as its own 'Verified' label, though geolocation and timing have some independent anchors.

  1. other [Section 5, Other Evidence & Findings, Supporting Sources]
    "Multiple video analysis YAML files (ID43-3_ID43-2, ID43-3_ID43-3, ID43-3_ID43-1) consistently report the explosion at a bridge in Dnipro, Ukraine..."

    These YAML files are outputs of the system's own Stage 1 video analysis pipeline (Gemini 2.0 Flash, Section 4.1.1), not independent external sources. The verification report lists them under 'Supporting Sources' as corroboration for its own conclusions, so the evidence chain loops back to the system itself: the system's intermediate outputs are used as evidence for the system's final verdict. The only clearly external item in the same list is the Twitter post, whose authenticity is itself part of what is being verified.

  2. self definitional [Section 5, paragraph after the verification report]
    "The verification process successfully confirmed the authenticity of the submitted content, classifying it as 'Verified', based on converging evidence from multiple sources."

    The paper's measure of 'successfully confirmed authenticity' is the system's own 'Verified' classification; no comparison against the challenge's ground-truth labels or an independent forensic detector is presented. The report is both the measuring instrument and the verdict, so this demonstration cannot fail by construction unless the pipeline itself emits an anomaly flag. The success claim is thus tied definitionally to the system's self-generated label rather than to an external check, making the central 'Verified' result self-confirmatory.

full rationale

The paper does not derive quantitative predictions, and the architecture is a tool-orchestration pipeline rather than a mathematical derivation, so most classical circularity patterns do not apply. However, the demonstration on ID43-3 contains a concrete self-referential evidence loop: the report's 'Supporting Sources' include the system's own YAML analysis files, so part of the 'converging evidence' is generated by the same pipeline whose output it is meant to validate. Additionally, the claim of forensic authenticity relies on unnamed 'state-of-the-art detection tools' and on overlay consistency, yet Section 4.4 lists no deepfake or manipulation detector among the system's four tools; this is a missing-support issue rather than a circular step, but it reinforces that the 'Verified' verdict is largely self-assessed. There is some independent grounding: geolocation is cross-referenced with satellite imagery and known landmarks, and the timestamp is tied to a Twitter post, so the central claim does not reduce entirely to the system's own output. On balance, the partial self-corroboration warrants a moderate circularity score rather than a clean bill; the paper would need ground-truth comparison or an externally validated forensic detector to remove the circularity. A score of 4 reflects that the geolocation/timing content has independent anchors while the authenticity conclusion is substantially self-confirmatory.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

No fitted parameters or invented entities appear. The verification conclusions rely on unstated trust in MLLM analysis, reverse-search coverage, overlay consistency as authenticity evidence, and social-media metadata.

assumptions (4)
  • domain assumption Gemini 2.0 Flash produces accurate frame descriptions and metadata extraction for verification purposes.
    Stage 1 relies on the MLLM's video analysis; no error analysis or ground-truth checks are provided (Section 4.1.1).
  • domain assumption Consistency of embedded overlays (timestamps, location, MAC address) across video segments indicates authenticity.
    Used as the main forensic evidence in the ID43-3 report; overlays could in principle be part of a manipulated composite (Section 5).
  • domain assumption Reverse image search and fact-checking database results provide comprehensive and trustworthy corroboration.
    The Deep Researcher Agent treats returned URLs and articles as evidence; no coverage or reliability analysis is given (Sections 4.1.2 and 4.4).
  • domain assumption The Twitter post identified as the original source is genuinely the first publication and was posted at the claimed time.
    Source attribution in the report rests on platform post metadata without independent verification (Section 5).

how reviews work

0 comments
Cite this review

Pith. "Pith review of Multimedia Verification Through Multi-Agent Deep Research Multimodal Large Language Models." pith.science (2026). https://pith.science/paper/BFTB57P7

@misc{pith2026250704410,
  author       = {Pith},
  title        = {Pith review of: Multimedia Verification Through Multi-Agent Deep Research Multimodal Large Language Models},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/BFTB57P7}},
  note         = {Machine review of arXiv:2507.04410}
}
read the original abstract

This paper presents our submission to the ACMMM25 - Grand Challenge on Multimedia Verification. We developed a multi-agent verification system that combines Multimodal Large Language Models (MLLMs) with specialized verification tools to detect multimedia misinformation. Our system operates through six stages: raw data processing, planning, information extraction, deep research, evidence collection, and report generation. The core Deep Researcher Agent employs four tools: reverse image search, metadata analysis, fact-checking databases, and verified news processing that extracts spatial, temporal, attribution, and motivational context. We demonstrate our approach on a challenge dataset sample involving complex multimedia content. Our system successfully verified content authenticity, extracted precise geolocation and timing information, and traced source attribution across multiple platforms, effectively addressing real-world multimedia verification scenarios.

Figures

Figures reproduced from arXiv: 2507.04410 by the authors.

Figure 1
Figure 1. ACMMM25 - Grand Challenge on Multimedia Verifi￾cation dataset for verifying the authenticity and context of multimedia content. Prior studies highlight two major challenges: (1) “deepfakes” (i.e., content tampering and synthesis, where images/videos are digitally altered or generated) and (2) “cheapfakes” (i.e., content miscon￾textualization, where genuine media is reused in a false context to mislead). Early effort… view at source ↗
Figure 2
Figure 2. Our proposed Multi-Agent Deep Research MLLMs architecture for the multimedia verification system. [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. Detected key frames from raw data ID43-3. detailed source attribution demonstrates the effectiveness of our multi-modal approach. The system’s comprehensive source analysis is particularly noteworthy, tracing the content’s origin to a specific Twitter account and documenting its subsequent distribution across multiple platforms. The forensic analysis component detected no signs of synthetic manipulation or deepfake … view at source ↗

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

35 extracted references · 25 canonical work pages

  1. [1]

    Kars Alfrink et al. 2023. Contestable AI by design: Towards a framework. Minds and Machines 33, 4 (2023), 613–639

  2. [2]

    Rajaa Alqudah, Mohammed Al-Qaisi, Rakan Ammari, and Yazan Abu Ta’a. 2023. OSINT-Based Tool for Social Media User Impersonation Detection Through Machine Learning. In 2023 International Conference on Information Technology (ICIT). IEEE, 752–757

  3. [3]

    Shivangi Aneja, Cise Midoglu, Duc-Tien Dang-Nguyen, Sohail Ahmed Khan, Michael Riegler, Pål Halvorsen, Chris Bregler, and Balu Adsumilli. 2022. Acm mul- timedia grand challenge on detecting cheapfakes. arXiv preprint arXiv:2207.14534 (2022)

  4. [4]

    Christina Boididou, Katerina Andreadou, Symeon Papadopoulos, Duc Tien Dang Nguyen, Giulia Boato, Michael Riegler, Yiannis Kompatsiaris, et al. 2015. Verifying multimedia use at mediaeval 2015. In MediaEval 2015. Vol. 1436. CEUR- WS

  5. [5]

    Tobias Braun, Mark Rothermel, Marcus Rohrbach, and Anna Rohrbach. 2024. DEFAME: Dynamic Evidence-based FAct-checking with Multimodal Experts. arXiv preprint arXiv:2412.10510 (2024)

  6. [6]

    Swathi Chundru. 2021. Leveraging AI for Data Provenance: Enhancing Tracking and Verification of Data Lineage in FATE Assessment. International Journal of Inventions in Engineering & Science Technology 7, 1 (2021), 87–104

  7. [7]

    Yandex Cloud. 2025. Yandex Cloud Documentation | Yandex Search API | Web Search API, gRPC: ImageSearchService . https://yandex.cloud/en/docs/search- api/api-ref/grpc/ImageSearch/

  8. [8]

    Duc-Tien Dang-Nguyen, Morten Langfeldt Dahlback, Henrik Vold, Silje Førsund, Minh-Son Dao, Kha-Luan Pham, Sohail Ahmed Khan, Marc Gallofré Ocaña, Minh-Triet Tran, and Anh-Duy Tran. 2025. The 2025 Grand Challenge on Multimedia Verification: Foundations and Overview. In Proceedings of the 33rd ACM International Conference on Multimedia . xxxx–yyyy

Show all 35 references
  1. [9]

    Minh-Son Dao and Koji Zettsu. 2023. Leveraging knowledge graphs for cheap- fakes detection: Beyond dataset evaluation. In 2023 IEEE International Conference on Multimedia and Expo Workshops (ICMEW) . IEEE, 99–104

  2. [10]

    Jessica Deuschel, Andreas Foltyn, Karsten Roscher, and Stephan Scheele. 2024. The role of uncertainty quantification for trustworthy AI. In Unlocking Artificial Intelligence: From Theory to Applications . Springer, 95–115

  3. [11]

    Sharad Duwal, Mir Nafis Sharear Shopnil, Abhishek Tyagi, and Adiba Mahbub Proma. 2025. Evidence-Grounded Multimodal Misinformation Detection with Attention-Based GNNs. arXiv preprint arXiv:2505.18221 (2025)

  4. [12]

    Dhanvi Ganti. 2022. A novel method for detecting misinformation in videos, utilizing reverse image search, semantic analysis, and sentiment comparison of metadata. Utilizing Reverse Image Search, Semantic Analysis, and Sentiment Comparison of Metadata (June 5, 2022) (2022)

  5. [13]

    Xingyu Gao, Xi Wang, Zhenyu Chen, Wei Zhou, and Steven CH Hoi. 2024. Knowledge enhanced vision and language model for multi-modal fake news detection. IEEE Transactions on Multimedia (2024)

  6. [14]

    Bishwamittra Ghosh, Sarah Hasan, Naheed Anjum Arafat, and Arijit Khan. 2024. Logical Consistency of Large Language Models in Fact-checking. arXiv preprint arXiv:2412.16100 (2024). arXiv:2412.16100

  7. [15]

    Sonal Goel, Niharika Sachdeva, Ponnurangam Kumaraguru, A. V. Subramanyam, and Divam Gupta. 2016. PicHunt: Social Media Image Retrieval for Improved Law Enforcement. In Social Informatics, Emma Spiro and Yong-Yeol Ahn (Eds.). Springer International Publishing, Cham, 206–223

  8. [16]

    Haiying Guan. 2025. NIST Open Media Forensics Challenge (OpenMFC Briefing for IIRD)

  9. [17]

    Arash Heidari, Nima Jafari Navimipour, Hasan Dag, and Mehmet Unal. 2024. Deepfake detection using deep learning methods: A systematic and compre- hensive review. Wiley Interdisciplinary Reviews: Data Mining and Knowledge Discovery 14, 2 (2024), e1520

  10. [18]

    Naveena Karusala, Sohini Upadhyay, Rajesh Veeraraghavan, and Krzysztof Z Gajos. 2024. Understanding Contestability on the Margins: Implications for the Design of Algorithmic Decision-making in Public Services. In Proceedings of the Multimedia Verification Through Multi-Agent D...

  11. [19]

    Kyungha Kim, Sangyun Lee, Kung-Hsiang Huang, Hou Pong Chan, Manling Li, and Heng Ji. 2024. Can llms produce faithful explanations for fact-checking? towards faithful explainable fact-checking via multi-agent debate. arXiv preprint arXiv:2402.07401 (2024)

  12. [20]

    Ashish Kumar, Divya Singh, Rachna Jain, Deepak Kumar Jain, Chenquan Gan, and Xudong Zhao. 2025. Advances in DeepFake detection algorithms: Exploring fusion techniques in single and multi-modal approach. Information Fusion (2025), 102993

  13. [21]

    Kumud Lakara, Juil Sock, Christian Rupprecht, Philip Torr, John Collomosse, and Christian Schroeder de Witt. 2024. MAD-Sherlock: Multi-Agent Debates for Out-of-Context Misinformation Detection. arXiv preprint arXiv:2410.20140 (2024)

  14. [22]

    Xuannan Liu, Peipei Li, Huaibo Huang, Zekun Li, Xing Cui, Jiahao Liang, Lixiong Qin, Weihong Deng, and Zhaofeng He. 2024. Fka-owl: Advancing multimodal fake news detection through knowledge-augmented lvlms. In Proceedings of the 32nd ACM International Conference on Multimedia ...

  15. [23]

    Dale Meredith. 2024. The OSINT Handbook: A practical guide to gathering and analyzing online information. Packt Publishing Ltd

  16. [24]

    Bao-Tin Nguyen, Van-Loc Nguyen, Thanh-Son Nguyen, Duc-Tien Dang-Nguyen, Trong-Le Do, and Minh-Triet Tran. 2024. A Hybrid Approach for Cheapfake Detection Using Reputation Checking and End-To-End Network. In Proceedings of the 1st Workshop on Security-Centric Strategies for Com...

  17. [25]

    Hung Nguyen et al. 2024. LangXAI: Integrating Large Vision Models for Gener- ating Textual Explanations to Enhance Explainability in Visual Perception Tasks. In Proceedings of the Thirty-Third International Joint Conference on Artificial Intel- ligence, IJCAI-24, Kate Larson (...

  18. [26]

    Hung Nguyen, Alireza Rahimi, Veronica Whitford, Hélène Fournier, Irina Kon- dratova, René Richard, and Hung Cao. 2025. Heart2Mind: Human-Centered Contestable Psychiatric Disorder Diagnosis System using Wearable ECG Moni- tors. arXiv preprint arXiv:2505.11612 (2025)

  19. [27]

    Minh-Tam Nguyen, Quynh T Nguyen, Minh Son Dao, and Binh T Nguyen. 2025. Multimodal scene-graph matching for cheapfakes detection.International Journal of Multimedia Information Retrieval 14, 2 (2025), 17

  20. [28]

    Thanh-Son Nguyen, Vinh Dang, Minh-Triet Tran, and Duc-Tien Dang-Nguyen

  21. [29]

    Thanh-Son Nguyen and Minh-Triet Tran. 2023. Multi-Models from Computer Vision to Natural Language Processing for Cheapfakes Detection. In 2023 IEEE International Conference on Multimedia and Expo Workshops (ICMEW) . IEEE, 93–98

  22. [30]

    Van-Hoang Phan, Long-Khanh Pham, Dang Vu, Anh-Duy Tran, and Minh-Son Dao. 2025. E-FreeM2: Efficient Training-Free Multi-Scale and Cross-Modal News Verification via MLLMs. arXiv preprint arXiv:2506.20944 (2025)

  23. [31]

    Andreas Rossler, Davide Cozzolino, Luisa Verdoliva, Christian Riess, Justus Thies, and Matthias Nießner. 2019. Faceforensics++: Learning to detect manipulated facial images. In Proceedings of the IEEE/CVF international conference on computer vision. 1–11

  24. [32]

    Timothée Schmude. 2025. Explainability and Contestability for the Responsible Use of Public Sector AI. In Proceedings of the Extended Abstracts of the CHI Conference on Human Factors in Computing Systems . 1–6

  25. [33]

    Quang-Tien Tran, Thanh-Phuc Tran, Minh-Son Dao, Tuan-Vinh La, Anh-Duy Tran, and Duc Tien Dang Nguyen. 2022. A textual-visual-entailment-based unsupervised algorithm for cheapfake detection. In Proceedings of the 30th ACM International Conference on Multimedia . 7145–7149

  26. [34]

    Yuxi Xie, Guanzhen Li, Xiao Xu, and Min-Yen Kan. 2024. V-DPO: Mitigat- ing Hallucination in Large Vision Language Models via Vision-Guided Direct Preference Optimization. In Findings of the Association for Computational Linguis- tics: EMNLP 2024, Yaser Al-Onaizan, Mohit Bansal...

  27. [2023]

    In Proceedings of the 4th ACM Workshop on Intelligent Cross-Data Analysis and Retrieval

    Leveraging cross-modals for cheapfakes detection. In Proceedings of the 4th ACM Workshop on Intelligent Cross-Data Analysis and Retrieval . 51–59

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.