REVIEW 4 major objections 5 minor 127 references
TRIED: Truly Innovative and Effective AI Detection Benchmark, developed by WITNESS
T0 review · 4 major / 5 minor · reviewed 2026-08-16 · deepseek-v4-flash
Pith's one-line read The paper claims a 177-point self-assessment checklist can classify AI detection tools by real-world effectiveness, with scores above 143 meaning 'truly effective'.
desk verdict The TRIED checklist is a solid sociotechnical contribution, but the 177-point scoring system and the '143 = truly effective' cutoff are arbitrary and need recalibration or reframing. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central mechanism is the TRIED Benchmark checklist, a 177-point assessment instrument whose items ask developers to answer 'Yes', 'No', 'To do', or 'N/A' with a one-sentence justification. Each 'Yes' earns 3 points, a justified 'To do' earns 2, a justified 'No' earns 1, and unjustified or empty answers earn 0. The checklist's work is to convert qualitative sociotechnical lessons, drawn from frontline cases involving low-resolution video, noisy audio, underrepresented languages, and false accusations of AI use, into a single numeric score with explicit bands for effectiveness, so that 'truly effective' has a public, checkable meaning.
What would settle it
Take a set of detection tools that score above 143 on the checklist, run them on a corpus of low-resolution, noisy, and non-English manipulated and authentic media, and compare their error rates with tools scoring below 71; if the high-scoring tools are not systematically more accurate on those real-world cases, the score-to-effectiveness mapping is contradicted.
Extended reading notes
Core claim
The paper's discovery is a definition of 'truly innovative and effective' that is grounded in six interconnected pillars: handling real-world media conditions, transparency and explainability, accessibility, fairness, durability, and integration into broader verification workflows. The TRIED Benchmark operationalizes these pillars as a scored checklist with 177 points, distributed across lifecycle stages: design, development, testing, implementation, and maintenance. A self-assessed score above 143 means the tool is 'truly effective'; scores of 107 to 142 are moderately effective, 72 to 106 somewhat effective, and below 71 not effective. The claim is that a tool passing this checklist is one that supports the people most exposed to deceptive AI rather than merely performing well on clean benchmark data.
Load-bearing premise
The load-bearing premise is that developers' self-reported answers, each backed by a single sentence, faithfully describe how a tool actually behaves, and that the chosen cutoff of 143 is a meaningful separation rather than an arbitrary number.
Editorial extensions
If this is right
- Detection developers can use the checklist to find gaps before release, such as missing training data for compressed social-media formats or missing multilingual support.
- Standards bodies and regulators could adopt the threshold bands as a baseline for procurement or certification of detection tools.
- Evaluation of detection tools would broaden from accuracy-only metrics to include explainability, accessibility, fairness, durability, and fit within verification workflows.
- Fact-checkers and civil-society users would gain a common vocabulary for comparing tools beyond vendor claims.
Reading between the lines
- Editorial inference: the 143-point threshold has not been calibrated against measured detection performance, so the same checklist could be tested by scoring a set of tools and then checking their actual outcomes on real-world cases.
- Editorial inference: rewarding 'to do with justification' almost as much as 'yes' means a roadmap can count nearly as much as a finished capability; a stricter scoring variant might separate planning from delivery.
- Editorial inference: the checklist could be extended to independent third-party assessment, where reviewers rather than developers answer the questions, or to weighted scoring for specific user groups such as human-rights defenders versus general audiences.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This WITNESS report argues that AI detection tools are typically evaluated on narrow technical metrics and therefore fail in real-world use, and it proposes the TRIED Benchmark as a remedy: a 177-point checklist organized into six pillars (real-world design, transparency/explainability, accessibility, fairness, durability, and integration with verification ecosystems). The qualitative argument is grounded in WITNESS's Deepfakes Rapid Response Force (DRRF) casework, global consultations, and external work such as Deepfake-Eval-2024. The report claims that a score above 143 points on the Annex A checklist indicates a tool is 'truly effective,' with bands for moderately, somewhat, and not effective. The main body also offers recommendations for developers, regulators, standards bodies, and governments. The central quantitative claim, however, rests entirely on Annex A's scoring rubric, which is introduced without calibration, pilot testing, inter-rater reliability analysis, or comparison with external detection outcomes, and which contains apparent internal inconsistencies in the point accounting.
Significance. If the qualitative framework were adopted, it would make a useful contribution by expanding AI-detection evaluation beyond accuracy metrics toward accessibility, fairness, explainability, and workflow integration. The report's documentation of low-quality media, language barriers, false-positive harms, and the importance of human-assisted verification is valuable and well illustrated by concrete DRRF cases. The alignment with existing trustworthy-AI frameworks (EU AI Act, NIST, PAI) provides useful grounding. However, the benchmark's quantitative scoring system is the paper's headline deliverable, and that system is currently unsupported by validation data and is internally inconsistent in its arithmetic. The 'truly effective' threshold is precisely the kind of actionable claim that stakeholders would use in procurement or policy decisions, so the lack of support is load-bearing rather than cosmetic.
major comments (4)
- [Annex A (scoring instructions and thresholds)] The claim that 'a score above 143 points indicates that your tool is truly effective' is unsupported by any calibration, pilot testing, inter-rater reliability analysis, or comparison against external measures of detection performance. The point weights (Yes=3, To Do with justification=2, No with justification=1, otherwise 0) and the 143/107/72 cutoffs are asserted rather than derived or fitted, so the report's central quantitative claim is not backed by evidence. Given that the abstract and Executive Summary present this threshold as actionable guidance, the authors should either validate the scale or explicitly reframe the score as an unvalidated self-assessment instrument.
- [Annex A (Development checklist, maximum points)] The Development section of the checklist contains 22 items, including conditional items 3.4, 3.5, 4.1, 4.2, and 4.3. If all applicable items are counted, the maximum for that section is 22×3=66 points, yet the text states a maximum of 57 points. This makes the stated total maximum of 177 points, and therefore the 143-point threshold, depend on an unexplained exclusion of applicable items; the numerical scale is internally inconsistent for a multimodal tool that honestly answers all conditionals.
- [Annex A (answer rubric)] The instructions allow N/A as a response, but the scoring table specifies points only for Yes, To Do with or without justification, and No with or without justification; it does not state how N/A is counted. If N/A receives 0 points, tools with a narrower intended scope are penalized relative to general-purpose tools. If N/A is excluded from the denominator, the maximum possible score varies by tool and the fixed thresholds 143/107/72 are not well-defined. Either interpretation breaks the comparability that the fixed thresholds presuppose.
- [Sections 4 and 7, Annex A] The benchmark is presented as grounded in the same DRRF casework and WITNESS consultations that are then used as the primary evidence for its validity; no independent validation of the checklist items' content validity or of the thresholds against external assessments is provided. This does not undermine the qualitative lessons from the case studies, but it means the paper does not demonstrate that the benchmark measures 'real-world effectiveness' rather than the authors' own design priorities.
minor comments (5)
- [Abstract] The abstract contains missing spaces in 'stakeholderscandriveinnovation, safeguardpublictrust, strengthenAIliteracy, andcontribute' and in 'consid erations'; these should be fixed.
- [Sections 3 and 4] Section 3 begins 'IntroductionDeceptiveAIcontent' with a missing space after the heading, and Section 4 uses 'Al' instead of 'AI' in 'evaluates Al through a sociotechnical lens.'
- [References] Several references are incomplete or malformed: [19] and [32] lack closing brackets in their arXiv identifiers, and references such as [10], [18], [35], [37], [39], [40], [46], [48], [52], and [70] are bare social-media URLs without author, title, or access-date information, which makes source verification difficult.
- [Section 6.4 and Section 2] Section 6.4 contains typos ('funs' for 'funds', 'prioritze' for 'prioritize'), and Section 2 defines TRIED as 'Truly Innovative and Effective Detection' while the title and abstract include 'AI Detection'; the acronym expansion should be consistent.
- [Annex A (scoring table)] The scoring table gives 0 points for both 'To do without justification' and 'No without justification,' which is equivalent to not answering at all; the intended distinction should be made explicit.
Circularity Check
The 'truly effective' cutoff is a self-imposed definition, not a derived result.
-
self definitional
[Annex A, 'A WITNESS’ TRIED Benchmark: A Checklist for Truly Innovative and Effective AI Detection', final score bands]
"The Total Maximum Number of Points is 177. A score above 143 points indicates that your tool is truly effective. A score between 107 and 142 points indicates that your tool is moderately effective. A score between 72 and 106 points indicates that your tool is somewhat effective. A score below 71 points indicates that your tool is not effective."
The predicate 'truly effective' is defined by the authors' own score bands in the same annex. The report provides no calibration, pilot testing, inter-rater reliability analysis, or outcome-data link connecting 143/177 (80.8%) to real-world effectiveness. Thus the central claim that a score above 143 indicates a tool is 'truly effective' is a restatement of the chosen threshold, not a result derived from evidence. The label is attached to the score by definition, so the 'indication' is tautological.
full rationale
The report's qualitative sociotechnical framework is not circular: it is grounded in external literature (EU AI Act, NIST, OECD, PAI, BetterBench, Deepfake-Eval-2024) and illustrated by concrete DRRF case examples. Those sections stand as independent content. The circularity is localized to the benchmark's quantitative scoring claim. The point weights (Yes=3, To do with justification=2, No with justification=1, otherwise 0) and the effectiveness bands (143/107/72) are asserted in Annex A without calibration, piloting, inter-rater agreement, or comparison against actual detection outcomes. The final statement that a score above 143 points indicates a tool is 'truly effective' reduces, by construction, to the authors' own stipulation in the same annex. Frequent self-citations to WITNESS reports and blogs support background motivation and are not load-bearing for this threshold, so they do not independently raise the circularity score beyond the definitional issue.
Assumptions & free parameters
free parameters (2)
- Point weights for checklist answers =
Yes=3, To do with justification=2, No with justification=1, all other responses=0
- Effectiveness thresholds =
144-177 truly effective, 107-142 moderately effective, 72-106 somewhat effective, below 72 not effective
assumptions (3)
- domain assumption The six pillars are the complete set of relevant effectiveness dimensions
- domain assumption Self-reported one-sentence justifications are reliable evidence of a tool's properties
- domain assumption Anecdotal DRRF cases generalize across global contexts
invented entities (1)
-
TRIED Benchmark
Cite this review
Pith. "Pith review of TRIED: Truly Innovative and Effective AI Detection Benchmark, developed by WITNESS." pith.science (2026). https://pith.science/paper/NENS4J4Y
@misc{pith2026250421489,
author = {Pith},
title = {Pith review of: TRIED: Truly Innovative and Effective AI Detection Benchmark, developed by WITNESS},
year = {2026},
howpublished = {\url{https://pith.science/paper/NENS4J4Y}},
note = {Machine review of arXiv:2504.21489}
}
read the original abstract
The proliferation of generative AI and deceptive synthetic media threatens the global information ecosystem, especially across the Global Majority. This report from WITNESS highlights the limitations of current AI detection tools, which often underperform in real-world scenarios due to challenges related to explainability, fairness, accessibility, and contextual relevance. In response, WITNESS introduces the Truly Innovative and Effective AI Detection (TRIED) Benchmark, a new framework for evaluating detection tools based on their real-world impact and capacity for innovation. Drawing on frontline experiences, deceptive AI cases, and global consultations, the report outlines how detection tools must evolve to become truly innovative and relevant by meeting diverse linguistic, cultural, and technological contexts. It offers practical guidance for developers, policy actors, and standards bodies to design accountable, transparent, and user-centered detection solutions, and incorporate sociotechnical considerations into future AI standards, procedures and evaluation frameworks. By adopting the TRIED Benchmark, stakeholders can drive innovation, safeguard public trust, strengthen AI literacy, and contribute to a more resilient global information credibility.
Reference graph
Works this paper leans on
-
[1]
Accessed: 7 April 2025
European Union Artificial Intelligence Act.https://artificialintelligenceact.eu/article/1/ . Accessed: 7 April 2025
2025
-
[2]
Audio atribuido a margarita gonzález sobre programas sociales en morelos tiene indicios de manipulación digital.Animal Politico
Samedi Aguirre. Audio atribuido a margarita gonzález sobre programas sociales en morelos tiene indicios de manipulación digital.Animal Politico. https://www.animalpolitico.com/verificacion-de-hec hos/desinformacion/audio-candidata-morena-margarita/, 2024. Accessed: 8 April 2025
2024
-
[3]
Keep your ai claims in check.Internet Archive
Michael Atleson. Keep your ai claims in check.Internet Archive. https://web.archive.org/web/20 250115031325/https://www.ftc.gov/business-guidance/blog/2023/02/keep-your-ai-claims-c heck/, 2023. Accessed: 8 April 2025
2023
-
[4]
Tomorrow’s great digital divide: Content with or without provenance.WITNESS Blog
Jacobo Castellanos. Tomorrow’s great digital divide: Content with or without provenance.WITNESS Blog. https://blog.witness.org/2025/03/tomorrows-great-digital-divide , 2025. Accessed: 7 April 2025
2025
-
[5]
Reducing risks posed by synthetic content an overview of technical approaches to digital content transparency
Bilva Chandra, Jesse Dunietz, and Kathleen Roberts. Reducing risks posed by synthetic content an overview of technical approaches to digital content transparency. Technical report, National Institute of Standards and Technology, 2024
2024
-
[6]
Deepfake-eval-2024: A multi-modal in-the-wild benchmark of deepfakes circulated in
Nuria Alina Chandra, Ryan Murtfeldt, Lin Qiu, Arnab Karmakar, Hannah Lee, Emmanuel Tanumihardja, Kevin Farhat, Ben Caffee, Sejin Paik, Changyeon Lee, Jongwook Choi, Aerin Kim, and Oren Etzioni. Deepfake-eval-2024: A multi-modal in-the-wild benchmark of deepfakes circulated in
2024
-
[7]
An Indian politician says scandalous audio clips are AI deepfakes
Nilesh Christopher. An Indian politician says scandalous audio clips are AI deepfakes. we had them tested. Rest of World. https://restofworld.org/2023/indian-politician-leaked-audio-ai-dee pfake/, 2023. Accessed: 8 April 2025
2023
-
[8]
Accessed: 7 April 2025
Shakti: India Fact-Checking Collective.https://projectshakti.in/. Accessed: 7 April 2025
2025
Show all 127 references
-
[9]
Audio-visual person-of- interest deepfake detection
Davide Cozzolino, Alessandro Pianese, Matthias Nießner, and Luisa Verdoliva. Audio-visual person-of- interest deepfake detection. arXiv. https://arxiv.org/abs/2204.03083/ , 2023. arXiv:2204.03083 [cs.CV]
2023 arXiv
-
[10]
@dahrinoor2. X. https://x.com/dahrinoor2/status/1771155215414067603/ , 2024. Accessed: 7 April 2025
2024
-
[11]
Accessed: 7 April 2025
Synthetic Media Deepfakes and Generative AI.https://www.gen-ai.witness.org/?pk_vid=e9bdde b5077398fa1745328087fd83d2/. Accessed: 7 April 2025
2025
-
[12]
Facebook
Facebook. Facebook. https://www.facebook.com/100002167747651/videos/457953553957488/?vh= e&extid=MSG-UNK-UNK-UNK-COM_GK0T-GK1C/ , 2024. Accessed: 8 April 2025
2024
-
[13]
Accessed: 8 April 2025
Coalition for Content Provenance and Authenticity.https://c2pa.org/. Accessed: 8 April 2025
2025
-
[14]
C2PA Harms Modelling 1.4
Coalition for Content Provenance and Authenticity. C2PA Harms Modelling 1.4. https://partne rshiponai.org/resource/glossary-for-synthetic-media-transparency-methods-part-1/ . Accessed: 8 April 2025
2025
-
[15]
https://www.gen-ai.witness.org/deepfakes-rapid-respons e-force
Deepfakes Rapid Response Force. https://www.gen-ai.witness.org/deepfakes-rapid-respons e-force. Accessed: 7 April 2025
2025
-
[16]
Getting to the source: Understanding metadata removal on social media.Magnet Forensics Blog
Magnet Forensics. Getting to the source: Understanding metadata removal on social media.Magnet Forensics Blog. https://www.magnetforensics.com/blog/getting-to-the-source-understanding -metadata-removal-on-social-media// , 2024. Accessed: 8 April 2025. 17
2024
-
[17]
Disconnected from reality: American voters grapple with ai and flawed osint strategies.Institute for Strategic Dialogue
Isabelle Frances-Wright, Ellen Jacobs, and Ella Meyer. Disconnected from reality: American voters grapple with ai and flawed osint strategies.Institute for Strategic Dialogue. https://www.isdglobal. org/digital_dispatches/disconnected-from-reality-american-voters-grapple-with-...
2024
-
[18]
@GazetteNGR. X. https://x.com/GazetteNGR/status/1642218315685699586/ , 2023. Accessed: 8 April 2025
2023
-
[19]
Datasheets for datasets.arXiv
Timnit Gebru, Jamie Morgenstern, Briana Vecchione, Jennifer Wortman Vaughan, Hanna Wallach, Hal Daumé III, and Kate Crawford. Datasheets for datasets.arXiv. https://arxiv.org/abs/1803.0 9010/, 2021. arXiv:1803.09010 [cs.DB
2021 arXiv
-
[20]
Posts misrepresent a photo of a ukrainian soldier balancing on his prosthetic limbs.AP News
Melissa Goldin. Posts misrepresent a photo of a ukrainian soldier balancing on his prosthetic limbs.AP News. https://apnews.com/article/fact-check-ukrainian-soldier-photo-nazi-salute-863 799988341?fbclid=IwAR0w4ci7kN4tp0RKyfbrj_xH7VrxJxvckoY7PeAQ67WMZslobZZ5OCpHHLU/ , 2024. Ac...
2024
-
[21]
Sam Gregory. What’s needed in deepfakes detection? insights from witness’ global preparedness work and the partnership on ai’s steerco on media integrity work on the deepfake detection challenge. WITNESS Blog. https://blog.witness.org/2020/04/whats-needed-deepfakes-detection/ ,
2020
-
[22]
Pre-empting a crisis: Deepfake detection skills + global access to media forensics tools
Sam Gregory. Pre-empting a crisis: Deepfake detection skills + global access to media forensics tools. WITNESS Blog. https://blog.witness.org/2021/07/deepfake-detection-skills-tools-acces s, 2021. Accessed: 7 April 2025
2021
-
[23]
The world needs deepfake experts to stem this chaos.Wired
Sam Gregory. The world needs deepfake experts to stem this chaos.Wired. https://www.wired.co m/story/opinion-the-world-needs-deepfake-experts-to-stem-this-chaos/ , 2021. Accessed: 8 April 2025
2021
-
[24]
Grother, Mei L
Patrick J. Grother, Mei L. Ngan, and Kayee K. Hanaoka. Face recognition vendor test part 3: Demographic effects, nist interagency.Internal Report, National Institute of Standards and Technology. https://doi.org/10.6028/NIST.IR.8280/, 2019. Accessed: 8 April 2025
-
[25]
Youtube data viewer.https://citizenevidence.org/2014/07/01/youtube -dataviewer/
Amnesty International. Youtube data viewer.https://citizenevidence.org/2014/07/01/youtube -dataviewer/. Accessed: 8 April 2025
2014
-
[26]
Ticks or it didn’t happen: Key dilemmas in building authenticity infrastructure for multimedia
Gabi Ivens and Sam Gregory. Ticks or it didn’t happen: Key dilemmas in building authenticity infrastructure for multimedia. WITNESS Report. https://lab.witness.org/ticks-or-it-did nt-happen/#references/, 2019. Accessed: 8 April 2025
2019
-
[27]
@johnnygould. X. https://x.com/jonnygould/status/1768345232184148139/ , 2024. Accessed: 8 April 2025
2024
-
[28]
Chen, and Siwei Lyu
Yan Ju, Shu Hu, Shan Jia, George H. Chen, and Siwei Lyu. Improving fairness in deepfake detection. preprint arXiv. https://arxiv.org/abs/2306.16635/, 2023. arXiv:2306.16635 [cs.CV]
2023 arXiv
-
[29]
Adversarial machine learning in the context of network security: Challenges and solutions.Journal of Computational Intelligence and Robotics, 4(1):51–63, march 2024
Muskan Khan and Laiba Ghafoor. Adversarial machine learning in the context of network security: Challenges and solutions.Journal of Computational Intelligence and Robotics, 4(1):51–63, march 2024
2024
-
[30]
Le, Jiwon Kim, Simon S
Binh M. Le, Jiwon Kim, Simon S. Woo, Kristen Moore, Alsharif Abuadbba, and Shahroz Tariq. Sok: Systematization and benchmarking of deepfake detectors in a unified framework.preprint arXiv. https: //arxiv.org/abs/2401.04364/, 2025. arXiv:2401.04364 [cs.CV]
2025 arXiv
-
[31]
What is facial recognition technology?https://www.ajl.org/facial-r ecognition-technology/
Algorithmic Justice League. What is facial recognition technology?https://www.ajl.org/facial-r ecognition-technology/. Accessed: 8 April 2025. 18
2025
-
[32]
Leibowicz
Claire R. Leibowicz. Regulating reality: Exploring synthetic media through multistakeholder ai governance. arXiv. https://arxiv.org/abs/2502.04526/, 2025. arXiv:2502.04526 [cs.CY
2025 arXiv
-
[33]
Fortifying the truth in the age of synthetic media and generative ai perspectives from africa.WITNESS Blog
Raquel Vazquez Llorente, Jacobo Castellanos, and Nkem Agunwa. Fortifying the truth in the age of synthetic media and generative ai perspectives from africa.WITNESS Blog. https://blog.witness .org/2023/05/generative-ai-africa/, 2023. Accessed: 7 April 2025
2023
-
[34]
Deepfake detection improves when using algorithms that are more aware of demographic diversity
Siwei Lyu and Yan Ju. Deepfake detection improves when using algorithms that are more aware of demographic diversity. NiemanLab. https://www.niemanlab.org/2024/04/deepfake-detection -improves-when-using-algorithms-that-are-more-aware-of-demographic-diversity/ , 2023. Accessed:...
2024
-
[35]
Archived from X
@mini_razdan10. Archived from X. https://drive.google.com/file/d/1XVRz3vVikXX9ttMTLZlQ7 JOeUYZJQn9k/view?usp=sharing/, 2024. Accessed: 7 April 2025
2024
-
[36]
Model cards for model reporting
Margaret Mitchell, Simone Wu, Andrew Zaldivar, Parker Barnes, Lucy Vasserman, Ben Hutchinson, Elena Spitzer, Inioluwa Deborah Raji, and Timnit Gebru. Model cards for model reporting. In Proceedings of the Conference on Fairness, Accountability, and Transparency, page 220–229. ...
2019
-
[37]
MPvenezolano. Youtube. https://www.youtube.com/watch?v=T8RE-8OfFp4, 2024. Accessed: 8 April 2025
2024
-
[38]
Daragh Murray. Police use of retrospective facial recognition technology: A step change in surveillance capability necessitating an evolution of the human rights law framework.The Modern Law Review, 87(4):833–863, dec 2023
2023
-
[39]
@mustafali2001. Tik Tok. https://www.tiktok.com/@muss.mil/video/7438149251449752850/ ,
-
[40]
Facebook
Solid PH News. Facebook. https://www.facebook.com/solidphnews1/videos/501528382261724/ ,
-
[41]
OECD AI Principles overview.https://oecd.ai/en/ai-principles/
OECD AI Policy Observatory. OECD AI Principles overview.https://oecd.ai/en/ai-principles/. Accessed: 8 April 2025
2025
-
[43]
Ethics guidelines for trustworthy ai.European Commission
High-Level Expert Group on AI. Ethics guidelines for trustworthy ai.European Commission. https: //digital-strategy.ec.europa.eu/en/library/ethics-guidelines-trustworthy-ai/ , 2019. Accessed: 8 April 2025
2019
-
[44]
Accessed: 7 April 2025
2025
-
[45]
Building a glossary for synthetic media transparency methods part 1: Indirect disclosure
Partnership on AI. Building a glossary for synthetic media transparency methods part 1: Indirect disclosure. Partnership on AI. https://partnershiponai.org/resource/glossary-for-synthetic -media-transparency-methods-part-1/ , 2023. Accessed: 8 April 2025
2023
-
[46]
AI Risks and Trustworthiness.https://airc.nist
National Institute of Standards and Technology. AI Risks and Trustworthiness.https://airc.nist. gov/airmf-resources/airmf/3-sec-characteristics/#:~:text=NIST%20has%20identified%20t hree%20major,statistical%2C%20and%20human%2Dcognitive/, 2019. Accessed: 8 April 2025
2019
-
[47]
Archived from X
@PedroKonductaz. Archived from X. https://drive.google.com/file/d/1sN9OYILspCUaMf83LroQq 19 suqgLW4GRhT/view?usp=sharing/, 2024. Accessed: 8 April 2025
2024
-
[48]
The deepfake detection challenge: Insights and recommendations for AI and media integrity
Partnership on AI. The deepfake detection challenge: Insights and recommendations for AI and media integrity. Partnership on AI Report. https://partnershiponai.org/a-report-on-the-deepfake-d etection-challenge//, 2020. Accessed: 8 April 2025
2020
-
[49]
Determining trustworthiness through provenance and context
Google Public Policy. Determining trustworthiness through provenance and context. Google Policy Paper. https://static.googleusercontent.com/media/publicpolicy.google/en//resources/d etermining_trustworthiness_en.pdf/, 2024. Accessed: 8 April 2025
2024
-
[50]
UTV Ghana Online. Youtube. https://www.youtube.com/watch?v=K1KYhenW3oM&t=596s , 2024. Accessed: 8 April 2024
2024
-
[52]
Archived from Tik Tok
@phenixsaqartvelo. Archived from Tik Tok. https://drive.google.com/file/d/1EaB__rClt_pQCjH 8M1StmNejNspWvSaA/view?usp=sharing/, 2024. Accessed: 7 April 2025
2024
-
[53]
Faceforensics++: Learning to detect manipulated facial images
Andreas Rössler, Davide Cozzolino, Luisa Verdoliva, Christian Riess, Justus Thies, and Matthias Nießner. Faceforensics++: Learning to detect manipulated facial images. arxiv. https://arxiv. org/abs/1901.08971/, 2019. arXiv:1901.08971 [cs.CV]
1901 arXiv
-
[54]
Adversarial deep learning against intrusion detection classifiers
Maria Rigaki. Adversarial deep learning against intrusion detection classifiers. Master’s thesis, Luleå University of Technology, 2017
2017
-
[55]
Spotting the deepfakes in this year of elections: How ai detection tools work and where they fail.Reuters Institute
shirin anlen and Raquel Vazquez Llorente. Spotting the deepfakes in this year of elections: How ai detection tools work and where they fail.Reuters Institute. https://reutersinstitute.politics.ox .ac.uk/news/spotting-deepfakes-year-elections-how-ai-detection-tools-work-and-whe...
2024
-
[56]
Nimisha Singh, Amita Kapoor, and Neha Soni. A sociotechnical perspective for explicit unfairness mitigation techniques for algorithm fairness.International Journal of Information Management Data Insights, 4(2):6073–6105, nov 2024
2024
-
[57]
@rwomchechen. X. https://x.com/rwomchechen/status/1860001051367342123/, 2024. Accessed: 8 April 2025
2024
-
[58]
Governing access to synthetic media detection technology
Jonathan Stray, Aviv Ovadya, Claire Leibowicz, and Sam Gregory. Governing access to synthetic media detection technology. Tech Policy Press. https://www.techpolicy.press/governing-access-to-s ynthetic-media-detection-technology/, 2021. Accessed: 8 April 2025
2021
-
[59]
Human rights can be the spark of ai innovation—not stifle it.Tech Policy Press
shirin anlen. Human rights can be the spark of ai innovation—not stifle it.Tech Policy Press. https: //www.techpolicy.press/human-rights-can-be-the-spark-of-ai-innovation-not-stifle-it/ ,
-
[60]
Dadzie TV. Youtube. https://www.youtube.com/watch?v=p6VnexYxwrU , 2024. Accessed: 8 April 2025
2024
-
[61]
https://www.youtube.com/watch?v=OrFXklQz6bQ&t=2702s
GHOne TV. https://www.youtube.com/watch?v=OrFXklQz6bQ&t=2702s. Accessed: 8 April 2025
2025
-
[62]
Ai-image-detector
Matthew Maybe (umm maybe). Ai-image-detector. Hugging Face. https://huggingface.co/umm-m 20 aybe/AI-image-detector, 2022. Accessed: 8 April 2025
2022
-
[63]
Sohrawardi, Sovantharith Seng, Akash Chintha, Bao Thai, Raymond Ptucha, Matthew Wright, and Andrea Hickerson
Saniat J. Sohrawardi, Sovantharith Seng, Akash Chintha, Bao Thai, Raymond Ptucha, Matthew Wright, and Andrea Hickerson. Defaking deepfakes: Understanding journalists’ needs for deepfake detection. In Proceedings of the USENIX Symposium on Usable Privacy and Security, 2020
2020
-
[64]
Big brother watch briefing on clause 21 of the criminal justice bill.https://bills
Big Brother Watch. Big brother watch briefing on clause 21 of the criminal justice bill.https://bills. parliament.uk/publications/53817/documents/4320#:~:text=Clause%2021%20also%20allows%2 0for,powers%2C%20with%20limited%20parliamentary%20oversight.&text=policing%20needs/ . Ac...
2021
-
[65]
Big brother watch responds to facial recognition powers in new crime bill
Big Brother Watch Team. Big brother watch responds to facial recognition powers in new crime bill. Big Brother Watch. https://bigbrotherwatch.org.uk/press-releases/big-brother-watch-res ponds-to-facial-recognition-powers-in-new-crime-bill/#:~:text=The%20Bill%20allows%2 0the%20...
2025
-
[66]
Sociotechnical safety evaluation of generative ai systems.arxiv
Laura Weidinger, Maribeth Rauh, Nahema Marchal, Arianna Manzini, Lisa Anne Hendricks, Juan Mateos-Garcia, Stevie Bergman, Jackie Kay, Conor Griffin, Ben Bariach, Iason Gabriel, Verena Rieser, and William Isaac. Sociotechnical safety evaluation of generative ai systems.arxiv. h...
-
[67]
https://www.witness.org/
WITNESS. https://www.witness.org/. Accessed: 7 April 2025
2025
-
[68]
Prepare, don’t panic: Synthetic media and deepfakes.https://lab.witness.org/projec ts/synthetic-media-and-deep-fakes
WITNESS. Prepare, don’t panic: Synthetic media and deepfakes.https://lab.witness.org/projec ts/synthetic-media-and-deep-fakes. Accessed: 8 April 2025
2025
-
[69]
Outgoing US President Biden did not confess to helping orchestrate pakistan ‘regime change’
Haseem uz Zaman. Outgoing US President Biden did not confess to helping orchestrate pakistan ‘regime change’. Soch Fact Check. https://www.sochfactcheck.com/us-president-joe-biden-did-not -confess-to-regime-change-conspiracy-pakistan-army-imran-khan/ , 2024. Accessed: 8 April 2025
2024
-
[70]
@zeltzinjuareze. X. https://x.com/zeltzinjuareze/status/1762878195265638428/ , 2024. Accessed: 8 April 2025. 21 A WITNESS’TRIEDBenchmark: AChecklistforTrulyInnovative and Effective AI Detection Instructions The checklist below is adapted from the benchmark quality assessment f...
2024
-
[71]
Burkina faso: Video shows soldiers disemboweling.Human Rights Watch
Human Rights Watch. Burkina faso: Video shows soldiers disemboweling.Human Rights Watch. https: //www.hrw.org/news/2024/07/26/burkina-faso-video-shows-soldiers-disemboweling-body ,
2024
-
[76]
A community-based approach to visual verification to fortify the truth.WITNESS Guide
WITNESS. A community-based approach to visual verification to fortify the truth.WITNESS Guide. https://library.witness.org/product/community-based-approaches-to-verification// ,
-
[77]
Accessed: 8 April 2025
2025
-
[79]
The goals, key concepts, primary features, and target audience for the detection tool are clearly outlined
-
[80]
The design process actively involves input from domain experts with relevant expertise
-
[81]
Relevant academic research, industry standards, and existing literature are thoroughly reviewed and integrated into the design
-
[82]
Real-world scenarios and practical use cases are incorporated to guide the tool’s development and application
-
[83]
The mechanisms to ensure the tool’s durability and adaptability to rapid development of synthetic media are defined and prioritized in the design process
-
[84]
Checklist
The tool’s funding allows for a responsible and sustainable development and operation of the tool. Checklist
-
[85]
1.1 The intended audience is clearly defined
The goals, concepts, characteristics, and audience of the detection tool are defined. 1.1 The intended audience is clearly defined. □TO DO □YES □NO □N/A Justification: 1.2 The use cases on which the tool should be used are clearly described. □TO DO □YES □NO □N/A Justification:...
-
[86]
□TO DO □YES □NO □N/A Justification:
The design process involved consultation with diverse domain experts, including those from global and underserved contexts. □TO DO □YES □NO □N/A Justification:
-
[87]
3.1 The research and literature were diverse and steps were taken to include knowledge produced 23 outside of Europe and the United States
Diverse relevant academic research, industry reports, and existing literature were integrated into the tool’s design. 3.1 The research and literature were diverse and steps were taken to include knowledge produced 23 outside of Europe and the United States. □TO DO □YES □NO □N/...
-
[88]
4.1 Real-life cases and use cases were incorporated to reflect practical challenges and expectations
Real-world scenarios and practical use cases are incorporated to guide the tool’s development and application. 4.1 Real-life cases and use cases were incorporated to reflect practical challenges and expectations. □TO DO □YES □NO □N/A Justification: 4.2 Stakeholders representat...
-
[89]
5.1 Specific steps are outlined to ensure that the tool is future-proof and guarantee that it will remain relevant in the long term
The mechanisms to ensure the tool’s durability and adaptability to rapid development of synthetic media are defined. 5.1 Specific steps are outlined to ensure that the tool is future-proof and guarantee that it will remain relevant in the long term. □TO DO □YES □NO □N/A Justif...
-
[90]
The development stage requires explicit consideration of underrepresented languages, cultural nuances, and geopolitical challenges in both training and testing
-
[91]
Trainingdataisdiverse, representativeofglobaldemographics, andcollectedethically, withtransparent documentation of the sourcing process
-
[92]
Comprehensiveandaccessibledocumentationisprovided, integratingrelevantcontexttoaidunderstanding and usability
-
[93]
The tool’s limitations are clearly identified and communicated to ensure realistic expectations of its capabilities
-
[94]
Mechanisms for explainability are incorporated, ensuring that results and detections are interpretable by both technical and non-technical users
-
[95]
The development team is diverse, reflecting the perspectives and needs of the tool’s intended audience. 24
-
[96]
Accessibility is prioritized, ensuring the tool is usable by the intended audience regardless of technical expertise or resource constraints
-
[97]
The tool’s capabilities for handling various file types and quality levels are explicitly detailed
-
[98]
The tool demonstrates consistent performance across diverse contexts and use cases it aims to serve
-
[99]
Development respects and integrates with existing verification techniques and skill sets to enhance reliability and usability
-
[100]
Checklist
Adaptability is a key focus, allowing the tool to evolve in response to new challenges, use cases, and technological advancements. Checklist
-
[101]
1.1 Data collection process complied with the local data protection regulations
Training data is diverse, representative of global demographics, and ethically sourced. 1.1 Data collection process complied with the local data protection regulations. □TO DO □YES □NO □N/A Justification: 1.2 The training dataset is representative of diverse demographics. □TO ...
-
[102]
□TO DO □YES □NO □N/A Justification:
The tool’s capabilities for handling various file types and quality levels are explicitly detailed. □TO DO □YES □NO □N/A Justification:
-
[103]
3.1 The tool was trained on low-resolution content
The training dataset included examples of corrupted synthetic content. 3.1 The tool was trained on low-resolution content. □TO DO □YES □NO □N/A Justification: 3.2 The tool was trained on social media compression standards. □TO DO □YES □NO □N/A Justification: 3.3 The tool was t...
-
[104]
4.1 If the tool is capable of detecting audio, the training dataset included data with a transmission similar to telephones
The detection tool was trained to deal with diverse types of content. 4.1 If the tool is capable of detecting audio, the training dataset included data with a transmission similar to telephones. □TO DO □YES □NO □N/A Justification: 4.2 If the tool is capable of detecting audio,...
-
[105]
5.1 A system is in place, and adaptable to adjust the detection model to continually update the training dataset to include new examples of AI-generated or manipulated content
The tool is developed to scale across different levels of content complexity, size, and volume. 5.1 A system is in place, and adaptable to adjust the detection model to continually update the training dataset to include new examples of AI-generated or manipulated content. □TO ...
-
[106]
The tool provides users with clear information on how detection results should be interpreted. 6.1 The tool does not use binary labels (such as ‘real’ or ‘fake’) when communicating the results but instead describes manipulation using accessible description and language of degr...
-
[107]
7.1 The information is provided with respect to other verification techniques and skill sets that could be integrated
The tool acknowledges other existing verification techniques and incorporates information about them into its workflow. 7.1 The information is provided with respect to other verification techniques and skill sets that could be integrated. □TO DO □YES □NO □N/A Justification: 7....
-
[108]
□TO DO □YES □NO □N/A Justification: 27 The maximum number of points is 57
The tool is adaptable with processes in place to evolve in response to new challenges, use cases, and technological advancements. □TO DO □YES □NO □N/A Justification: 27 The maximum number of points is 57. Your score is: Benchmark Testing Testing Criteria
-
[109]
Proactive measurements are being taken to test the tool’s blind spots
-
[110]
The tool undergoes regular and systematic testing to maintain reliability and effectiveness
-
[111]
Diverse stakeholders and domain experts, including independent external groups, actively participate in the testing process to provide varied perspectives and expertise
-
[112]
The tool is rigorously tested on challenging edge cases, including adversarial content, false claims of AI-generated media, and heavily manipulated content
-
[113]
The tool demonstrates resilience against evasion techniques, maintaining its accuracy and reliability even under deliberate attempts to bypass detection
-
[114]
Checklist
The tool is evaluated from a human-detection perspective. Checklist
-
[115]
1.1 The tool is proactively tested to identify blind spots and potential failures
The tool was tested in an exhaustive and inclusive manner. 1.1 The tool is proactively tested to identify blind spots and potential failures. □TO DO □YES □NO □N/A Justification: 1.2 The tool’s robustness was tested against adversarial attacks and evasion techniques. □TO DO □YE...
-
[116]
2.1 Such testing included assessing how much time it took to receive the result and how difficult the process was from the human perspective
The tool is evaluated from a human-assisted detection perspective. 2.1 Such testing included assessing how much time it took to receive the result and how difficult the process was from the human perspective. □TO DO □YES □NO □N/A Justification: 2.2 Such testing included a comp...
-
[117]
The tool is designed and implemented to uphold human rights, ensuring it does not infringe on privacy, freedom of expression, or other fundamental rights
-
[118]
Accessibility is prioritized, ensuring the tool is usable by its intended audience, regardless of technical expertise or resource availability
-
[119]
Checklist
Documentation is clear and comprehensive, and top level version is easily accessible, providing users with the necessary guidance to operate the tool effectively and understand its limitations. Checklist
-
[120]
1.1 Thetoolexplicitlyavoidsoutputsorrecommendationsthatcouldharmindividualsorcommunities, aligning with ethical AI principles
Ensure the tool does not inadvertently violate human rights, such as privacy or freedom of expression. 1.1 Thetoolexplicitlyavoidsoutputsorrecommendationsthatcouldharmindividualsorcommunities, aligning with ethical AI principles. □TO DO □YES □NO □N/A Justification: 1.2 The too...
-
[121]
2.1 The tool includes educational resources or tutorials to help users understand its functionality and limitations
The tool is accessible to a diverse group of targeted users. 2.1 The tool includes educational resources or tutorials to help users understand its functionality and limitations. □TO DO □YES □NO □N/A Justification: 2.2 The tool does not require advanced technical knowledge or s...
-
[122]
3.1 Thetoolcommunicateserrorsoruncertaintiesclearlytousers, avoidingoverconfidenceinambiguous cases
The tool’s performance metrics are accurate and not exaggerated. 3.1 Thetoolcommunicateserrorsoruncertaintiesclearlytousers, avoidingoverconfidenceinambiguous cases. □TO DO □YES □NO □N/A Justification: 30 The maximum number of points is 33. Your score is: Benchmark Maintenance...
-
[123]
The tool undergoes regular performance evaluations to ensure its detection accuracy remains reliable across evolving datasets and content types
-
[124]
Resources are actively being distributed for updates and maintenance processes
-
[125]
Updatesareroutinelyimplementedtoaddressemergingchallenges, suchasadversarialattacks, advancements in deepfake technologies, and new forms of synthetic content
-
[126]
User feedback is actively collected, documented, and incorporated into updates, with a transparent and accessible communication channel available for issue reporting and suggestions
-
[127]
Clear policies are established regarding the support duration for older versions of the tool following updates or new releases
-
[128]
Periodic audits are conducted to identify and mitigate any biases introduced during updates or revealed by newer datasets
-
[129]
Checklist
Developers continuously monitor advancements in AI and related technologies, ensuring the tool remains aligned with the latest innovations and best practices. Checklist
-
[130]
1.1 The tool’s accuracy is evaluated against previously verified cases
The tool undergoes regular updates, and changes are documented transparently. 1.1 The tool’s accuracy is evaluated against previously verified cases. □TO DO □YES □NO □N/A Justification: 1.2 Clear policies are in place for how regularly updates will be conducted. □TO DO □YES □N...
-
[131]
2.1 Feedback channels for users are actively maintained, and updates incorporate user feedback
An accessible channel of communication is provided. 2.1 Feedback channels for users are actively maintained, and updates incorporate user feedback. □TO DO □YES □NO □N/A Justification: 2.2 A contact person or team for support and inquiries is clearly listed. □TO DO □YES □NO □N/...
-
[132]
3.1 The tool will be retired if it has not been updated for a specified time following clear policies
The tool has End-of-Life procedures in place. 3.1 The tool will be retired if it has not been updated for a specified time following clear policies. □TO DO □YES □NO □N/A Justification: 3.2 If the tool is to be discontinued, clear communication is provided to users, along with ...
-
[2021]
arXiv:2007.02407 [cs.LG]
2007 arXiv
-
[2024]
https://arxiv.org/abs/2503.02857/, 2025
preprint arXiv. https://arxiv.org/abs/2503.02857/, 2025. arXiv:2503.02857 [cs.CV]
2025 arXiv
Reviewed August 16, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.