Pith. sign in

REVIEW 4 major objections 5 minor 127 references

TRIED: Truly Innovative and Effective AI Detection Benchmark, developed by WITNESS

T0 review · 4 major / 5 minor · reviewed 2026-08-16 · deepseek-v4-flash

Pith's one-line read The paper claims a 177-point self-assessment checklist can classify AI detection tools by real-world effectiveness, with scores above 143 meaning 'truly effective'.

desk verdict The TRIED checklist is a solid sociotechnical contribution, but the 177-point scoring system and the '143 = truly effective' cutoff are arbitrary and need recalibration or reframing. read the letter →

arxiv 2504.21489 v2 pith:NENS4J4Y submitted 2025-04-30 cs.CY cs.AI

classification cs.CYcs.AI
keywords AIdetectiondeepfakessyntheticmediasociotechnicalevaluationTRIEDBenchmarkinformationintegrityexplainabilityalgorithmicfairness
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

AI detection tools are usually judged on accuracy and speed, but this report argues those metrics miss why tools fail in practice: poor-quality recordings, unfamiliar languages, unclear results, cost, and unfair performance. The paper's central claim is that a tool's real-world effectiveness can be assessed by the TRIED Benchmark, a 177-point checklist covering design, development, testing, implementation, and maintenance. On that scale, a score above 143 is said to indicate that a tool is 'truly effective'; lower bands mark it moderately, somewhat, or not effective. A sympathetic reader should care because the checklist turns frontline experience with deepfake detection into concrete, auditable criteria that developers, regulators, and fact-checkers could use.

What carries the argument

The central mechanism is the TRIED Benchmark checklist, a 177-point assessment instrument whose items ask developers to answer 'Yes', 'No', 'To do', or 'N/A' with a one-sentence justification. Each 'Yes' earns 3 points, a justified 'To do' earns 2, a justified 'No' earns 1, and unjustified or empty answers earn 0. The checklist's work is to convert qualitative sociotechnical lessons, drawn from frontline cases involving low-resolution video, noisy audio, underrepresented languages, and false accusations of AI use, into a single numeric score with explicit bands for effectiveness, so that 'truly effective' has a public, checkable meaning.

What would settle it

Take a set of detection tools that score above 143 on the checklist, run them on a corpus of low-resolution, noisy, and non-English manipulated and authentic media, and compare their error rates with tools scoring below 71; if the high-scoring tools are not systematically more accurate on those real-world cases, the score-to-effectiveness mapping is contradicted.

Watch

Extended reading notes

Core claim

The paper's discovery is a definition of 'truly innovative and effective' that is grounded in six interconnected pillars: handling real-world media conditions, transparency and explainability, accessibility, fairness, durability, and integration into broader verification workflows. The TRIED Benchmark operationalizes these pillars as a scored checklist with 177 points, distributed across lifecycle stages: design, development, testing, implementation, and maintenance. A self-assessed score above 143 means the tool is 'truly effective'; scores of 107 to 142 are moderately effective, 72 to 106 somewhat effective, and below 71 not effective. The claim is that a tool passing this checklist is one that supports the people most exposed to deceptive AI rather than merely performing well on clean benchmark data.

Load-bearing premise

The load-bearing premise is that developers' self-reported answers, each backed by a single sentence, faithfully describe how a tool actually behaves, and that the chosen cutoff of 143 is a meaningful separation rather than an arbitrary number.

Editorial extensions

If this is right

  • Detection developers can use the checklist to find gaps before release, such as missing training data for compressed social-media formats or missing multilingual support.
  • Standards bodies and regulators could adopt the threshold bands as a baseline for procurement or certification of detection tools.
  • Evaluation of detection tools would broaden from accuracy-only metrics to include explainability, accessibility, fairness, durability, and fit within verification workflows.
  • Fact-checkers and civil-society users would gain a common vocabulary for comparing tools beyond vendor claims.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Editorial inference: the 143-point threshold has not been calibrated against measured detection performance, so the same checklist could be tested by scoring a set of tools and then checking their actual outcomes on real-world cases.
  • Editorial inference: rewarding 'to do with justification' almost as much as 'yes' means a roadmap can count nearly as much as a finished capability; a stricter scoring variant might separate planning from delivery.
  • Editorial inference: the checklist could be extended to independent third-party assessment, where reviewers rather than developers answer the questions, or to weighted scoring for specific user groups such as human-rights defenders versus general audiences.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. This WITNESS report argues that AI detection tools are typically evaluated on narrow technical metrics and therefore fail in real-world use, and it proposes the TRIED Benchmark as a remedy: a 177-point checklist organized into six pillars (real-world design, transparency/explainability, accessibility, fairness, durability, and integration with verification ecosystems). The qualitative argument is grounded in WITNESS's Deepfakes Rapid Response Force (DRRF) casework, global consultations, and external work such as Deepfake-Eval-2024. The report claims that a score above 143 points on the Annex A checklist indicates a tool is 'truly effective,' with bands for moderately, somewhat, and not effective. The main body also offers recommendations for developers, regulators, standards bodies, and governments. The central quantitative claim, however, rests entirely on Annex A's scoring rubric, which is introduced without calibration, pilot testing, inter-rater reliability analysis, or comparison with external detection outcomes, and which contains apparent internal inconsistencies in the point accounting.

Significance. If the qualitative framework were adopted, it would make a useful contribution by expanding AI-detection evaluation beyond accuracy metrics toward accessibility, fairness, explainability, and workflow integration. The report's documentation of low-quality media, language barriers, false-positive harms, and the importance of human-assisted verification is valuable and well illustrated by concrete DRRF cases. The alignment with existing trustworthy-AI frameworks (EU AI Act, NIST, PAI) provides useful grounding. However, the benchmark's quantitative scoring system is the paper's headline deliverable, and that system is currently unsupported by validation data and is internally inconsistent in its arithmetic. The 'truly effective' threshold is precisely the kind of actionable claim that stakeholders would use in procurement or policy decisions, so the lack of support is load-bearing rather than cosmetic.

major comments (4)
  1. [Annex A (scoring instructions and thresholds)] The claim that 'a score above 143 points indicates that your tool is truly effective' is unsupported by any calibration, pilot testing, inter-rater reliability analysis, or comparison against external measures of detection performance. The point weights (Yes=3, To Do with justification=2, No with justification=1, otherwise 0) and the 143/107/72 cutoffs are asserted rather than derived or fitted, so the report's central quantitative claim is not backed by evidence. Given that the abstract and Executive Summary present this threshold as actionable guidance, the authors should either validate the scale or explicitly reframe the score as an unvalidated self-assessment instrument.
  2. [Annex A (Development checklist, maximum points)] The Development section of the checklist contains 22 items, including conditional items 3.4, 3.5, 4.1, 4.2, and 4.3. If all applicable items are counted, the maximum for that section is 22×3=66 points, yet the text states a maximum of 57 points. This makes the stated total maximum of 177 points, and therefore the 143-point threshold, depend on an unexplained exclusion of applicable items; the numerical scale is internally inconsistent for a multimodal tool that honestly answers all conditionals.
  3. [Annex A (answer rubric)] The instructions allow N/A as a response, but the scoring table specifies points only for Yes, To Do with or without justification, and No with or without justification; it does not state how N/A is counted. If N/A receives 0 points, tools with a narrower intended scope are penalized relative to general-purpose tools. If N/A is excluded from the denominator, the maximum possible score varies by tool and the fixed thresholds 143/107/72 are not well-defined. Either interpretation breaks the comparability that the fixed thresholds presuppose.
  4. [Sections 4 and 7, Annex A] The benchmark is presented as grounded in the same DRRF casework and WITNESS consultations that are then used as the primary evidence for its validity; no independent validation of the checklist items' content validity or of the thresholds against external assessments is provided. This does not undermine the qualitative lessons from the case studies, but it means the paper does not demonstrate that the benchmark measures 'real-world effectiveness' rather than the authors' own design priorities.
minor comments (5)
  1. [Abstract] The abstract contains missing spaces in 'stakeholderscandriveinnovation, safeguardpublictrust, strengthenAIliteracy, andcontribute' and in 'consid erations'; these should be fixed.
  2. [Sections 3 and 4] Section 3 begins 'IntroductionDeceptiveAIcontent' with a missing space after the heading, and Section 4 uses 'Al' instead of 'AI' in 'evaluates Al through a sociotechnical lens.'
  3. [References] Several references are incomplete or malformed: [19] and [32] lack closing brackets in their arXiv identifiers, and references such as [10], [18], [35], [37], [39], [40], [46], [48], [52], and [70] are bare social-media URLs without author, title, or access-date information, which makes source verification difficult.
  4. [Section 6.4 and Section 2] Section 6.4 contains typos ('funs' for 'funds', 'prioritze' for 'prioritize'), and Section 2 defines TRIED as 'Truly Innovative and Effective Detection' while the title and abstract include 'AI Detection'; the acronym expansion should be consistent.
  5. [Annex A (scoring table)] The scoring table gives 0 points for both 'To do without justification' and 'No without justification,' which is equivalent to not answering at all; the intended distinction should be made explicit.

Circularity Check

1 steps flagged · score 6.0 of 10

The 'truly effective' cutoff is a self-imposed definition, not a derived result.

  1. self definitional [Annex A, 'A WITNESS’ TRIED Benchmark: A Checklist for Truly Innovative and Effective AI Detection', final score bands]
    "The Total Maximum Number of Points is 177. A score above 143 points indicates that your tool is truly effective. A score between 107 and 142 points indicates that your tool is moderately effective. A score between 72 and 106 points indicates that your tool is somewhat effective. A score below 71 points indicates that your tool is not effective."

    The predicate 'truly effective' is defined by the authors' own score bands in the same annex. The report provides no calibration, pilot testing, inter-rater reliability analysis, or outcome-data link connecting 143/177 (80.8%) to real-world effectiveness. Thus the central claim that a score above 143 indicates a tool is 'truly effective' is a restatement of the chosen threshold, not a result derived from evidence. The label is attached to the score by definition, so the 'indication' is tautological.

full rationale

The report's qualitative sociotechnical framework is not circular: it is grounded in external literature (EU AI Act, NIST, OECD, PAI, BetterBench, Deepfake-Eval-2024) and illustrated by concrete DRRF case examples. Those sections stand as independent content. The circularity is localized to the benchmark's quantitative scoring claim. The point weights (Yes=3, To do with justification=2, No with justification=1, otherwise 0) and the effectiveness bands (143/107/72) are asserted in Annex A without calibration, piloting, inter-rater agreement, or comparison against actual detection outcomes. The final statement that a score above 143 points indicates a tool is 'truly effective' reduces, by construction, to the authors' own stipulation in the same annex. Frequent self-citations to WITNESS reports and blogs support background motivation and are not load-bearing for this threshold, so they do not independently raise the circularity score beyond the definitional issue.

Assumptions & free parameters 2 free parameters · 3 assumptions · 1 invented entities

The central claim rests on three kinds of pulled-in support: hand-chosen scoring weights and thresholds, an unproven assumption that self-reporting is reliable, and an assumption that WITNESS's own case experience is representative. The TRIED Benchmark is an invented artifact with no independent validation.

free parameters (2)
  • Point weights for checklist answers = Yes=3, To do with justification=2, No with justification=1, all other responses=0
    Introduced in Annex A scoring instructions; no empirical calibration or theoretical derivation is provided for these weights.
  • Effectiveness thresholds = 144-177 truly effective, 107-142 moderately effective, 72-106 somewhat effective, below 72 not effective
    Asserted at the end of Annex A; there is no validation against observed tool performance, no pilot data, and no sensitivity analysis.
assumptions (3)
  • domain assumption The six pillars are the complete set of relevant effectiveness dimensions
    Sections 5.1 through 5.6 define effectiveness solely via these six pillars, without a formal derivation or an empirically grounded argument for exhaustiveness.
  • domain assumption Self-reported one-sentence justifications are reliable evidence of a tool's properties
    Annex A instructs that each question 'should be answered with a Yes, No, TODO or N/A, followed by a concise justification', and these self-assessments directly determine the score.
  • domain assumption Anecdotal DRRF cases generalize across global contexts
    The benchmark items are derived from WITNESS's Deepfakes Rapid Response Force casework and consultations (Sections 5.1-5.6); no systematic sampling or external replication is presented.
invented entities (1)
  • TRIED Benchmark
    purpose: A checklist-based instrument to evaluate AI detection tools on six sociotechnical pillars and assign an effectiveness category
    Introduced in this paper; no independent validation of its predictive or concurrent validity is provided outside the authors' own casework and self-assessment instructions.

how reviews work

0 comments
Cite this review

Pith. "Pith review of TRIED: Truly Innovative and Effective AI Detection Benchmark, developed by WITNESS." pith.science (2026). https://pith.science/paper/NENS4J4Y

@misc{pith2026250421489,
  author       = {Pith},
  title        = {Pith review of: TRIED: Truly Innovative and Effective AI Detection Benchmark, developed by WITNESS},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/NENS4J4Y}},
  note         = {Machine review of arXiv:2504.21489}
}
read the original abstract

The proliferation of generative AI and deceptive synthetic media threatens the global information ecosystem, especially across the Global Majority. This report from WITNESS highlights the limitations of current AI detection tools, which often underperform in real-world scenarios due to challenges related to explainability, fairness, accessibility, and contextual relevance. In response, WITNESS introduces the Truly Innovative and Effective AI Detection (TRIED) Benchmark, a new framework for evaluating detection tools based on their real-world impact and capacity for innovation. Drawing on frontline experiences, deceptive AI cases, and global consultations, the report outlines how detection tools must evolve to become truly innovative and relevant by meeting diverse linguistic, cultural, and technological contexts. It offers practical guidance for developers, policy actors, and standards bodies to design accountable, transparent, and user-centered detection solutions, and incorporate sociotechnical considerations into future AI standards, procedures and evaluation frameworks. By adopting the TRIED Benchmark, stakeholders can drive innovation, safeguard public trust, strengthen AI literacy, and contribute to a more resilient global information credibility.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

127 extracted references · 73 canonical work pages

  1. [1]

    Accessed: 7 April 2025

    European Union Artificial Intelligence Act.https://artificialintelligenceact.eu/article/1/ . Accessed: 7 April 2025

  2. [2]

    Audio atribuido a margarita gonzález sobre programas sociales en morelos tiene indicios de manipulación digital.Animal Politico

    Samedi Aguirre. Audio atribuido a margarita gonzález sobre programas sociales en morelos tiene indicios de manipulación digital.Animal Politico. https://www.animalpolitico.com/verificacion-de-hec hos/desinformacion/audio-candidata-morena-margarita/, 2024. Accessed: 8 April 2025

  3. [3]

    Keep your ai claims in check.Internet Archive

    Michael Atleson. Keep your ai claims in check.Internet Archive. https://web.archive.org/web/20 250115031325/https://www.ftc.gov/business-guidance/blog/2023/02/keep-your-ai-claims-c heck/, 2023. Accessed: 8 April 2025

  4. [4]

    Tomorrow’s great digital divide: Content with or without provenance.WITNESS Blog

    Jacobo Castellanos. Tomorrow’s great digital divide: Content with or without provenance.WITNESS Blog. https://blog.witness.org/2025/03/tomorrows-great-digital-divide , 2025. Accessed: 7 April 2025

  5. [5]

    Reducing risks posed by synthetic content an overview of technical approaches to digital content transparency

    Bilva Chandra, Jesse Dunietz, and Kathleen Roberts. Reducing risks posed by synthetic content an overview of technical approaches to digital content transparency. Technical report, National Institute of Standards and Technology, 2024

  6. [6]

    Deepfake-eval-2024: A multi-modal in-the-wild benchmark of deepfakes circulated in

    Nuria Alina Chandra, Ryan Murtfeldt, Lin Qiu, Arnab Karmakar, Hannah Lee, Emmanuel Tanumihardja, Kevin Farhat, Ben Caffee, Sejin Paik, Changyeon Lee, Jongwook Choi, Aerin Kim, and Oren Etzioni. Deepfake-eval-2024: A multi-modal in-the-wild benchmark of deepfakes circulated in

  7. [7]

    An Indian politician says scandalous audio clips are AI deepfakes

    Nilesh Christopher. An Indian politician says scandalous audio clips are AI deepfakes. we had them tested. Rest of World. https://restofworld.org/2023/indian-politician-leaked-audio-ai-dee pfake/, 2023. Accessed: 8 April 2025

  8. [8]

    Accessed: 7 April 2025

    Shakti: India Fact-Checking Collective.https://projectshakti.in/. Accessed: 7 April 2025

Show all 127 references
  1. [9]

    Audio-visual person-of- interest deepfake detection

    Davide Cozzolino, Alessandro Pianese, Matthias Nießner, and Luisa Verdoliva. Audio-visual person-of- interest deepfake detection. arXiv. https://arxiv.org/abs/2204.03083/ , 2023. arXiv:2204.03083 [cs.CV]

  2. [10]

    @dahrinoor2. X. https://x.com/dahrinoor2/status/1771155215414067603/ , 2024. Accessed: 7 April 2025

  3. [11]

    Accessed: 7 April 2025

    Synthetic Media Deepfakes and Generative AI.https://www.gen-ai.witness.org/?pk_vid=e9bdde b5077398fa1745328087fd83d2/. Accessed: 7 April 2025

  4. [12]

    Facebook

    Facebook. Facebook. https://www.facebook.com/100002167747651/videos/457953553957488/?vh= e&extid=MSG-UNK-UNK-UNK-COM_GK0T-GK1C/ , 2024. Accessed: 8 April 2025

  5. [13]

    Accessed: 8 April 2025

    Coalition for Content Provenance and Authenticity.https://c2pa.org/. Accessed: 8 April 2025

  6. [14]

    C2PA Harms Modelling 1.4

    Coalition for Content Provenance and Authenticity. C2PA Harms Modelling 1.4. https://partne rshiponai.org/resource/glossary-for-synthetic-media-transparency-methods-part-1/ . Accessed: 8 April 2025

  7. [15]

    https://www.gen-ai.witness.org/deepfakes-rapid-respons e-force

    Deepfakes Rapid Response Force. https://www.gen-ai.witness.org/deepfakes-rapid-respons e-force. Accessed: 7 April 2025

  8. [16]

    Getting to the source: Understanding metadata removal on social media.Magnet Forensics Blog

    Magnet Forensics. Getting to the source: Understanding metadata removal on social media.Magnet Forensics Blog. https://www.magnetforensics.com/blog/getting-to-the-source-understanding -metadata-removal-on-social-media// , 2024. Accessed: 8 April 2025. 17

  9. [17]

    Disconnected from reality: American voters grapple with ai and flawed osint strategies.Institute for Strategic Dialogue

    Isabelle Frances-Wright, Ellen Jacobs, and Ella Meyer. Disconnected from reality: American voters grapple with ai and flawed osint strategies.Institute for Strategic Dialogue. https://www.isdglobal. org/digital_dispatches/disconnected-from-reality-american-voters-grapple-with-...

  10. [18]

    @GazetteNGR. X. https://x.com/GazetteNGR/status/1642218315685699586/ , 2023. Accessed: 8 April 2025

  11. [19]

    Datasheets for datasets.arXiv

    Timnit Gebru, Jamie Morgenstern, Briana Vecchione, Jennifer Wortman Vaughan, Hanna Wallach, Hal Daumé III, and Kate Crawford. Datasheets for datasets.arXiv. https://arxiv.org/abs/1803.0 9010/, 2021. arXiv:1803.09010 [cs.DB

  12. [20]

    Posts misrepresent a photo of a ukrainian soldier balancing on his prosthetic limbs.AP News

    Melissa Goldin. Posts misrepresent a photo of a ukrainian soldier balancing on his prosthetic limbs.AP News. https://apnews.com/article/fact-check-ukrainian-soldier-photo-nazi-salute-863 799988341?fbclid=IwAR0w4ci7kN4tp0RKyfbrj_xH7VrxJxvckoY7PeAQ67WMZslobZZ5OCpHHLU/ , 2024. Ac...

  13. [21]

    Sam Gregory. What’s needed in deepfakes detection? insights from witness’ global preparedness work and the partnership on ai’s steerco on media integrity work on the deepfake detection challenge. WITNESS Blog. https://blog.witness.org/2020/04/whats-needed-deepfakes-detection/ ,

  14. [22]

    Pre-empting a crisis: Deepfake detection skills + global access to media forensics tools

    Sam Gregory. Pre-empting a crisis: Deepfake detection skills + global access to media forensics tools. WITNESS Blog. https://blog.witness.org/2021/07/deepfake-detection-skills-tools-acces s, 2021. Accessed: 7 April 2025

  15. [23]

    The world needs deepfake experts to stem this chaos.Wired

    Sam Gregory. The world needs deepfake experts to stem this chaos.Wired. https://www.wired.co m/story/opinion-the-world-needs-deepfake-experts-to-stem-this-chaos/ , 2021. Accessed: 8 April 2025

  16. [24]

    Grother, Mei L

    Patrick J. Grother, Mei L. Ngan, and Kayee K. Hanaoka. Face recognition vendor test part 3: Demographic effects, nist interagency.Internal Report, National Institute of Standards and Technology. https://doi.org/10.6028/NIST.IR.8280/, 2019. Accessed: 8 April 2025

  17. [25]

    Youtube data viewer.https://citizenevidence.org/2014/07/01/youtube -dataviewer/

    Amnesty International. Youtube data viewer.https://citizenevidence.org/2014/07/01/youtube -dataviewer/. Accessed: 8 April 2025

  18. [26]

    Ticks or it didn’t happen: Key dilemmas in building authenticity infrastructure for multimedia

    Gabi Ivens and Sam Gregory. Ticks or it didn’t happen: Key dilemmas in building authenticity infrastructure for multimedia. WITNESS Report. https://lab.witness.org/ticks-or-it-did nt-happen/#references/, 2019. Accessed: 8 April 2025

  19. [27]

    @johnnygould. X. https://x.com/jonnygould/status/1768345232184148139/ , 2024. Accessed: 8 April 2025

  20. [28]

    Chen, and Siwei Lyu

    Yan Ju, Shu Hu, Shan Jia, George H. Chen, and Siwei Lyu. Improving fairness in deepfake detection. preprint arXiv. https://arxiv.org/abs/2306.16635/, 2023. arXiv:2306.16635 [cs.CV]

  21. [29]

    Adversarial machine learning in the context of network security: Challenges and solutions.Journal of Computational Intelligence and Robotics, 4(1):51–63, march 2024

    Muskan Khan and Laiba Ghafoor. Adversarial machine learning in the context of network security: Challenges and solutions.Journal of Computational Intelligence and Robotics, 4(1):51–63, march 2024

  22. [30]

    Le, Jiwon Kim, Simon S

    Binh M. Le, Jiwon Kim, Simon S. Woo, Kristen Moore, Alsharif Abuadbba, and Shahroz Tariq. Sok: Systematization and benchmarking of deepfake detectors in a unified framework.preprint arXiv. https: //arxiv.org/abs/2401.04364/, 2025. arXiv:2401.04364 [cs.CV]

  23. [31]

    What is facial recognition technology?https://www.ajl.org/facial-r ecognition-technology/

    Algorithmic Justice League. What is facial recognition technology?https://www.ajl.org/facial-r ecognition-technology/. Accessed: 8 April 2025. 18

  24. [32]

    Leibowicz

    Claire R. Leibowicz. Regulating reality: Exploring synthetic media through multistakeholder ai governance. arXiv. https://arxiv.org/abs/2502.04526/, 2025. arXiv:2502.04526 [cs.CY

  25. [33]

    Fortifying the truth in the age of synthetic media and generative ai perspectives from africa.WITNESS Blog

    Raquel Vazquez Llorente, Jacobo Castellanos, and Nkem Agunwa. Fortifying the truth in the age of synthetic media and generative ai perspectives from africa.WITNESS Blog. https://blog.witness .org/2023/05/generative-ai-africa/, 2023. Accessed: 7 April 2025

  26. [34]

    Deepfake detection improves when using algorithms that are more aware of demographic diversity

    Siwei Lyu and Yan Ju. Deepfake detection improves when using algorithms that are more aware of demographic diversity. NiemanLab. https://www.niemanlab.org/2024/04/deepfake-detection -improves-when-using-algorithms-that-are-more-aware-of-demographic-diversity/ , 2023. Accessed:...

  27. [35]

    Archived from X

    @mini_razdan10. Archived from X. https://drive.google.com/file/d/1XVRz3vVikXX9ttMTLZlQ7 JOeUYZJQn9k/view?usp=sharing/, 2024. Accessed: 7 April 2025

  28. [36]

    Model cards for model reporting

    Margaret Mitchell, Simone Wu, Andrew Zaldivar, Parker Barnes, Lucy Vasserman, Ben Hutchinson, Elena Spitzer, Inioluwa Deborah Raji, and Timnit Gebru. Model cards for model reporting. In Proceedings of the Conference on Fairness, Accountability, and Transparency, page 220–229. ...

  29. [37]

    MPvenezolano. Youtube. https://www.youtube.com/watch?v=T8RE-8OfFp4, 2024. Accessed: 8 April 2025

  30. [38]

    Daragh Murray. Police use of retrospective facial recognition technology: A step change in surveillance capability necessitating an evolution of the human rights law framework.The Modern Law Review, 87(4):833–863, dec 2023

  31. [39]

    @mustafali2001. Tik Tok. https://www.tiktok.com/@muss.mil/video/7438149251449752850/ ,

  32. [40]

    Facebook

    Solid PH News. Facebook. https://www.facebook.com/solidphnews1/videos/501528382261724/ ,

  33. [41]

    OECD AI Principles overview.https://oecd.ai/en/ai-principles/

    OECD AI Policy Observatory. OECD AI Principles overview.https://oecd.ai/en/ai-principles/. Accessed: 8 April 2025

  34. [43]

    Ethics guidelines for trustworthy ai.European Commission

    High-Level Expert Group on AI. Ethics guidelines for trustworthy ai.European Commission. https: //digital-strategy.ec.europa.eu/en/library/ethics-guidelines-trustworthy-ai/ , 2019. Accessed: 8 April 2025

  35. [44]

    Accessed: 7 April 2025

  36. [45]

    Building a glossary for synthetic media transparency methods part 1: Indirect disclosure

    Partnership on AI. Building a glossary for synthetic media transparency methods part 1: Indirect disclosure. Partnership on AI. https://partnershiponai.org/resource/glossary-for-synthetic -media-transparency-methods-part-1/ , 2023. Accessed: 8 April 2025

  37. [46]

    AI Risks and Trustworthiness.https://airc.nist

    National Institute of Standards and Technology. AI Risks and Trustworthiness.https://airc.nist. gov/airmf-resources/airmf/3-sec-characteristics/#:~:text=NIST%20has%20identified%20t hree%20major,statistical%2C%20and%20human%2Dcognitive/, 2019. Accessed: 8 April 2025

  38. [47]

    Archived from X

    @PedroKonductaz. Archived from X. https://drive.google.com/file/d/1sN9OYILspCUaMf83LroQq 19 suqgLW4GRhT/view?usp=sharing/, 2024. Accessed: 8 April 2025

  39. [48]

    The deepfake detection challenge: Insights and recommendations for AI and media integrity

    Partnership on AI. The deepfake detection challenge: Insights and recommendations for AI and media integrity. Partnership on AI Report. https://partnershiponai.org/a-report-on-the-deepfake-d etection-challenge//, 2020. Accessed: 8 April 2025

  40. [49]

    Determining trustworthiness through provenance and context

    Google Public Policy. Determining trustworthiness through provenance and context. Google Policy Paper. https://static.googleusercontent.com/media/publicpolicy.google/en//resources/d etermining_trustworthiness_en.pdf/, 2024. Accessed: 8 April 2025

  41. [50]

    UTV Ghana Online. Youtube. https://www.youtube.com/watch?v=K1KYhenW3oM&t=596s , 2024. Accessed: 8 April 2024

  42. [52]

    Archived from Tik Tok

    @phenixsaqartvelo. Archived from Tik Tok. https://drive.google.com/file/d/1EaB__rClt_pQCjH 8M1StmNejNspWvSaA/view?usp=sharing/, 2024. Accessed: 7 April 2025

  43. [53]

    Faceforensics++: Learning to detect manipulated facial images

    Andreas Rössler, Davide Cozzolino, Luisa Verdoliva, Christian Riess, Justus Thies, and Matthias Nießner. Faceforensics++: Learning to detect manipulated facial images. arxiv. https://arxiv. org/abs/1901.08971/, 2019. arXiv:1901.08971 [cs.CV]

  44. [54]

    Adversarial deep learning against intrusion detection classifiers

    Maria Rigaki. Adversarial deep learning against intrusion detection classifiers. Master’s thesis, Luleå University of Technology, 2017

  45. [55]

    Spotting the deepfakes in this year of elections: How ai detection tools work and where they fail.Reuters Institute

    shirin anlen and Raquel Vazquez Llorente. Spotting the deepfakes in this year of elections: How ai detection tools work and where they fail.Reuters Institute. https://reutersinstitute.politics.ox .ac.uk/news/spotting-deepfakes-year-elections-how-ai-detection-tools-work-and-whe...

  46. [56]

    Nimisha Singh, Amita Kapoor, and Neha Soni. A sociotechnical perspective for explicit unfairness mitigation techniques for algorithm fairness.International Journal of Information Management Data Insights, 4(2):6073–6105, nov 2024

  47. [57]

    @rwomchechen. X. https://x.com/rwomchechen/status/1860001051367342123/, 2024. Accessed: 8 April 2025

  48. [58]

    Governing access to synthetic media detection technology

    Jonathan Stray, Aviv Ovadya, Claire Leibowicz, and Sam Gregory. Governing access to synthetic media detection technology. Tech Policy Press. https://www.techpolicy.press/governing-access-to-s ynthetic-media-detection-technology/, 2021. Accessed: 8 April 2025

  49. [59]

    Human rights can be the spark of ai innovation—not stifle it.Tech Policy Press

    shirin anlen. Human rights can be the spark of ai innovation—not stifle it.Tech Policy Press. https: //www.techpolicy.press/human-rights-can-be-the-spark-of-ai-innovation-not-stifle-it/ ,

  50. [60]

    Dadzie TV. Youtube. https://www.youtube.com/watch?v=p6VnexYxwrU , 2024. Accessed: 8 April 2025

  51. [61]

    https://www.youtube.com/watch?v=OrFXklQz6bQ&t=2702s

    GHOne TV. https://www.youtube.com/watch?v=OrFXklQz6bQ&t=2702s. Accessed: 8 April 2025

  52. [62]

    Ai-image-detector

    Matthew Maybe (umm maybe). Ai-image-detector. Hugging Face. https://huggingface.co/umm-m 20 aybe/AI-image-detector, 2022. Accessed: 8 April 2025

  53. [63]

    Sohrawardi, Sovantharith Seng, Akash Chintha, Bao Thai, Raymond Ptucha, Matthew Wright, and Andrea Hickerson

    Saniat J. Sohrawardi, Sovantharith Seng, Akash Chintha, Bao Thai, Raymond Ptucha, Matthew Wright, and Andrea Hickerson. Defaking deepfakes: Understanding journalists’ needs for deepfake detection. In Proceedings of the USENIX Symposium on Usable Privacy and Security, 2020

  54. [64]

    Big brother watch briefing on clause 21 of the criminal justice bill.https://bills

    Big Brother Watch. Big brother watch briefing on clause 21 of the criminal justice bill.https://bills. parliament.uk/publications/53817/documents/4320#:~:text=Clause%2021%20also%20allows%2 0for,powers%2C%20with%20limited%20parliamentary%20oversight.&text=policing%20needs/ . Ac...

  55. [65]

    Big brother watch responds to facial recognition powers in new crime bill

    Big Brother Watch Team. Big brother watch responds to facial recognition powers in new crime bill. Big Brother Watch. https://bigbrotherwatch.org.uk/press-releases/big-brother-watch-res ponds-to-facial-recognition-powers-in-new-crime-bill/#:~:text=The%20Bill%20allows%2 0the%20...

  56. [66]

    Sociotechnical safety evaluation of generative ai systems.arxiv

    Laura Weidinger, Maribeth Rauh, Nahema Marchal, Arianna Manzini, Lisa Anne Hendricks, Juan Mateos-Garcia, Stevie Bergman, Jackie Kay, Conor Griffin, Ben Bariach, Iason Gabriel, Verena Rieser, and William Isaac. Sociotechnical safety evaluation of generative ai systems.arxiv. h...

  57. [67]

    https://www.witness.org/

    WITNESS. https://www.witness.org/. Accessed: 7 April 2025

  58. [68]

    Prepare, don’t panic: Synthetic media and deepfakes.https://lab.witness.org/projec ts/synthetic-media-and-deep-fakes

    WITNESS. Prepare, don’t panic: Synthetic media and deepfakes.https://lab.witness.org/projec ts/synthetic-media-and-deep-fakes. Accessed: 8 April 2025

  59. [69]

    Outgoing US President Biden did not confess to helping orchestrate pakistan ‘regime change’

    Haseem uz Zaman. Outgoing US President Biden did not confess to helping orchestrate pakistan ‘regime change’. Soch Fact Check. https://www.sochfactcheck.com/us-president-joe-biden-did-not -confess-to-regime-change-conspiracy-pakistan-army-imran-khan/ , 2024. Accessed: 8 April 2025

  60. [70]

    @zeltzinjuareze. X. https://x.com/zeltzinjuareze/status/1762878195265638428/ , 2024. Accessed: 8 April 2025. 21 A WITNESS’TRIEDBenchmark: AChecklistforTrulyInnovative and Effective AI Detection Instructions The checklist below is adapted from the benchmark quality assessment f...

  61. [71]

    Burkina faso: Video shows soldiers disemboweling.Human Rights Watch

    Human Rights Watch. Burkina faso: Video shows soldiers disemboweling.Human Rights Watch. https: //www.hrw.org/news/2024/07/26/burkina-faso-video-shows-soldiers-disemboweling-body ,

  62. [76]

    A community-based approach to visual verification to fortify the truth.WITNESS Guide

    WITNESS. A community-based approach to visual verification to fortify the truth.WITNESS Guide. https://library.witness.org/product/community-based-approaches-to-verification// ,

  63. [77]

    Accessed: 8 April 2025

  64. [79]

    The goals, key concepts, primary features, and target audience for the detection tool are clearly outlined

  65. [80]

    The design process actively involves input from domain experts with relevant expertise

  66. [81]

    Relevant academic research, industry standards, and existing literature are thoroughly reviewed and integrated into the design

  67. [82]

    Real-world scenarios and practical use cases are incorporated to guide the tool’s development and application

  68. [83]

    The mechanisms to ensure the tool’s durability and adaptability to rapid development of synthetic media are defined and prioritized in the design process

  69. [84]

    Checklist

    The tool’s funding allows for a responsible and sustainable development and operation of the tool. Checklist

  70. [85]

    1.1 The intended audience is clearly defined

    The goals, concepts, characteristics, and audience of the detection tool are defined. 1.1 The intended audience is clearly defined. □TO DO □YES □NO □N/A Justification: 1.2 The use cases on which the tool should be used are clearly described. □TO DO □YES □NO □N/A Justification:...

  71. [86]

    □TO DO □YES □NO □N/A Justification:

    The design process involved consultation with diverse domain experts, including those from global and underserved contexts. □TO DO □YES □NO □N/A Justification:

  72. [87]

    3.1 The research and literature were diverse and steps were taken to include knowledge produced 23 outside of Europe and the United States

    Diverse relevant academic research, industry reports, and existing literature were integrated into the tool’s design. 3.1 The research and literature were diverse and steps were taken to include knowledge produced 23 outside of Europe and the United States. □TO DO □YES □NO □N/...

  73. [88]

    4.1 Real-life cases and use cases were incorporated to reflect practical challenges and expectations

    Real-world scenarios and practical use cases are incorporated to guide the tool’s development and application. 4.1 Real-life cases and use cases were incorporated to reflect practical challenges and expectations. □TO DO □YES □NO □N/A Justification: 4.2 Stakeholders representat...

  74. [89]

    5.1 Specific steps are outlined to ensure that the tool is future-proof and guarantee that it will remain relevant in the long term

    The mechanisms to ensure the tool’s durability and adaptability to rapid development of synthetic media are defined. 5.1 Specific steps are outlined to ensure that the tool is future-proof and guarantee that it will remain relevant in the long term. □TO DO □YES □NO □N/A Justif...

  75. [90]

    The development stage requires explicit consideration of underrepresented languages, cultural nuances, and geopolitical challenges in both training and testing

  76. [91]

    Trainingdataisdiverse, representativeofglobaldemographics, andcollectedethically, withtransparent documentation of the sourcing process

  77. [92]

    Comprehensiveandaccessibledocumentationisprovided, integratingrelevantcontexttoaidunderstanding and usability

  78. [93]

    The tool’s limitations are clearly identified and communicated to ensure realistic expectations of its capabilities

  79. [94]

    Mechanisms for explainability are incorporated, ensuring that results and detections are interpretable by both technical and non-technical users

  80. [95]

    The development team is diverse, reflecting the perspectives and needs of the tool’s intended audience. 24

  81. [96]

    Accessibility is prioritized, ensuring the tool is usable by the intended audience regardless of technical expertise or resource constraints

  82. [97]

    The tool’s capabilities for handling various file types and quality levels are explicitly detailed

  83. [98]

    The tool demonstrates consistent performance across diverse contexts and use cases it aims to serve

  84. [99]

    Development respects and integrates with existing verification techniques and skill sets to enhance reliability and usability

  85. [100]

    Checklist

    Adaptability is a key focus, allowing the tool to evolve in response to new challenges, use cases, and technological advancements. Checklist

  86. [101]

    1.1 Data collection process complied with the local data protection regulations

    Training data is diverse, representative of global demographics, and ethically sourced. 1.1 Data collection process complied with the local data protection regulations. □TO DO □YES □NO □N/A Justification: 1.2 The training dataset is representative of diverse demographics. □TO ...

  87. [102]

    □TO DO □YES □NO □N/A Justification:

    The tool’s capabilities for handling various file types and quality levels are explicitly detailed. □TO DO □YES □NO □N/A Justification:

  88. [103]

    3.1 The tool was trained on low-resolution content

    The training dataset included examples of corrupted synthetic content. 3.1 The tool was trained on low-resolution content. □TO DO □YES □NO □N/A Justification: 3.2 The tool was trained on social media compression standards. □TO DO □YES □NO □N/A Justification: 3.3 The tool was t...

  89. [104]

    4.1 If the tool is capable of detecting audio, the training dataset included data with a transmission similar to telephones

    The detection tool was trained to deal with diverse types of content. 4.1 If the tool is capable of detecting audio, the training dataset included data with a transmission similar to telephones. □TO DO □YES □NO □N/A Justification: 4.2 If the tool is capable of detecting audio,...

  90. [105]

    5.1 A system is in place, and adaptable to adjust the detection model to continually update the training dataset to include new examples of AI-generated or manipulated content

    The tool is developed to scale across different levels of content complexity, size, and volume. 5.1 A system is in place, and adaptable to adjust the detection model to continually update the training dataset to include new examples of AI-generated or manipulated content. □TO ...

  91. [106]

    The tool provides users with clear information on how detection results should be interpreted. 6.1 The tool does not use binary labels (such as ‘real’ or ‘fake’) when communicating the results but instead describes manipulation using accessible description and language of degr...

  92. [107]

    7.1 The information is provided with respect to other verification techniques and skill sets that could be integrated

    The tool acknowledges other existing verification techniques and incorporates information about them into its workflow. 7.1 The information is provided with respect to other verification techniques and skill sets that could be integrated. □TO DO □YES □NO □N/A Justification: 7....

  93. [108]

    □TO DO □YES □NO □N/A Justification: 27 The maximum number of points is 57

    The tool is adaptable with processes in place to evolve in response to new challenges, use cases, and technological advancements. □TO DO □YES □NO □N/A Justification: 27 The maximum number of points is 57. Your score is: Benchmark Testing Testing Criteria

  94. [109]

    Proactive measurements are being taken to test the tool’s blind spots

  95. [110]

    The tool undergoes regular and systematic testing to maintain reliability and effectiveness

  96. [111]

    Diverse stakeholders and domain experts, including independent external groups, actively participate in the testing process to provide varied perspectives and expertise

  97. [112]

    The tool is rigorously tested on challenging edge cases, including adversarial content, false claims of AI-generated media, and heavily manipulated content

  98. [113]

    The tool demonstrates resilience against evasion techniques, maintaining its accuracy and reliability even under deliberate attempts to bypass detection

  99. [114]

    Checklist

    The tool is evaluated from a human-detection perspective. Checklist

  100. [115]

    1.1 The tool is proactively tested to identify blind spots and potential failures

    The tool was tested in an exhaustive and inclusive manner. 1.1 The tool is proactively tested to identify blind spots and potential failures. □TO DO □YES □NO □N/A Justification: 1.2 The tool’s robustness was tested against adversarial attacks and evasion techniques. □TO DO □YE...

  101. [116]

    2.1 Such testing included assessing how much time it took to receive the result and how difficult the process was from the human perspective

    The tool is evaluated from a human-assisted detection perspective. 2.1 Such testing included assessing how much time it took to receive the result and how difficult the process was from the human perspective. □TO DO □YES □NO □N/A Justification: 2.2 Such testing included a comp...

  102. [117]

    The tool is designed and implemented to uphold human rights, ensuring it does not infringe on privacy, freedom of expression, or other fundamental rights

  103. [118]

    Accessibility is prioritized, ensuring the tool is usable by its intended audience, regardless of technical expertise or resource availability

  104. [119]

    Checklist

    Documentation is clear and comprehensive, and top level version is easily accessible, providing users with the necessary guidance to operate the tool effectively and understand its limitations. Checklist

  105. [120]

    1.1 Thetoolexplicitlyavoidsoutputsorrecommendationsthatcouldharmindividualsorcommunities, aligning with ethical AI principles

    Ensure the tool does not inadvertently violate human rights, such as privacy or freedom of expression. 1.1 Thetoolexplicitlyavoidsoutputsorrecommendationsthatcouldharmindividualsorcommunities, aligning with ethical AI principles. □TO DO □YES □NO □N/A Justification: 1.2 The too...

  106. [121]

    2.1 The tool includes educational resources or tutorials to help users understand its functionality and limitations

    The tool is accessible to a diverse group of targeted users. 2.1 The tool includes educational resources or tutorials to help users understand its functionality and limitations. □TO DO □YES □NO □N/A Justification: 2.2 The tool does not require advanced technical knowledge or s...

  107. [122]

    3.1 Thetoolcommunicateserrorsoruncertaintiesclearlytousers, avoidingoverconfidenceinambiguous cases

    The tool’s performance metrics are accurate and not exaggerated. 3.1 Thetoolcommunicateserrorsoruncertaintiesclearlytousers, avoidingoverconfidenceinambiguous cases. □TO DO □YES □NO □N/A Justification: 30 The maximum number of points is 33. Your score is: Benchmark Maintenance...

  108. [123]

    The tool undergoes regular performance evaluations to ensure its detection accuracy remains reliable across evolving datasets and content types

  109. [124]

    Resources are actively being distributed for updates and maintenance processes

  110. [125]

    Updatesareroutinelyimplementedtoaddressemergingchallenges, suchasadversarialattacks, advancements in deepfake technologies, and new forms of synthetic content

  111. [126]

    User feedback is actively collected, documented, and incorporated into updates, with a transparent and accessible communication channel available for issue reporting and suggestions

  112. [127]

    Clear policies are established regarding the support duration for older versions of the tool following updates or new releases

  113. [128]

    Periodic audits are conducted to identify and mitigate any biases introduced during updates or revealed by newer datasets

  114. [129]

    Checklist

    Developers continuously monitor advancements in AI and related technologies, ensuring the tool remains aligned with the latest innovations and best practices. Checklist

  115. [130]

    1.1 The tool’s accuracy is evaluated against previously verified cases

    The tool undergoes regular updates, and changes are documented transparently. 1.1 The tool’s accuracy is evaluated against previously verified cases. □TO DO □YES □NO □N/A Justification: 1.2 Clear policies are in place for how regularly updates will be conducted. □TO DO □YES □N...

  116. [131]

    2.1 Feedback channels for users are actively maintained, and updates incorporate user feedback

    An accessible channel of communication is provided. 2.1 Feedback channels for users are actively maintained, and updates incorporate user feedback. □TO DO □YES □NO □N/A Justification: 2.2 A contact person or team for support and inquiries is clearly listed. □TO DO □YES □NO □N/...

  117. [132]

    3.1 The tool will be retired if it has not been updated for a specified time following clear policies

    The tool has End-of-Life procedures in place. 3.1 The tool will be retired if it has not been updated for a specified time following clear policies. □TO DO □YES □NO □N/A Justification: 3.2 If the tool is to be discontinued, clear communication is provided to users, along with ...

  118. [2021]

    arXiv:2007.02407 [cs.LG]

  119. [2024]

    https://arxiv.org/abs/2503.02857/, 2025

    preprint arXiv. https://arxiv.org/abs/2503.02857/, 2025. arXiv:2503.02857 [cs.CV]

Pith tools

Reviewed August 16, 2026 · model on record in the stance chip above.