REVIEW 4 major objections 3 minor 90 references
People spot AI-generated content more reliably when a post combines text and an image, and most reliably when the two clash, according to a 154,552-post human study and the LLM-agent system built on it.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
A 154K-post study reports that humans identify AI content best when text and images are both present and inconsistent, and offers metrics plus an LLM agent for human-aligned responses.
T0 review reviewed 2026-08-05 challenge →
load-bearing objection Plausible human-centered contribution with a load-bearing label-provenance question that the abstract doesn't answer; worth peer review if the full text checks the boxes. the 4 major comments →
Modeling Human Responses to Multimodal AI Content
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
Core claim
The central discovery is that the human mind treats a multimodal post as a composite: when the text and image are present together, people are better at identifying AI-generated content than when either appears alone, and the effect is strongest when the two modalities disagree. The paper reports this through a human study on the MhAIM dataset and then operationalizes the finding in three metrics—trustworthiness, impact, and openness—that turn subjective judgments into numeric scores for engagement. These predicted human responses are wrapped in HR-MCP (Human Response Model Context Protocol), a component built on the standardized Model Context Protocol, so that any LLM using T-Lens can answe
What carries the argument
The load-bearing machinery is the MhAIM dataset: a corpus of 154,552 real online posts with provenance labels for AI generation, which makes the human-perception claim measurable at scale. Inside that setting, the operative mechanism is cross-modal inconsistency—when the text and the visual in a post do not agree, that contradiction functions as a visible cue that the content is AI-generated. On the system side, HR-MCP is the transducer that converts predicted human judgments (trustworthiness, impact, openness) into a standard protocol any LLM can call, which is what lets T-Lens anticipate human reactions rather than merely classify authenticity.
Load-bearing premise
The paper's conclusions stand on the accuracy of its AI-generated labels: if a meaningful share of the 111,153 AI-labeled posts are actually human-written, or if the human-written set is contaminated with AI content, the observed perceptual advantage for text-image inconsistency could shrink or disappear.
What would settle it
Take a random sample of the 111,153 AI-labeled posts and of the human-authored posts, verify provenance by independent annotators or platform metadata, and re-run the human-detection comparison; separately, run a controlled web experiment where the same text is paired with a consistent image and with a deliberately inconsistent image and measure human accuracy. If accuracy is no higher for inconsistent pairs, or if verified labels erase the effect, the central claim is false.
If this is right
- If text-image inconsistency is a reliable cue, then AI-generated misinformation can be made easier to spot simply by preserving and surfacing the mismatch between a post's text and its image.
- The MhAIM dataset, with 154,552 posts and human-response labels, can support large-scale studies of what makes content spread, not just whether it is true.
- The trustworthiness, impact, and openness metrics give content moderators and social platforms a shared vocabulary for user judgment that goes beyond binary authentic/not-authentic labels.
- Because T-Lens consumes predicted human responses through a standard protocol, LLM agents can be tuned to answer in ways that anticipate what a human reader would believe, which may reduce the spread of AI-driven misinformation.
- A direct corollary: detection systems should treat text-image consistency as a feature, not as noise.
Where Pith is reading between the lines
- An untested extension: the same text-image inconsistency cue could be built into automated moderation as a nudge that asks a reader to scrutinize before sharing; a controlled field test would show whether exposing the mismatch changes sharing behavior.
- If the MhAIM AI labels were assigned by platform provenance or generator metadata, the dataset may under-represent human-AI hybrid posts; a follow-up annotation study on mixed-authored content would show whether the finding persists when only part of a post is machine-generated.
- The three metrics could be repurposed as lightweight engagement predictors: treating trustworthiness and impact as regression targets might outperform authenticity classifiers for predicting virality, which is the practical motivation the paper opens with.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper introduces the MhAIM dataset (154,552 posts, 111,153 labeled AI-generated) and reports a human study whose headline finding is that people are better at identifying AI content when posts contain both text and visuals, especially when text and visuals are inconsistent. The paper also proposes three metrics (trustworthiness, impact, openness), and presents T-Lens, an LLM-based agent system built on HR-MCP, a Model Context Protocol-based component designed to incorporate predicted human responses into LLM queries. The claims are plausible and potentially useful, but the provided manuscript does not allow verification of the human study, the dataset label construction, or the T-Lens evaluation; much of the body text is unreadable due to encoding corruption.
Significance. If the findings hold, MhAIM would be a valuable large-scale resource for human-centered AI-content research, and the perceptual finding would provide a concrete, falsifiable design cue for AI-detection and misinformation-mitigation interfaces. The three metrics and the T-Lens/HR-MCP system are interesting system-level contributions, and the paper's emphasis on human perception rather than factual verification alone is timely. However, the MhAIM dataset's validity is currently not established because the provenance of the AI/human labels is not reported, and the human-study claim is presented without the methodological details needed for assessment. The paper ships no machine-checked proofs or reproducible code in the readable portions, and no precise numerical predictions beyond the abstract-level claim. The stress-test concern lands: label provenance is a load-bearing input-labeling premise, and it is missing.
major comments (4)
- [Abstract / MhAIM dataset description] The central quantity is the binary AI/human label on the 111,153 'AI-generated' posts, but the manuscript never states how these labels were obtained. The body text is unreadable in the submitted file, so no dataset-construction or label-validation section can be inspected. If labels come from platform provenance, the human study may be detecting account/format cues rather than content; if from an automatic detector, the 'text-image inconsistency' cue may reflect detector error patterns. Contamination in either direction would shrink or invert the reported perceptual difference. Since every downstream claim inherits this label quality, the provenance must be reported and a validation sub-study (e.g., human review of a random sample, inter-rater agreement on labels) provided.
- [Abstract, human study claim] The headline result — people are better at identifying AI content in text-plus-visual posts, particularly under text-image inconsistency — is presented without any of the study details needed to evaluate it. I could not find participant counts, recruitment procedure, stimulus selection, modality-balancing, controls for prior exposure or platform familiarity, inter-annotator agreement, or significance tests anywhere in the available text. The claim needs condition-level accuracy (or d') with confidence intervals and a test for the moderation effect of inconsistency; as written it is an assertion rather than a reportable result.
- [Abstract, T-Lens and HR-MCP paragraph] The evaluation of T-Lens appears circular as described. The abstract states that T-Lens aligns with human reactions by consuming 'predicted human responses' from HR-MCP; if those predictions are trained on the same human labels that ground the paper's empirical findings, then measuring T-Lens against human reactions is partly testing its own training signal. The text does not state how HR-MCP was trained, whether T-Lens evaluation uses held-out posts/participants, or what baseline (e.g., an LLM agent without HR-MCP) was used. This must be specified before the claim 'better align with human reactions' can be interpreted.
- [Appendix metrics (trustworthiness/impact/openness)] The three metrics are named in the abstract and appear as garbled table entries in the appendix, but no readable definition, normalization, or validation is provided. As they are purported new measurement instruments, the paper should give exact formulas, annotation scales, and reliability/validity evidence (e.g., inter-rater reliability, convergent associations with behavioral outcomes). Without this, the MhAIM analysis built on these metrics is not reproducible.
minor comments (3)
- [Full text / rendering] The entire body text is mojibake; equations, tables, references, and section headings are unreadable. Please resubmit a clean PDF/source; this is a prerequisite for any technical review.
- [References] No verifiable references are readable in the provided text; related-work positioning and comparisons to prior datasets cannot be checked.
- [Tables] Partially legible tables cannot be matched to conditions or metrics; add clear captions and readable numbers so that reported counts and effect sizes can be verified.
Circularity Check
No significant circularity found: the abstract reports an empirical human-study result and a supervised agent, and no step in the available text reduces to its own input by construction.
full rationale
The load-bearing claim is a reported experimental finding: 'our human study reveals that people are better at identifying AI content when posts include both text and visuals, particularly when inconsistencies exist between the two.' The abstract does not state that the AI-generated labels were produced by the human study itself or by T-Lens; they are an input premise. The human detection finding is a measured outcome conditional on those labels, not a fitted parameter renamed as a prediction. T-Lens/HR-MCP is described as 'incorporating predicted human responses,' but the available text does not show that the same human labels used for training are also used as the evaluation target in a way that makes agreement tautological, nor that the proposed metrics (trustworthiness, impact, openness) are defined in terms of the model's own outputs. Because the full body text is encoding-corrupted beyond reliable recovery, no equation-level reduction (e.g., Eq. X = Eq. Y by construction) can be exhibited. Under the hard rule requiring a quotable reduction, no circularity is established. Concerns about label provenance are validation/correctness risks rather than circularity. The self-citation chain is not visible in the readable content, so no load-bearing self-citation can be charged.
Axiom & Free-Parameter Ledger
axioms (3)
- domain assumption Dataset labels of AI-generation status are accurate
- domain assumption Study participants' detection behavior generalizes to the broader population
- domain assumption Sampled online posts are representative of real-world AI-content exposure
invented entities (5)
-
Trustworthiness metric
no independent evidence
-
Impact metric
no independent evidence
-
Openness metric
no independent evidence
-
HR-MCP (Human Response Model Context Protocol)
no independent evidence
-
T-Lens agent
no independent evidence
Cite this review
Pith. "Pith review of Modeling Human Responses to Multimodal AI Content." pith.science (2026). https://pith.science/paper/SKA3YXQP
@misc{pith2026250810769,
author = {Pith},
title = {Pith review of: Modeling Human Responses to Multimodal AI Content},
year = {2026},
howpublished = {\url{https://pith.science/paper/SKA3YXQP}},
note = {Machine review of arXiv:2508.10769}
}
read the original abstract
As AI-generated content becomes widespread, so does the risk of misinformation. While prior research has primarily focused on identifying whether content is authentic, much less is known about how such content influences human perception and behavior. In domains like trading or the stock market, predicting how people react (e.g., whether a news post will go viral), can be more critical than verifying its factual accuracy. To address this, we take a human-centered approach and introduce the MhAIM Dataset, which contains 154,552 online posts (111,153 of them AI-generated), enabling large-scale analysis of how people respond to AI-generated content. Our human study reveals that people are better at identifying AI content when posts include both text and visuals, particularly when inconsistencies exist between the two. We propose three new metrics: trustworthiness, impact, and openness, to quantify how users judge and engage with online content. We present T-Lens, an LLM-based agent system designed to answer user queries by incorporating predicted human responses to multimodal information. At its core is HR-MCP (Human Response Model Context Protocol), built on the standardized Model Context Protocol (MCP), enabling seamless integration with any LLM. This integration allows T-Lens to better align with human reactions, enhancing both interpretability and interaction capabilities. Our work provides empirical insights and practical tools to equip LLMs with human-awareness capabilities. By highlighting the complex interplay among AI, human cognition, and information reception, our findings suggest actionable strategies for mitigating the risks of AI-driven misinformation.
Reference graph
Works this paper leans on
-
[1]
LangChain
2022. LangChain. https://github.com/langchain-ai/langchain/. Accessed: 2023-10-01
2022
-
[2]
Online crowd-sourcing platform Toluna
2023. Online crowd-sourcing platform Toluna. www.toluna-group.com. Accessed: 2023-10-01
2023
-
[3]
ChatGPT Large Language Model
2024. ChatGPT Large Language Model. https://chat.openai.com/. Accessed: 2024-08-10
2024
-
[4]
Snopes fact checking website
2024. Snopes fact checking website. https://www.snopes.com/. Accessed: 2023-10-01
2024
-
[5]
Stable Diffusion Online
2024. Stable Diffusion Online. https://stablediffusionweb.com/. Accessed: 2024-08-10
2024
-
[6]
A \" meur, E.; Amri, S.; and Brassard, G. 2023. Fake news, disinformation and misinformation in social media: a review. Social Network Analysis and Mining, 13(1): 30
2023
-
[7]
Amoroso, R.; Morelli, D.; Cornia, M.; Baraldi, L.; Del Bimbo, A.; and Cucchiara, R. 2023. Parents and Children: Distinguishing Multimodal DeepFakes from Natural Images. arXiv preprint arXiv:2304.00500
Pith/arXiv arXiv 2023
-
[8]
Aneja, S.; Bregler, C.; and Nie ner, M. 2021. Cosmos: Catching out-of-context misinformation with self-supervised learning. arXiv preprint arXiv:2101.06278
Pith/arXiv arXiv 2021
-
[9]
Anthropic . 2024. Introducing the Model Context Protocol. https://www.anthropic.com/news/model-context-protocol. Accessed: 2025-06-30
2024
-
[10]
Anthropic . 2025. Claude 3.7 Sonnet and Claude Code. https://www.anthropic.com/news/claude-3-7-sonnet. Accessed: 2025-07-28
2025
-
[11]
M.; Nayak, V.; Dinkov, Y.; Zlatkova, D.; Dent, K.; Bhatawdekar, A.; Bouchard, G.; et al
Arora, A.; Nakov, P.; Hardalov, M.; Sarwar, S. M.; Nayak, V.; Dinkov, Y.; Zlatkova, D.; Dent, K.; Bhatawdekar, A.; Bouchard, G.; et al. 2021. Detecting Harmful Content on Online Platforms: What Platforms Need vs. Where Research Efforts Go. ACM Computing Surveys
2021
-
[12]
Aslett, K.; Sanderson, Z.; Godel, W.; Persily, N.; Nagler, J.; and Tucker, J. A. 2024. Online searches to evaluate misinformation can increase its perceived veracity. Nature, 625(7995): 548--556
2024
-
[13]
Bai, J.; Bai, S.; Chu, Y.; Cui, Z.; Dang, K.; Deng, X.; Fan, Y.; Ge, W.; Han, Y.; Huang, F.; et al. 2023. Qwen technical report. arXiv preprint arXiv:2309.16609
Pith/arXiv arXiv 2023
-
[14]
Bailey, R. A. 2008. Design of comparative experiments, volume 25. Cambridge University Press
work page 2008
-
[15]
Bandi, A.; Adapa, P. V. S. R.; and Kuchi, Y. E. V. P. K. 2023. The Power of Generative AI: A Review of Requirements, Models, Input--Output Formats, Evaluation Metrics, and Challenges. Future Internet, 15(8): 260
work page 2023
-
[16]
Batailler, C.; Brannon, S. M.; Teas, P. E.; and Gawronski, B. 2022. A signal detection approach to understanding the identification of fake news. Perspectives on Psychological Science, 17(1): 78--98
work page 2022
-
[17]
H.; Ragnhildstveit, A.; Sprockett, S.; Barr, N.; Christensen, A.; and Seli, P
Bellaiche, L.; Shahi, R.; Turpin, M. H.; Ragnhildstveit, A.; Sprockett, S.; Barr, N.; Christensen, A.; and Seli, P. 2023. Humans versus AI: whether and why we prefer human-created compared to AI-created artwork. Cognitive Research: Principles and Implications, 8(1): 1--22
work page 2023
-
[18]
Boididou, C.; Papadopoulos, S.; Zampoglou, M.; Apostolidis, L.; Papadopoulou, O.; and Kompatsiaris, Y. 2018. Detection and visualization of misleading content on Twitter. International Journal of Multimedia Information Retrieval, 7(1): 71--86
work page 2018
-
[19]
Budak, C.; Nyhan, B.; Rothschild, D. M.; Thorson, E.; and Watts, D. J. 2024. Misunderstanding the harms of online misinformation. Nature, 630(8015): 45--53
work page 2024
-
[20]
Cao, Y.; Li, S.; Liu, Y.; Yan, Z.; Dai, Y.; Yu, P. S.; and Sun, L. 2023. A comprehensive survey of ai-generated content (aigc): A history of generative ai from gan to chatgpt. arXiv preprint arXiv:2303.04226
Pith/arXiv arXiv 2023
-
[21]
Chaka, C. 2023. Detecting AI content in responses generated by ChatGPT, YouChat, and Chatsonic: The case of five AI content detection tools. Journal of Applied Learning and Teaching, 6(2)
work page 2023
-
[22]
Chen, Y.; Li, D.; Zhang, P.; Sui, J.; Lv, Q.; Tun, L.; and Shang, L. 2022. Cross-modal ambiguity learning for multimodal fake news detection. In Proceedings of the ACM Web Conference 2022, 2897--2905
work page 2022
-
[23]
Comanici, G.; Bieber, E.; Schaekermann, M.; Pasupat, I.; Sachdeva, N.; Dhillon, I.; Blistein, M.; Ram, O.; Zhang, D.; Rosen, E.; et al. 2025. Gemini 2.5: Pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities. arXiv preprint arXiv:2507.06261
Pith/arXiv arXiv 2025
-
[24]
Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; et al. 2020. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929
Pith/arXiv arXiv 2020
-
[25]
Du, H.; Zhang, R.; Niyato, D.; Kang, J.; Xiong, Z.; Kim, D. I.; Shen, X. S.; and Poor, H. V. 2023 a . Exploring collaborative distributed diffusion-based AI-generated content (AIGC) in wireless networks. IEEE Network, (99): 1--8
work page 2023
-
[26]
Du, W.; Li, Q.; Zhou, J.; Ding, X.; Wang, X.; Zhou, Z.; and Liu, J. 2023 b . FinGuard: A Multimodal AIGC Guardrail in Financial Scenarios. In Proceedings of the 5th ACM International Conference on Multimedia in Asia, 1--3
work page 2023
-
[27]
K.; Lewandowsky, S.; Cook, J.; Schmid, P.; Fazio, L
Ecker, U. K.; Lewandowsky, S.; Cook, J.; Schmid, P.; Fazio, L. K.; Brashier, N.; Kendeou, P.; Vraga, E. K.; and Amazeen, M. A. 2022. The psychological drivers of misinformation belief and its resistance to correction. Nature Reviews Psychology, 1(1): 13--29
work page 2022
-
[28]
Edwin, L. 2025. Model Context Protocol (MCP): Solution to AI Integration Bottlenecks. https://addepto.com/blog/model-context-protocol-mcp-solution-to-ai-integration-bottlenecks/. Accessed: 2025-06-30
work page 2025
-
[29]
R.; Groh, M.; Herman, L.; Leach, N.; et al
Epstein, Z.; Hertzmann, A.; of Human Creativity, I.; Akten, M.; Farid, H.; Fjeld, J.; Frank, M. R.; Groh, M.; Herman, L.; Leach, N.; et al. 2023. Art and the science of generative AI. Science, 380(6650): 1110--1111
work page 2023
-
[30]
Fan, D.-P.; Ji, G.-P.; Xu, P.; Cheng, M.-M.; Sakaridis, C.; and Van Gool, L. 2023. Advances in deep concealed scene understanding. Visual Intelligence, 1(1): 16
work page 2023
-
[31]
Fan, S.; Shen, Z.; Jiang, M.; Koenig, B. L.; Xu, J.; Kankanhalli, M. S.; and Zhao, Q. 2018. Emotional attention: A study of image sentiment and visual attention. In Proceedings of the IEEE Conference on computer vision and pattern recognition, 7521--7531
work page 2018
-
[32]
L.; Ng, T.-T.; and Kankanhalli, M
Fan, S.; Shen, Z.; Koenig, B. L.; Ng, T.-T.; and Kankanhalli, M. S. 2020. When and why static images are more effective than videos. IEEE Transactions on Affective Computing
work page 2020
-
[33]
Ferrara, E. 2024. GenAI against humanity: Nefarious applications of generative artificial intelligence and large language models. Journal of Computational Social Science, 1--21
work page 2024
-
[34]
Ghorbanpour, F.; Ramezani, M.; Fazli, M. A.; and Rabiee, H. R. 2023. FNR: a similarity and transformer-based approach to detect multi-modal fake news in social media. Social Network Analysis and Mining, 13(1): 56
work page 2023
-
[35]
Gong, D.; Goh, O. S.; Kumar, Y. J.; Ye, Z.; and Chi, W. 2020. Deepfake forensics, an ai-synthesized detection with deep convolutional generative adversarial networks. Int J, 9(3): 2861--2870
work page 2020
-
[36]
Hangloo, S.; and Arora, B. 2023. Evidence-Aware Fake News Detection: A Review. In 2023 International Conference on Advanced Computing & Communication Technologies (ICACCTech), 81--86. IEEE
work page 2023
-
[37]
Hartwig, K.; Doell, F.; and Reuter, C. 2024. The Landscape of User-centered Misinformation Interventions-A Systematic Literature Review. ACM Computing Surveys, 56(11): 1--36
work page 2024
-
[38]
He, B.; Ahamad, M.; and Kumar, S. 2023. Reinforcement learning-based counter-misinformation response generation: a case study of COVID-19 vaccine misinformation. In Proceedings of the ACM Web Conference 2023, 2698--2709
work page 2023
-
[39]
Hermann, E. 2022. Artificial intelligence and mass personalization of communication content—An ethical and literacy perspective. New Media & Society, 24(5): 1258--1277
work page 2022
-
[40]
Hill, K. M. 2025. The rising threat of fake news in financial markets. CU Boulder Today. Accessed: 2025-08-01
work page 2025
-
[41]
Hou, X.; Zhao, Y.; Wang, S.; and Wang, H. 2025. Model context protocol (mcp): Landscape, security threats, and future research directions. arXiv preprint arXiv:2503.23278
Pith/arXiv arXiv 2025
-
[42]
Hu, X.; Chen, P.-Y.; and Ho, T.-Y. 2023. Radar: Robust ai-text detection via adversarial learning. Advances in Neural Information Processing Systems, 36: 15077--15095
work page 2023
-
[43]
Jo, A. 2023. The promise and peril of generative AI. Nature, 614(1): 214--216
work page 2023
-
[44]
Kaate, I.; Salminen, J.; Jung, S.-G.; Almerekhi, H.; and Jansen, B. J. 2023. How Do Users Perceive Deepfake Personas? Investigating the Deepfake User Perception and Its Implications for Human-Computer Interaction. In Proceedings of the 15th Biannual Conference of the Italian SIGCHI Chapter, 1--12
work page 2023
-
[45]
S.; Torralba, A.; and Oliva, A
Khosla, A.; Raju, A. S.; Torralba, A.; and Oliva, A. 2015. Understanding and Predicting Image Memorability at a Large Scale. In International Conference on Computer Vision (ICCV)
work page 2015
-
[46]
Kramer, M. A.; Hebart, M. N.; Baker, C. I.; and Bainbridge, W. A. 2023. The features underlying the memorability of objects. Science advances, 9(17): eadd2981
work page 2023
-
[47]
Kreps, S.; McCain, R. M.; and Brundage, M. 2022. All the news that’s fit to fabricate: AI-generated text as a tool of media misinformation. Journal of experimental political science, 9(1): 104--117
work page 2022
-
[48]
Li, J.; Li, D.; Xiong, C.; and Hoi, S. 2022. Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation. In International Conference on Machine Learning, 12888--12900. PMLR
work page 2022
-
[49]
Lin, L.; Gupta, N.; Zhang, Y.; Ren, H.; Liu, C.-H.; Ding, F.; Wang, X.; Li, X.; Verdoliva, L.; and Hu, S. 2024. Detecting Multimedia Generated by Large AI Models: A Survey. arXiv preprint arXiv:2402.00045
Pith/arXiv arXiv 2024
-
[50]
Liu, A.; Feng, B.; Xue, B.; Wang, B.; Wu, B.; Lu, C.; Zhao, C.; Deng, C.; Zhang, C.; Ruan, C.; et al. 2024. Deepseek-v3 technical report. arXiv preprint arXiv:2412.19437
Pith/arXiv arXiv 2024
-
[51]
Liu, H. 2024. ‘Worldview’of the AIGC systems: stability, tendency and polarization. AI & SOCIETY, 1--14
work page 2024
-
[52]
Lu, Z.; Huang, D.; Bai, L.; Liu, X.; Qu, J.; and Ouyang, W. 2023. Seeing is not always believing: A Quantitative Study on Human Perception of AI-Generated Images. arXiv preprint arXiv:2304.13023
Pith/arXiv arXiv 2023
-
[53]
Malik, A.; Kuribayashi, M.; Abdullahi, S. M.; and Khan, A. N. 2022. DeepFake detection for human face images and videos: A survey. Ieee Access, 10: 18757--18775
work page 2022
-
[54]
M.; Javed, A.; Irtaza, A.; and Malik, H
Masood, M.; Nawaz, M.; Malik, K. M.; Javed, A.; Irtaza, A.; and Malik, H. 2023. Deepfakes generation and detection: State-of-the-art, open challenges, countermeasures, and way forward. Applied intelligence, 53(4): 3974--4026
work page 2023
-
[55]
M.; Figueira, O.; Wang, Y.; and Wang, G
Mink, J.; Luo, L.; Barbosa, N. M.; Figueira, O.; Wang, Y.; and Wang, G. 2022. \ DeepPhish \ : Understanding User Trust Towards Artificially Generated Profiles in Online Social Networks. In 31st USENIX Security Symposium (USENIX Security 22), 1669--1686
work page 2022
-
[56]
Mirsky, Y.; and Lee, W. 2021. The creation and detection of deepfakes: A survey. ACM Computing Surveys (CSUR), 54(1): 1--41
work page 2021
-
[57]
Mittal, G.; Yenphraphai, J.; Hegde, C.; and Memon, N. 2022. Gotcha: A Challenge-Response System for Real-Time Deepfake Detection. arXiv preprint arXiv:2210.06186
work page internal anchor Pith review Pith/arXiv arXiv 2022
-
[58]
Mundra, S.; Porcile, G. J. A.; Marvaniya, S.; Verbus, J. R.; and Farid, H. 2023. Exposing GAN-Generated Profile Photos From Compact Embeddings. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 884--892
work page 2023
-
[59]
Nakamura, K.; Levy, S.; and Wang, W. Y. 2020. Fakeddit: A new multimodal benchmark dataset for fine-grained fake news detection. Conference on Language Resources and Evaluation (LREC 2020), 6149--6157
work page 2020
-
[60]
OpenAI. 2024. GPT-4o Technical Report. Accessed: 2025-08-01
work page 2024
-
[61]
Papadopoulou, O.; Zampoglou, M.; Papadopoulos, S.; and Kompatsiaris, I. 2019. A corpus of debunked and verified user-generated videos. Online information review, 43(1): 72--88
work page 2019
-
[62]
Pu, J.; Mangaokar, N.; Kelly, L.; Bhattacharya, P.; Sundaram, K.; Javed, M.; Wang, B.; and Viswanath, B. 2021. Deepfake videos in the wild: Analysis and detection. In Proceedings of the Web Conference 2021, 981--992
work page 2021
-
[63]
Qi, P.; Bu, Y.; Cao, J.; Ji, W.; Shui, R.; Xiao, J.; Wang, D.; and Chua, T.-S. 2023. FakeSV: A multimodal benchmark with rich social context for fake news detection on short video platforms. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 37, 14444--14452
work page 2023
-
[64]
Qi, P.; Yan, Z.; Hsu, W.; and Lee, M. L. 2024. SNIFFER: Multimodal Large Language Model for Explainable Out-of-Context Misinformation Detection. In IEEE Conference on Computer Vision and Patten Recognition (CVPR)
work page 2024
-
[65]
W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al
Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al. 2021. Learning transferable visual models from natural language supervision. In International conference on machine learning, 8748--8763. PMLR
2021
-
[66]
Robertson, C.; and Ridge-Newman, A. 2022. The Potential of Artificial Intelligence to Rejuvenate Public Trust in Journalism. In Futures of Journalism: Technology-stimulated Evolution in the Audience-News Media Relationship, 127--142. Springer
work page 2022
-
[67]
Rossler, A.; Cozzolino, D.; Verdoliva, L.; Riess, C.; Thies, J.; and Nie ner, M. 2019. Faceforensics++: Learning to detect manipulated facial images. In Proceedings of the IEEE/CVF international conference on computer vision, 1--11
work page 2019
-
[68]
Seo, H.; Xiong, A.; and Lee, D. 2019. Trust it or not: Effects of machine-learning warnings in helping individuals mitigate misinformation. In Proceedings of the 10th ACM Conference on Web Science, 265--274
work page 2019
-
[69]
Shao, R.; Wu, T.; Wu, J.; Nie, L.; and Liu, Z. 2024. Detecting and Grounding Multi-Modal Media Manipulation and Beyond. IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI)
work page 2024
-
[70]
K.; Bhattacharyya, A.; Baths, V.; Chen, C.; Ratn Shah, R.; Krishnamurthy, B.; et al
Singh, S.; Singla, Y. K.; Bhattacharyya, A.; Baths, V.; Chen, C.; Ratn Shah, R.; Krishnamurthy, B.; et al. 2023. Long-Term Memorability On Advertisements. arXiv e-prints, arXiv--2309
work page 2023
-
[71]
L.; Daum \'e III, H.; Dodge, J.; Evans, E.; Hooker, S.; et al
Solaiman, I.; Talat, Z.; Agnew, W.; Ahmad, L.; Baker, D.; Blodgett, S. L.; Daum \'e III, H.; Dodge, J.; Evans, E.; Hooker, S.; et al. 2023. Evaluating the Social Impact of Generative AI Systems in Systems and Society. arXiv preprint arXiv:2306.05949
Pith/arXiv arXiv 2023
-
[72]
St \"o ckl, A. 2023. Evaluating a synthetic image dataset generated with stable diffusion. In International Congress on Information and Communication Technology, 805--818. Springer
work page 2023
-
[73]
Sun, M.; Zhang, X.; Ma, J.; Xie, S.; Liu, Y.; and Philip, S. Y. 2023. Inconsistent Matters: A Knowledge-guided Dual-consistency Network for Multi-modal Rumor Detection. IEEE Transactions on Knowledge and Data Engineering
work page 2023
-
[74]
Tong, Z.; Song, Y.; Wang, J.; and Wang, L. 2022. Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training. Advances in neural information processing systems, 35: 10078--10093
2022
-
[75]
Touvron, H.; Lavril, T.; Izacard, G.; Martinet, X.; Lachaux, M.-A.; Lacroix, T.; Rozi \`e re, B.; Goyal, N.; Hambro, E.; Azhar, F.; Rodriguez, A.; Joulin, A.; Grave, E.; and Lample, G. 2023. LLaMA: Open and Efficient Foundation Language Models. ArXiv, abs/2302.13971
Pith/arXiv arXiv 2023
-
[76]
Uzun, L. 2023. ChatGPT and academic integrity concerns: Detecting artificial intelligence generated content. Language Education and Technology, 3(1)
work page 2023
-
[77]
von der Weth, C.; Abdul, A.; Fan, S.; and Kankanhalli, M. 2020. Helping Users Tackle Algorithmic Threats on Social Media: A Multimedia Research Agenda. In Proceedings of the 28th ACM International Conference on Multimedia, 4425--4434
work page 2020
-
[78]
Wang, Y.; Ma, F.; Jin, Z.; Yuan, Y.; Xun, G.; Jha, K.; Su, L.; and Gao, J. 2018. Eann: Event adversarial neural networks for multi-modal fake news detection. In Proceedings of the 24th acm sigkdd international conference on knowledge discovery & data mining, 849--857
work page 2018
-
[79]
Wang, Z.; Shan, X.; Zhang, X.; and Yang, J. 2022. N24News: A New Dataset for Multimodal News Classification. In Proceedings of the Language Resources and Evaluation Conference, 6768--6775. Marseille, France: European Language Resources Association
work page 2022
-
[80]
Wickens, T. D. 2001. Elementary signal detection theory. Oxford university press
work page 2001
This paper was first reviewed by deepseek-v4-flash on August 5, 2026.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.