Pith. sign in

REVIEW 3 major objections 5 minor 26 references

Do people rely on ChatGPT more than their peers to detect deepfake news?

T0 review · 3 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read People rely more on ChatGPT than on peers to detect deepfake news.

desk verdict A careful, preregistered study with a credible AI-over-peer reliance result that needs one more robustness check before the headline is bulletproof. read the letter →

arxiv 2608.01540 v1 pith:FR5XGDKP submitted 2026-08-02 econ.GN cs.CYq-fin.EC

classification econ.GNcs.CYq-fin.EC
keywords ChatGPTAIreliancedeepfakedetectionadvicetakingweightofhuman-AIinteractionmisinformationexperimentaleconomics
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper tries to establish that when people judge whether a news article is partly AI-written, they adjust their answer more toward advice labeled as coming from ChatGPT than toward a peer's answer. In the main experiment the mean weight of advice was 0.592 for ChatGPT advice versus 0.326 for peer advice, a gap that persists when the peer advice comes from earlier sessions. The paper also argues that the benefit of AI advice is conditional on its realized quality: higher-quality advice produces larger accuracy gains, and the AI label by itself does not. A 2025 session with the same procedure adds linguistic experts as a source, and participants then relied more on experts (WOA 0.551) than on ChatGPT (WOA 0.505), so the paper's overall ordering is experts > ChatGPT > peers in that wave. A sympathetic reader would care because whether people over- or under-trust AI detectors determines whether 'fight AI with AI' tools can actually reduce deepfake harm.

What carries the argument

The carrying object is the weight of advice (WOA), defined as $(Final - Initial)/(Advice - Initial)$, with undefined cases (initial equal to advice, 4.4% of observations) excluded. WOA converts each round into a scalar showing how far a participant moved toward the advice; 0.5 means equal weight. The paper compares WOA across between-subject treatments whose only difference is the stated source of the single piece of advice—GPT-4, a same-session peer, a peer from earlier sessions, or a linguistic expert—and pairs it with a researcher-side advice quality measure $Advq = 1 - |Advice - HMpro^*|/100$ to separate reliance from benefit. The activation–integration decomposition (two-stage Heckman m

What would settle it

Run the identical experiment but swap labels, showing ChatGPT-drawn advice as coming from a peer and peer-drawn advice as coming from ChatGPT. If WOA tracks the label, the source effect is real; if WOA tracks the values, the gap is mechanical. A cheaper check is comparing the distributions of |advice - initial| and advice dispersion across the AI and Human treatments.

Watch

Extended reading notes

Core claim

The central claim is an ordering of reliance by advice source in a controlled deepfake detection task. Participants first estimated the percentage of human-written content in each article, then saw one piece of advice and submitted a revised estimate. Mean WOA was 0.592 in the ChatGPT treatment and 0.326 in the peer treatment (Holm-adjusted p < 0.001), meaning participants moved more than halfway from their own estimate toward the AI advice but only about a third of the way toward the peer advice. The ordering is not simply 'AI over humans': in the 2025 within-year comparison, linguistic experts (0.551) were trusted more than ChatGPT (0.505), and the preregistered hypothesis that ChatGPT wou

Load-bearing premise

The reliance comparison assumes that the gap in weight of advice is caused by the label 'ChatGPT' versus 'peer', not by differences in the advice-value distributions generated by the two sources.

Editorial extensions

If this is right

  • If the reliance ordering holds, deploying ChatGPT-style detectors in news environments can improve people's final judgments, because participants actually move toward the AI estimate.
  • The same results imply this improvement is fragile: when AI advice is low quality, it can underperform peer advice; performance gains appear only when reliance and advice quality coincide.
  • The persistence of the peer gap in the preHuman treatment suggests the ChatGPT-over-peers effect is not an artifact of same-day social pressure or concern about being observed.
  • The 2025 expert-vs-ChatGPT reversal implies that 'AI versus human' is too coarse a frame; the relevant comparison is among advisor types, and relative trust can shift over time.
  • Since the human-written proportion does not systematically affect reliance or accuracy, the results apply roughly uniformly across fully fake, fully real, and mixed articles in this task.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • One extension the paper does not run is a label swap: presenting identical advice values under the two labels to test whether the WOA gap is driven by the source label or by distributional differences in the advice pools; this would directly isolate the mechanism.
  • If the ordering generalizes beyond the lab, content platforms could expect users to over-adjust to AI flags even when the detector is unreliable; a field test varying detector accuracy while holding the interface constant would test this.
  • The 2025 reversal suggests reliance may track perceived model competence rather than a fixed attitude toward AI; a natural follow-up is varying the stated model version (e.g., GPT-4 vs a newer model) while holding advice values identical.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. This paper reports a laboratory experiment (N=87 in the main study, plus N=133 in an additional study) on how people weigh advice from ChatGPT (GPT-4), human peers, and linguistic experts when detecting GPT-2-generated deepfake news. Participants first estimate the proportion of human-written content in each article, then receive advice labeled as coming from one of these sources, and then submit a revised estimate. Reliance is measured by weight of advice (WOA), defined as (Final – Initial)/(Advice – Initial). The central claim is Result 2: participants rely more on ChatGPT than on human peers (mean WOA 0.592 vs. 0.326, Holm-adjusted p<0.001). Secondary results concern the role of advice quality in driving performance improvements, the rejection of H2 (no effect of human-written proportion on reliance), and a mixed expert-vs-AI comparison across waves, with a within-2025 comparison favoring experts over ChatGPT (WOA 0.551 vs. 0.505, Holm-adjusted p=0.013). The paper is preregistered, reports Holm-adjusted p-values, and includes multiple robustness checks, but the main reliance comparison is vulnerable to a distributional confound between the AI and Human advice streams.

Significance. If the central claim holds, the paper makes a useful contribution to the behavioral literature on human-AI interaction and misinformation detection: it directly compares reliance on ChatGPT versus human peers in a controlled, incentivized task, something existing work has rarely done. The preregistration, Holm corrections, multiple waves, and the preHuman control condition are genuine methodological strengths. The paper also transparently reports that the preregistered hypothesis H3 (ChatGPT over linguistic experts) was not supported in the 2025 comparison. However, the significance of the main claim is currently weakened by a potential mechanical artifact: WOA is denominator-sensitive, and the AI and Human advice streams are constructed differently. Because the paper does not show that the source-label effect survives controls for the distribution of advice values and the advice–initial gap, the headline ordering cannot yet be attributed to the source label with confidence.

major comments (3)
  1. [§3.2.3 and §4.2] The central comparison in Result 2 is confounded by differences in the advice streams. In the AI treatment, advice is a random draw from a fixed pool of 24 GPT-4 responses per article; in the Human treatment, advice is a single peer's first identification. WOA = (Final − Initial)/(Advice − Initial) is mechanically sensitive to the distance between the advice and the participant's initial response. If ChatGPT's advice values happen to sit closer to participants' initial estimates (e.g., because GPT-4 uses the same surface cues as participants), a larger WOA could arise even with identical willingness to follow advice. The paper reports nearly identical mean advice quality (0.721 vs. 0.719) but does not report the dispersion of advice values, the distribution of |Advice − Initial| by treatment, or WOA conditional on that gap. The preHuman control matches the Human advice distribution but n
  2. [§6.2, Figure 12] The claim that performance improvement reflects the joint role of reliance and advice quality is partly mechanical. Because Imp = Accu2 − Accu1 and Accu = 1 − |Response − Truth|/100, moving from the initial response toward an advice value that is closer to the truth necessarily increases measured improvement. Thus, the significant Advq × WOA interaction in Figure 12 and the associated OLS regression (β=0.513) could arise by construction: high WOA means the final response is near the advice, and high Advq means that advice is near the truth. The paper should explicitly separate the behavioral mechanism (reliance as a decision process) from the arithmetic fact that using accurate advice improves accuracy. At minimum, this mechanical relationship should be acknowledged and, ideally, benchmarked against a predicted no-choice baseline where participants move toward advice with a fixed probabi
  3. [§5.2.3, Result 5] The ordering in Figure 10 (WOA_AI > WOA_Expert > WOA_AIadd > WOA_Human ≈ WOA_preHuman) is presented as a single ranking, but the AI and Expert treatments are not directly comparable: AI was run in 2023–2024, while Expert was run in 2025. The only clean within-year comparison is Expert vs. AIadd (p=0.013), and there the result is opposite to the preregistered H3. This is not itself an error—the paper acknowledges the cross-wave caveat—but Result 5's summary sentence ('participants assign the greatest reliance to linguistic experts, followed by ChatGPT') overstates the evidence. The summary should be restricted to the within-year comparison or presented as a mixed pattern, consistent with the abstract.
minor comments (5)
  1. [§3.2.3] The AI treatment draws advice from a 24-response pool, while the Human treatment draws from another participant's single first identification. The paper should state whether the AI advice values are i.i.d. draws from this pool, how the 24 responses were generated (e.g., temperature), and whether any of them were excluded for being identical or extreme. This information is relevant to the distributional concern above.
  2. [§4.1.1, Figure 4] The paper says AI advice quality is 'statistically significant' at 0.721 vs. 0.719 with Holm-adjusted p=0.0022. This is a numerically tiny difference; the authors should be careful not to overinterpret practical significance. The later statement that AI 'provides higher-quality advice' is technically correct but the magnitude is negligible.
  3. [Table 2] The table does not define the variable HMpro, although it is used in later regression tables. Also, the variable name freqGPT is defined as 'Average days per week using ChatGPT' but the range 0–7 is inconsistent with SQ6's wording. Clarify.
  4. [Figure 3 caption] The caption says 'red points indicate the original AI advice ... and the 95% CI', but no confidence interval is visible in the figure. Either the figure or the caption is incomplete.
  5. [References] Several URLs are inline (e.g., the Japanese fake-news dataset in §3.2.1, and OpenAI Help Center in §2.1.2) rather than in the reference list. For a journal submission, these should be formal references or footnotes.

Circularity Check

1 steps flagged · score 4.0 of 10

Central reliance result is empirical, but Section 6.2's performance-improvement mechanism is an algebraic artifact of the WOA/Imp definitions.

  1. self definitional [Section 6.2 (with WOA definition in Section 4.2 and Imp definition in Section 4.1.2)]
    "W OA_i,r = (Final Response_i,r − Initial Response_i,r)/(Advice_i,r − Initial Response_i,r) ... Imp = Accu2 − Accu1 ... An OLS regression of Imp on Advq, WOA, and Advq×WOA yields a positive and significant interaction (β= 0.513, p <0.001; participant-clustered standard errors)."

    By definition, Final = Initial + WOA·(Advice − Initial). Substituting into Imp = (|Initial − Truth| − |Final − Truth|)/100 gives, in the common case where the final response lies between initial and advice, Imp = WOA·(Advq − Accu1) up to the 1/100 scaling. Thus a positive Advq×WOA interaction is forced by the definitions of WOA, Advq, and Imp: weighting advice that is closer to the truth improves accuracy algebraically, and weighting poor advice worsens it by the same algebra. The paper interprets this as a behavioral 'gating role of reliance,' but it is a restatement of the measurement equations, not an independent empirical discovery.

full rationale

The paper's central claim (Result 2: WOA 0.592 in AI treatment vs 0.326 in Human treatment) is a genuine between-subject comparison that could have gone either way; it does not reduce to a fitted parameter or to a self-citation. The preHuman control further strengthens the empirical content. The only self-citation (footnote 18, Fu and Hanaki 2025) is supplementary and not load-bearing. The skeptic's concern about advice-stream distributional differences affecting WOA comparability is a measurement-validity issue, not a circularity, because the paper does not define the treatment effect in terms of the outcome. However, the Section 6.2 mechanism claim — that performance improvement arises from the joint role of reliance and advice quality — is partially circular: the Advq×WOA interaction in predicting Imp is a mathematical consequence of the definitions of WOA, Advq, and Imp. This is a secondary interpretive result rather than the main reliance ordering, so the overall circularity is moderate rather than severe.

Assumptions & free parameters 3 free parameters · 5 assumptions · 0 invented entities

The paper introduces no fitted physical or economic parameters and no invented constructs: WOA and Advq are computed from observed responses and true proportions, and the payoff rule is a standard quadratic scoring rule. The main unverified inputs are behavioral assumptions (source-label credibility, validity of WOA, unobserved perceived quality) and the representativeness of GPT-2 stimuli, all disclosed. The effect-size and mechanism interpretations rest on these assumptions plus the unpublished appendix regressions.

free parameters (3)
  • Uninformative baseline value = 50
    Accuracy benchmark predicting a fixed HM pro of 50 for every round (red dashed line, Figures 4 and 7). Chosen by hand as a benchmark, not fitted; it is not part of the central claim.
  • Payoff function constants = 2300 JPY fixed, 0.3 JPY penalty per squared error
    Quadratic scoring rule parameters in Section 3.4. Design parameters that incentivize accurate reporting; their specific values are not claimed to drive any qualitative result.
  • WOA tercile cutoffs = not reported in text
    Section 6.2 and Figure 12 classify reliance into low/medium/high WOA groups by terciles of the sample WOA distribution; the cut points are sample-dependent.
assumptions (5)
  • domain assumption WOA measures reliance validly in the judge-advisor paradigm
    Central measure throughout (Section 4.2): WOA = (Final - Initial) / (Advice - Initial). Assumes movement toward advice reflects reliance and that undefined cases (4.4%) can be excluded without biasing treatment comparisons; robustness is delegated to the Heckman activation-integration model in Online Appendix D.
  • domain assumption Participants believed the advice source labels (ChatGPT, peer, linguistic expert)
    Sections 3.2.3 and 5.1: the treatment is the source label. If participants did not credit the labels, the WOA differences would not identify source effects. The AIadd and preHuman controls partially support this assumption.
  • domain assumption Advice quality can be measured against the true HM pro, which participants never observe
    Section 4.1.1: Advq = 1 - |Advice - truth| / 100 uses the true value; the paper assumes ex post quality can proxy for the unobserved perceived quality that shapes behavior.
  • ad hoc to paper GPT-2-generated Japanese articles represent deepfake news for the reliance question
    Sections 3.2.1 and 7: stimuli are GPT-2-generated or mixed (human-written first part, AI-generated second part). The generality to frontier-model deepfakes is explicitly acknowledged as a limitation, since GPT-2 output is far easier to detect than current models.
  • standard math Standard econometric assumptions for OLS, probit, Mann-Whitney, Wilcoxon, and Heckman models
    Used throughout Sections 4-6 and Online Appendices A-D; standard assumptions, not examined in the main text.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Do people rely on ChatGPT more than their peers to detect deepfake news?." pith.science (2026). https://pith.science/paper/FR5XGDKP

@misc{pith2026260801540,
  author       = {Pith},
  title        = {Pith review of: Do people rely on ChatGPT more than their peers to detect deepfake news?},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/FR5XGDKP}},
  note         = {Machine review of arXiv:2608.01540}
}
read the original abstract

This experimental study investigates how people rely on different sources of advice when detecting AI-generated fake news (deepfake news). In a laboratory deepfake detection task, student participants identified the proportion of human-written (non-AI-generated) content in synthetic deepfake news articles and received advice from ChatGPT (GPT-4), human peers, or linguistic experts. The results show that participants rely more on ChatGPT than on human peers when detecting GPT-2-generated deepfake news. Participants also rely more on linguistic experts than on peers, while the relative reliance on experts versus ChatGPT is mixed across experimental waves, potentially reflecting time trends in beliefs about AI-based detection. Importantly, in the additional experiment conducted in 2025 under the same experimental procedure, participants relied more on linguistic experts than on ChatGPT. Moreover, performance improvements reflect the joint role of reliance and advice quality, arising primarily when participants rely on high-quality advice. Overall, relying on AI to detect AI-generated deepfakes can improve detection outcomes, but only when AI-based detection tools are of sufficiently high quality. These findings highlight the dual role of GAI as both a source of deepfakes and a tool for mitigating related risks.

Figures

Figures reproduced from arXiv: 2608.01540 by the authors.

Figure 1
Figure 1. Experimental Procedure: (a) Overall Procedure and (b) Main Task [PITH_FULL_IMAGE:figures/full_fig_p016_1.png] view at source ↗
Figure 2
Figure 2. Instructional Illustration of News Composition Shown to Participants [PITH_FULL_IMAGE:figures/full_fig_p018_2.png] view at source ↗
Figure 3
Figure 3. HMpro∗ , the original AI advice set and Round Number Note: The points marked with “▲” represent the true HMpro values in the 30-round tasks (HMpro∗ ). The red points indicate the original AI advice (24 data points in each round) generated before the experiment, and the 95% CI. 3.2.2 First Identification After each participant had read the news, they were asked to report a number between 0 and 100 to represent their … view at source ↗
Figures from the paper (9 more)
Figure 4
Figure 4. Figure 4: Overall Performance Note: + p < 0.1, * p < 0.05, ** p < 0.01, *** p < 0.001. “n.s.” means that the difference is not statistically significant at the 0.1 level. Error bars denote 95% confidence intervals across participants. Accu1, Advq, and Accu2 are compared within t…
Figure 5
Figure 5. Figure 5: Performance Improvement Note: ImpUP is compared across treatments using Fisher’s exact test, and Imp and PRE are compared across treatments using the Mann–Whitney U test. For PRE, 124 observations are excluded because participants with Accu1 = 1 have undefined PRE valu…
Figure 6
Figure 6. Figure 6: WOA Across Treatments Note: + p < 0.1, * p < 0.05, ** p < 0.01, *** p < 0.001. “n.s.” means that the difference is not statistically significant at the 0.1 level. Error bars denote 95% confidence intervals across participants. WOA is compared across treatments using th…
Figure 7
Figure 7. Figure 7: Mean Advq Across Treatments (Excluding Expert) Note: + p < 0.1, * p < 0.05, ** p < 0.01, *** p < 0.001. “n.s.” means that the difference is not statistically significant at the 0.1 level. Error bars denote 95% confidence intervals across participants. The Expert treatm…
Figure 8
Figure 8. Figure 8: Mean Accu1 Across Treatments Note: + p < 0.1, * p < 0.05, ** p < 0.01, *** p < 0.001. “n.s.” means that the difference is not statistically significant at the 0.1 level. Error bars denote 95% confidence intervals across participants. Accu1 is compared across treatments…
Figure 9
Figure 9. Figure 9: Mean Accu2 Across Treatments Note: + p < 0.1, * p < 0.05, ** p < 0.01, *** p < 0.001. “n.s.” means that the difference is not statistically significant at the 0.1 level. Error bars denote 95% confidence intervals across participants. Accu2 is compared across treatments…
Figure 10
Figure 10. Figure 10: WOA Across Treatments Note: + p < 0.1, * p < 0.05, ** p < 0.01, *** p < 0.001. “n.s.” means that the difference is not statistically significant at the 0.1 level. Error bars denote 95% confidence intervals across participants. WOA is compared across treatments using t…
Figure 11
Figure 11. Figure 11: Mean prefAdvSrc Across Treatments Note: +p < 0.1, ∗p < 0.05, ∗∗p < 0.01, ∗∗∗p < 0.001. “n.s.” indicates that the difference is not statistically significant at the 10% level. Error bars denote 95% confidence intervals across participants. Pairwise comparisons across t…
Figure 12
Figure 12. Figure 12: Advice Quality, WOA & Performance Improvement [PITH_FULL_IMAGE:figures/full_fig_p041_12.png]

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

26 extracted references · 20 canonical work pages

  1. [2]

    Can GPT-4 Models Detect Misleading Visualizations?

    J. Alexander, P. Nanda, K.-C. Yang, and A. Sarvghad. Can gpt-4 models detect mis- leading visualizations?arXiv preprint arXiv:2408.12617,

  2. [4]

    Accessed: 2025-11-01

    URLhttps://www.abc.net.au/news/2025-10-20/universities-using-ai-t o-detect-students-cheating/105905804. Accessed: 2025-11-01. A. Bhattacharjee and H. Liu. Fighting fire with fire: can chatgpt detect ai-generated text?ACM SIGKDD Explorations Newsletter, 25(2):14–21,

  3. [7]

    Updated: 2023-02-02

    URLhttps://www.indiatoday.in/technology/new s/story/oh-god-open-ai-tool-that-identifies-text-written-chatgpt-bel ieves-bible-was-written-by-ai-2329163-2023-02-01. Updated: 2023-02-02. C. Chauhan and G. Currie. The impact of generative artificial intelligence on research integrity in scholarly publishing.The American journal of pathology, 194(12):2234– 2238,

  4. [8]

    Chen and K

    C. Chen and K. Shu. Can llm-generated misinformation be detected?arXiv preprint arXiv:2309.13788,

  5. [9]

    A. T. Y. Chong, H. N. Chua, M. B. Jasser, and R. T. Wong. Bot or human? de- tection of deepfake text with semantic, emoji, sentiment and linguistic features. In 2023 IEEE 13th International Conference on System Engineering and Technology (ICSET), pages 205–210. IEEE,

  6. [10]

    Explains false positives such as the U.S

    URLhttps://arstechn ica.com/information-technology/2023/07/why-ai-detectors-think-the-u s-constitution-was-written-by-ai/. Explains false positives such as the U.S. Constitution being flagged as AI-generated. A. Friggeri, L. Adamic, D. Eckles, and J. Cheng. Rumor cascades. Inproceedings of the international AAAI conference on web and social media, volume ...

  7. [11]

    Accessed: 2025-11-01

    URLhttps://spectrumlocalnews.com/nys/central-ny/news/2025/05 /14/ub-student-says-false-ai-use-accusation-caused-stress--inspired-p etition. Accessed: 2025-11-01. N. Harvey and I. Fischer. Taking advice: Accepting help, improving judgment, and sharing responsibility.Organizational behavior and human decision processes, 70(2): 117–133,

  8. [12]

    Trade Press Services blog post. S. Koka, A. Vuong, and A. Kataria. Evaluating the efficacy of large language models in detecting fake news: A comparative analysis.arXiv preprint arXiv:2406.06584,

Show all 26 references
  1. [15]

    Sallami, Y.-C

    D. Sallami, Y.-C. Chang, and E. A ¨ ımeur. From deception to detection: The dual roles of large language models in fake news.arXiv preprint arXiv:2409.17416,

  2. [18]

    Accessed: 2025-11-01

    URLhttps://www.adelaidenow.com.au/education/h igher-education/robocheating-fiasco-saw-australian-catholic-universit y-students-falsely-accused-of-using-ai-by-an-unreliable-ai-tool/news -story/4a08732c84499263a709ec3bb1980802. Accessed: 2025-11-01. The Courier-Mail. ‘impossible...

  3. [19]

    Accessed: 2025-11-01

    URLhttps://www.couriermail.com.au/queensland-education/impossible-n ew-ai-detection-tool-slammed-by-experts/news-story/c2443d79a81ff2fea 705b6e8baf8377a. Accessed: 2025-11-01. T. T. K. Tse, N. Hanaki, and B. Mao. Beware the performance of an algorithm be- fore relying on it: E...

  4. [20]

    Uchendu, S

    A. Uchendu, S. Venkatraman, T. Le, and D. Lee. Catch me if you gpt: Tutorial on deepfake texts. InProceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 5: Tutorial Abstracts), pages 1–7,

  5. [21]

    Umbach, N

    R. Umbach, N. Henry, G. F. Beard, and C. M. Berryessa. Non-consensual synthetic intimate imagery: Prevalence, attitudes, and knowledge in 10 countries. InProceed- ings of the 2024 CHI Conference on Human Factors in Computing Systems, pages 1–20,

  6. [22]

    Vodrahalli, R

    59 K. Vodrahalli, R. Daneshjou, T. Gerstenberg, and J. Zou. Do humans trust advice more if it comes from ai? an analysis of human-ai interactions. InProceedings of the 2022 AAAI/ACM Conference on AI, Ethics, and Society, pages 763–777,

  7. [23]

    Workshop, T

    B. Workshop, T. L. Scao, A. Fan, C. Akiki, E. Pavlick, S. Ili´ c, D. Hesslow, R. Castagn´ e, A. S. Luccioni, F. Yvon, et al. Bloom: A 176b-parameter open-access multilingual language model.arXiv preprint arXiv:2211.05100,

  8. [24]

    Z. Zeng, S. Liu, L. Sha, Z. Li, K. Yang, S. Liu, D. Gaˇ sevi´ c, and G. Chen. Detecting ai-generated sentences in human-ai collaborative hybrid texts: Challenges, strategies, and insights.arXiv preprint arXiv:2403.03506,

  9. [25]

    60 P. Zhang. Taking advice from chatgpt.arXiv preprint arXiv:2305.11888,

  10. [26]

    Zhang, Y

    Y. Zhang, Y. Ma, J. Liu, X. Liu, X. Wang, and W. Lu. Detection vs. anti-detection: Is text generated by ai detectable? InInternational Conference on Information, pages 209–222. Springer, 2024a. Z. Zhang, W. Qin, and B. A. Plummer. Machine-generated text localization.arXiv prep...

  11. [1989]

    Somoray, D

    K. Somoray, D. J. Miller, and M. Holmes. Human performance in deepfake detection: A systematic review.Human Behavior and Emerging Technologies, 2025(1):1833228,

  12. [2019]

    Chadha, V

    A. Chadha, V. Kumar, S. Kashyap, and M. Gupta. Deepfake: an overview. InProceed- ings of second international conference on computing, communications, and cyber- security: IC4S 2020, pages 557–566. Springer,

  13. [2021]

    doi: 10.2 4251/hicss.2021.496. Y. Mirsky and W. Lee. The creation and detection of deepfakes: A survey.ACM computing surveys (CSUR), 54(1):1–41,

  14. [2022]

    Laurier, A

    L. Laurier, A. Giulietta, A. Octavia, and M. Cleti. The cat and mouse game: The ongoing arms race between diffusion models and detection methods.arXiv preprint arXiv:2410.18866,

  15. [2023]

    Agrawal, S

    V. Agrawal, S. Kandul, M. Kneer, and M. Christen. From oecd to india: Explor- ing cross-cultural differences in perceived trust, responsibility and reliance of ai and human experts.arXiv preprint arXiv:2307.15452,

  16. [2024]

    J. Y. Bo, S. Wan, and A. Anderson. To rely or not to rely? evaluating interventions for appropriate reliance on large language models. InProceedings of the 2025 CHI Conference on Human Factors in Computing Systems, pages 1–23,

  17. [2025]

    E. M. Bender, T. Gebru, A. McMillan-Major, and S. Shmitchell. On the dangers of stochastic parrots: Can language models be too big? InProceedings of the 2021 ACM conference on fairness, accountability, and transparency, pages 610–623,

  18. [2026]

    X. Sun, R. Ma, X. Zhao, Z. Li, J. Lindqvist, A. E. Ali, and J. A. Bosch. Trusting the search: unraveling human trust in health information from google and chatgpt. arXiv preprint arXiv:2403.09987,

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.