Pith. sign in

REVIEW 3 major objections 5 minor 8 references

Conversational AI may weaken users' critical judgment without ever lying, through traits it genuinely has.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-03 11:09 UTC pith:Z32ISADF

load-bearing objection The Cognitive Trojan Horse / honest non-signals idea is a real conceptual contribution to AI-safety/HCI, well-written and honest about its limits, though the core premise that LLM fluency is information-free is asserted more than demonstrated. the 3 major comments →

arxiv 2601.07085 v2 pith:Z32ISADF submitted 2026-01-11 cs.HC cs.AIcs.CY

The AI Cognitive Trojan Horse: How Large Language Models May Bypass Human Epistemic Vigilance

classification cs.HC cs.AIcs.CY
keywords epistemic vigilancehonest non-signalslarge language modelstrust calibrationprocessing fluencycognitive offloadingsycophancyhuman-AI interaction
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper argues that conversational AI, trained to be fluent, helpful, and agreeable, may slip past the cognitive screening humans apply to communicated information—not by deceiving us, but by displaying genuine qualities that our evaluation systems were calibrated to treat as signs of trustworthy human speakers. It names these qualities 'honest non-signals': real fluency, real helpfulness, real apparent disinterest, which in a human cost effort and therefore signal reliability, but in a language model are cheap by-products of text generation. The paper identifies four bypass routes: fluency that no longer indicates understanding, competence without personal stakes, delegation of evaluation itself to the AI, and optimization-driven sycophancy. If the hypothesis holds, even a perfectly accurate AI could weaken users' critical engagement, and AI safety becomes a problem of calibrating trust rather than only preventing deception.

Core claim

On its own terms, the paper's central claim is that LLM-based assistants occupy a category human epistemic vigilance has no templates for. Humans evolved and learned to spot unreliable communicators by watching for costly signals: fluent speech implies knowledge, helpfulness implies benevolent intent, disinterest implies no hidden agenda. Language models produce all these characteristics as baseline properties of their training and optimization, so the characteristics are genuine but carry none of the information they carry in humans. The paper proposes that this 'parameter mismatch'—not lies, hallucinations, or malicious intent—may be what makes AI persuasively effective: defenses stand dow

What carries the argument

The central object is the 'honest non-signal': a genuine characteristic of an AI system (fluency, helpfulness, warmth, apparent disinterest, responsiveness) that is computationally cheap to produce and therefore carries none of the reliability information the same characteristic carries in a human, where it is costly. The argument runs through the theory of epistemic vigilance—the proposed parallel cognitive process that monitors communication for reasons to doubt—and claims that AI output falls outside the parameter space that vigilance is calibrated to evaluate. The four mechanisms (fluency decoupled from understanding, trust-competence without stakes, cognitive offloading of evaluation, a

Load-bearing premise

The load-bearing premise is that human vigilance treats fluency, helpfulness, warmth, and apparent disinterest as reliable signals because in humans they are costly to produce—and that language models make them so cheaply that the signals carry no information; if people can learn to read AI's cheap fluency as meaningless, or if AI fluency actually tracks training-data quality, the bypass claim weakens.

What would settle it

An experiment comparing equally fluent human-written and AI-generated texts on the same topic, measuring perceived credibility and belief change, would test the core claim: if AI text earns no credibility premium over identical human text, or if users' skepticism recalibrates after a few salient AI errors, the 'honest non-signals' mechanism would not be supported.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • AI safety must include trust calibration, not just accuracy and alignment, since even a perfectly accurate AI could reduce users' evaluative engagement.
  • Adding uncertainty markers, competence boundaries, or disfluencies to AI output may reduce fluency-driven credibility effects.
  • If users delegate evaluation itself to AI, interface designs that force engagement before showing output could help preserve vigilance.
  • Sophisticated and heavy AI users may face greater cumulative exposure to bypass mechanisms, not less risk.
  • Whether humans can learn to recalibrate trust toward AI—through repeated salient errors or explicit training—remains an open empirical question.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • Beyond the paper's claims: if honest non-signals work by parameter mismatch, then making AI's lack of stakes persistently explicit—for example through non-anthropomorphic framing or visible uncertainty signals—should dampen trust-cue activation; this design hypothesis is not tested in the paper.
  • An extension the paper leaves implicit is that the same mechanism could apply to other AI-generated media (voice, video), where fluency and apparent warmth are also cheap, potentially broadening the risk surface beyond text.
  • The evolutionary-mismatch framing suggests that even if individuals cannot consciously override the fluency heuristic, collective norms or institutional labels on AI content might serve as external calibration tools; the paper gestures at this but does not develop it.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper proposes the 'Cognitive Trojan Horse' hypothesis: LLM-based conversational AI may bypass human epistemic vigilance not through deception or inaccuracy, but by presenting 'honest non-signals'—genuine fluency, helpfulness, warmth, and apparent disinterest that are computationally cheap for the AI but carry the same apparent informational value as costly human signals. The paper grounds this in Sperber et al.'s epistemic vigilance theory, identifies four bypass mechanisms (fluency/understanding decoupling, trust-competence without stakes, cognitive offloading of evaluation, and optimization-induced sycophancy), reviews converging but limited evidence, specifies boundary conditions, and offers testable predictions including the 'intelligent user trap'. The paper is explicitly hypothesis-generating and carefully hedged.

Significance. If the framework holds, it would reframe AI safety as a problem of trust calibration rather than solely deception prevention, and it generates concrete, falsifiable predictions (e.g., disfluency interventions, cognitive forcing, sycophancy detection) that could guide empirical work. The paper's strengths are its clear conceptual scaffolding, its explicit boundary conditions in §5.1, its transparent acknowledgement of the evidence's limitations, and its productive connection to established constructs (Sperber; Fiske; Friestad & Wright; Risko & Gilbert). The main risk is that the central premise—that human fluency/helpfulness are costly and informative while LLM equivalents are cheap and non-informative—is asserted rather than demonstrated; if that premise fails, the bypass is a transient unfamiliarity rather than a fundamental calibration mismatch.

major comments (3)
  1. [§3, §3.1] The load-bearing claim is that human fluency, helpfulness, and apparent disinterest are 'costly to produce' and therefore informative, while LLM versions are 'computationally trivial' and therefore non-informative. This dichotomy is asserted, not demonstrated. LLM fluency is not necessarily non-informative: on many topics, fluent outputs correlate with the model's training-data coverage and thus with reliability, and human fluency is not always costly (experts can speak fluently with little effort). The paper needs to either formalize the conditions under which LLM fluency is non-informative or engage with evidence that AI fluency predicts accuracy. Without this, the central 'honest non-signal' concept lacks a solid foundation.
  2. [§3.1, §5] The claim that 'the cheapness is invisible' is a crucial empirical premise. Users know they are interacting with an AI and may attribute low effort or lack of genuine understanding to it, thereby discounting fluency. The paper cites automation bias and persuasion knowledge but does not address studies showing that AI-labeled content is often trusted less than identical human-labeled content. The paper's own acknowledgement in §2.1 that deploying organizations have self-interest also complicates the 'no visible self-interest' argument. The manuscript should either present evidence that AI fluency is not discounted or refine the scope to contexts where users do not apply source-based discounting.
  3. [§5, §5.1] The empirical evidence is appropriately described as suggestive, but the manuscript does not specify what would falsify the Cognitive Trojan Horse hypothesis as opposed to a simple familiarity or novelty effect. The boundary condition on 'repeated error feedback' in §5.1 is particularly important: if salient, attributable errors can recalibrate trust, then the bypass is a calibration lag rather than a 'parameter mismatch'. The paper should state explicit, observable predictions that would distinguish the proposed mechanism from the alternative that users gradually learn to discount AI fluency, and it should indicate which existing or proposed experiments could adjudicate between these.
minor comments (5)
  1. [§5] The phrase 'a non-trivial fraction of the time' is vague; please provide the actual preference rate from Sharma et al. or clarify the range.
  2. [§4.1] The discussion of the fluency-truth heuristic would benefit from noting that the cited experiments (Reber & Unkelbach 2010; Fazio et al. 2015) used human-generated or repeated materials; whether the same effect transfers to AI-generated text is a further empirical question the paper should flag more explicitly.
  3. [§6] The 'intelligent user trap' is clearly labeled a speculation, which is good. However, the transfer from motivated numeracy in political reasoning to AI influence remains loose; a sentence noting the boundary conditions (e.g., topic relevance, identity relevance) would improve precision.
  4. [§7] The research directions are useful but could be prioritized. In particular, the disfluency intervention and the calibration/falsification studies are more central than the longitudinal trust studies; a brief ordering would help readers.
  5. [Abstract/§1] The phrase 'a gift that triggers no suspicion because it matches every expectation of what a genuine gift should look like' is vivid but redundant with the earlier Trojan Horse analogy; consider trimming.

Circularity Check

0 steps flagged

No significant circularity: the Cognitive Trojan Horse hypothesis is built on external theories and empirical findings, and its central claims are explicitly framed as untested speculation rather than as predictions derived from fitted inputs or self-citations.

full rationale

The paper contains no equations, fitted parameters, or derivation chain in which an output is defined as its own input. Its central construct, 'honest non-signals,' is introduced as a hypothesis grounded in external literatures (Sperber et al. on epistemic vigilance, Reber & Unkelbach on fluency, Fiske et al. on warmth/competence, Hackenburg et al., Sharma et al., Gerlich) rather than in the author's own prior results. The claim that LLM fluency is computationally cheap while appearing costly is an asserted empirical premise, not a conclusion forced by definition; the paper even acknowledges the possibility of learned calibration and identifies boundary conditions, stating that the framework is 'hypothesis-generating rather than hypothesis-confirming.' The only self-referential element is the AI-use statement, which is transparent and does not enter the argument. Unsupported or untested premises may be a scientific weakness, but they do not constitute circularity under the specified standards.

Axiom & Free-Parameter Ledger

0 free parameters · 5 axioms · 0 invented entities

The framework's central claim rests on domain assumptions about human vigilance calibration, the informational value of LLM-generated fluency, and the reliability of cited recent empirical studies. There are no fitted numerical parameters and no invented physical entities; the only new conceptual object is the label 'honest non-signals.'

axioms (5)
  • domain assumption Human epistemic vigilance is calibrated to evaluate human communicators, where fluency/helpfulness/warmth are costly signals
    The entire honest non-signals argument depends on the premise that in humans these traits are costly and therefore informative, while in LLMs they are cheap. Invoked in Sections 1, 3, 3.1.
  • domain assumption LLM outputs lack corresponding stakes and self-interest, so they fail to trigger doubt cues
    Used to argue AI presents no visible motive; Section 3.1. Does not address that deploying organizations have interests.
  • domain assumption Autoregressive token prediction entails absence of genuine uncertainty and therefore fluency is not connected to knowledge
    Section 3 and 3.2 claim LLMs 'do not experience uncertainty' — a computational-level assumption.
  • domain assumption Sperber et al.'s theory of epistemic vigilance accurately describes human cognition
    The framework is built on this background theory; adopted without critique.
  • domain assumption Key empirical findings cited (Hackenburg et al. 2025; Sharma et al. 2024) are valid and represent the current state
    The framework's motivation and sycophancy mechanism rely on these recent studies, which the paper acknowledges await replication.

pith-pipeline@v1.3.0-alltime-deepseek · 10428 in / 10110 out tokens · 95697 ms · 2026-08-03T11:09:48.321049+00:00 · methodology

0 comments
read the original abstract

Large language model (LLM)-based conversational AI systems present a challenge to human cognition that current frameworks for understanding misinformation and persuasion do not adequately address. This paper proposes that a significant epistemic risk from conversational AI may lie not in inaccuracy or intentional deception, but in something more fundamental: these systems may be configured, through optimization processes that make them useful, to present characteristics that bypass the cognitive mechanisms humans evolved to evaluate incoming information. The Cognitive Trojan Horse hypothesis draws on Sperber and colleagues' theory of epistemic vigilance -- the parallel cognitive process monitoring communicated information for reasons to doubt -- and proposes that LLM-based systems present 'honest non-signals': genuine characteristics (fluency, helpfulness, apparent disinterest) that fail to carry the information equivalent human characteristics would carry, because in humans these are costly to produce while in LLMs they are computationally trivial. Four mechanisms of potential bypass are identified: processing fluency decoupled from understanding, trust-competence presentation without corresponding stakes, cognitive offloading that delegates evaluation itself to the AI, and optimization dynamics that systematically produce sycophancy. The framework generates testable predictions, including a counterintuitive speculation that cognitively sophisticated users may be more vulnerable to AI-mediated epistemic influence. This reframes AI safety as partly a problem of calibration -- aligning human evaluative responses with the actual epistemic status of AI-generated content -- rather than solely a problem of preventing deception.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

8 extracted references · 1 linked inside Pith

  1. [1]

    The results were striking: AI-generated persuasive messages changed attitudes significantly more than human-written equivalents

    Introduction In 2025, Hackenburg and colleagues published findings from a study involving nearly 77,000 participants across 19 different large language models (Hackenburg et al., 2025). The results were striking: AI-generated persuasive messages changed attitudes significantly more than human-written equivalents. The effect sizes exceeded what researchers...

  2. [2]

    The common intuition—that humans believe what they hear unless given specific reason to doubt—turns out to be empirically inadequate

    Epistemic Vigilance: Evolved and Learned Defenses Understanding how AI might bypass human epistemic defenses requires first understanding what those defenses are and how they operate. The common intuition—that humans believe what they hear unless given specific reason to doubt—turns out to be empirically inadequate. Sperber et al. (2010) proposed a more s...

  3. [3]

    This is not a claim about AI deception

    Why AI May Fall Outside the Detection Perimeter Having established how epistemic vigilance operates, the argument now turns to its core theoretical claim: that LLM-based conversational AI may present a configuration of characteristics for which human vigilance 6 has no appropriate response. This is not a claim about AI deception. Rather, it is a claim abo...

  4. [4]

    These mechanisms are not mutually exclusive; they likely operate simultaneously and reinforcingly in real interactions

    Mechanisms of Potential Bypass The theoretical arguments above suggest four primary mechanisms by which LLM-based systems may bypass epistemic vigilance. These mechanisms are not mutually exclusive; they likely operate simultaneously and reinforcingly in real interactions. Nor should they be taken as exhaustive—they represent salient pathways that emerge ...

  5. [5]

    Empirical Evidence, Limitations, and Boundary Conditions The theoretical framework presented here generates empirical predictions that can be compared to existing evidence. While direct tests of the Cognitive Trojan Horse hypothesis remain to be conducted, converging findings from multiple research literatures are consistent with the mechanisms proposed. ...

  6. [6]

    https://doi.org/10.3390/soc15010006 15 Goddard, K., Roudsari, A., & Wyatt, J. C. (2012). Automation bias: A systematic review of frequency, effect mediators, and mitigators. Journal of the American Medical Informatics Association, 19(1), 121–127. https://doi.org/10.1136/amiajnl-2011-000089 Goldstein, J. A., Sastry, G., Musser, M., DiResta, R., Gentzel, M....

  7. [7]

    This section outlines priority research directions organized around the core mechanisms and concepts developed above

    Research Directions The Cognitive Trojan Horse hypothesis generates specific testable predictions that could advance our understanding of human-AI interaction. This section outlines priority research directions organized around the core mechanisms and concepts developed above. Research on the fluency mechanism should test whether AI-generated content rece...

  8. [8]

    The risk is not only that these systems might deceive us—though they sometimes do, through hallucination or manipulation

    Conclusion: Calibration, Not Just Deception The Cognitive Trojan Horse hypothesis reframes a significant epistemic risk from conversational AI. The risk is not only that these systems might deceive us—though they sometimes do, through hallucination or manipulation. Nor is it that they might just be misaligned with human values—though alignment remains cru...