{"id":"49e090f9-c29a-4efc-8064-4b9dcb94352b","arxiv_id":"2505.13418","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Non-expert humans and LLMs both often miss dementia in transcribed picture descriptions, but LLMs use a wider set of language cues closer to clinical markers, while humans rely on a few simple and sometimes misleading signals.","lead":"Researchers asked 27 non-expert people and three AI chatbots to read 514 short descriptions of a picture and guess whether each speaker had dementia. They then labeled 38 language features in each text and built simple statistical models to see which cues actually drove each group's guesses.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"GPT-4o both annotates the features and contributes to the LLM perception vote, so the 'richer LLM cues' claim may be a self-consistency artifact rather than a real perception difference.","rationale":"The reader's weakest assumption was that the 38 expert-guided binary features annotated by GPT-4o faithfully capture the cues driving human and LLM judgments, and that low human R2 could reflect bad measurement. I share that concern but emphasize a sharper and more specific version: GPT-4o is both the feature annotator and one of the three LLMs whose majority vote defines LLM perception. This creates a direct measurement-dependence channel that can inflate the LLM model's fit even if the features are individually valid. The reader noted this confound in the rationale but did not make it the primary weakest assumption. My proposed check—re-annotating features with an independent LLM and/or excluding GPT-4o from the perception majority—directly tests whether the 'richer LLM feature set' survives without GPT-4o on both sides. Since the paper is otherwise transparent, the limitations section acknowledges LLM-annotation sensitivity, and the central comparison is exploratory rather than causal, the appropriate verdict remains CONDITIONAL. My read does not move the reader's verdict; it sharpens the condition under which the headline should be accepted.","tokens_in":1796,"tokens_out":836,"duration_ms":44887,"concrete_test":"Re-run the full pipeline with Gemini-1.5-Pro as the feature annotator for all 514 transcripts, using the same 38 prompts and keeping the original three-model majority as the LLM perception target; then refit the stepwise logistic regression and compare McFadden R2 and the significant feature set against Figure 3 and Table 3. For a stricter version, also redefine LLM perception as the majority of LLaMA-3 and Gemini-1.5-Pro only, excluding GPT-4o from the perception label. If the alternative-annotator R2 remains near 0.5 with the same four feature categories represented, the concern is resolved; if it drops substantially or the significant features narrow, the claim that LLMs draw on a richer, more nuanced feature set needs qualification.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central asymmetry—humans use a narrow, sometimes misleading cue set while LLMs use a richer set that 'aligns more closely with clinical patterns'—is established by comparing logistic-regression fits over 38 binary features annotated by GPT-4o (Section 5.2) to perception labels that include GPT-4o's own majority vote (Section 4.2). Because the same model constructs the feature representation and participates in the LLM perception label, the LLM model's very high McFadden R2 (0.527, vs 0.058 for humans; Section 7.1, Table 3) can reflect within-model consistency rather than a genuinely broader or more clinically aligned cue set. The Alternative Annotator Test validates GPT-4o's feature labels against human labels on only 10 transcripts (380 values) and establishes agreement, not independence from GPT-4o's perception; it also does not show that the 38-feature vocabulary is the right causal vocabulary for what humans actually use. The low human R2 could therefore reflect that the GPT-4o-generated features are semantically aligned with its own perception while missing the cues humans rely on, making the headline 'LLMs are richer, humans are narrow' an artifact of measurement rather than a substantive finding about perception.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper investigates how non-expert humans and LLMs perceive dementia from language, using 514 Cookie Theft picture descriptions from the Pitt corpus. The authors collect intuitive healthy/dementia judgments from 27 non-expert annotators and from three LLMs (GPT-4o, LLaMA 3, Gemini-1.5-Pro), define 38 expert-guided binary features, annotate those features with GPT-4o, and fit stepwise logistic regressions to explain human perception, LLM majority perception, and clinical diagnosis. The central claims are that human perception is inconsistent and relies on a narrow and sometimes misleading cue set, that LLMs draw on a richer feature set more aligned with clinical patterns, and that both groups show a tendency toward false negatives. The paper also analyzes misperceptions and self-reported human rationales.","tokens_in":29297,"tokens_out":6099,"duration_ms":60696,"significance":"If the central claims were fully supported, the paper would make a valuable contribution to dementia awareness and to the design of LLM-based monitoring tools. The research question is timely and understudied, the use of a clinically grounded corpus is appropriate, and the authors are transparent about several limitations. They also apply a statistical test for LLM-as-annotator quality and present inherently interpretable models. However, the main human-vs-LLM asymmetry currently rests on a confounded measurement setup and on an interpretation of a very low human model fit that may not be unique. The paper's headline conclusion is therefore not yet established, although it is plausibly fixable with additional analyses.","major_comments":[{"comment":"The main evidence for the claim that LLMs rely on a richer feature set is confounded. GPT-4o is both the annotator that produced all 38 feature values for all 514 transcripts (Section 5.2) and one of the three LLMs whose majority vote defines the LLM perception label (Section 4.2). The very high McFadden R² of 0.527 in Table 3 may therefore reflect GPT-4o's internal consistency rather than a genuinely broader or more clinically aligned cue set. I request that the authors re-estimate the LLM perception model using only the LLaMA-3 and Gemini-1.5-Pro majority, and ideally with features annotated by a model that is not part of the perception panel, and report whether the R² and the coefficient pattern survive.","section":"§4.2, §5.2, Table 3"},{"comment":"The low human-perception fit (McFadden R²=0.058) is interpreted as evidence of human inconsistency, but it is equally consistent with the alternative that the 38 binary features annotated by GPT-4o do not capture the cues humans actually use. The Alternative Annotator Test in Section 5.2 validates agreement between GPT-4o and human annotators on the feature values for 10 transcripts (380 values); it does not establish that this feature vocabulary is complete or causally aligned with human perception. Without a larger human-annotated feature validation, or an analysis using an expanded feature set derived from the human rationales in Section 7.3, the central asymmetry between narrow human cues and rich LLM cues may be an artifact of measurement rather than a substantive finding.","section":"§7.1, Table 3"},{"comment":"The set of significant features is obtained through stepwise logistic regression on the full dataset, and the p-values in Table 3 are not adjusted for the 38 candidate predictors. Under this procedure, the number of significant features per judgment (4 for humans, 13 for LLMs, 8 for clinical diagnosis) is not a clean measure of cue 'richness,' since stepwise selection is known to be unstable and to produce inflated significance. I ask for a robustness check, such as bootstrap stability of selected features or a regularized alternative like LASSO, before the richness comparison is used as a central result. In addition, the claim that LLMs align 'more closely' with clinical patterns is not quantified; a formal comparison, such as correlation or overlap of coefficient vectors including signs, would make this claim testable.","section":"§6, Table 3"}],"minor_comments":[{"comment":"The statement that 65% of self-reported cues align with the predefined feature set is based on a manual review of responses from 18 of 27 annotators, but the coding procedure and coder agreement are not described. Please provide the coding protocol and report inter-coder reliability.","section":"§7.3"},{"comment":"Please state explicitly in the modeling section that all logistic regressions use the majority-vote perception labels, and clarify whether the clinical diagnosis model is fit on the same feature matrix as the perception models.","section":"§6"},{"comment":"Five features are removed because they are positive in fewer than 5% of samples. Since some of these features (e.g., empathy, irritability) are clinically interesting, please report whether the main conclusions change if these features are merged with related categories or analyzed descriptively.","section":"Appendix C.1"}],"recommendation":"major_revision","confidential_remarks":"The manuscript addresses an important question and reports a carefully collected dataset, but the main quantitative comparison is currently vulnerable to the annotation-perception confound. I believe this is fixable: the authors should show that the LLM richness result survives when GPT-4o is excluded from the perception majority and when features are annotated by an independent model, and they should provide a more direct validation of the feature vocabulary for human perception. There is no indication of questionable research practices; the issue is methodological and visible in the manuscript itself."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Nick,\n\nHere's my take on the dementia perception paper. The genuinely new piece is that they treat dementia assessment as a perception problem: they collect judgments from 27 non-experts and three LLMs on the same 514 Cookie Theft transcripts, then try to explain those judgments with a shared set of 38 expert-guided binary features. That framing is useful and I haven't seen it before. The feature design is careful, with literature sources and example prompts. The use of the Alternative Annotator Test to justify GPT-4o as an annotator is a step up from the usual hand-waving. And the human findings—low inter-annotator agreement, significant coefficients for only a few features, a consistent false-negative bias—look credible and match common sense.\n\nBut the main comparative claim in the abstract—that LLMs rely on a richer, more clinically aligned feature set than humans—is weakened by a design confound. GPT-4o wrote all the feature labels, and GPT-4o is also one of the three models whose majority vote is the LLM perception label. The logistic regression then predicts that majority using features produced by one of the voters. A model that is internally consistent will naturally look 'richer' when its own features predict its own vote. The R2 of 0.527 for LLM perception is almost certainly inflated by this overlap. The Alternative Annotator Test on 10 transcripts validates that GPT-4o's feature labels match human labels, but it doesn't break the circularity when the outcome variable also comes from GPT-4o. The fix is straightforward—extract features with a different model, or use human feature annotations on a larger sample—but as it stands the headline asymmetry is not proven.\n\nThe other soft spot is the misperception analysis (Section 7.2). They fit stepwise logistic regressions on subsets of 121 and 88 cases with no cross-validation, then report McFadden R2=0.6. With that many features and small n, those coefficients are unstable. A cross-validated or regularized version would be more honest. The low human R2 (0.058) is less troubling—it's consistent with the low inter-annotator agreement and the authors are appropriately cautious—but they could be more explicit that it might also reflect measurement error in the features.\n\nOverall, I'd send this to review. The perception angle is a real contribution, the data collection effort is substantial, and the Limitations section shows the authors are thinking carefully. But referees need to push on the circularity and on the overinterpretation of the LLM findings. With those fixed, this could be a nice paper for a clinical NLP audience or an interpretability workshop. I would not cite the LLM result as it stands, though I might cite the human perception result after revision.","headline":"A genuinely new perception-angle paper whose central human-vs-LLM comparison is compromised by GPT-4o serving as both feature extractor and one of the perception voters.","tokens_in":29993,"tokens_out":3759,"would_cite":false,"duration_ms":35952,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper shows that non-expert humans and large language models both miss many clinically diagnosed dementia cases, while the cues they rely on differ: humans use a narrow, sometimes misleading set, and LLMs use a richer set closer to…","keywords":["dementia perception","large language models","explainable NLP","Cookie Theft picture description","logistic regression","LLM as annotator","false negatives"],"falsifier":"Have a fresh group of non-expert humans annotate the same 38 features on a random sample of transcripts, then fit the same logistic regression using those human feature annotations to predict human perception. If the model fit rises well above the reported McFadden's $R^2=0.058$, the conclusion that human perception is inherently inconsistent would be undermined, because the noisy part would be the original LLM feature annotations rather than human judgment.","tokens_in":28847,"feed_emoji":"🧠","tokens_out":9276,"duration_ms":80580,"temperature":0.7,"pith_summary":"This paper tries to establish that dementia is perceived differently by non-expert humans and by large language models when both read the same transcribed speech, and that neither group perceives it the way clinicians diagnose it. Using 514 Cookie Theft picture descriptions, the authors show that human judgments are inconsistent and hinge on a few simple cues, some of which point opposite to clinical diagnosis, while LLMs draw on a richer feature set that aligns more closely with clinical patterns. Still, both humans and LLMs miss a large share of clinically diagnosed dementia cases, mostly by labeling them healthy. This matters because early detection often begins with non-experts, and increasingly with LLM assistants, so knowing which cues each group actually uses is a step toward better public awareness and safer automated monitoring.","feed_headline":"Both humans and LLMs miss about 4 in 10 dementia cases","feed_subtitle":"Non-experts lean on a few misleading cues; LLMs use a broader set closer to clinical patterns.","key_machinery":"The load-bearing mechanism is a four-step explainable pipeline. First, 38 binary features are defined in consultation with a neurologist and a neuropsychologist, grouped into five categories: Objective Interpretation, Subjective Interpretation, Linguistic, Human Experience, and Interview Context. Second, a large language model (GPT-4o) annotates every transcript for all 38 features, after an Alternative Annotator Test showed a 90% chance its annotations were as good as or better than human annotations. Third, stepwise logistic regression models are fit separately to human perception, LLM perception, and clinical diagnosis using the same features. Fourth, the coefficients are compared to see which cues each judgment type actually relies on; the same feature representation across all three makes the comparison interpretable.","core_discovery":"The paper's central claim is that non-expert human perception of dementia in language is inconsistent and driven by a narrow, partly misleading set of cues, whereas LLM perception is driven by a broader set that overlaps more with clinical diagnosis; nevertheless both are biased toward false negatives. Concretely, human annotators agreed only weakly (Fleiss' $\\kappa=0.28$), a logistic regression on the 38 expert-guided features explained almost none of their judgments (McFadden's $R^2=0.058$), and the few cues that mattered included reading short sentences as a sign of health even though clinicians associate short sentences with dementia. LLM judgments were better explained by the same features (McFadden's $R^2=0.527$) and drew on subjective and emotional cues such as Theory of Mind, lightheartedness, and sadness. With 283 clinically diagnosed dementia cases, humans correctly identified 57% and LLMs 60%, and among their errors both groups predominantly missed dementia cases (65% and 70% false negatives, respectively). The authors also report that annotators' self-described reasoning did not match the features that actually predicted their judgments.","pith_inferences":["The single-transcript design likely understates how well humans detect decline in real life, where repeated exposure gives a baseline; the narrow-cue result may not transfer to longitudinal monitoring.","A practical screening system would need to track changes over time rather than classify individual transcripts, because the LLM false-negative pattern shows a transcript without overt linguistic problems is almost always labeled healthy.","If the 38-feature set omits cues humans actually use, the low human model fit could reflect incomplete measurement rather than human inconsistency; re-annotating a larger sample with human feature labels would settle this.","The paper does not test whether LLM explanations improve human accuracy, but its results make that a natural next experiment for public-awareness tools."],"forward_implications":["If non-experts read short sentences as a sign of health, awareness materials can be aimed directly at that specific misreading.","Any LLM-based early-warning tool should be designed around the demonstrated false-negative tendency, since 70% of LLM errors were missed dementia cases.","If people's self-reported cues do not match their modeled decision patterns, asking users to reflect on their own reasoning is not a reliable way to audit or improve their judgments.","If LLMs already use a feature set closer to clinical diagnosis, their explanations could serve as coaching signals to help non-experts attend to clinically relevant cues."],"supporting_citations":[{"why":"Supplies the Pitt corpus transcripts and clinical diagnoses used throughout the study.","marker":"Becker et al., 1994"},{"why":"Provides the Alternative Annotator Test that justifies replacing human annotators with GPT-4o for the 38 features.","marker":"Calderon et al., 2025"},{"why":"Demonstrates the LLM-as-annotator feature extraction approach the pipeline builds on.","marker":"Badian et al., 2023"},{"why":"Grounds several linguistic features, such as disfluencies and vocabulary richness, in established dementia markers.","marker":"Fraser et al., 2016a"},{"why":"Defines the pseudo-$R^2$ used to compare how well the feature model explains each judgment type.","marker":"McFadden, 1972"}],"fun_headline_variants":["Humans and LLMs both miss ~40% of dementia cases","Dementia detection: humans and LLMs share a 40% miss rate","Why humans and LLMs both overlook dementia in speech","Both humans and LLMs fail to catch 40% of dementia signs","Non-experts and LLMs miss dementia in language at similar rates"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the 38 binary features extracted by a single LLM (GPT-4o) faithfully capture the linguistic cues that actually drive human and LLM judgments; the validation of this feature set was carried out on only 10 transcripts.","fun_headline_variants_meta":{"raw":{"variants":["Humans and LLMs both miss ~40% of dementia cases","Dementia detection: humans and LLMs share a 40% miss rate","Why humans and LLMs both overlook dementia in speech","Both humans and LLMs fail to catch 40% of dementia signs","Non-experts and LLMs miss dementia in language at similar rates"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000735,"raw_usage":{"total_tokens":3318,"prompt_tokens":1013,"completion_tokens":2305,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":629,"completion_tokens_details":{"reasoning_tokens":2214}},"tokens_in":629,"tokens_out":2305,"duration_ms":15849,"temperature":1.0,"reasoning_tokens":2214,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T20:13:57.048025+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Have a fresh group of non-expert humans annotate the same 38 features on a random sample of transcripts, then fit the same logistic regression using those human feature annotations to predict human perception. If the model fit rises well above the reported McFadden's $R^2=0.058$, the conclusion that human perception is inherently inconsistent would be undermined, because the noisy part would be the original LLM feature annotations rather than human judgment.","supporting_citations":[],"review_version":1}