{"id":"160df7dd-e800-4aa3-9391-3d8a88e4e5d8","arxiv_id":"2608.05630","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Some open-weight LLMs show human-like sensitivity to distance and discourse prominence in anaphor resolution, but weaker sensitivity to semantic interference.","lead":"This paper tests whether five large language models respond to the same discourse, distance, and semantic factors that make anaphors (words like 'it' or 'she') harder or easier for humans to resolve. It finds that some models, especially Mistral-7B and GPT-2-XL, partly mimic human patterns on distance and prominence effects, but rarely on semantic interference.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Load-bearing gap: the selective-alignment trends are descriptive only; with 16–19 items and no inferential statistics, chance alone could produce the reported ordinal patterns.","rationale":"The reader's verdict is CONDITIONAL and centers on the automatic judge as the weakest assumption, while also noting the absence of inferential statistics in the rationale. I agree that the judge is a real fragility for the accuracy measures, but I see the more load-bearing concern as the reliability of the descriptive ordinal patterns themselves, particularly in surprisal, which carries most of the selective-alignment claim. The paper is transparent about this limitation and even labels it the largest one, which is credit to the authors, but transparency does not establish the claim. The concern is concrete and addressable: item-level sign tests and permutation tests can be run immediately on the released code and stimuli without collecting new materials or human data. Because the authors deliberately avoided simulated participants and offer cautious language, their conclusion is appropriately provisional; the existing CONDITIONAL verdict is the right one, and no verdict change is needed. I therefore recommend UNCHANGED, with the condition that the authors add the proposed item-level inferential tests or explicitly soften the central claim to 'descriptive trends consistent with human-like sensitivity' until such tests are run.","tokens_in":10004,"tokens_out":5298,"duration_ms":74460,"concrete_test":"Using the released code and data, recover per-text, per-condition normalized surprisals and per-text accuracy labels for all five models. For each model, experiment, and measure, compute (i) a two-sided sign test on the within-text A-versus-D contrast and (ii) a permutation test of the full predicted ordering A < {B,C} < D for surprisal (A > {B,C} > D for accuracy), permuting condition labels within each text. Apply Benjamini-Hochberg correction across the 5 models x 3 experiments x 2 measures. If the A/D sign test fails at p<0.05 for the models highlighted in Table 1 (e.g., all models for Experiment 1 surprisal; GPT-2-XL and Mistral-7B for alignment), the selective-alignment conclusion is not supported beyond chance.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim of selective cognitive alignment rests on the ordinal patterns in Figures 1–6 being genuine effects of the manipulated factors. The paper's own General Discussion identifies the lack of items as 'the largest limitation' and reports only descriptive means with standard errors, without any inferential statistics. With 16 texts in Experiment 1 and 19 in Experiments 2–3, across 5 models x 2 measures x 3 experiments, there are many opportunities for chance to yield a predicted A/D ordering or a partial ordinal pattern. Several error bars in the figures appear to overlap (notably in the accuracy panels), so the visually 'human-like' patterns may reflect sampling noise rather than factor sensitivity. This concern applies most directly to the surprisal results, which carry most of the headline generalization. The automatic judge is an additional uncontrolled measurement layer for the accuracy measures, but the more load-bearing gap is that the surprisal trends have no reported reliability estimate. This is not an internal inconsistency; it is an unestablished empirical claim, and it is directly testable with item-level statistics on the already-released data.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper evaluates whether five open-weight LLMs (GPT-2-XL, Llama-3.1-8B, Pythia-12B, Mistral-7B, Mistral-24B) are sensitive to six factors that cognitive science has shown to affect human anaphor resolution: antecedent topicality, sentential distance, spatial distance, temporal duration, semantic similarity, and semantic interference. Three experiments use published human materials, and each compares four text versions with model surprisal at the anaphor as a reading-time proxy and model accuracy on a comprehension question as a success measure, with accuracy scored by an LLM-as-a-judge. The paper reports qualitative ordinal patterns and concludes that LLMs show selective cognitive alignment: more consistent human-like sensitivity to discourse-prominence and distance-based factors than to semantic interference, suggesting that LLMs can serve as partial cognitive models of discourse processing while localizing some divergences.","tokens_in":10319,"tokens_out":6310,"duration_ms":74182,"significance":"If the reported selective-alignment pattern were statistically established, this would be a useful contribution to the cognitive-science evaluation of LLMs: it moves beyond coreference-accuracy benchmarks, uses classic human experimental materials, tests predictions derived from external cognitive science findings rather than fitting model parameters, and publicly releases code and data. The asymmetry between distance/discourse factors and semantic-interference factors is also a concrete, falsifiable claim about where current LLM architectures diverge from human discourse comprehension. However, the current evidentiary basis is almost entirely descriptive, so the significance of the contribution is conditional on the missing inferential support and on validation of the automatic judge.","major_comments":[{"comment":"The central claim of selective cognitive alignment rests on ordinal patterns of descriptive means computed over 16 texts in Experiment 1 and 19 texts in Experiments 2 and 3. The General Discussion itself states that the item count 'precluded running statistical analyses' and describes the trends as 'informal.' With five models x two measures x three experiments, there are many opportunities for chance to yield the predicted A/D ordering or a full predicted pattern, and several error bars visibly overlap, especially in the accuracy panels (e.g., Figures 2, 4, and 6). The surprisal results, which carry the headline generalization, are reported without any reliability estimate. This is not an internal inconsistency, but it means the empirical claim that 'some LLMs exhibit human-like sensitivity' is currently unestablished. The released data would permit item-level mixed-effects models or bootstrap confidence intervals for each model-by-factor contrast; these should be added, with appropriate correction for multiple comparisons.","section":"General Discussion, limitations; Figures 1-6"},{"comment":"All comprehension-accuracy results pass through an automatic judge (Gemini-2.5-flash-preview) that the authors acknowledge 'introduces known biases and can be different from human judgments.' No human-agreement sample, judge reliability statistic, or error analysis is reported. Because accuracy is one of only two dependent measures and is used in every experiment to support the selective-alignment narrative, the paper needs at least a validation subset scored by human annotators, or an error analysis showing that judge disagreements do not systematically favor the predicted conditions.","section":"Experiment 1, Comprehension Question Answering; General Discussion limitations"},{"comment":"The summary table does not state what count as 'human-like' for each cell, and it is not always consistent with the text. In Experiment 1, the text reports that GPT-2-XL, Llama-3.1-8B, Pythia-12B, and Mistral-7B showed the predicted surprisal pattern, whereas Table 1 codes 'All models'; in Experiment 2, the text identifies only Llama-3.1-8B and Mistral-7B as showing the predicted pattern, while Table 1's 'All models except Pythia-12B' implies additional models did so. Similar ambiguities appear in the accuracy columns, where 'None/weak effects' coexists with 'all of the other models correctly order versions A and D.' Since the main conclusion about selective alignment is a meta-level summary of these table cells, the coding criteria need to be explicit and applied consistently.","section":"Table 1 and Results sections"}],"minor_comments":[{"comment":"The normalization equation is unnumbered and its notation is loose; please label it as Equation (1), define whether the min and max are taken over the four version-level surprisals for each text, and justify min-max normalization over z-scoring, especially given that it can change the relative weighting of texts with different within-text variance.","section":"Experiment 1, Procedure and Dependent Measures"},{"comment":"The model name 'Pythia' is used instead of 'Pythia-12B', and the paper alternates between 'LLaMa' and 'Llama'; please harmonize names across the text, tables, and figures.","section":"General Discussion, Table 1"},{"comment":"The description says the base texts are 'the version D (far spatial distance, long temporal duration) texts from Experiment 2'; please clarify whether the surrounding discourse and the anaphor tokens are identical to Experiment 2 or were modified, and how the semantic manipulations were crossed with the pre-existing spatial/temporal properties.","section":"Experiment 3, Design and Materials"},{"comment":"The date-stamped judge name 'Gemini-2.5-flash-preview-09-2025' will age quickly; please report the exact model version used and consider adding a note about reproducibility as the judge model changes.","section":"Experiment 1, Comprehension Question Answering"},{"comment":"The figures report only means and standard errors; adding item-level points or box plots would make the degree of overlap and the ordinal claims much easier to assess visually.","section":"Figures 1-6"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is transparent about its main limitation, and the missing inferential analysis is directly addressable with the released data, so I do not see this as a reject. However, the current wording overstates descriptive trends, and the automatic judge needs validation before the accuracy results can carry weight. I would advise the editor to require item-level statistics and a human-annotation check in the revision, and to ask the authors to reconcile Table 1 with the text."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick note on arXiv:2608.05630. Read it this morning.\n\nWhat you should know: this is a systematic, transparent look at whether five open-weight LLMs show the same six-factor sensitivity as humans in anaphor resolution (topicality, sentential/spatial/temporal distance, semantic similarity, interference). Used classic stimuli from O'Brien (1987) and Varma & Janssen (2019), measured both surprisal at the anaphor and comprehension-answer accuracy, and released the code. That is genuinely new; nobody had run all six factors through a modern LLM panel with both measures.\n\nIt also does something right: no parameter fitting, no post-hoc cherry-picking of predictions. Predictions come from the cognitive science results, and the authors report all conditions, including the failures (e.g., semantic interference effects are mostly absent). The General Discussion is candid, admitting both the item-count limitation and concerns about the LLM-as-a-judge scoring.\n\nThe soft spots are real, and the main one is load-bearing. With 16 or 19 texts per experiment, and only descriptive means with overlapping error bars, the ordering patterns the headline rests on could easily be noise. The authors flag this as the largest limitation, but the abstract still says the results 'show selective cognitive alignment' as if established. The stress-test concern about missing inferential statistics is correct, and it applies most to the surprisal results, which carry the generalization.\n\nSecond issue: the abstract promises 'compare model accuracy to human accuracy,' but I did not find any human accuracy data in the paper. The comparisons are to directional predictions from prior work, not to actual human participants or even published human accuracy values. That is a billing problem, not fatal, but it should be fixed.\n\nThird, the comprehension accuracy results sit behind a Gemini judge, another uncontrolled layer. The authors acknowledge it, but it makes the accuracy half of the study hard to interpret.\n\nWho gets value: cognitive scientists evaluating LLMs as process models, and NLP people designing coreference benchmarks. The paper is a useful descriptive exploration, and the authors are honest about limits. If a referee asks for item-level statistics (mixed-effects models, or at least a permutation test on the ordinal pattern) plus a direct human comparison or a revised abstract, it could become a solid contribution.\n\nSend it to peer review. It deserves referee time; the claims need tempering, not abandoning. I would take the review and would cite it as a cautionary example of promising trends that need statistical backing.","headline":"A transparent, novel evaluation of six classic cognitive factors in anaphor resolution across five LLMs, but the central claim of selective alignment rests on descriptive bar-chart patterns with no inferential statistics; worth peer review as a promising descriptive study that needs statistical backing.","tokens_in":10697,"tokens_out":2830,"would_cite":true,"duration_ms":34952,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Large language models show human-like sensitivity to discourse prominence and distance in anaphor resolution, but largely miss semantic interference effects.","keywords":["anaphor resolution","large language models","surprisal","cognitive alignment","situation models","discourse prominence","semantic interference","comprehension accuracy"],"falsifier":"Collect human ratings on the same model-generated comprehension answers for all text versions; if the condition ordering in accuracy does not reproduce for any model, the accuracy-side evidence for cognitive alignment collapses.","tokens_in":9825,"feed_emoji":"🔗","tokens_out":6138,"duration_ms":69919,"temperature":0.7,"pith_summary":"The paper asks whether large language models resolve anaphors the way people do, driven by the same discourse, situational, and semantic factors. It tests five open-weight models on standard cognitive science materials, using token surprisal at the anaphor as a stand-in for reading time and comprehension-question accuracy as a stand-in for resolution success. The central finding is selective alignment: several models mirror human sensitivity to discourse prominence and to sentential, spatial, and temporal distance, while sensitivity to semantic similarity and interference is weaker or absent. The authors take this as evidence that LLMs can serve as partial cognitive models of human discourse processing, with the divergence localized to semantic interference.","feed_headline":"LLMs mirror human anaphor resolution, except for semantics","feed_subtitle":"Surprisal tracks prominence and distance cues, but semantic interference barely registers in current models.","key_machinery":"The load-bearing machinery is the surprisal linking hypothesis—the assumption that a model's $-\\log_2 p$ for the anaphor token indexes processing difficulty the way reading time does for humans—combined with comprehension questions scored by an automatic judge. These two measures are applied to factor-manipulated texts drawn from prior cognitive science experiments, so each text's four versions isolate one accessibility factor at a time. What carries the argument is the pattern across conditions: if models show the same ordinal ordering as human reading-time and accuracy findings, they exhibit cognitive alignment.","core_discovery":"Using materials from earlier studies of anaphor resolution, the paper manipulates six factors orthogonally: antecedent topicality, sentential distance, spatial distance, temporal duration, semantic overlap between anaphor and antecedent, and semantic interference from a non-antecedent distractor. For each of five open-weight LLMs, it measures mean surprisal on the anaphor and accuracy on a comprehension question about the antecedent. The predicted human pattern is that resolution is fastest and most accurate when the antecedent is topicalized, close in surface or situational distance, semantically similar to the anaphor, and free of competing typical distractors. The results show that several models track the topicality and sentential-distance pattern in surprisal, that two models track spatial and temporal distance, and that only isolated models show the predicted semantic patterns. The authors conclude that some LLMs, especially Mistral-7B and GPT-2-XL, approximate human anaphor resolution on prominence and distance factors but diverge on semantic interference.","pith_inferences":["An extension the paper does not run: replacing the automatic judge with human raters on the same model outputs would show whether the reported accuracy orderings are artifacts of judge bias.","If the absence of semantic interference is genuine, it points to a structural difference—current LLMs lack the competitive memory retrieval that cognitive theories posit—which would recommend memory-augmented architectures, a prediction the paper does not make.","The small number of texts makes the conclusions descriptive; a larger-item replication with inferential statistics is the natural next step, which the authors explicitly defer.","The asymmetry between prominence and distance effects on the one hand and semantic effects on the other suggests a possible ordering of difficulty for cognitive alignment in models: surface accessibility first, situation-model distance second, semantic competition last."],"forward_implications":["Surprisal at the anaphor can serve as a process-level measure for discourse phenomena, not just word- and sentence-level effects.","Models like Mistral-7B and GPT-2-XL become candidate tools for generating behavioral predictions about human anaphor resolution.","Benchmarks for coreference and anaphor resolution should manipulate the cognitive factors that drive human resolution, since raw accuracy alone cannot distinguish human-like processing from shallow heuristics.","Semantic interference is a reliable diagnostic dimension on which current LLMs diverge from human comprehension.","The documented recency and ceiling effects in larger models can mask distance-based effects, so model size alone does not determine cognitive alignment."],"supporting_citations":[{"why":"Supplies the Experiment 1 materials and the topicality and sentential-distance manipulation.","marker":"O'Brien (1987)"},{"why":"Supplies the Experiment 2 and Experiment 3 materials and the spatial, temporal, and semantic manipulations.","marker":"Varma & Janssen (2019)"},{"why":"Introduces the surprisal-based linking hypothesis that connects model probabilities to reading times.","marker":"Hale (2001)"},{"why":"Formalizes expectation-based comprehension, grounding the surprisal–reading time link.","marker":"Levy (2008)"},{"why":"Provides evidence that neural language model surprisal predicts human real-time comprehension.","marker":"Wilcox et al. (2020)"},{"why":"Shows that temporal duration affects anaphor accessibility, motivating the factor tested in Experiment 2.","marker":"Anderson et al. (1983)"},{"why":"Establishes that semantic typicality between anaphor and antecedent affects resolution, motivating Experiment 3.","marker":"Garrod & Sanford (1977)"},{"why":"Provides the semantic interference manipulation used in Experiment 3.","marker":"Corbett (1984)"}],"fun_headline_variants":["LLMs mirror human anaphor cues, miss semantic traps","LLMs align with humans on anaphor prominence, not semantics","Anaphor resolution: LLMs match humans on prominence, not semantics","LLMs show humanlike anaphor cues, skip semantic interference"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The reported comprehension-accuracy results depend on an automatic judge scoring model responses correctly enough to preserve the ordering across conditions, and the paper itself flags that this judge can differ from human judgments.","fun_headline_variants_meta":{"raw":{"variants":["LLMs mirror human anaphor cues, miss semantic traps","LLMs align with humans on anaphor prominence, not semantics","Anaphor resolution: LLMs match humans on prominence, not semantics","LLMs show humanlike anaphor cues, skip semantic interference"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000681,"raw_usage":{"total_tokens":3089,"prompt_tokens":937,"completion_tokens":2152,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":553,"completion_tokens_details":{"reasoning_tokens":2077}},"tokens_in":553,"tokens_out":2152,"duration_ms":19670,"temperature":1.0,"reasoning_tokens":2077,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T00:53:30.631068+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Collect human ratings on the same model-generated comprehension answers for all text versions; if the condition ordering in accuracy does not reproduce for any model, the accuracy-side evidence for cognitive alignment collapses.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the Experiment 2 and Experiment 3 materials and the spatial, temporal, and semantic manipulations."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduces the surprisal-based linking hypothesis that connects model probabilities to reading times."},{"cited_title":"C., & Sanford, A","cited_arxiv_id":null,"evidence_quote":"Shows that temporal duration affects anaphor accessibility, motivating the factor tested in Experiment 2."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Establishes that semantic typicality between anaphor and antecedent affects resolution, motivating Experiment 3."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the semantic interference manipulation used in Experiment 3."}],"review_version":1}