{"id":"bbf7cbb5-12d2-4b66-8f4a-8951c7957f5f","arxiv_id":"2607.20208","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"LLM-based surprisal is not interchangeable across models: different architectures compute word probabilities via visibly different internal computations, undermining representation-agnostic claims for Surprisal Theory.","lead":"This paper argues that surprisal—the negative log probability of a word—is not a single theory once it is computed by different large language models, because different models use different internal representations to arrive at their probabilities. It demonstrates, on three popular models, that the internal trajectories of probability estimates and the word-likeness of top predictions differ dramatically across architectures.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Experiments show internal layerwise divergence, not that final-layer surprisal values differ in behaviorally consequential ways; the central inference from algorithmic differences to non-interchangeable surprisal theories is unsupported.","rationale":"The reader identified the same load-bearing weakness: the paper relies on a functional-equivalence premise without showing that differences in internal computation translate into behaviorally relevant differences in the final surprisal values. My stress-test sharpens this by noting that the paper's own reported 0.92 final-layer correlation between GPT-2 and Pythia is direct evidence against a large behavioral difference, and by observing that no reading-time or neural-effort measure is used anywhere in the experiments. The conceptual argument about representational commitments and Marr's levels is coherent and worth making, but the empirical support is descriptive and lacks inferential statistics; the layerwise curves are interesting but do not bear on the computational-level interchangeability question unless final outputs diverge in behaviorally predictive ways. I am not claiming the paper is wrong—it may be right—but the central inference is underdetermined by the data as presented. The reader's CONDITIONAL verdict already captures this, so no change is needed. My concrete test would settle the question by directly comparing final-layer surprisal values from the three models on a behavioral benchmark.","tokens_in":21661,"tokens_out":3024,"duration_ms":34920,"concrete_test":"On a reading-time corpus (e.g., the realigned Natural Stories or Dundee), fit mixed-effects models with GPT-2, Pythia-160M, and RoBERTa final-layer surprisal as predictors of word reading times. Test whether model identity significantly moderates the surprisal coefficient (interaction term) or whether replacing one model's surprisal with another changes the qualitative conclusion (e.g., log-linear vs. linear fit). Additionally, compute the item-level distribution of absolute final-layer surprisal differences between GPT-2 and Pythia; if the largest divergences are concentrated among rare/low-cloze words, check whether those items drive any differential behavioral prediction. If final-layer surprisal values from different models yield statistically indistinguishable fits to behavioral data, the paper's central empirical claim fails, regardless of layerwise trajectory differences.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim is that LLM-derived surprisal values are not interchangeable because model architecture and algorithm affect the computation. But the operationalized surprisal metric in the psycholinguistic literature is the final-layer negative log probability, and the experiments directly report that final-layer GPT-2 and Pythia probabilities correlate at 0.92 and that final-layer correlations to cloze are similar across models. The observed divergence is in layerwise trajectories (Figures 1-3) and in layerwise lexicality (Figure 4), i.e., in internal algorithmic detail. The paper itself acknowledges in §4.1 that multiple realizability is intrinsic to computational-level theories and that such theories abstract over algorithmic specification. Therefore, to show that these models instantiate different computational-level linking hypotheses, the authors must show that final surprisal values differ in ways that change predictions about behavior (e.g., reading times or N400). They do not: no behavioral effort measure is used in Experiments 1 or 2. The conditional statement in §4.1 ('If different LLMs change the linking hypothesis between probabilities and effort...') is exactly the untested premise. The Natural Stories misalignment discussion is a separate methodological point and does not fill this gap. Thus the most load-bearing inference—from different internal computations to non-interchangeable surprisal-based theories—is unsupported by the presented evidence.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper argues that the common psycholinguistic practice of treating surprisal values extracted from different large language models as interchangeable is theoretically unjustified. It reports two analyses using GPT-2, Pythia-160M, and RoBERTa: first, layerwise logit-lens probabilities are correlated with human cloze probabilities (Experiment 1); second, the 'lexicality' of each model's top-k predictions is traced across layers (Experiment 2). The authors find that layerwise trajectories differ markedly across models and conclude that different architectures and training objectives instantiate different algorithmic and representational commitments, undermining representation-agnosticism in Surprisal Theory. The paper then develops a conceptual critique of using LLM-derived surprisal as a computational-level explanation and offers recommendations for theory-driven model selection.","tokens_in":21991,"tokens_out":4282,"duration_ms":42187,"significance":"The paper addresses a timely and important methodological question: whether results obtained with one LLM's surprisal values can be taken to test Surprisal Theory independently of the model that generated them. Its strengths include public data and scripts (OSF), an emphasis on falsifiable practice rather than fitted parameters, and a clear engagement with Marr's levels. If the central claim were established, it would make a meaningful contribution to computational psycholinguistics. However, the significance is conditional: the empirical results demonstrate differences in internal algorithmic trajectories, but the operationalized surprisal in the literature is the final-layer probability, and the paper itself reports that final-layer GPT-2 and Pythia probabilities correlate at 0.92. The gap between algorithmic divergence and non-interchangeable linking hypotheses is the load-bearing issue that must be addressed.","major_comments":[{"comment":"The central inference from divergent layerwise computations to non-interchangeable surprisal-based theories is not supported by the presented evidence. Study 1.1 reports that final-layer GPT-2 and Pythia probabilities correlate at 0.92, and final-layer correlations to cloze are similar across models. Surprisal as commonly operationalized is the final-layer negative log probability, and Marr's computational level explicitly abstracts over algorithmic implementation—a point the paper itself acknowledges in §4.1 ('multiple-realizability is intrinsic to computational-level theories'). The statement at the end of §4.1, 'If different LLMs change the linking hypothesis between probabilities and effort...', is exactly the untested premise. To support the paper's headline claim, the authors need to show that final surprisal values differ across models in ways that change behavioral predictions (e","section":"§2.3 and §4.1"},{"comment":"The claims of 'starkly different' trajectories (Figure 3) and 'massive fluctuations' in lexicality (Figure 4) rest on visual inspection of only three models, with no error bars, confidence intervals, or inferential statistics on the layerwise curves. This is particularly consequential because the qualitative reading of the curves is a key part of the argument that the models are not interchangeable. Additionally, the lexicality metric depends on the choice of k=10 in §3.2, but no sensitivity analysis is reported; different tokenizers across models could confound the comparison. I recommend quantifying trajectory differences (e.g., bootstrap CIs, pairwise curve-distance measures) and testing whether the lexicality patterns are robust to k and to tokenization.","section":"§2.4–2.5 and §3.2"},{"comment":"The discussion of the Natural Stories misalignment error is presented as evidence that the field's reliance on surprisal is theoretically fragile. This is an interesting argument, but the inference is not logically entailed: a correlation can survive a one-position shift for many reasons (e.g., autocorrelation in surprisal, smoothness of the cost function). The claim that the error 'should have led researchers to reconsider why exactly their surprisal analyses had been so successful' is speculative. This is not the load-bearing part of the paper, but it should be framed as an open question or a testable prediction rather than a necessary consequence.","section":"§4.4"}],"minor_comments":[{"comment":"Typo: 'effors' should be 'efforts'. Also, the phrase 'these assumptions are is in fact not wholly valid' has a grammar error.","section":"§1"},{"comment":"The title 'surprisal is Not a Theory' overstates the paper's scope; the argument is specifically about LLM-derived surprisal and its use as a computational-level proxy. A more precise title would help manage reader expectations.","section":"Title"},{"comment":"The figure would be more informative with confidence bands or per-context variability; as published, the eye is drawn to mean trajectories that may mask high variance.","section":"Figure 3"},{"comment":"The tokenization differences across models make the qualitative examples in Table 2 hard to interpret; a brief note on subword segmentation would improve clarity.","section":"§3.2 / Table 2"}],"recommendation":"major_revision","confidential_remarks":"The skeptical reading is largely correct: the empirical work shows algorithmic and representational divergence, but the paper's central claim about non-interchangeable theories requires demonstrating consequences for the operationalized surprisal measure (final-layer probabilities) or for behavioral predictions. I regard this as fixable within the paper's scope, e.g., by adding a final-layer probability divergence analysis or a re-analysis of existing reading-time data with multiple models. The paper is likely best positioned as a position paper with illustrative empirical support rather than as a decisive falsification of interchangeability. I would not reject, but the current version does not fully support the abstract's strong conclusion."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this paper makes a point that needed making, and illustrates it with genuinely new layerwise analyses, but the empirical support doesn't quite carry the central inference.\n\nWhat's new is the logit-lens comparison of GPT-2, Pythia-160M, and RoBERTa on cloze norms, plus a lexicality metric showing the three models traverse very different internal trajectories. That's a concrete way to show model internals are not interchangeable, and it goes beyond previous critiques that stayed at the philosophical level. The paper also situates itself well relative to the multiple-realizability literature and aims a fair critique at predictive-power-driven model selection.\n\nThe soft spot is the jump from internal divergence to non-interchangeability of surprisal as a test of Surprisal Theory. The experiments show layerwise trajectories and lexicality curves differ, but the final-layer probabilities correlate at 0.92 between GPT-2 and Pythia, and layerwise correlations to cloze are highest at the final layer. If Surprisal Theory is a computational-level account, it should abstract over internal algorithms; the paper's own §4.1 concedes that multiple realizability is intrinsic to such accounts. To show different models instantiate different linking hypotheses, you'd need to show their final surprisal values make different predictions for reading times or N400. The paper doesn't. The Natural Stories misalignment discussion is a separate methodological issue and does not fill that gap.\n\nThe analyses are also purely descriptive — no error bars, no inferential tests on the trajectory differences, three models only, and the lexicality metric uses k=10 with no sensitivity check. Not fatal, but the evidence is illustrative rather than probative.\n\nVerdict: worth a serious referee. The critique matters and the layerwise analyses are worth building on, but the authors should either soften the central inference or show that model-induced surprisal differences actually change behavioral predictions. I'd bring it to our reading group; it will generate useful argument.","headline":"Useful critique of interchangeable surprisal, but the layerwise evidence doesn't close the gap to behavioral non-interchangeability.","tokens_in":22457,"tokens_out":3706,"would_cite":true,"duration_ms":35803,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper argues that surprisal scores drawn from different large language models encode divergent internal computations, so treating them as interchangeable evidence for Surprisal Theory is unsound.","keywords":["surprisal","large language models","psycholinguistics","multiple realizability","computational-level theory","representation-agnosticism","reading times","model interchangeability"],"falsifier":"A decisive test would be to fit the same reading-time regression twice, once with final-layer surprisals from two models whose internal trajectories diverge (e.g., GPT-2 and Pythia) and once with the same final values but shuffled layerwise trajectories; if the coefficients and model fit are statistically indistinguishable, the paper's claim that different internal computations constitute different linking hypotheses would be undermined. Alternatively, showing that the divergence in layerwise trajectories never translates into differences in predicted reading times for any naturalistic corpus","tokens_in":21595,"feed_emoji":"🧠","tokens_out":5418,"duration_ms":51869,"temperature":0.7,"pith_summary":"The authors aim to show that LLM-derived surprisal is not a single, theory-neutral metric despite its apparent uniformity at the output layer. They argue that each language model instantiates a distinct linking hypothesis between probabilities and processing effort, because models differ in training objective, architecture, and the representations built along the way. Two experiments—tracing next-word probability trajectories across layers and measuring the word-likeness ('lexicality') of models' top predictions—find strikingly different internal behaviors across GPT-2, Pythia, and RoBERTa, even where final correlations to human cloze or reading behavior look similar. If the paper is right, pooling surprisal from many models or choosing models by predictive power alone conflates different algorithmic commitments and undermines claims that Surprisal Theory is a computational-level explanation.","feed_headline":"Surprisal from different LLMs is not interchangeable","feed_subtitle":"Layer-by-layer probes show the same metric arises from different computations, so pooling these scores is unsound.","key_machinery":"The key mechanism is the 'logit lens' and the companion 'lexical lens': probes that read off a model's next-word probability distribution and its top token guesses at each hidden layer, making internal computation visible. The paper uses these to compare GPT-2, Pythia, and RoBERTa, which are similar in parameter count but differ in training objective (decoder vs. masked encoder), tied vs. untied embeddings, and context directionality. The interpretive frame is multiple realizability combined with functional equivalence: if a computational-level theory is allowed to abstract over algorithms, then internal differences matter only when they change downstream processing. The paper argues they do","core_discovery":"The central claim is that the 'representation-agnosticism' often invoked to justify using LLM surprisal as a computational-level measure is untenable. Different LLMs compute next-word probabilities through different latent structures and algorithms; the probabilities themselves are therefore not implementations of one same linking hypothesis. The authors introduce the criterion of functional equivalence: two models' surprisal values are interchangeable only if all subsequent processing is identical. Using logit-lens and lexical-lens probes, they show that although models like GPT-2 and Pythia correlate at 0.92 in final-layer probabilities, the layerwise evolution of those probabilities and t","pith_inferences":["If the authors' premise is right, meta-analyses pooling surprisal across model families may be averaging over qualitatively different computations, which could explain some conflicting results in the surprisal-reading-time literature.","The same critique likely applies to other LLM-extracted psycholinguistic indices (e.g., attention weights, entropy), not just surprisal.","A concrete extension: researchers could vary architecture while matching final surprisal values, to test whether layerwise trajectory differences actually change reading-time predictions—if they don't, the practical significance of this argument would shrink.","The argument could be transformed into a positive methodological program: use layerwise probes to determine which representational commitments are necessary to replicate human processing, rather than treating any LLM as neutral."],"forward_implications":["Researchers should stop treating surprisal estimates from different LLMs as interchangeable in reading-time and neurobehavioral analyses.","Model selection in psycholinguistics must be justified by explicit cognitive commitments, not by predictive power alone.","Surprisal Theory cannot claim to be a computational-level account while leaving representations and algorithms unexamined.","Published comparisons across dozens of models may be conflating distinct linking hypotheses rather than measuring the same construct.","The field should move toward explainable, cognitively motivated models, and reviewers should not routinely demand adding LLM surprisal as a covariate."],"fun_headline_variants":["LLM surprisal scores hide divergent computations","Same surprisal, different algorithms: a warning","LLM probabilities are not a single computational measure","Don't pool LLM surprisal scores"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The argument rests on the premise that two models' surprisal values are interchangeable only when every step of their internal computation is functionally equivalent; if a computational-level theory is allowed to abstract away internal algorithms, then the observed layerwise differences are irrelevant to the interchangeability of the metric.","fun_headline_variants_meta":{"raw":{"variants":["LLM surprisal scores hide divergent computations","Same surprisal, different algorithms: a warning","LLM probabilities are not a single computational measure","Don't pool LLM surprisal scores"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000895,"raw_usage":{"total_tokens":3648,"prompt_tokens":655,"completion_tokens":2993,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":399,"completion_tokens_details":{"reasoning_tokens":2933}},"tokens_in":399,"tokens_out":2993,"duration_ms":21850,"temperature":1.0,"reasoning_tokens":2933,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T10:27:03.487327+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A decisive test would be to fit the same reading-time regression twice, once with final-layer surprisals from two models whose internal trajectories diverge (e.g., GPT-2 and Pythia) and once with the same final values but shuffled layerwise trajectories; if the coefficients and model fit are statistically indistinguishable, the paper's claim that different internal computations constitute different linking hypotheses would be undermined. Alternatively, showing that the divergence in layerwise trajectories never translates into differences in predicted reading times for any naturalistic corpus","supporting_citations":[],"review_version":1}