REVIEW 4 major objections 5 minor 21 references
Focus particles and scalar inferences across humans and language models
T0 review · 4 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read A large language model reproduces the average human pattern of scalar judgments for 'even' and 'only', yet its complete lack of response variability reveals that matching aggregates is not evidence for shared cognitive representations.
desk verdict A compact extended abstract with a clean aggregate result and a real confound in the LLM variability claim; worth a round of revision, not a desk reject. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The study's machinery is a paired contrast between two focus particles—'even,' which highlights an unexpected or low-probability alternative, and 'only,' which restricts the set of valid alternatives—used to elicit scalar ability ratings across four response-scale configurations (horizontal or vertical layout, standard or reversed label mapping). The argument separates two rival accounts: spatial coding, which predicts ratings should shift with the physical arrangement of the scale, and the polarity correspondence principle, which predicts stable particle differences because conceptual polarity aligns with response polarity. The decisive diagnostic is response variability: human ratings are graded across individuals, whereas the LLM's repeated samples are nearly constant, so the spread of responses—not the mean—carries the claim that the model's mechanism differs.
What would settle it
Check the model's token-level probability distribution over the five response integers before any forced decoding: if 'only' items place substantial probability mass on ratings other than 5 even though every sampled integer is 5, then the categorical rating is an artifact of the response format, and the paper's mechanism-difference conclusion loses its main evidence.
Extended reading notes
Core claim
The central claim is that sensitivity to the scalar implications of 'even' and 'only' is stable across response-scale configurations in both humans and an LLM, but that the two systems arrive at this stability through different mechanisms. Human participants rated people described with 'even' at 2.07 on average and with 'only' at 4.31, while Llama 3.3 70B rated them at 1.41 and 5.00, respectively; the differential was larger for the model (3.58 vs. 2.20) and every 'only' trial received the scale maximum. Reversing or rotating the response scale did not reverse the pattern, which the paper reads as evidence that scalar judgments are driven by the evaluative content of the particles rather than by spatial magnitude coding. The absence of response variability—confirmed by a temperature sweep from 0.0 to 2.0—leads the authors to conclude that repeated samples from a single model reflect noise around one fixed judgment, not the genuine individual differences seen across people, and therefore that aggregate human-model agreement is not sufficient evidence for shared representations.
Load-bearing premise
The conclusion rests on the assumption that the instruction to 'respond with a single integer from 1 to 5' captures the same judgment humans made; if the discrete-response prompt itself collapsed the model's uncertainty onto scale endpoints, then the observed mechanism difference would be an artifact of task framing rather than a genuine representational difference.
Editorial extensions
If this is right
- If the paper's reading is right, scalar judgments about 'even' and 'only' are governed by semantic/evaluative structure, not by the spatial layout of a response scale.
- Aggregate benchmarks that compare only average human and LLM ratings can overstate cognitive similarity; response distributions must be compared as well.
- Sampling one language model many times at high temperature does not reproduce human individual differences, so such sampling is not a substitute for collecting human data.
- Temperature adjustment alone will not make current LLMs produce human-like graded judgments when the model's probability mass is already concentrated on one response.
- Claims that LLMs mirror human semantic representations need to be checked against human-like variability, not just matched means.
Reading between the lines
- A direct extension the paper does not run: ask the model to output a full probability distribution over the 1-5 scale, or to answer with a continuous slider, and compare the spread to human ratings. If spread appears under those formats, the zero-variability result would be a response-format artifact rather than a representational difference.
- The stability of LLM judgments across spatial layouts is plausibly trivial given transformers lack embodied spatial history; a sharper test would manipulate the set of alternatives in the linguistic context (e.g., replacing 'only' with 'even' in identical frames) and see whether the model's categorical response shifts.
- The paper's logic implies that human gradedness is itself informative: individual differences in how strongly 'only' raises ability ratings are real variation, so future model evaluations should target the distribution of human responses, not a normative single judgment.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This extended abstract reports a cross-system comparison of scalar ability judgments for sentences containing the focus particles 'even' and 'only'. Human participants (N=108) and Llama-3.3-70B rated 60 sentences on 5-point Likert scales under four response-scale configurations that varied spatial format (horizontal/vertical) and label mapping (standard/reversed). Both humans and the LLM rated 'only' sentences higher than 'even' sentences, with the LLM producing more extreme and less variable responses. The authors interpret the stable particle effect as evidence for semantic/evaluative rather than spatial-format-driven judgments, and the LLM's lack of response variability as evidence that its underlying mechanism differs from humans'.
Significance. If the central claim is established, the paper contributes a useful cautionary result: aggregate matching of average human and LLM ratings does not entail shared cognitive representations. The repeated-sampling design and the temperature-sweep robustness check are sensible probes of model variability, and the explicit reporting of a null result for spatial-format reversal is a strength in transparency. However, the significance is moderate because the key inference about mechanism difference currently rests on a single, potentially confounded observation, and because several statistical and design details are insufficiently specified.
major comments (4)
- [Method / Results] The inference that the LLM's lack of variability reflects a different underlying mechanism is not yet established because the only response format used was a prompt asking the model to 'respond with a single integer from 1 to 5.' A discrete-output instruction can collapse a graded internal distribution onto a canonical label, even when the model's token probabilities are diffuse. The temperature sweep (in the Discussion) does not address this concern because it varies sampling randomness while keeping the same single-integer output constraint; it cannot distinguish a genuinely concentrated distribution from a response-format artifact. The authors should either report the token-level probabilities of the rating options from the model's softmax (e.g., for '1' through '5' in each condition) or add a control condition that allows graded responses (e.g., a continuous slider or a distribution output). Without such evidence, the mechanism-difference claim is unsupported.
- [Results] The paper claims that spatial configuration did not affect scalar judgments, but it never reports a statistical test of the Format or Mapping factors or their interactions with Particle. The only reported effects are the main effect of Particle and the Particle×Source interaction. The statement that changing orientation or label mapping 'did not eliminate or reverse this pattern' is not a substitute for a formal test; an absence of reversal is compatible with a small but real spatial effect. The authors should report the full fixed-effects table for the mixed model, including main effects of Format and Mapping and all two- and three-way interactions, and ideally model comparisons (e.g., with vs. without spatial terms) to support the claim that scalar judgments are driven primarily by semantic structure.
- [Abstract / Method] The abstract states that the dataset contains 'approximately 100 items,' but the Method section says 60 sentences (30 with 'even' and 30 with 'only'). This discrepancy is not cosmetic: the item sample size directly affects statistical power and the generality of the item-level random effects. The authors must reconcile the count and report the exact number of items used in each condition.
- [Discussion] The temperature sweep is described only in the Discussion ('a follow-up temperature sweep on a subset of items') and no results are reported for it. Since this negative result is load-bearing for the claim that sampling noise cannot produce human-like variability, the authors should move the description to the Method and provide the number of items, the temperatures tested, the variability metric (e.g., standard deviation or entropy of sampled responses), and the actual stability values across temperatures. Without these details, the claim that 'increasing temperature alone was insufficient' is not verifiable.
minor comments (5)
- [Method] The design lacks a neutral control condition without 'even' or 'only' (e.g., 'Mike can bake a cake' with no particle). Such a baseline would help interpret the absolute ratings and confirm that the effect is specifically due to the particles rather than to some other aspect of the ability statements.
- [References] The reference for Llama 3.3 70B is given as Touvron et al. (2023), which describes the original Llama architecture, not Llama 3.3. The authors should cite the correct model card or a technical report for Llama 3.3.
- [Introduction] The sentence about transformer-based language models processing 'textual input tokens simultaneously through bidirectional attention' is inaccurate for Llama, which is an autoregressive, left-to-right model. The discussion of spatial orientation may still hold, but the architectural description should be corrected.
- [Method] The assignment of the 108 human participants to the four response-scale conditions is not specified; it should state whether participants were between-subjects assigned to one condition each and how items were randomized across trials.
- [Results / Figure 1] The figure caption states that error bars are 95% CIs and are absent for the LLM in 'only' conditions, but the figure itself is not included in the manuscript. Please ensure the figure is present and that the caption describes the aggregation units (participant-level vs. item-level CIs).
Circularity Check
No significant circularity: the study is an empirical comparison of human and LLM ratings, with no fitted parameters or self-citation chain that reduces its conclusions to its inputs.
full rationale
The paper reports direct empirical measurements: humans and Llama 3.3 70B produced scalar ratings for sentences with 'even' and 'only' across four response-scale configurations. The main comparisons (Particle effect, Particle × Source interaction, extremeness and lack of variability in LLM responses) are descriptive and inferential results from those measurements, not predictions derived from fitted parameters. The only post hoc element is the follow-up temperature sweep, but it is explicitly presented as a robustness check, not as a fitted or predicted outcome; it does not define any quantity in terms of another quantity from the same data. The claim that the LLM's mechanism differs from humans is an interpretation of observed zero response variability, and while one could criticize it as confounded by the single-integer response prompt, that is a validity or alternative-explanation concern, not a circularity concern. The paper also does not rely on load-bearing self-citations: the cited works are external, and none is invoked as a uniqueness theorem or as the source of an ansatz that smuggles in the conclusion. Therefore no circular step can be exhibited, and the derivation chain is self-contained with respect to circularity.
Assumptions & free parameters
free parameters (1)
- sampling temperature =
1.0 (swept 0.0 to 2.0 in follow-up)
assumptions (5)
- domain assumption Likert ratings of implied ability are a valid operationalization of scalar judgments for both humans and LLMs.
- domain assumption The 60 sentence stimuli (30 per particle) are representative of the broader class of even/only constructions.
- domain assumption The four response-scale configurations are a valid test of the spatial coding account.
- domain assumption LLM sampling at temperature 1.0, repeated 20 times, is an appropriate way to approximate a response distribution.
- standard math Linear mixed-effects model with random intercepts for participants and items adequately accounts for the data structure.
Cite this review
Pith. "Pith review of Focus particles and scalar inferences across humans and language models." pith.science (2026). https://pith.science/paper/CI4ZTW5H
@misc{pith2026260808227,
author = {Pith},
title = {Pith review of: Focus particles and scalar inferences across humans and language models},
year = {2026},
howpublished = {\url{https://pith.science/paper/CI4ZTW5H}},
note = {Machine review of arXiv:2608.08227}
}
read the original abstract
Focus particles such as "even" and "only" are central to formal semantic theories that posit structured representations over sets of alternatives. "Even" highlights unexpected or extreme alternatives, while "only" enforces exclusivity. If such scalar representations are robust and generalizable, they should give rise to consistent judgments across contexts and systems. In this work, we test whether humans and large language models (LLMs) construct stable scalar representations from sentences containing these particles. Using a dataset of approximately 100 items, participants and models were asked to make scalar judgments. Preliminary results suggest that similar outputs across humans and LLMs may arise from different underlying mechanisms.
Figures
Reference graph
Works this paper leans on
-
[1]
arXiv preprint arXiv:2603.18171 , year=
Modeling the human lexicon under temperature variations: linguistic factors, diversity and typicality in LLM word associations , author=. arXiv preprint arXiv:2603.18171 , year=
-
[2]
arXiv preprint arXiv:2505.16164 , year=
Can LLMs Simulate Human Behavioral Variability? A Case Study in the Phonemic Fluency Task , author=. arXiv preprint arXiv:2505.16164 , year=
-
[3]
arXiv preprint arXiv:2502.05234 , year=
Optimizing temperature for language models with multi-sample inference , author=. arXiv preprint arXiv:2502.05234 , year=
-
[4]
The Necessity of Setting Temperature in LLM-as-a-Judge
The Necessity of Setting Temperature in LLM-as-a-Judge , author=. arXiv preprint arXiv:2603.28304 , year=
-
[5]
arXiv preprint arXiv:2302.13971 , year=
Llama: Open and efficient foundation language models , author=. arXiv preprint arXiv:2302.13971 , year=
-
[6]
Advances in neural information processing systems , volume=
Attention is all you need , author=. Advances in neural information processing systems , volume=
-
[7]
arXiv preprint arXiv:2512.03676 , year=
Different types of syntactic agreement recruit the same units within large language models , author=. arXiv preprint arXiv:2512.03676 , year=
-
[8]
Proceedings of the National Academy of Sciences , volume=
Using cognitive psychology to understand GPT-3 , author=. Proceedings of the National Academy of Sciences , volume=. 2023 , publisher=
2023
Show all 21 references
-
[9]
The generic book , pages=
Focus and the interpretation of generic sentences , author=. The generic book , pages=. 1995 , publisher=
1995
-
[10]
Findings of the Association for Computational Linguistics: ACL 2025 , pages=
Is it JUST semantics? a case study of discourse particle understanding in LLMs , author=. Findings of the Association for Computational Linguistics: ACL 2025 , pages=
2025
-
[11]
, author=
The mental representation of parity and number magnitude. , author=. Journal of experimental psychology: General , volume=. 1993 , publisher=
1993
-
[12]
Mathematical cognition , volume=
The importance of magnitude information in numerical processing: Evidence from the SNARC effect , author=. Mathematical cognition , volume=. 1996 , publisher=
1996
-
[13]
Cognition , volume=
The mental representation of ordinal sequences is spatially organized , author=. Cognition , volume=. 2003 , publisher=
2003
-
[14]
Cognition , volume=
The SNARC effect does not imply a mental number line , author=. Cognition , volume=. 2008 , publisher=
2008
-
[15]
Memory & cognition , volume=
Spatial structure of quantitative representation of numbers: Evidence from the SNARC effect , author=. Memory & cognition , volume=. 2004 , publisher=
2004
-
[16]
, author=
Polarity correspondence: A general principle for performance of speeded binary classification tasks. , author=. Psychological bulletin , volume=. 2006 , publisher=
2006
-
[17]
Cognitive Effects in Large Language Models , ISBN=
Shaki, Jonathan and Kraus, Sarit and Wooldridge, Michael , year=. Cognitive Effects in Large Language Models , ISBN=. doi:10.3233/faia230505 , booktitle=
-
[18]
Newell and H
A. Newell and H. A. Simon , title =
-
[19]
Computational models of scientific discovery and theory formation , publisher =
-
[20]
Natural language semantics , volume=
A theory of focus interpretation , author=. Natural language semantics , volume=. 1992 , publisher=
1992
-
[21]
Semantics and Linguistic Theory , pages=
A compositional semantics for multiple focus constructions , author=. Semantics and Linguistic Theory , pages=
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.