REVIEW 6 minor 12 references
Language models agree with each other far more than with human readers, even when the human baseline is naturalistic and uninstructed.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-03 10:12 UTC pith:INX4YHXL
load-bearing objection A rigorous, unusually candid measurement of model-model vs. reader-reader agreement, but the headline claim overreaches: the model side gets a task and the human side doesn't, and the paper says so itself.
Language Models Agree With Each Other, Not With Readers
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
Measured on a scale where chance agreement from sentence position and length is subtracted out, two language models share about 8.7 of 14 chosen sentences on a median document while two readers share about 4.1, an excess of 2.8 sentences against 0.6. Across 153 model pairs from 18 models spanning vendors, countries, sizes, and weight regimes, the median excess agreement is +0.093 versus +0.040 for reader pairs, with 99 pairs above the human interval. The two frontier models from rival labs reach +0.203, more than twice what a leading frontier model agrees with itself on a second call. The effect is graded: the smallest models agree at the human level. Agreement with readers rises with model
What carries the argument
The central instrument is a chance-adjusted excess-agreement estimator. For two sets of sentences, each of size b (20% of the document), excess agreement is the observed overlap minus the overlap expected when each set is independently resampled within narrow depth-and-length bands. This removes the dominant confound that raw overlap is mostly position. The null is calibrated by showing that random baselines land within 0.006 of zero and classical extractive algorithms near zero. The human reference is a corpus of 2,523 reader mark sets on 120 public web documents, produced by people highlighting for their own purposes on a platform where the overlay of others' marks is off by default.
Load-bearing premise
The load-bearing premise is that the readers whose marks form the human baseline made their highlights independently, without seeing others' marks; the platform setting that hides other readers' highlights was off by default and rarely enabled, but actual take-up is not verifiable from the data.
What would settle it
A pre-registered replication using documents with verified reader independence (for example, server logs of whether the overlay was enabled) that finds human-human excess agreement matching or exceeding model-model agreement would refute the central claim. So would a future model whose agreement with the crowd-consensus ceiling clears the human interval in a pre-registered out-of-sample test.
If this is right
- If valid, simulated populations built from several language models are not several populations; on this task they agree several times more than human readers do.
- Consensus among capable models cannot serve as evidence that a selection is what a reader would choose, since no model in the panel agrees with readers detectably more than a reader does.
- The claim that capability buys machine agreement and not human agreement is wrong: agreement with readers rises with generation, but it saturates near the human-human level while machine-machine agreement keeps rising.
- Because models and readers pick different sentences with the same surface properties, efforts to align models to reader taste will need to target content-level salience, not stylistic register.
- The ordering (models above humans) is robust to procedure, but the magnitude is not; a defensible alternative that blunts models like readers halves the gap, so the size of the effect depends on the comparison protocol.
Where Pith is reading between the lines
- One testable extension would apply the same position-controlled estimator to other naturalistic choice traces, such as bookmarks, annotations in e-readers, or search-click data, to see whether the model-human gap generalizes beyond highlighting.
- If the recency trend holds, agreement between models may continue to rise while agreement with readers plateaus, implying a growing divergence between model taste and human taste that could surface in recommendation, summarization, and content-curation systems.
- The paper's inability to verify reader independence suggests a measurement experiment: compare human-human agreement on documents where the overlay was demonstrably enabled versus disabled; if visible highlights raise agreement to model levels, the naturalistic baseline is inflated.
- A further consequence the authors leave implicit: the same machinery could be used to audit 'diversity' claims about model outputs, since it gives a scale on which chance, human, and machine agreement are directly comparable.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper measures agreement between language models' sentence selections and between human readers' highlights, using a naturalistic corpus of 2,523 mark sets from a social highlighting platform where readers are uninstructed and uncompensated. The authors propose a chance-adjusted overlap estimator that controls for sentence depth and length, demonstrate its calibration with random baselines, and conduct a pre-registered 18-arm model panel spanning vendors, sizes, and generations. The central finding is that model-model agreement (median +0.093) substantially exceeds human-human agreement (+0.040), with 99 of 153 pairs entirely above the human interval, and no model agreeing with readers more than a typical reader does. The ordering is shown to be robust to sensitivity analyses, a blunting control that applies the readers' procedure to models, sign tests, and out-of-sample predictions on four later models.
Significance. Provided the result holds, the paper makes a valuable methodological contribution by constructing a human baseline that is not an artifact of the study design. The calibration of the null, the pre-registration of panel and predictions, the transparent reporting of failed null versions and of the one unremovable confound (Section 8), and the sign test robustness are all exemplary. The finding that model-model agreement grows with recency while reader agreement saturates is an important nuance. This work strengthens the evidence for model homogenization by measuring it against a naturalistic reference, and it gives future studies a template for avoiding the 'instructed human' artifact.
minor comments (6)
- [Abstract] Typographical errors: 'T ested' should be 'Tested', and 'sharpestb' appears to be a rendering artifact. Please proofread the abstract and body for similar formatting slips.
- [5.6] The out-of-sample test covers only four arms from a single vendor. While the paper acknowledges this, it should be stated in the abstract or conclusion that the out-of-sample result is vendor-specific.
- [5.4, Table 2] In Table 2, the row for 'same country' has a dash in the range column; consider providing the range or a note for consistency.
- [6] In the blunting table, the bottom row lists 8 domains; the text correctly says 'on 8 domains', but the table header could be more explicit that 'docs' and 'domains' refer to the subset used in that row.
- [8] The self-referential narrative 'an earlier version of this paper...' appears several times. While transparency is valuable, condensing these remarks into footnotes or an appendix would improve readability without sacrificing disclosure.
- [4] In the definition of excess agreement, it would help to state explicitly that the expectation is taken over the null distribution, and to define the depth-length band in words once in addition to the formal statement.
Circularity Check
No circularity: the central comparison is self-contained, externally calibrated, and its admitted confounds are disclosed limitations rather than construction steps.
full rationale
The paper's derivation chain does not reduce its target result to its inputs. The human reference is a pre-existing, uninstructed corpus (2,523 mark sets across 120 documents) external to the study; model-model and human-human agreement are computed with the same estimator on different pairs. The null is not fitted to the outcome: its one free parameter (band tolerance) is chosen on grounds of null degeneracy, not to maximize the gap, and Section 5.7 shows the ordering survives all three tolerances. Calibration is checked against random baselines and classical extractive pairs, which are independent of the substantive claim. The main asymmetry (models receive a ranking task; readers receive none) is explicitly acknowledged in Section 8 as an unremovable confound and is a scope limitation, not a construction that makes model agreement equal to the instruction. The post-hoc items (shared-goal form, routing as DEV-2) are disclosed and do not enter the central ordering. The self-citation to earlier work [12] supplies the estimator, corpus and prompt, but the estimator's validity is demonstrated internally (random-baseline and classical-pair checks, plus the reported null-correction failures), and out-of-sample predictions were fixed before calls. No fitted parameter is renamed as a prediction. The robustness section addresses the truncation asymmetry by blunting models to the reader procedure and shows the gap remains above zero where the measurement has power. Therefore no circular step can be exhibited.
Axiom & Free-Parameter Ledger
free parameters (2)
- band_width_tolerance =
0.10 (relative depth and length rank)
- keep_set_budget_ratio =
0.2 (round(0.2n))
axioms (5)
- domain assumption Readers' highlights are produced independently; the on-page overlay of other readers' marks is off by default and rarely enabled.
- domain assumption The 2,523 mark sets approximate independent reader samples; uid-free shuffling prevents tracking individuals.
- domain assumption The depth-and-length resampling null captures chance overlap due to position/length.
- domain assumption Models returning document-order rankings still provide usable arms after neutralization by the null.
- domain assumption Selecting the top 20% of sentences by importance is a meaningful operationalization of highlighting behavior.
read the original abstract
Claims that language models homogenise are usually measured against human judgements collected for the study, which makes the human side an artifact of the design: a crowdworker given the model's instruction is running the model's prompt. We measure convergence against a human reference nobody built for the purpose -- 2,523 reader mark sets across 120 web documents, produced by people highlighting for their own reasons on a platform where the overlay of others' marks is off by default. Agreement is the overlap between two size-matched sentence sets minus the overlap expected when each is resampled within its own depth-and-length bands. The null's calibration is demonstrated, not asserted: every pair involving a random baseline lands within 0.006 of zero. On the median document each party names 14 sentences of 70; two readers share 4.1 and two models 8.7. Across 18 model arms spanning 11 vendors, 3 countries and both weight regimes, the median of 153 model pairs is +0.093 against a human yardstick of +0.040, and 99 sit entirely above the human interval. Two frontier models from rival labs reach +0.203, twice what GPT-4o agrees with itself on a second call. The effect is not determinism, prompt wording, procedure, vendor or routing, and it is graded: the smallest models agree at the human level. No model agrees with readers detectably more than a reader does, and at equal depth and length no surface feature separates their choices. The multiples are procedure-dependent and the ordering is not: models are cut to their sharpest set while a reader's is a random draw from what they marked, and blunting the models alike halves the gap without closing it. Tested out of sample on four models released after this analysis, against predictions fixed beforehand, none clears the human interval. A population simulated from several models is not several populations.
Reference graph
Works this paper leans on
-
[1]
A. R. Doshi and O. P. Hauser. Generative AI enhances individual creativity but reduces the collective diversity of novel content.Science Advances, 10(28):eadn5290, 2024
2024
-
[2]
V. Padmakumar and H. He. Does writing with language models reduce content diversity? In International Conference on Learning Representations (ICLR), 2024. arXiv:2309.05196
Pith/arXiv arXiv 2024
-
[3]
Gilardi, M
F. Gilardi, M. Alizadeh, and M. Kubli. ChatGPT outperforms crowd workers for text- annotation tasks.Proceedings of the National Academy of Sciences, 120(30):e2305016120, 2023
2023
-
[4]
S. Goel, J. Str¨ uber, I. A. Auzina, K. K. Chandra, P. Kumaraguru, D. Kiela, A. Prabhu, M. Bethge, and J. Geiping. Great models think alike and this undermines AI oversight. In International Conference on Machine Learning (ICML), 2025. arXiv:2502.04313
Pith/arXiv arXiv 2025
-
[5]
J. Trienes, J. Schl¨ otterer, J. J. Li, and C. Seifert. Behavioral analysis of information salience in large language models. InFindings of the Association for Computational Linguistics: ACL 2025, 2025. arXiv:2502.14613
Pith/arXiv arXiv 2025
-
[6]
M. V. Reiss. Testing the reliability of ChatGPT for text annotation and classification: A cautionary remark. arXiv:2304.11085, 2023. 17
Pith/arXiv arXiv 2023
-
[7]
G. J. Rath, A. Resnick, and T. R. Savage. The formation of abstracts by the selection of sentences.American Documentation, 12(2):139–143, 1961
1961
-
[8]
Kleinberg and M
J. Kleinberg and M. Raghavan. Algorithmic monoculture and social welfare.Proceedings of the National Academy of Sciences, 118(22):e2018340118, 2021
2021
-
[9]
R. Bommasani, K. A. Creel, A. Kumar, D. Jurafsky, and P. Liang. Picking on the same person: Does algorithmic monoculture lead to outcome homogenization? InAdvances in Neural Information Processing Systems (NeurIPS), 2022. arXiv:2211.13972
Pith/arXiv arXiv 2022
-
[10]
L. P. Argyle, E. C. Busby, N. Fulda, J. R. Gubler, C. Rytting, and D. Wingate. Out of one, many: Using language models to simulate human samples.Political Analysis, 31(3):337–351, 2023
2023
-
[11]
H. P. Luhn. The automatic creation of literature abstracts.IBM Journal of Research and Development, 2(2):159–165, 1958
1958
-
[12]
K. Nakayashiki and K. Watanabe. Measuring Alignment With Reader Highlights Net of Position and Length. arXiv:2607.27739 [cs.IR], 2026. 18
Pith/arXiv arXiv 2026
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.