Pith. sign in

REVIEW 6 minor 12 references

Language models agree with each other far more than with human readers, even when the human baseline is naturalistic and uninstructed.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-03 10:12 UTC pith:INX4YHXL

load-bearing objection A rigorous, unusually candid measurement of model-model vs. reader-reader agreement, but the headline claim overreaches: the model side gets a task and the human side doesn't, and the paper says so itself.

arxiv 2607.29274 v1 pith:INX4YHXL submitted 2026-07-31 cs.IR cs.CLcs.CYcs.HC

Language Models Agree With Each Other, Not With Readers

classification cs.IR cs.CLcs.CYcs.HC
keywords language modelshomogenizationconvergencereader highlightsnaturalistic baselineexcess agreementposition-controlled nullmodel-human agreement
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper claims that the well-known homogenization of language models is not an artifact of how human comparisons are usually collected. Using 2,523 sets of highlights made by real readers on 120 web documents, where readers marked text for their own reasons and did not see others' marks, it measures how often two models pick the same sentences versus how often two readers do. After correcting for position and length, the median model pair agrees 2.3 times more than two readers, and the strongest model pair agrees more than a model agrees with itself on a second call. No model in an 18-model panel agrees with readers detectably more than one reader agrees with another. The paper argues convergence is graded, growing with scale and recency, and is not explained by determinism, prompt wording, vendor, or procedure.

Core claim

Measured on a scale where chance agreement from sentence position and length is subtracted out, two language models share about 8.7 of 14 chosen sentences on a median document while two readers share about 4.1, an excess of 2.8 sentences against 0.6. Across 153 model pairs from 18 models spanning vendors, countries, sizes, and weight regimes, the median excess agreement is +0.093 versus +0.040 for reader pairs, with 99 pairs above the human interval. The two frontier models from rival labs reach +0.203, more than twice what a leading frontier model agrees with itself on a second call. The effect is graded: the smallest models agree at the human level. Agreement with readers rises with model

What carries the argument

The central instrument is a chance-adjusted excess-agreement estimator. For two sets of sentences, each of size b (20% of the document), excess agreement is the observed overlap minus the overlap expected when each set is independently resampled within narrow depth-and-length bands. This removes the dominant confound that raw overlap is mostly position. The null is calibrated by showing that random baselines land within 0.006 of zero and classical extractive algorithms near zero. The human reference is a corpus of 2,523 reader mark sets on 120 public web documents, produced by people highlighting for their own purposes on a platform where the overlay of others' marks is off by default.

Load-bearing premise

The load-bearing premise is that the readers whose marks form the human baseline made their highlights independently, without seeing others' marks; the platform setting that hides other readers' highlights was off by default and rarely enabled, but actual take-up is not verifiable from the data.

What would settle it

A pre-registered replication using documents with verified reader independence (for example, server logs of whether the overlay was enabled) that finds human-human excess agreement matching or exceeding model-model agreement would refute the central claim. So would a future model whose agreement with the crowd-consensus ceiling clears the human interval in a pre-registered out-of-sample test.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • If valid, simulated populations built from several language models are not several populations; on this task they agree several times more than human readers do.
  • Consensus among capable models cannot serve as evidence that a selection is what a reader would choose, since no model in the panel agrees with readers detectably more than a reader does.
  • The claim that capability buys machine agreement and not human agreement is wrong: agreement with readers rises with generation, but it saturates near the human-human level while machine-machine agreement keeps rising.
  • Because models and readers pick different sentences with the same surface properties, efforts to align models to reader taste will need to target content-level salience, not stylistic register.
  • The ordering (models above humans) is robust to procedure, but the magnitude is not; a defensible alternative that blunts models like readers halves the gap, so the size of the effect depends on the comparison protocol.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • One testable extension would apply the same position-controlled estimator to other naturalistic choice traces, such as bookmarks, annotations in e-readers, or search-click data, to see whether the model-human gap generalizes beyond highlighting.
  • If the recency trend holds, agreement between models may continue to rise while agreement with readers plateaus, implying a growing divergence between model taste and human taste that could surface in recommendation, summarization, and content-curation systems.
  • The paper's inability to verify reader independence suggests a measurement experiment: compare human-human agreement on documents where the overlay was demonstrably enabled versus disabled; if visible highlights raise agreement to model levels, the naturalistic baseline is inflated.
  • A further consequence the authors leave implicit: the same machinery could be used to audit 'diversity' claims about model outputs, since it gives a scale on which chance, human, and machine agreement are directly comparable.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

0 major / 6 minor

Summary. This paper measures agreement between language models' sentence selections and between human readers' highlights, using a naturalistic corpus of 2,523 mark sets from a social highlighting platform where readers are uninstructed and uncompensated. The authors propose a chance-adjusted overlap estimator that controls for sentence depth and length, demonstrate its calibration with random baselines, and conduct a pre-registered 18-arm model panel spanning vendors, sizes, and generations. The central finding is that model-model agreement (median +0.093) substantially exceeds human-human agreement (+0.040), with 99 of 153 pairs entirely above the human interval, and no model agreeing with readers more than a typical reader does. The ordering is shown to be robust to sensitivity analyses, a blunting control that applies the readers' procedure to models, sign tests, and out-of-sample predictions on four later models.

Significance. Provided the result holds, the paper makes a valuable methodological contribution by constructing a human baseline that is not an artifact of the study design. The calibration of the null, the pre-registration of panel and predictions, the transparent reporting of failed null versions and of the one unremovable confound (Section 8), and the sign test robustness are all exemplary. The finding that model-model agreement grows with recency while reader agreement saturates is an important nuance. This work strengthens the evidence for model homogenization by measuring it against a naturalistic reference, and it gives future studies a template for avoiding the 'instructed human' artifact.

minor comments (6)
  1. [Abstract] Typographical errors: 'T ested' should be 'Tested', and 'sharpestb' appears to be a rendering artifact. Please proofread the abstract and body for similar formatting slips.
  2. [5.6] The out-of-sample test covers only four arms from a single vendor. While the paper acknowledges this, it should be stated in the abstract or conclusion that the out-of-sample result is vendor-specific.
  3. [5.4, Table 2] In Table 2, the row for 'same country' has a dash in the range column; consider providing the range or a note for consistency.
  4. [6] In the blunting table, the bottom row lists 8 domains; the text correctly says 'on 8 domains', but the table header could be more explicit that 'docs' and 'domains' refer to the subset used in that row.
  5. [8] The self-referential narrative 'an earlier version of this paper...' appears several times. While transparency is valuable, condensing these remarks into footnotes or an appendix would improve readability without sacrificing disclosure.
  6. [4] In the definition of excess agreement, it would help to state explicitly that the expectation is taken over the null distribution, and to define the depth-length band in words once in addition to the formal statement.

Circularity Check

0 steps flagged

No circularity: the central comparison is self-contained, externally calibrated, and its admitted confounds are disclosed limitations rather than construction steps.

full rationale

The paper's derivation chain does not reduce its target result to its inputs. The human reference is a pre-existing, uninstructed corpus (2,523 mark sets across 120 documents) external to the study; model-model and human-human agreement are computed with the same estimator on different pairs. The null is not fitted to the outcome: its one free parameter (band tolerance) is chosen on grounds of null degeneracy, not to maximize the gap, and Section 5.7 shows the ordering survives all three tolerances. Calibration is checked against random baselines and classical extractive pairs, which are independent of the substantive claim. The main asymmetry (models receive a ranking task; readers receive none) is explicitly acknowledged in Section 8 as an unremovable confound and is a scope limitation, not a construction that makes model agreement equal to the instruction. The post-hoc items (shared-goal form, routing as DEV-2) are disclosed and do not enter the central ordering. The self-citation to earlier work [12] supplies the estimator, corpus and prompt, but the estimator's validity is demonstrated internally (random-baseline and classical-pair checks, plus the reported null-correction failures), and out-of-sample predictions were fixed before calls. No fitted parameter is renamed as a prediction. The robustness section addresses the truncation asymmetry by blunting models to the reader procedure and shows the gap remains above zero where the measurement has power. Therefore no circular step can be exhibited.

Axiom & Free-Parameter Ledger

2 free parameters · 5 axioms · 0 invented entities

The central claim rests on two unverified domain assumptions (reader independence and non-identifiability of readers), one data-informed modeling choice (0.10 band width), and a task-design assumption that models were given an instruction while readers were not. The paper discloses each of these and provides sensitivity or bounding analyses. No invented entities are introduced.

free parameters (2)
  • band_width_tolerance = 0.10 (relative depth and length rank)
    The null's only free parameter; changed from 0.05 in prior work because 27.5% of sentences had no permissible replacement at 0.05. Ordering is robust across 0.05/0.10/0.20 but magnitudes are procedure-dependent (Sections 4 and 5.7).
  • keep_set_budget_ratio = 0.2 (round(0.2n))
    Models and readers are cut to the top 20% of sentences; the gap persists at 0.10 and 0.30 budget ratios but the magnitude changes (Section 6).
axioms (5)
  • domain assumption Readers' highlights are produced independently; the on-page overlay of other readers' marks is off by default and rarely enabled.
    Section 3 (Independence): explicitly an assumption, not a finding; take-up of the overlay setting is not verifiable from the artifacts. Load-bearing for the naturalistic baseline.
  • domain assumption The 2,523 mark sets approximate independent reader samples; uid-free shuffling prevents tracking individuals.
    Section 8 (Limitations): heavy users could account for many sets, narrowing human-side intervals; the sign test on 90 documents is unaffected.
  • domain assumption The depth-and-length resampling null captures chance overlap due to position/length.
    Section 4: calibration is demonstrated by random baselines landing within 0.006 of zero; standard statistical model assumption.
  • domain assumption Models returning document-order rankings still provide usable arms after neutralization by the null.
    Section 8 (Limitations): 9 of 18 arms show identity rankings on some documents; the depth-matched null neutralizes them, and the algorithm control (naive truncation ≈ 0) confirms.
  • domain assumption Selecting the top 20% of sentences by importance is a meaningful operationalization of highlighting behavior.
    Section 3; the task-vs-no-task confound is acknowledged as unremovable in Section 8.

pith-pipeline@v1.3.0-daily-deepseek · 14301 in / 17276 out tokens · 164539 ms · 2026-08-03T10:12:01.206417+00:00 · methodology

0 comments
read the original abstract

Claims that language models homogenise are usually measured against human judgements collected for the study, which makes the human side an artifact of the design: a crowdworker given the model's instruction is running the model's prompt. We measure convergence against a human reference nobody built for the purpose -- 2,523 reader mark sets across 120 web documents, produced by people highlighting for their own reasons on a platform where the overlay of others' marks is off by default. Agreement is the overlap between two size-matched sentence sets minus the overlap expected when each is resampled within its own depth-and-length bands. The null's calibration is demonstrated, not asserted: every pair involving a random baseline lands within 0.006 of zero. On the median document each party names 14 sentences of 70; two readers share 4.1 and two models 8.7. Across 18 model arms spanning 11 vendors, 3 countries and both weight regimes, the median of 153 model pairs is +0.093 against a human yardstick of +0.040, and 99 sit entirely above the human interval. Two frontier models from rival labs reach +0.203, twice what GPT-4o agrees with itself on a second call. The effect is not determinism, prompt wording, procedure, vendor or routing, and it is graded: the smallest models agree at the human level. No model agrees with readers detectably more than a reader does, and at equal depth and length no surface feature separates their choices. The multiples are procedure-dependent and the ordering is not: models are cut to their sharpest set while a reader's is a random draw from what they marked, and blunting the models alike halves the gap without closing it. Tested out of sample on four models released after this analysis, against predictions fixed beforehand, none clears the human interval. A population simulated from several models is not several populations.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

12 extracted references · 6 linked inside Pith

  1. [1]

    A. R. Doshi and O. P. Hauser. Generative AI enhances individual creativity but reduces the collective diversity of novel content.Science Advances, 10(28):eadn5290, 2024

  2. [2]

    Padmakumar and H

    V. Padmakumar and H. He. Does writing with language models reduce content diversity? In International Conference on Learning Representations (ICLR), 2024. arXiv:2309.05196

  3. [3]

    Gilardi, M

    F. Gilardi, M. Alizadeh, and M. Kubli. ChatGPT outperforms crowd workers for text- annotation tasks.Proceedings of the National Academy of Sciences, 120(30):e2305016120, 2023

  4. [4]

    S. Goel, J. Str¨ uber, I. A. Auzina, K. K. Chandra, P. Kumaraguru, D. Kiela, A. Prabhu, M. Bethge, and J. Geiping. Great models think alike and this undermines AI oversight. In International Conference on Machine Learning (ICML), 2025. arXiv:2502.04313

  5. [5]

    Trienes, J

    J. Trienes, J. Schl¨ otterer, J. J. Li, and C. Seifert. Behavioral analysis of information salience in large language models. InFindings of the Association for Computational Linguistics: ACL 2025, 2025. arXiv:2502.14613

  6. [6]

    M. V. Reiss. Testing the reliability of ChatGPT for text annotation and classification: A cautionary remark. arXiv:2304.11085, 2023. 17

  7. [7]

    G. J. Rath, A. Resnick, and T. R. Savage. The formation of abstracts by the selection of sentences.American Documentation, 12(2):139–143, 1961

  8. [8]

    Kleinberg and M

    J. Kleinberg and M. Raghavan. Algorithmic monoculture and social welfare.Proceedings of the National Academy of Sciences, 118(22):e2018340118, 2021

  9. [9]

    Bommasani, K

    R. Bommasani, K. A. Creel, A. Kumar, D. Jurafsky, and P. Liang. Picking on the same person: Does algorithmic monoculture lead to outcome homogenization? InAdvances in Neural Information Processing Systems (NeurIPS), 2022. arXiv:2211.13972

  10. [10]

    L. P. Argyle, E. C. Busby, N. Fulda, J. R. Gubler, C. Rytting, and D. Wingate. Out of one, many: Using language models to simulate human samples.Political Analysis, 31(3):337–351, 2023

  11. [11]

    H. P. Luhn. The automatic creation of literature abstracts.IBM Journal of Research and Development, 2(2):159–165, 1958

  12. [12]

    Nakayashiki and K

    K. Nakayashiki and K. Watanabe. Measuring Alignment With Reader Highlights Net of Position and Length. arXiv:2607.27739 [cs.IR], 2026. 18