Pith. sign in

REVIEW 4 major objections 4 minor

A one-word census of 44 language models finds extreme answer-choice convergence, with newest mainline flagships most conformist and persona-tuned models most divergent.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-15 03:16 UTC pith:YQN6KZ6Q

load-bearing objection Clean, cheap instrument for measuring answer-choice conformity across models; the rankings and generational claims rest on only four draws and need error bars. the 4 major comments →

arxiv 2607.12796 v2 pith:YQN6KZ6Q submitted 2026-07-14 cs.CL cs.AIcs.CY

The One-Word Census: Answer-Choice Conformity Across 44 Language Models

classification cs.CL cs.AIcs.CY
keywords language modelsanswer-choice conformityone-word censusleave-one-out surprisalmodel diversitycategory productionLLM evaluationhomogenization
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

When language models must freely pick one word from a large space of equally valid options, they converge far more than chance would suggest: asked to "pick a word—any word," 44 models chose "serendipity" 41 percent of the time, and in seven of 31 categories a single answer took more than 80 percent of all responses. The paper introduces a deliberately minimal instrument—the One-Word Census—31 single-turn category prompts asked four times each with no system prompt, scored by exact-match on normalized tokens. It then ranks every model by leave-one-out answer-choice surprisal, the average surprise of its answers under the pooled answers of every other model. The ranking is structured: persona- and community-tuned models diverge most, while the newest flagships produce almost no unique answers; conformity also rises across successive generations inside four major lineages before a possible reversal at the latest Claude and GPT flagships. The field is more concentrated than human category-production norms in 18 of 20 shared categories, and the rankings remain stable when any one model family is held out.

Core claim

The One-Word Census shows that free one-word answer choice is extremely convergent across 44 language models, yet the degree of conformity varies more than fourfold and is systematically ordered: persona- and community-tuned models are most divergent while newest mainline flagships are most conformist, with leave-one-out answer-choice surprisal producing rankings that are robust to leave-one-family-out (rho = 0.985).

What carries the argument

Answer-choice surprisal: the average −log₂ probability of a model’s answers under the pooled answer distribution of all other models (leave-one-out). This single scalar, computed from exact-match normalized tokens on 31 minimal category prompts, carries the ranking and the structural claims.

Load-bearing premise

That four exact-match normalized answers to each of 31 deliberately minimal single-turn prompts, with no system prompt, give a stable and representative measure of a model’s answer-choice distribution and of field-wide conformity.

What would settle it

Re-running the identical 31 prompts with substantially more samples per model (or a larger independent set of categories) and finding either that leave-one-out surprisal rankings reverse or that peak category concentrations fall to or below human norms would falsify the claims of extreme, structured conformity.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • A sub-dollar, exact-match census can rank any new model’s conformity relative to the existing field without embeddings or judges.
  • Newest mainline flagships currently occupy the most conformist end of the spectrum and almost never emit answers unseen in the rest of the field.
  • Within Claude, GPT, Qwen and Grok lineages, conformity has risen with each generation, with a possible early reversal only at the latest top-tier Claude and GPT models.
  • The aggregate model field is more concentrated than human category norms on the large majority of shared categories, implying reduced answer diversity relative to people.
  • Rankings remain essentially unchanged under leave-one-family-out, so the conformity order is not an artifact of any single model family dominating the pool.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The same leave-one-out surprisal could serve as a lightweight monitor for whether future post-training (RLHF, preference tuning, or community fine-tunes) is increasing or decreasing homogenization.
  • If the generational rise-then-reversal pattern continues, it may indicate deliberate repositioning by frontier labs away from pure consensus answers once a conformity ceiling is reached.
  • Because the instrument uses only exact token match and four samples, it can be re-run continuously as new models appear, turning the public release into an ongoing conformity index rather than a one-shot snapshot.
  • The excess concentration relative to humans raises a testable question about whether training-data overlap, shared preference models, or decoding defaults are the dominant driver; ablating any one of those factors on open models would distinguish them.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. The paper proposes the One-Word Census: 31 deliberately minimal single-turn category prompts (e.g., “Name a tree,” “pick a word—any word”), each asked four times with no system prompt to 44 language models. Analysis is exact-match on normalized single tokens. Models are scored by leave-one-out answer-choice surprisal (average −log2 probability of a model’s answers under the pooled answers of all other models). The abstract reports extreme convergence (e.g., “serendipity” 41% of free-word answers; 7/31 categories with one answer >80%), more than fourfold variation in conformity (persona/community-tuned models most divergent; newest mainline flagships most conformist), within-lineage generational rise with a late Claude/GPT reversal, leave-one-family-out ranking robustness (ρ=0.985), and greater concentration than human norms in 18/20 shared categories. Prompts, transcripts, and code are stated to be public.

Significance. If the instrument is statistically stable and the structured conformity patterns hold, this is a cheap, reproducible, judge-free diagnostic of field-wide answer-choice homogenization that could track training-data overlap and product positioning over time. Explicit strengths include the minimal exact-match design (no embeddings or LLM judges), leave-one-out surprisal without fitted free parameters, the leave-one-family-out robustness check, the human category-norm comparison, and the commitment to full public release of prompts, transcripts, and code. Those are real methodological assets for a conformity census.

major comments (4)
  1. Abstract: each of the 31 prompts is “asked four times per model.” With n=4 multinomial draws, modal frequencies (3/4 vs 4/4) and the presence/absence of a single unique token can materially swing a model’s average leave-one-out −log2 p and its rank. The fine-grained claims—fourfold conformity range, persona models most divergent, newest flagships most conformist, and the within-lineage Claude/GPT reversal—are load-bearing and rest on these high-variance estimates. The manuscript needs either substantially more draws per prompt or bootstrap/standard-error quantification showing that the reported ranking structure exceeds sampling noise.
  2. Abstract: “Rankings are robust to roster composition (leave-one-family-out rho = 0.985).” Because the leave-one-family-out correlation is computed on the same n=4 exact-match estimates, it does not independently validate ranking stability against sampling variance; it mainly shows that no single family dominates the pool. A proper stability check would resample answers within models (or increase n) and recompute ranks.
  3. Abstract: “Analysis is exact-match on normalized tokens.” Exact-match collapses near-synonyms and morphological variants, which can inflate apparent concentration and suppress genuine diversity. For the central claim of extreme convergence (7 categories >80% modal share; field more concentrated than humans in 18/20 categories), the paper should quantify sensitivity to synonym grouping or report token-level entropy alongside exact-match rates so readers can separate true mode collapse from label granularity.
  4. Abstract: decoding settings (temperature, top-p, seed policy, API defaults) are not stated. Answer-choice diversity is highly sensitive to sampling temperature; if models were run at different defaults or at temperature 0, the conformity ranking and the “almost no answer no other model gave” claim for flagships could be partly an artifact of decoding rather than of model preference. This must be fixed and held constant (or ablated) for the comparative claims to be interpretable.
minor comments (4)
  1. Abstract only is available for this review; the full methods section should list the exact 31 prompts, the 44 models with family labels, normalization rules, and any temperature/API settings so the instrument is fully reproducible from the paper alone.
  2. Clarify how “normalized tokens” handle multi-token answers, capitalization, punctuation, and pluralization; a short appendix table of normalization examples would prevent ambiguity.
  3. The human comparison (“18 of 20 shared categories”) should name the human norm source and the matching procedure so the concentration claim can be audited.
  4. Report per-category and per-model answer entropies (or unique-answer counts) alongside surprisal so readers can see whether high conformity is driven by a few ultra-concentrated categories.

Circularity Check

0 steps flagged

No significant circularity: leave-one-out surprisal is an empirical measurement against other models, not forced by definition or self-fit.

full rationale

The paper's central quantities (cross-model answer concentration, leave-one-out answer-choice surprisal, and the resulting conformity rankings) are defined from raw exact-match token counts on a fixed set of 31 minimal prompts asked four times each. Surprisal for model M is the average -log2 probability of M's answers under the empirical distribution of all models except M; a model's own answers therefore do not enter its score, so the ranking is not self-definitional. No free parameters are fitted to a subset of the data and then re-presented as predictions; no uniqueness theorems or ansatzes are imported via self-citation; and the instrument is not a renaming of a prior closed-form result. Extreme modal frequencies (e.g., 'serendipity' 41 %, seven categories >80 %) and the leave-one-family-out robustness (rho=0.985) are direct empirical observations on the collected transcripts. Shared training-data modes that produce the observed convergence are domain structure being measured, not a circular reduction of the derivation chain. With only the abstract available, no load-bearing self-citation chain or definitional identity appears. Score 0 is therefore the correct, non-manufactured finding.

Axiom & Free-Parameter Ledger

0 free parameters · 3 axioms · 0 invented entities

Abstract-only: no free parameters are fitted; scoring is leave-one-out surprisal on exact-match tokens. Domain assumptions are the usual ones for single-turn LM sampling and category norms. No new physical or theoretical entities are invented; "answer-choice surprisal" is a standard information-theoretic score applied to the pooled answer distribution.

axioms (3)
  • domain assumption Exact-match on normalized single tokens is a sufficient and unbiased measure of answer choice for open categories.
    Stated as the analysis method; ignores near-synonyms, multi-token answers, and sampling temperature effects.
  • domain assumption Four independent single-turn draws with no system prompt adequately sample each model's answer distribution.
    Sample size and prompt regime are fixed by the instrument description; stability under more draws or system prompts is assumed.
  • domain assumption Leave-one-out pooled answers of the other 43 models form a meaningful reference distribution for "field" conformity.
    Defines the surprisal score; roster composition is claimed robust (rho=0.985) but still an operational choice.

pith-pipeline@v1.1.0-grok45 · 6233 in / 2484 out tokens · 19166 ms · 2026-07-15T03:16:30.910203+00:00 · methodology

0 comments
read the original abstract

When a language model must pick one answer from a large space of equally valid options, which does it pick -- and how often is it the same answer every other model picks? Asked to "pick a word -- any word," 44 models chose "serendipity" 41% of the time. We characterize this convergence with a deliberately minimal instrument: 31 single-turn prompts, each naming a category with many valid one-word answers ("Name a tree."), asked four times per model with no system prompt. Analysis is exact-match on normalized tokens -- no embeddings, no judge -- at about a dollar per model. That models converge is well documented; our contribution is the instrument itself -- the One-Word Census -- and what it reveals about the structure of the convergence. We score each model by answer-choice surprisal: the average $-\log2$ probability of its answers under the pooled answers of all other models, leave-one-out. Convergence is extreme -- in 7 of 31 categories one answer takes over 80% of all answers -- yet conformity varies more than fourfold across models, and the variation is structured. Persona- and community-tuned models are the most divergent; the newest mainline flagships are the most conformist, producing almost no answer no other model gave. Within four lineages (Claude, GPT, Qwen, Grok) conformity rises with each generation -- but reverses for the latest flagship Claude and GPT models, a possible early signal of repositioning at the top tier. Rankings are robust to roster composition (leave-one-family-out rho = 0.985). Against human category-production norms, the field is more concentrated than people in 18 of 20 shared categories. All prompts, transcripts, and code are public.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.