Pith. sign in

REVIEW 2 major objections 1 minor 13 references

Computational conceptual history of scientific concepts: From early digital methods to LLMs

T0 review · 2 major / 1 minor · reviewed 2026-06-28 · grok-4.3

Pith's one-line read LLMs extend computational methods for tracking scientific concept change but repeat earlier problems with data selection, modeling choices, and result validation.

desk verdict A clear survey that links pre-LLM computational work in HPSS to LLM applications and flags the same old methodological issues, without adding new findings. read the letter →

arxiv 2606.04118 v1 pith:NBNVGG2O submitted 2026-06-02 cs.CL

classification cs.CL
keywords computationalconceptualhistorylargelanguagemodelslexicalsemanticchangedetectionofsciencedigitalhumanitiesdistributionalsemantics
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper reconstructs the pre-LLM history of computational approaches to concept analysis in the history, philosophy, and sociology of science by combining three strands of earlier work. It then reviews how LLMs are being applied to lexical semantic change detection and specific HPSS case studies, showing that familiar difficulties around corpus construction, operationalization, and evaluation persist. A reader would care because this places current LLM experiments in a longer methodological conversation rather than treating them as an entirely new start. The account makes clear that scale and new architectures do not automatically remove the need for careful decisions about sources and interpretation.

What carries the argument

Three strands of prior work—early digital methods in HPSS, distributional approaches from digital history, and lexical semantic change detection—used to frame how LLMs handle corpus construction, operationalization, and evaluation.

What would settle it

An LLM-based study in HPSS that produces stable, reproducible results on conceptual change without requiring manual corpus curation or expert interpretation of outputs would undermine the claim that the longstanding problems are inherited.

Watch

Extended reading notes

Core claim

The review reconstructs computational conceptual history before LLMs by bringing together early digital methods in HPSS, distributional approaches from digital history, and lexical semantic change detection. It then examines LLM-based work on lexical semantic change detection and relevant HPSS case studies, revisiting the methodological questions of corpus construction, model choice and training data, operationalization trade-offs, and evaluation and interpretation to show how these issues continue to shape LLM workflows.

Load-bearing premise

The three strands of earlier work adequately capture the main challenges and opportunities in computational conceptual history prior to LLMs.

Editorial extensions

If this is right

  • LLMs allow processing of larger and more heterogeneous corpora than earlier distributional methods for detecting shifts in scientific terminology.
  • Prompting and fine-tuning decisions in LLM workflows introduce operationalization trade-offs comparable to those in pre-LLM vector-space models.
  • Evaluation of LLM outputs for historical concepts continues to depend on alignment with expert historical knowledge rather than purely quantitative metrics.
  • Corpus construction choices remain decisive even when models can ingest raw text at scale.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The same pattern of inherited methodological limits may appear when LLMs are applied to conceptual analysis outside HPSS, such as in legal or literary history.
  • One testable extension would be to compare stability of LLM-derived concept representations against earlier methods across the same long historical text collections.
  • The review implies that progress may lie more in designing transparent workflows that document data and modeling choices than in adopting ever-larger models.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 1 minor

Summary. The paper situates LLMs within the history of computational concept analysis in HPSS. Part 1 reconstructs pre-LLM work by synthesizing three strands—early digital methods in HPSS, distributional approaches from digital history, and lexical semantic change detection—while cataloguing challenges in corpus construction, operationalization/modelling choices, and evaluation/interpretation. Part 2 introduces LLMs, reviews their use in lexical semantic change detection and HPSS case studies, and revisits the same three problem areas to argue that LLMs add capabilities but inherit the longstanding issues.

Significance. If the three-strand reconstruction is representative, the paper supplies a timely synthesis that clarifies continuities between pre-LLM and LLM-based workflows in conceptual history. Explicit credit is due for the structured two-part organization and the attempt to link methodological questions across eras via recent case studies.

major comments (2)
  1. [Introduction / first part framing] Introduction and the opening of the first part: the central claim that LLMs inherit problems around corpus construction, operationalization, and evaluation rests on the reconstruction via exactly three strands. The manuscript does not supply an explicit rationale for why these strands (rather than, e.g., topic-modeling pipelines or citation-network concept mapping common in science studies) suffice to identify the full set of longstanding challenges; without such justification the inheritance diagnosis is under-supported.
  2. [Second part, LLM case studies and revisit of methodological questions] The LLM section (second part): the assertion that the same three problem areas “play out” in LLM workflows is illustrated by case studies, yet the manuscript provides no systematic side-by-side comparison (e.g., how a specific operationalization choice in a pre-LLM distributional model maps onto an LLM prompt or fine-tuning decision) that would make the inheritance claim load-bearing rather than suggestive.
minor comments (1)
  1. [Abstract] The abstract states the two-part structure clearly but does not foreground the paper’s distinctive contribution (the explicit mapping of inherited problems onto LLM practice); a single sentence to this effect would improve reader orientation.

Simulated Author's Rebuttal

2 responses · 0 unresolved

We thank the referee for the constructive report and the recommendation for major revision. The two major comments identify opportunities to strengthen the framing of our three-strand reconstruction and the explicitness of continuities between pre-LLM and LLM workflows. We address each point below and commit to revisions that make the inheritance argument more robust while preserving the survey character of the paper.

read point-by-point responses
  1. Referee: [Introduction / first part framing] Introduction and the opening of the first part: the central claim that LLMs inherit problems around corpus construction, operationalization, and evaluation rests on the reconstruction via exactly three strands. The manuscript does not supply an explicit rationale for why these strands (rather than, e.g., topic-modeling pipelines or citation-network concept mapping common in science studies) suffice to identify the full set of longstanding challenges; without such justification the inheritance diagnosis is under-supported.

    Authors: We accept that an explicit rationale for the three strands is currently implicit rather than stated. The strands were selected because they constitute the primary text-based computational approaches that have been applied to the semantic content and historical evolution of scientific concepts within HPSS: (1) early digital methods supply the disciplinary context and initial operationalizations of concepts; (2) distributional methods from digital history introduce scalable vector representations of meaning; and (3) lexical semantic change detection supplies the diachronic modeling techniques most directly relevant to conceptual history. Topic modeling and citation-network approaches, while valuable in science studies, primarily capture thematic distributions or relational structures rather than the fine-grained semantic trajectories of individual concepts that are the focus of our survey. We will insert a short subsection (or expanded paragraph) at the end of the introduction that articulates this selection criterion and notes the boundaries of the reconstruction, thereby supporting the claim that the identified challenges are representative for computational conceptual history. revision: yes

  2. Referee: [Second part, LLM case studies and revisit of methodological questions] The LLM section (second part): the assertion that the same three problem areas “play out” in LLM workflows is illustrated by case studies, yet the manuscript provides no systematic side-by-side comparison (e.g., how a specific operationalization choice in a pre-LLM distributional model maps onto an LLM prompt or fine-tuning decision) that would make the inheritance claim load-bearing rather than suggestive.

    Authors: We agree that the current revisit section relies on illustrative discussion rather than a systematic mapping, which leaves the inheritance claim more suggestive than demonstrated. While the case studies already show how corpus, modeling, and evaluation issues recur, we did not provide explicit correspondences (e.g., between static word embeddings and contextual LLM representations, or between manual seed-word selection and prompt design). We will add a concise comparative table in the second part that pairs representative pre-LLM choices with their LLM analogues and indicates where the same underlying problems persist or are transformed. This addition will make the continuities load-bearing without expanding the paper beyond its survey scope. revision: yes

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: survey paper with no derivations or fitted quantities

full rationale

The paper is a historical and methodological survey that reconstructs pre-LLM work by organizing existing literature into three strands and then compares LLM applications against the same problem areas (corpus construction, operationalization, evaluation). No equations, parameters, predictions, or derivations are present. The selection of strands is an explicit framing choice in the abstract and introduction, not a self-referential reduction or self-citation chain that forces the central claim. The analysis therefore contains no load-bearing steps that reduce to inputs by construction.

Assumptions & free parameters 0 free parameters · 0 assumptions · 0 invented entities

This is a survey paper with no new mathematical derivations, fitted parameters, or postulated entities. No free parameters, axioms, or invented entities are introduced.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Computational conceptual history of scientific concepts: From early digital methods to LLMs." pith.science (2026). https://pith.science/paper/NBNVGG2O

@misc{pith2026260604118,
  author       = {Pith},
  title        = {Pith review of: Computational conceptual history of scientific concepts: From early digital methods to LLMs},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/NBNVGG2O}},
  note         = {Machine review of arXiv:2606.04118}
}
read the original abstract

This article situates large language models (LLMs) within the longer history of computational approaches to concept analysis in the history, philosophy, and sociology of science (HPSS). We examine what LLMs add to existing methods, how they inherit longstanding problems, and review recent case studies that employ them. In the first part, we reconstruct computational conceptual history before LLMs by bringing together three strands of work: early digital methods in HPSS, distributional approaches from digital history and related research, and lexical semantic change detection. We provide an overview of the main challenges and opportunities, focusing on corpus construction, operationalization and modelling choices, and evaluation and interpretation. In the second part, we turn to the era of LLMs, starting with a short introduction to LLMs before reviewing LLM-based work on lexical semantic change detection and relevant case studies in HPSS. We then revisit the earlier methodological questions, showing how issues of corpus construction, model choice and training data, operationalization trade-offs, and evaluation and interpretation play out in LLM-based workflows.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

13 extracted references · 7 canonical work pages

  1. [1]

    (2026) Discursive parallels of the chem- ical revolution

    Aguilar-Valdez S, Phan-Tất B, Speelman D, et al. (2026) Discursive parallels of the chem- ical revolution. Topic modelling and distributional analysis. In: Simons A, Wüthrich A, Zichert M, et al. (eds) Understanding Science with Large Language Models? Potentials for the History, Philosophy, and Sociology of Science. Bielefeld: transcript, part-3. Ahmadi E...

  2. [2]

    Journal of machine Learn- ing research 3(Jan): 993–1022

    Blei DM, Ng AY and Jordan MI (2003) Latent Dirichlet allocation. Journal of machine Learn- ing research 3(Jan): 993–1022. Blei DM and Lafferty JD (2006) Dynamic topic models. In: Proceedings of the 23rd interna- tional conference on Machine learning, New York, NY, USA, 25 June 2006, pp. 113–120. ICML ’06. Association for Computing Machinery. Available at:...

  3. [3]

    Transactions of the Association for Computational Linguistics 13: 690–708

    Cassotti P and Tahmasebi N (2025a) Sense-specific Historical Word Usage Generation. Transactions of the Association for Computational Linguistics 13: 690–708. Available at: https://doi.org/10.1162/tacl_a_00761. Cassotti P and Tahmasebi N (2025) A Hypothesis-Driven Framework for Detecting Lexi- cal Semantic Change. In: Proceedings of the Eleventh Italian C...

  4. [4]

    CEUR Workshop Proceedings. CEUR. Available at: https://ceur-ws .org/Vol-4112/#18_main_long Chang J, Gerrish S, Wang C, et al. (2009) Reading Tea Leaves: How Humans Interpret Topic Models. In: Advances in Neural Information Processing Systems,

  5. [5]

    BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

    Curran As- sociates. Callon M, Courtial J-P , Turner WA, et al. (1983) From translations to problematic net- works: An introduction to co-word analysis. Social Science Information 22(2). SAGE Publications Ltd: 191–235. Callon M, Law J and Rip A (1986) Qualitative Scientometrics. In: Callon M, Law J, and Rip A (eds) Mapping the Dynamics of Science and Tech...

  6. [6]

    Information Processing & Management 37(6): 817–842

    Ding Y, Chowdhury GG and Foo S (2001) Bibliometric cartography of information re- trieval research by using co-word analysis. Information Processing & Management 37(6): 817–842. Dubossarsky H, Weinshall D and Grossman E (2017) Outta Control: Laws of Semantic Change and Inherent Biases in Word Representation Models. In: Proceedings of the 2017 Conference o...

  7. [7]

    Retrieval-Augmented Generation for Large Language Models: A Survey

    Gao Y, Xiong Y, Gao X, et al. (2024) Retrieval-Augmented Generation for Large Language Models: A Survey. arXiv:2312.10997. arXiv. Available at: http://arxiv.org/abs/2312.109

  8. [8]

    (2019) Spaces of Meaning: Conceptual History, Vec- tor Semantics, and Close Reading

    Gavin M, Jennings C, Kersey L, et al. (2019) Spaces of Meaning: Conceptual History, Vec- tor Semantics, and Close Reading. In: Gold MK and Klein LF (eds) Debates in the Digital Humanities

Show all 13 references
  1. [9]

    Minneapolis: University of Minnesota Press, pp. 243–267. Available at: https://www.jstor.org/stable/10.5749/j.ctvg251hk.24. Griffiths TL and Steyvers M (2004) Finding scientific topics. Proceedings of the National Academy of Sciences

  2. [10]

    Gulordava K and Baroni M (2011) A distributional similarity approach to the detection of semantic change in the Google Books Ngram corpus

    Proceedings of the National Academy of Sciences: 5228–5235. Gulordava K and Baroni M (2011) A distributional similarity approach to the detection of semantic change in the Google Books Ngram corpus. In: Proceedings of the GEMS 2011 Workshop on GEometrical Models of Natural Lan...

  3. [11]

    Müller E and Schmieder F (2018) Begriffsgeschichte und Wissenschaftsgeschichte: Bestandsaufnahme und Forschungsperspektiven

    Suhrkamp. Müller E and Schmieder F (2018) Begriffsgeschichte und Wissenschaftsgeschichte: Bestandsaufnahme und Forschungsperspektiven. Geschichte und Gesellschaft 44(1): 79–106. Pennington J, Socher R and Manning C (2014) GloVe: Global Vectors for Word Representa- tion. In: Pr...

  4. [12]

    Writing science

    Rheinberger H-J (1997) Toward a History of Epistemic Things: Synthesizing Proteins in the Test Tube. Writing science. Stanford, Calif.: Stanford Univ. Press. Rip A and Courtial JP (1984) Co-word maps of biotechnology: An example of cognitive scientometrics. Scientometrics 6(6)...

  5. [13]

    Historical Methods: A Journal of Quantitative and Interdisciplinary History 53(4)

    Wevers M and Koolen M (2020) Digital begriffsgeschichte: Tracing semantic change us- ing word embeddings. Historical Methods: A Journal of Quantitative and Interdisciplinary History 53(4). Routledge: 226–243. Wu J, Gan W , Chen Z, et al. (2023) Multimodal Large Language Models...

Pith tools

Reviewed June 28, 2026 · model on record in the stance chip above.