REVIEW 2 major objections 1 minor 13 references
Computational conceptual history of scientific concepts: From early digital methods to LLMs
T0 review · 2 major / 1 minor · reviewed 2026-06-28 · grok-4.3
Pith's one-line read LLMs extend computational methods for tracking scientific concept change but repeat earlier problems with data selection, modeling choices, and result validation.
desk verdict A clear survey that links pre-LLM computational work in HPSS to LLM applications and flags the same old methodological issues, without adding new findings. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
Three strands of prior work—early digital methods in HPSS, distributional approaches from digital history, and lexical semantic change detection—used to frame how LLMs handle corpus construction, operationalization, and evaluation.
What would settle it
An LLM-based study in HPSS that produces stable, reproducible results on conceptual change without requiring manual corpus curation or expert interpretation of outputs would undermine the claim that the longstanding problems are inherited.
Extended reading notes
Core claim
The review reconstructs computational conceptual history before LLMs by bringing together early digital methods in HPSS, distributional approaches from digital history, and lexical semantic change detection. It then examines LLM-based work on lexical semantic change detection and relevant HPSS case studies, revisiting the methodological questions of corpus construction, model choice and training data, operationalization trade-offs, and evaluation and interpretation to show how these issues continue to shape LLM workflows.
Load-bearing premise
The three strands of earlier work adequately capture the main challenges and opportunities in computational conceptual history prior to LLMs.
Editorial extensions
If this is right
- LLMs allow processing of larger and more heterogeneous corpora than earlier distributional methods for detecting shifts in scientific terminology.
- Prompting and fine-tuning decisions in LLM workflows introduce operationalization trade-offs comparable to those in pre-LLM vector-space models.
- Evaluation of LLM outputs for historical concepts continues to depend on alignment with expert historical knowledge rather than purely quantitative metrics.
- Corpus construction choices remain decisive even when models can ingest raw text at scale.
Reading between the lines
- The same pattern of inherited methodological limits may appear when LLMs are applied to conceptual analysis outside HPSS, such as in legal or literary history.
- One testable extension would be to compare stability of LLM-derived concept representations against earlier methods across the same long historical text collections.
- The review implies that progress may lie more in designing transparent workflows that document data and modeling choices than in adopting ever-larger models.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper situates LLMs within the history of computational concept analysis in HPSS. Part 1 reconstructs pre-LLM work by synthesizing three strands—early digital methods in HPSS, distributional approaches from digital history, and lexical semantic change detection—while cataloguing challenges in corpus construction, operationalization/modelling choices, and evaluation/interpretation. Part 2 introduces LLMs, reviews their use in lexical semantic change detection and HPSS case studies, and revisits the same three problem areas to argue that LLMs add capabilities but inherit the longstanding issues.
Significance. If the three-strand reconstruction is representative, the paper supplies a timely synthesis that clarifies continuities between pre-LLM and LLM-based workflows in conceptual history. Explicit credit is due for the structured two-part organization and the attempt to link methodological questions across eras via recent case studies.
major comments (2)
- [Introduction / first part framing] Introduction and the opening of the first part: the central claim that LLMs inherit problems around corpus construction, operationalization, and evaluation rests on the reconstruction via exactly three strands. The manuscript does not supply an explicit rationale for why these strands (rather than, e.g., topic-modeling pipelines or citation-network concept mapping common in science studies) suffice to identify the full set of longstanding challenges; without such justification the inheritance diagnosis is under-supported.
- [Second part, LLM case studies and revisit of methodological questions] The LLM section (second part): the assertion that the same three problem areas “play out” in LLM workflows is illustrated by case studies, yet the manuscript provides no systematic side-by-side comparison (e.g., how a specific operationalization choice in a pre-LLM distributional model maps onto an LLM prompt or fine-tuning decision) that would make the inheritance claim load-bearing rather than suggestive.
minor comments (1)
- [Abstract] The abstract states the two-part structure clearly but does not foreground the paper’s distinctive contribution (the explicit mapping of inherited problems onto LLM practice); a single sentence to this effect would improve reader orientation.
Simulated Author's Rebuttal
We thank the referee for the constructive report and the recommendation for major revision. The two major comments identify opportunities to strengthen the framing of our three-strand reconstruction and the explicitness of continuities between pre-LLM and LLM workflows. We address each point below and commit to revisions that make the inheritance argument more robust while preserving the survey character of the paper.
read point-by-point responses
-
Referee: [Introduction / first part framing] Introduction and the opening of the first part: the central claim that LLMs inherit problems around corpus construction, operationalization, and evaluation rests on the reconstruction via exactly three strands. The manuscript does not supply an explicit rationale for why these strands (rather than, e.g., topic-modeling pipelines or citation-network concept mapping common in science studies) suffice to identify the full set of longstanding challenges; without such justification the inheritance diagnosis is under-supported.
Authors: We accept that an explicit rationale for the three strands is currently implicit rather than stated. The strands were selected because they constitute the primary text-based computational approaches that have been applied to the semantic content and historical evolution of scientific concepts within HPSS: (1) early digital methods supply the disciplinary context and initial operationalizations of concepts; (2) distributional methods from digital history introduce scalable vector representations of meaning; and (3) lexical semantic change detection supplies the diachronic modeling techniques most directly relevant to conceptual history. Topic modeling and citation-network approaches, while valuable in science studies, primarily capture thematic distributions or relational structures rather than the fine-grained semantic trajectories of individual concepts that are the focus of our survey. We will insert a short subsection (or expanded paragraph) at the end of the introduction that articulates this selection criterion and notes the boundaries of the reconstruction, thereby supporting the claim that the identified challenges are representative for computational conceptual history. revision: yes
-
Referee: [Second part, LLM case studies and revisit of methodological questions] The LLM section (second part): the assertion that the same three problem areas “play out” in LLM workflows is illustrated by case studies, yet the manuscript provides no systematic side-by-side comparison (e.g., how a specific operationalization choice in a pre-LLM distributional model maps onto an LLM prompt or fine-tuning decision) that would make the inheritance claim load-bearing rather than suggestive.
Authors: We agree that the current revisit section relies on illustrative discussion rather than a systematic mapping, which leaves the inheritance claim more suggestive than demonstrated. While the case studies already show how corpus, modeling, and evaluation issues recur, we did not provide explicit correspondences (e.g., between static word embeddings and contextual LLM representations, or between manual seed-word selection and prompt design). We will add a concise comparative table in the second part that pairs representative pre-LLM choices with their LLM analogues and indicates where the same underlying problems persist or are transformed. This addition will make the continuities load-bearing without expanding the paper beyond its survey scope. revision: yes
Circularity Check
No circularity: survey paper with no derivations or fitted quantities
full rationale
The paper is a historical and methodological survey that reconstructs pre-LLM work by organizing existing literature into three strands and then compares LLM applications against the same problem areas (corpus construction, operationalization, evaluation). No equations, parameters, predictions, or derivations are present. The selection of strands is an explicit framing choice in the abstract and introduction, not a self-referential reduction or self-citation chain that forces the central claim. The analysis therefore contains no load-bearing steps that reduce to inputs by construction.
Assumptions & free parameters
Cite this review
Pith. "Pith review of Computational conceptual history of scientific concepts: From early digital methods to LLMs." pith.science (2026). https://pith.science/paper/NBNVGG2O
@misc{pith2026260604118,
author = {Pith},
title = {Pith review of: Computational conceptual history of scientific concepts: From early digital methods to LLMs},
year = {2026},
howpublished = {\url{https://pith.science/paper/NBNVGG2O}},
note = {Machine review of arXiv:2606.04118}
}
read the original abstract
This article situates large language models (LLMs) within the longer history of computational approaches to concept analysis in the history, philosophy, and sociology of science (HPSS). We examine what LLMs add to existing methods, how they inherit longstanding problems, and review recent case studies that employ them. In the first part, we reconstruct computational conceptual history before LLMs by bringing together three strands of work: early digital methods in HPSS, distributional approaches from digital history and related research, and lexical semantic change detection. We provide an overview of the main challenges and opportunities, focusing on corpus construction, operationalization and modelling choices, and evaluation and interpretation. In the second part, we turn to the era of LLMs, starting with a short introduction to LLMs before reviewing LLM-based work on lexical semantic change detection and relevant case studies in HPSS. We then revisit the earlier methodological questions, showing how issues of corpus construction, model choice and training data, operationalization trade-offs, and evaluation and interpretation play out in LLM-based workflows.
Reference graph
Works this paper leans on
-
[1]
(2026) Discursive parallels of the chem- ical revolution
Aguilar-Valdez S, Phan-Tất B, Speelman D, et al. (2026) Discursive parallels of the chem- ical revolution. Topic modelling and distributional analysis. In: Simons A, Wüthrich A, Zichert M, et al. (eds) Understanding Science with Large Language Models? Potentials for the History, Philosophy, and Sociology of Science. Bielefeld: transcript, part-3. Ahmadi E...
2026
-
[2]
Journal of machine Learn- ing research 3(Jan): 993–1022
Blei DM, Ng AY and Jordan MI (2003) Latent Dirichlet allocation. Journal of machine Learn- ing research 3(Jan): 993–1022. Blei DM and Lafferty JD (2006) Dynamic topic models. In: Proceedings of the 23rd interna- tional conference on Machine learning, New York, NY, USA, 25 June 2006, pp. 113–120. ICML ’06. Association for Computing Machinery. Available at:...
-
[3]
Transactions of the Association for Computational Linguistics 13: 690–708
Cassotti P and Tahmasebi N (2025a) Sense-specific Historical Word Usage Generation. Transactions of the Association for Computational Linguistics 13: 690–708. Available at: https://doi.org/10.1162/tacl_a_00761. Cassotti P and Tahmasebi N (2025) A Hypothesis-Driven Framework for Detecting Lexi- cal Semantic Change. In: Proceedings of the Eleventh Italian C...
-
[4]
CEUR Workshop Proceedings. CEUR. Available at: https://ceur-ws .org/Vol-4112/#18_main_long Chang J, Gerrish S, Wang C, et al. (2009) Reading Tea Leaves: How Humans Interpret Topic Models. In: Advances in Neural Information Processing Systems,
2009
-
[5]
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Curran As- sociates. Callon M, Courtial J-P , Turner WA, et al. (1983) From translations to problematic net- works: An introduction to co-word analysis. Social Science Information 22(2). SAGE Publications Ltd: 191–235. Callon M, Law J and Rip A (1986) Qualitative Scientometrics. In: Callon M, Law J, and Rip A (eds) Mapping the Dynamics of Science and Tech...
work page Pith review arXiv 1983
-
[6]
Information Processing & Management 37(6): 817–842
Ding Y, Chowdhury GG and Foo S (2001) Bibliometric cartography of information re- trieval research by using co-word analysis. Information Processing & Management 37(6): 817–842. Dubossarsky H, Weinshall D and Grossman E (2017) Outta Control: Laws of Semantic Change and Inherent Biases in Word Representation Models. In: Proceedings of the 2017 Conference o...
2001
-
[7]
Retrieval-Augmented Generation for Large Language Models: A Survey
Gao Y, Xiong Y, Gao X, et al. (2024) Retrieval-Augmented Generation for Large Language Models: A Survey. arXiv:2312.10997. arXiv. Available at: http://arxiv.org/abs/2312.109
work page Pith review arXiv 2024
-
[8]
(2019) Spaces of Meaning: Conceptual History, Vec- tor Semantics, and Close Reading
Gavin M, Jennings C, Kersey L, et al. (2019) Spaces of Meaning: Conceptual History, Vec- tor Semantics, and Close Reading. In: Gold MK and Klein LF (eds) Debates in the Digital Humanities
2019
Show all 13 references
-
[9]
Minneapolis: University of Minnesota Press, pp. 243–267. Available at: https://www.jstor.org/stable/10.5749/j.ctvg251hk.24. Griffiths TL and Steyvers M (2004) Finding scientific topics. Proceedings of the National Academy of Sciences
2004 doi
-
[10]
Gulordava K and Baroni M (2011) A distributional similarity approach to the detection of semantic change in the Google Books Ngram corpus
Proceedings of the National Academy of Sciences: 5228–5235. Gulordava K and Baroni M (2011) A distributional similarity approach to the detection of semantic change in the Google Books Ngram corpus. In: Proceedings of the GEMS 2011 Workshop on GEometrical Models of Natural Lan...
2011 arXiv
-
[11]
Müller E and Schmieder F (2018) Begriffsgeschichte und Wissenschaftsgeschichte: Bestandsaufnahme und Forschungsperspektiven
Suhrkamp. Müller E and Schmieder F (2018) Begriffsgeschichte und Wissenschaftsgeschichte: Bestandsaufnahme und Forschungsperspektiven. Geschichte und Gesellschaft 44(1): 79–106. Pennington J, Socher R and Manning C (2014) GloVe: Global Vectors for Word Representa- tion. In: Pr...
2018
-
[12]
Writing science
Rheinberger H-J (1997) Toward a History of Epistemic Things: Synthesizing Proteins in the Test Tube. Writing science. Stanford, Calif.: Stanford Univ. Press. Rip A and Courtial JP (1984) Co-word maps of biotechnology: An example of cognitive scientometrics. Scientometrics 6(6)...
1997 doi
-
[13]
Historical Methods: A Journal of Quantitative and Interdisciplinary History 53(4)
Wevers M and Koolen M (2020) Digital begriffsgeschichte: Tracing semantic change us- ing word embeddings. Historical Methods: A Journal of Quantitative and Interdisciplinary History 53(4). Routledge: 226–243. Wu J, Gan W , Chen Z, et al. (2023) Multimodal Large Language Models...
2020
Reviewed June 28, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.