REVIEW 6 minor 7 references
Large Language Models for History, Philosophy, and Sociology of Science: Interpretive Uses, Methodological Challenges, and Critical Perspectives
T0 review · 0 major / 6 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read LLMs may mark an inflection point in computational and interpretive studies of science.
desk verdict A useful, carefully hedged survey that is more optimistic about LLM-based workflows than its own caveats support; deserves peer review. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The carrying mechanism is the contextualized word embedding (CWE), produced by transformer architectures through self-attention. In this setup each token becomes a vector whose values depend on the surrounding text, so meaning is learned as position in an embedding space where semantic similarity is spatial proximity. The paper contrasts BERT-style full-context models, whose bidirectional attention makes CWEs well suited to classification, retrieval, and word sense modeling, with GPT-style generative models that predict tokens left to right and are better suited to fluent generation and in-context learning. This machinery carries the argument because all three HPSS challenges—data structuring, pattern detection, and dynamic modeling—are reframed as operations on these contextualized representations.
What would settle it
Take a set of historical passages with expert-validated sense shifts (e.g., 19th-century uses of "energy" or "objectivity") and check whether a state-of-the-art LLM's nearest-neighbor embeddings or prompted paraphrases align with the period senses or with modern senses. If modern senses systematically win, the claim that LLMs can model historically situated meaning fails for the core diachronic use case.
Extended reading notes
Core claim
The central claim is that LLMs may mark an inflection point in the long-standing tension between computational and interpretive research on science. The core discovery is that contextualized word embeddings operationalize the distributional hypothesis of meaning—meaning as position in a high-dimensional semantic space—so that semantic similarity, conceptual variation, and diachronic shift can be measured at scale while still remaining tied to context. On this basis, the paper claims that HPSS researchers can trace conceptual change (e.g., shifts in terms like "energy", "evolution", or "objectivity"), map argumentative structures, and study boundary-work and scientific authority with methods that scale interpretive depth rather than substituting for it. Crucially, the paper frames LLMs as epistemic infrastructures that encode assumptions about meaning and relevance, so the same reflexive skills HPSS applies to science must be applied to the models themselves.
Load-bearing premise
The load-bearing premise is that LLMs trained largely on contemporary text can faithfully represent the historically situated, context-dependent meaning of older scientific texts; if that fails, the proposed workflows for conceptual history and discourse analysis lose their ground.
Editorial extensions
If this is right
- Historians of science could trace shifts in terms like "energy", "evolution", and "objectivity" across large corpora, making conceptual history scalable.
- Philosophers of science could map conceptual structures, analyze argumentative patterns, and simulate rival positions to test their coherence.
- Sociologists of science could study how scientific authority, credibility, and dissent are constructed in journals, public discourse, and policy texts.
- Hybrid workflows pairing a full-context model for pattern detection with a generative model for interpretation become the recommended path, since no single model type serves all HPSS tasks.
- HPSS must define its own benchmarks, corpora, and annotation strategies because standard NLP evaluation metrics are not designed for historical ambiguity and shifting vocabularies.
Reading between the lines
- An untested implication is that continued pretraining on period-specific corpora will matter more for HPSS than architecture choice; the cited studies of "Planck" and "virtual" suggest domain adaptation improves sense distinctions, but the survey does not demonstrate this across fields.
- If the temporal mismatch between training data and historical sources proves intractable, LLM-assisted HPSS may systematically favor research questions aligned with modern text distributions, pushing archives toward what models handle well rather than what history needs.
- The infrastructure critique could be turned into a study design: HPSS could use its own conceptual-history methods to trace how "meaning", "context", and "similarity" are being redefined inside LLM training pipelines and benchmark communities.
- A testable extension would be a shared benchmark of diachronic concept shifts across multiple languages (for example, early quantum mechanics in German and English), which the paper identifies as an evaluation gap but does not build.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper is a methodological and conceptual survey of how large language models (LLMs) can support interpretive research in the history, philosophy, and sociology of science (HPSS). It offers a non-technical primer on transformer architectures, then maps LLM-based methods onto three challenges identified by Laubichler et al. (2019): structuring data, detecting patterns, and modeling dynamics over time. The central claim is that LLMs may mark an inflection point by offering improved semantic modeling, accessibility, and a bridge between close and distant reading, while also embedding assumptions that require critical scrutiny. The paper closes with four lessons: model selection is a trade-off, LLM literacy is foundational, HPSS must build its own benchmarks and corpora, and LLMs should enhance rather than replace interpretive methods. The manuscript does not present new empirical results; its contribution is a synthesis, a conceptual frame, and practical guidance.
Significance. If accepted as a programmatic contribution, this survey could help orient HPSS researchers entering the LLM space and stimulate discipline-specific evaluation standards. Its main strengths are the balanced cataloguing of affordances and risks, the explicit recognition that embeddings are not neutral representations of meaning, and the consistent call for qualitative validation and critical infrastructure awareness. The paper gives appropriate weight to known failure modes, such as the Kleymann et al. (2022) null result and the contextual-noise caveat of Kutuzov et al. (2022). Its main limitation is that the positive case studies for token-level dynamics (Simons 2024; Zichert et al. 2025) are author self-citations without independent gold standards, and the paper sometimes moves from hedged possibility to affirmative 'can' in summarizing these results. Overall, the central claim is defensible because it is framed as a conditional promise with explicit caveats; the survey does not overclaim certainty about the historical alignment of embeddings.
minor comments (6)
- [Section 4.2] Two in-text citations read 'Ji et al., 2000', but the reference list gives Ji, Wei, and Xu (2020); correct the year in both places.
- [Section 5.2] The in-text citation 'Gorour et al. (2024)' in Section 5.2 is spelled 'Gorur et al.' in the reference list; unify the spelling.
- [References] In the reference for Song et al. (2023), 'sematic analysis' should be 'semantic analysis'.
- [Section 5.1] The sentence 'These studies show how LLMs can support historically grounded concept tracing' overstates the evidence, because Simons (2024) and Zichert et al. (2025) rely on the authors' own interpretive judgment and lack external replication; consider changing 'show' to 'suggest' or adding a caveat about the absence of independent validation.
- [Section 6.1] The two workflow examples would be strengthened by an explicit sentence stipulating a validation step (for example, comparing embedding-based sense clusters against expert-annotated target concepts) before the pipeline is adopted, consistent with the temporal-mismatch caveats in Sections 3.1 and 3.3.
- [Section 1] The paper says it is 'structured in five parts', but the body contains more than five sections; rephrase to 'five parts beyond this introduction' or explicitly list the parts.
Circularity Check
No significant circularity: the paper is an interpretive survey whose central claim is argued from external evidence and explicitly acknowledged limitations, not derived from its own inputs.
full rationale
This paper is a survey and conceptual argument, not a derivation; there is no equation, fitted parameter, or uniqueness theorem whose output is fed back as input. The central claim that LLMs 'may mark an inflection point' (Section 1) is supported by external NLP results (e.g., BERT, SciBERT, BERTopic), by non-author applications (Kleymann et al. 2022; Lucy et al. 2023), and by the authors' own prior case studies (Simons 2024; Zichert et al. 2025). Those self-citations are illustrative exemplars of a proposed workflow, not load-bearing premises that force the conclusion; the paper itself flags the key vulnerability in Section 3.1 ('LLMs trained mainly on contemporary data may flatten or misrepresent historically specific language') and Section 5.1 (embedding shifts 'can reflect contextual noise rather than true semantic drift', citing Kutuzov et al. 2022). The skeptical concern that the temporal-alignment assumption is unvalidated is a correctness/evidentiary risk, not a circularity: the paper does not define or fit its conclusion in terms of that assumption. No circular step can be exhibited.
Assumptions & free parameters
assumptions (3)
- domain assumption The distributional hypothesis of meaning (Harris 1954) is a sufficient basis for computational analysis of historically situated meaning.
- domain assumption Domain-adapted models trained on historical corpora can adequately mitigate the temporal mismatch between training data and historical sources.
- domain assumption The cited applications in HPSS are representative and successful examples of LLM-based interpretive research.
Cite this review
Pith. "Pith review of Large Language Models for History, Philosophy, and Sociology of Science: Interpretive Uses, Methodological Challenges, and Critical Perspectives." pith.science (2026). https://pith.science/paper/ZE76FABF
@misc{pith2026250612242,
author = {Pith},
title = {Pith review of: Large Language Models for History, Philosophy, and Sociology of Science: Interpretive Uses, Methodological Challenges, and Critical Perspectives},
year = {2026},
howpublished = {\url{https://pith.science/paper/ZE76FABF}},
note = {Machine review of arXiv:2506.12242}
}
read the original abstract
This paper explores the use of large language models (LLMs) as research tools in the history, philosophy, and sociology of science (HPSS). LLMs are remarkably effective at processing unstructured text and inferring meaning from context, offering new affordances that challenge long-standing divides between computational and interpretive methods. This raises both opportunities and challenges for HPSS, which emphasizes interpretive methodologies and understands meaning as context-dependent, ambiguous, and historically situated. We argue that HPSS is uniquely positioned not only to benefit from LLMs' capabilities but also to interrogate their epistemic assumptions and infrastructural implications. To this end, we first offer a concise primer on LLM architectures and training paradigms tailored to non-technical readers. We frame LLMs not as neutral tools but as epistemic infrastructures that encode assumptions about meaning, context, and similarity, conditioned by their training data, architecture, and patterns of use. We then examine how computational techniques enhanced by LLMs, such as structuring data, detecting patterns, and modeling dynamic processes, can be applied to support interpretive research in HPSS. Our analysis compares full-context and generative models, outlines strategies for domain and task adaptation (e.g., continued pretraining, fine-tuning, and retrieval-augmented generation), and evaluates their respective strengths and limitations for interpretive inquiry in HPSS. We conclude with four lessons for integrating LLMs into HPSS: (1) model selection involves interpretive trade-offs; (2) LLM literacy is foundational; (3) HPSS must define its own benchmarks and corpora; and (4) LLMs should enhance, not replace, interpretive methods.
Reference graph
Works this paper leans on
-
[1]
Amrami, A., & Goldberg, Y. (2019). Towards better substitution-based word sense induction. arXiv Preprint arXiv:1905.12598 . Arnaout, H., Sternlicht, N., Hope, T., & Gurevych, I. (2025). In-depth Research Impact Summarization through Fine-Grained Temporal Citation Analysis . arXiv:2505.14838. Beltagy, I., Lo, K., & Cohan, A. (2019). SciBERT: A Pretrained ...
arXiv 2019
-
[4]
Larooij, M., & Törnberg, P. (2025). Do Large Language Models Solve the Problems of Agent-Based Modeling? A Critical Review of Generative Social Simulations . arXiv:2504.03274. Latour, B. (1987). Science in action: How to follow scientists and engineers through society . Harvard University Press. Laubichler, M., Maienschein, J., & Renn, J. (2019). Computat...
arXiv 2025
-
[58]
Zhu, Y., Yuan, H., Wang, S., et al. (2024b). Large Language Models for Information Retrieval: A Survey . arXiv:2308.07107. Zichert, M., Simons, A., & Wüthrich, A. (2025). Expanding Conceptual Histories: Using Contextualized Word Embeddings for the History and Philosophy of the Virtual Particle Concept. Computational Humanities Research . Ziems, C., Held, ...
arXiv 2024
-
[269]
Jiang, C., Xu, W., & Stevens, S. (2022). Edits: Understanding the Human Revision Process in Scientific Writing. In Y. Goldberg, Z. Kozareva, & Y. Zhang (Eds.), Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing (pp. 9420–9435). Association for Computational Linguistics. Jiang, X., & Chen, J. (2023). Contextualised segme...
work page 2022
-
[275]
Wang, W., Downey, J., & Yang, F. (2023a). AI anxiety? Comparing the sociotechnical imaginaries of artificial intelligence in UK, Chinese and Indian newspapers. Global Media and China . Wang, Z., Chen, J., Chen, J., & Chen, H. (2023b). Identifying interdisciplinary topics and their evolution based on BERTopic. Scientometrics , 1–26. Wang, X., Chen, G., Qia...
work page Pith review arXiv 2023
-
[1418]
Davidson, T., & Karell, D. (2025). Integrating Generative Artificial Intelligence into Social Science Research: Measurement, Prompting, and Simulation. Sociological Methods & Research . de Winter, J. (2024). Can ChatGPT be used to predict citation counts, readership, and social media interaction? An exploration among 2222 scientific abstracts. Scientometr...
arXiv 2025
-
[2024]
(pp. 144–158). Association for Computational Linguistics. Lewis, P., Perez, E., Piktus, A., et al. (2020). Retrieval-augmented generation for knowledge-intensive nlp tasks. Advances in Neural Information Processing Systems , 33 , 9459–9474. 24 Large Language Models for History, Philosophy, and Sociology of Science Leydesdorff, L., Ràfols, I., & Milojević,...
arXiv 2020
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.