{"id":"b89ce323-7d5b-41d8-8568-d73b01f11ce6","arxiv_id":"2506.12242","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":3.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A position paper arguing that LLMs can enhance interpretive research in history, philosophy, and sociology of science, but only with model literacy, domain-specific benchmarks, and critical reflection.","lead":"This paper examines how large language models can be used as research tools in the history, philosophy, and sociology of science, arguing they can support interpretive analysis when used critically. It offers a primer on LLMs, reviews computational strategies, and concludes with four lessons for integrating such models into the field.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The survey's central promise that LLM embeddings can scale interpretive HPSS work rests on an unvalidated temporal-alignment assumption; the paper's own cited case study shows the failure mode, so the claim needs an explicit validation check before workflows are adopted.","rationale":"The reader's weakest-assumption diagnosis is essentially correct: temporal mismatch is a load-bearing premise. I partially agree because the same concern is already acknowledged in the paper; the sharper issue is that the paper's own cited evidence contains a documented failure mode (Kleymann et al. 2022) without a mitigation or validation protocol. That said, this is a survey, and its contribution is conceptual rather than empirical. The central claim is plausible, the limitations are discussed candidly, and the four lessons are sensible. The concern does not justify rejection; it justifies maintaining the CONDITIONAL verdict and asking for an explicit validation check before the proposed workflows are treated as ready-to-use. I credit the paper for its balanced treatment of model trade-offs, its recognition of epistemic opacity, and its insistence that LLMs supplement rather than replace interpretive judgment. No ad hominem or methodological fraud is involved; this is an ordinary evidential gap in a forward-looking survey.","tokens_in":33847,"tokens_out":2578,"duration_ms":32875,"concrete_test":"Run a diachronic concept-tracing study on a historical corpus (e.g., Physical Review 1893-2020 for 'virtual') with two models: a general BERT (BERT-base) and a domain-adapted BERT (as in Zichert et al. 2025). Have historians blind-annotate a stratified sample of 500 usages into conceptual senses (virtual particle, virtual work, virtual displacement, etc.). Compare time-sliced CWE clusters against these expert labels using adjusted Rand index and cluster purity, and compare detected semantic drift against a null model where time slices are shuffled. If the domain-adapted model does not significantly outperform the general model on expert agreement, or if drift signals do not exceed the shuffled baseline, Sections 5-6's embedding-based workflows are not supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The strongest claim (Section 1) is that LLMs 'may mark an inflection point' by scaling interpretive depth without losing nuance. That claim stands or falls on whether contextual embeddings trained on modern corpora can represent historically situated meaning. The paper acknowledges the risk (Section 3.1: LLMs trained mainly on contemporary data 'may flatten or misrepresent historically specific language') and even notes the need to account for temporal mismatch (Section 3.3), but Sections 5-6 then build recommended workflows directly on time-sliced CWE clusters, BERTopic, and semantic-similarity measures, as if the risk were manageable. The internal evidence is mixed in a telling way: Kleymann et al. (2022), cited approvingly in Section 4.1, found that CWE clustering of 'theory' largely captured syntactic variation rather than meaningful semantic distinctions; the positive counterexamples (Simons 2024; Zichert et al. 2025) are author self-citations with no independent gold standard. The paper also cites Kutuzov et al. (2022) for the fact that embedding shifts can reflect contextual noise rather than true semantic drift. Without a validation step that ties embedding-space shifts to expert-annotated conceptual senses, the 'inflection point' claim is an extrapolation, not an established result. This is a correctness risk in the paper's central argument, not a disagreement with consensus.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper is a methodological and conceptual survey of how large language models (LLMs) can support interpretive research in the history, philosophy, and sociology of science (HPSS). It offers a non-technical primer on transformer architectures, then maps LLM-based methods onto three challenges identified by Laubichler et al. (2019): structuring data, detecting patterns, and modeling dynamics over time. The central claim is that LLMs may mark an inflection point by offering improved semantic modeling, accessibility, and a bridge between close and distant reading, while also embedding assumptions that require critical scrutiny. The paper closes with four lessons: model selection is a trade-off, LLM literacy is foundational, HPSS must build its own benchmarks and corpora, and LLMs should enhance rather than replace interpretive methods. The manuscript does not present new empirical results; its contribution is a synthesis, a conceptual frame, and practical guidance.","tokens_in":34077,"tokens_out":6276,"duration_ms":78168,"significance":"If accepted as a programmatic contribution, this survey could help orient HPSS researchers entering the LLM space and stimulate discipline-specific evaluation standards. Its main strengths are the balanced cataloguing of affordances and risks, the explicit recognition that embeddings are not neutral representations of meaning, and the consistent call for qualitative validation and critical infrastructure awareness. The paper gives appropriate weight to known failure modes, such as the Kleymann et al. (2022) null result and the contextual-noise caveat of Kutuzov et al. (2022). Its main limitation is that the positive case studies for token-level dynamics (Simons 2024; Zichert et al. 2025) are author self-citations without independent gold standards, and the paper sometimes moves from hedged possibility to affirmative 'can' in summarizing these results. Overall, the central claim is defensible because it is framed as a conditional promise with explicit caveats; the survey does not overclaim certainty about the historical alignment of embeddings.","major_comments":[],"minor_comments":[{"comment":"Two in-text citations read 'Ji et al., 2000', but the reference list gives Ji, Wei, and Xu (2020); correct the year in both places.","section":"Section 4.2"},{"comment":"The in-text citation 'Gorour et al. (2024)' in Section 5.2 is spelled 'Gorur et al.' in the reference list; unify the spelling.","section":"Section 5.2"},{"comment":"In the reference for Song et al. (2023), 'sematic analysis' should be 'semantic analysis'.","section":"References"},{"comment":"The sentence 'These studies show how LLMs can support historically grounded concept tracing' overstates the evidence, because Simons (2024) and Zichert et al. (2025) rely on the authors' own interpretive judgment and lack external replication; consider changing 'show' to 'suggest' or adding a caveat about the absence of independent validation.","section":"Section 5.1"},{"comment":"The two workflow examples would be strengthened by an explicit sentence stipulating a validation step (for example, comparing embedding-based sense clusters against expert-annotated target concepts) before the pipeline is adopted, consistent with the temporal-mismatch caveats in Sections 3.1 and 3.3.","section":"Section 6.1"},{"comment":"The paper says it is 'structured in five parts', but the body contains more than five sections; rephrase to 'five parts beyond this introduction' or explicitly list the parts.","section":"Section 1"}],"recommendation":"minor_revision","confidential_remarks":"This is a competent and carefully hedged survey that should be useful to the HPSS community. One editorial concern is the disproportionate reliance on the authors' own preprints in Section 5.1; adding a neutral third-party example or a clearer disclaimer would strengthen the appearance of objectivity. The journal should also consider whether it wants a survey without new empirical content; if so, the didactic primer and comparison tables are valuable justifications for publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nWhat you should know: this is a competent, well-cautious survey of LLM methods for history, philosophy, and sociology of science, not a research paper with new empirical results. If you need one place to point a graduate student for the current methodological landscape, this works.\n\nThe paper does several things well. The primer on BERT-style vs generative models is accessible and accurate. The 'epistemic infrastructures' framing, while not new to STS, is used productively to organize the discussion. The four lessons at the end are sensible and appropriately hedged. The authors repeatedly flag the key risks: temporal mismatch between training data and historical sources, embeddings capturing noise rather than semantic change, and the need for qualitative validation. They even cite Kleymann et al.'s negative result, which suggests the 'theory' clusters were mostly syntactic, and take it seriously rather than explaining it away.\n\nThe soft spots are proportionate to the genre. The positive examples that underwrite the 'inflection point' claim are mostly the authors' own previous studies (Simons 2024; Zichert et al. 2025), which are not validated against an independent gold standard. The stress-test correctly identifies the temporal alignment assumption as load-bearing, but I think it slightly overstates the case: the paper does not present these workflows as established, and it repeatedly says interpretive reading is essential. Still, the overall tone in Sections 5 and 6 is more optimistic than the caveats justify, and a revised version should make expert validation a named first step in the proposed workflows rather than a cautionary note. That would strengthen the paper without changing its argument.\n\nThe reference list is broad and citation patterns look fair. No red flags. The paper is a bit long for its message, but that is a minor editorial issue.\n\nThis deserves a serious referee as a review/perspective piece. I would publish it after light revision, mainly to tighten the validation message. It will be useful for HPSS scholars entering this area.\n\nBest.","headline":"A useful, carefully hedged survey that is more optimistic about LLM-based workflows than its own caveats support; deserves peer review.","tokens_in":34618,"tokens_out":3591,"would_cite":true,"duration_ms":43454,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"LLMs may mark an inflection point in computational and interpretive studies of science.","keywords":["large language models","history and philosophy of science","sociology of science","contextualized word embeddings","computational methods","conceptual change","digital humanities","epistemic infrastructure"],"falsifier":"Take a set of historical passages with expert-validated sense shifts (e.g., 19th-century uses of \"energy\" or \"objectivity\") and check whether a state-of-the-art LLM's nearest-neighbor embeddings or prompted paraphrases align with the period senses or with modern senses. If modern senses systematically win, the claim that LLMs can model historically situated meaning fails for the core diachronic use case.","tokens_in":33620,"feed_emoji":"📜","tokens_out":5409,"duration_ms":58882,"temperature":0.7,"pith_summary":"The paper argues that large language models are not just bigger search or coding tools but a genuine methodological turn for the history, philosophy, and sociology of science (HPSS). Their contextualized representations of meaning, the authors claim, offer three advances over earlier computational methods: better semantic modeling, lower technical barriers, and a bridge between close reading of individual texts and distant reading of large corpora. Because HPSS treats meaning as historically situated and context-dependent, the authors see the field as both a prime beneficiary of LLMs and the right community to scrutinize what these models assume about meaning, similarity, and relevance. The paper's contribution is a methodological survey organized around three challenges—structuring data, detecting patterns, and explaining scientific change—plus a set of lessons for using LLMs without surrendering interpretive control.","feed_headline":"Language models can end the close-reading vs. scale divide","feed_subtitle":"Contextualized embeddings can scale interpretive depth without flattening historical meaning, the survey argues.","key_machinery":"The carrying mechanism is the contextualized word embedding (CWE), produced by transformer architectures through self-attention. In this setup each token becomes a vector whose values depend on the surrounding text, so meaning is learned as position in an embedding space where semantic similarity is spatial proximity. The paper contrasts BERT-style full-context models, whose bidirectional attention makes CWEs well suited to classification, retrieval, and word sense modeling, with GPT-style generative models that predict tokens left to right and are better suited to fluent generation and in-context learning. This machinery carries the argument because all three HPSS challenges—data structuring, pattern detection, and dynamic modeling—are reframed as operations on these contextualized representations.","core_discovery":"The central claim is that LLMs may mark an inflection point in the long-standing tension between computational and interpretive research on science. The core discovery is that contextualized word embeddings operationalize the distributional hypothesis of meaning—meaning as position in a high-dimensional semantic space—so that semantic similarity, conceptual variation, and diachronic shift can be measured at scale while still remaining tied to context. On this basis, the paper claims that HPSS researchers can trace conceptual change (e.g., shifts in terms like \"energy\", \"evolution\", or \"objectivity\"), map argumentative structures, and study boundary-work and scientific authority with methods that scale interpretive depth rather than substituting for it. Crucially, the paper frames LLMs as epistemic infrastructures that encode assumptions about meaning and relevance, so the same reflexive skills HPSS applies to science must be applied to the models themselves.","pith_inferences":["An untested implication is that continued pretraining on period-specific corpora will matter more for HPSS than architecture choice; the cited studies of \"Planck\" and \"virtual\" suggest domain adaptation improves sense distinctions, but the survey does not demonstrate this across fields.","If the temporal mismatch between training data and historical sources proves intractable, LLM-assisted HPSS may systematically favor research questions aligned with modern text distributions, pushing archives toward what models handle well rather than what history needs.","The infrastructure critique could be turned into a study design: HPSS could use its own conceptual-history methods to trace how \"meaning\", \"context\", and \"similarity\" are being redefined inside LLM training pipelines and benchmark communities.","A testable extension would be a shared benchmark of diachronic concept shifts across multiple languages (for example, early quantum mechanics in German and English), which the paper identifies as an evaluation gap but does not build."],"forward_implications":["Historians of science could trace shifts in terms like \"energy\", \"evolution\", and \"objectivity\" across large corpora, making conceptual history scalable.","Philosophers of science could map conceptual structures, analyze argumentative patterns, and simulate rival positions to test their coherence.","Sociologists of science could study how scientific authority, credibility, and dissent are constructed in journals, public discourse, and policy texts.","Hybrid workflows pairing a full-context model for pattern detection with a generative model for interpretation become the recommended path, since no single model type serves all HPSS tasks.","HPSS must define its own benchmarks, corpora, and annotation strategies because standard NLP evaluation metrics are not designed for historical ambiguity and shifting vocabularies."],"supporting_citations":[{"why":"Supplies the three-challenge framework—structuring data, detecting patterns, explaining change—that organizes the paper's survey.","marker":"Laubichler et al. (2019)"},{"why":"Gives the distributional hypothesis of meaning on which the embedding-based methods rest.","marker":"Harris (1954)"},{"why":"Provides the transformer and self-attention architecture that produces contextualized word embeddings.","marker":"Vaswani et al. (2017)"},{"why":"Defines BERT, the representative full-context model used throughout the methodological comparisons.","marker":"Devlin et al. (2018)"},{"why":"Defines GPT, the representative generative model used for in-context learning and generation tasks.","marker":"Radford et al. (2018)"},{"why":"Presents SciBERT, the key evidence that domain-specific pretraining improves performance on scientific text tasks.","marker":"Beltagy et al. (2019)"},{"why":"Introduces retrieval-augmented generation, the mechanism behind the \"chatting with papers\" workflows discussed.","marker":"Lewis et al. (2020)"},{"why":"Demonstrates CWE-based conceptual history on \"Planck\" in physics, showing domain-adapted models capture fine-grained sense distinctions.","marker":"Simons (2024)"},{"why":"Applies time-sliced embeddings to trace the evolution and polysemy of \"virtual\" in nearly a century of physics publications.","marker":"Zichert et al. (2025)"}],"fun_headline_variants":["LLMs can bridge interpretive and computational science studies","Scaling interpretive depth without flattening historical meaning","Language models as epistemic infrastructure for science studies","Close reading meets scale: LLMs for history and philosophy of science","Contextual embeddings track conceptual change at scale"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that LLMs trained largely on contemporary text can faithfully represent the historically situated, context-dependent meaning of older scientific texts; if that fails, the proposed workflows for conceptual history and discourse analysis lose their ground.","fun_headline_variants_meta":{"raw":{"variants":["LLMs can bridge interpretive and computational science studies","Scaling interpretive depth without flattening historical meaning","Language models as epistemic infrastructure for science studies","Close reading meets scale: LLMs for history and philosophy of science","Contextual embeddings track conceptual change at scale"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000242,"raw_usage":{"total_tokens":1560,"prompt_tokens":1012,"completion_tokens":548,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":628,"completion_tokens_details":{"reasoning_tokens":475}},"tokens_in":628,"tokens_out":548,"duration_ms":6626,"temperature":1.0,"reasoning_tokens":475,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T00:54:15.888938+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a set of historical passages with expert-validated sense shifts (e.g., 19th-century uses of \"energy\" or \"objectivity\") and check whether a state-of-the-art LLM's nearest-neighbor embeddings or prompted paraphrases align with the period senses or with modern senses. If modern senses systematically win, the claim that LLMs can model historically situated meaning fails for the core diachronic use case.","supporting_citations":[],"review_version":1}