{"id":"ff34c4bb-6595-4aa9-971d-793e2ab4ff86","arxiv_id":"2606.04118","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":2.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"A survey paper that reconstructs computational conceptual history in HPSS before and after LLMs, revisiting issues of corpus construction, modeling, and interpretation.","lead":"This paper reviews the evolution of computational methods for tracking how scientific concepts change, from early digital tools through statistical word models to large language models. It highlights what LLMs add and the persistent challenges they share with older approaches in history and philosophy of science.","discovery_kind":"review","skeptic_critique":{"model":"grok-4.3","headline":"Adequacy of the three pre-LLM strands for identifying inherited problems","rationale":"The reader's weakest assumption directly identifies the same framing dependency that makes the inheritance claim vulnerable. Because the paper offers no new quantitative results or derivations, this is the primary point where the argument could fail to generalize; the concern is therefore internal to the paper's own structure rather than an external consensus issue.","tokens_in":1681,"tokens_out":347,"duration_ms":16967,"concrete_test":"In the full manuscript section that reconstructs the three strands, extract every cited work and method mentioned; cross-check against a standard HPSS computational bibliography (e.g., via keyword search for 'topic model' or 'scientometrics' in the same period); if ≥3 additional distinct approaches addressing corpus/operationalization issues are absent and no rationale is given for their exclusion, the strand selection is incomplete.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that LLMs add capabilities while inheriting problems around corpus construction, operationalization, and evaluation—rests on the reconstruction of pre-LLM work via exactly three strands (early digital methods in HPSS, distributional approaches from digital history, and lexical semantic change detection). This selection must capture the main longstanding challenges for the inheritance diagnosis to hold; if other established lines (e.g., topic modeling or citation-network concept mapping in science studies) are omitted without explicit justification, the problems attributed to LLMs may be incomplete or unrepresentative. The abstract frames the reconstruction as bringing these strands together to focus on the three problem areas, but the load-bearing step is whether that framing is exhaustive enough to ground the later LLM comparison.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper situates LLMs within the history of computational concept analysis in HPSS. Part 1 reconstructs pre-LLM work by synthesizing three strands—early digital methods in HPSS, distributional approaches from digital history, and lexical semantic change detection—while cataloguing challenges in corpus construction, operationalization/modelling choices, and evaluation/interpretation. Part 2 introduces LLMs, reviews their use in lexical semantic change detection and HPSS case studies, and revisits the same three problem areas to argue that LLMs add capabilities but inherit the longstanding issues.","tokens_in":1821,"tokens_out":478,"duration_ms":16599,"significance":"If the three-strand reconstruction is representative, the paper supplies a timely synthesis that clarifies continuities between pre-LLM and LLM-based workflows in conceptual history. Explicit credit is due for the structured two-part organization and the attempt to link methodological questions across eras via recent case studies.","major_comments":[{"comment":"Introduction and the opening of the first part: the central claim that LLMs inherit problems around corpus construction, operationalization, and evaluation rests on the reconstruction via exactly three strands. The manuscript does not supply an explicit rationale for why these strands (rather than, e.g., topic-modeling pipelines or citation-network concept mapping common in science studies) suffice to identify the full set of longstanding challenges; without such justification the inheritance diagnosis is under-supported.","section":"Introduction / first part framing"},{"comment":"The LLM section (second part): the assertion that the same three problem areas “play out” in LLM workflows is illustrated by case studies, yet the manuscript provides no systematic side-by-side comparison (e.g., how a specific operationalization choice in a pre-LLM distributional model maps onto an LLM prompt or fine-tuning decision) that would make the inheritance claim load-bearing rather than suggestive.","section":"Second part, LLM case studies and revisit of methodological questions"}],"minor_comments":[{"comment":"The abstract states the two-part structure clearly but does not foreground the paper’s distinctive contribution (the explicit mapping of inherited problems onto LLM practice); a single sentence to this effect would improve reader orientation.","section":"Abstract"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the constructive report and the recommendation for major revision. The two major comments identify opportunities to strengthen the framing of our three-strand reconstruction and the explicitness of continuities between pre-LLM and LLM workflows. We address each point below and commit to revisions that make the inheritance argument more robust while preserving the survey character of the paper.","responses":[{"response":"We accept that an explicit rationale for the three strands is currently implicit rather than stated. The strands were selected because they constitute the primary text-based computational approaches that have been applied to the semantic content and historical evolution of scientific concepts within HPSS: (1) early digital methods supply the disciplinary context and initial operationalizations of concepts; (2) distributional methods from digital history introduce scalable vector representations of meaning; and (3) lexical semantic change detection supplies the diachronic modeling techniques most directly relevant to conceptual history. Topic modeling and citation-network approaches, while valuable in science studies, primarily capture thematic distributions or relational structures rather than the fine-grained semantic trajectories of individual concepts that are the focus of our survey. We will insert a short subsection (or expanded paragraph) at the end of the introduction that articulates this selection criterion and notes the boundaries of the reconstruction, thereby supporting the claim that the identified challenges are representative for computational conceptual history.","revision_made":"yes","referee_comment":"[Introduction / first part framing] Introduction and the opening of the first part: the central claim that LLMs inherit problems around corpus construction, operationalization, and evaluation rests on the reconstruction via exactly three strands. The manuscript does not supply an explicit rationale for why these strands (rather than, e.g., topic-modeling pipelines or citation-network concept mapping common in science studies) suffice to identify the full set of longstanding challenges; without such justification the inheritance diagnosis is under-supported."},{"response":"We agree that the current revisit section relies on illustrative discussion rather than a systematic mapping, which leaves the inheritance claim more suggestive than demonstrated. While the case studies already show how corpus, modeling, and evaluation issues recur, we did not provide explicit correspondences (e.g., between static word embeddings and contextual LLM representations, or between manual seed-word selection and prompt design). We will add a concise comparative table in the second part that pairs representative pre-LLM choices with their LLM analogues and indicates where the same underlying problems persist or are transformed. This addition will make the continuities load-bearing without expanding the paper beyond its survey scope.","revision_made":"yes","referee_comment":"[Second part, LLM case studies and revisit of methodological questions] The LLM section (second part): the assertion that the same three problem areas “play out” in LLM workflows is illustrated by case studies, yet the manuscript provides no systematic side-by-side comparison (e.g., how a specific operationalization choice in a pre-LLM distributional model maps onto an LLM prompt or fine-tuning decision) that would make the inheritance claim load-bearing rather than suggestive."}],"tokens_in":1371,"tokens_out":634,"duration_ms":17602,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The paper organizes existing strands of computational concept analysis into a two-part history: it first reconstructs three lines of pre-LLM work (early digital HPSS methods, distributional approaches from digital history, and lexical semantic change detection) and then checks how LLMs handle the same questions of corpus building, operationalization, and evaluation. That framing is the main contribution. It gives a reader a compact way to see continuity rather than rupture when moving from older tools to LLMs, and the case-study reviews in the second half make the inheritance claim concrete.\n\nThe structure is straightforward and the focus on those three problem areas is consistent. The paper does not claim new techniques or data, so its value sits in synthesis and in showing that familiar difficulties reappear in LLM workflows.\n\nThe choice of exactly those three strands is the soft spot. Topic modeling and citation-network methods have long been used in science studies for concept mapping, yet they receive little attention here. If the paper does not justify why they fall outside the scope, the list of inherited problems risks looking incomplete. The review also stays at the level of summarizing case studies rather than testing any of the claims against fresh examples, which keeps the argument at a high level.\n\nThis is useful background for researchers already working in digital humanities or computational HPSS who want a quick map of the terrain. It is not the sort of paper one would cite for a new result or method. A methods-oriented journal or a science-studies venue could reasonably send it to referees, provided the strand selection is tightened or explicitly defended.","headline":"A clear survey that links pre-LLM computational work in HPSS to LLM applications and flags the same old methodological issues, without adding new findings.","tokens_in":2283,"tokens_out":390,"would_cite":false,"duration_ms":17490,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"LLMs extend computational methods for tracking scientific concept change but repeat earlier problems with data selection, modeling choices, and result validation.","keywords":["computational conceptual history","large language models","lexical semantic change detection","history of science","digital humanities","conceptual change","distributional semantics"],"falsifier":"An LLM-based study in HPSS that produces stable, reproducible results on conceptual change without requiring manual corpus curation or expert interpretation of outputs would undermine the claim that the longstanding problems are inherited.","tokens_in":2573,"feed_emoji":"","tokens_out":685,"duration_ms":21434,"temperature":0.7,"pith_summary":"The paper reconstructs the pre-LLM history of computational approaches to concept analysis in the history, philosophy, and sociology of science by combining three strands of earlier work. It then reviews how LLMs are being applied to lexical semantic change detection and specific HPSS case studies, showing that familiar difficulties around corpus construction, operationalization, and evaluation persist. A reader would care because this places current LLM experiments in a longer methodological conversation rather than treating them as an entirely new start. The account makes clear that scale and new architectures do not automatically remove the need for careful decisions about sources and interpretation.","feed_headline":"LLMs extend concept tracking methods but keep old corpus and evaluation limits","feed_subtitle":"Review places current LLM work in a longer line of digital history methods where data selection and validation choices remain central.","key_machinery":"Three strands of prior work—early digital methods in HPSS, distributional approaches from digital history, and lexical semantic change detection—used to frame how LLMs handle corpus construction, operationalization, and evaluation.","core_discovery":"The review reconstructs computational conceptual history before LLMs by bringing together early digital methods in HPSS, distributional approaches from digital history, and lexical semantic change detection. It then examines LLM-based work on lexical semantic change detection and relevant HPSS case studies, revisiting the methodological questions of corpus construction, model choice and training data, operationalization trade-offs, and evaluation and interpretation to show how these issues continue to shape LLM workflows.","pith_inferences":["The same pattern of inherited methodological limits may appear when LLMs are applied to conceptual analysis outside HPSS, such as in legal or literary history.","One testable extension would be to compare stability of LLM-derived concept representations against earlier methods across the same long historical text collections.","The review implies that progress may lie more in designing transparent workflows that document data and modeling choices than in adopting ever-larger models."],"forward_implications":["LLMs allow processing of larger and more heterogeneous corpora than earlier distributional methods for detecting shifts in scientific terminology.","Prompting and fine-tuning decisions in LLM workflows introduce operationalization trade-offs comparable to those in pre-LLM vector-space models.","Evaluation of LLM outputs for historical concepts continues to depend on alignment with expert historical knowledge rather than purely quantitative metrics.","Corpus construction choices remain decisive even when models can ingest raw text at scale."],"fun_headline_variants":["LLMs track scientific concepts but retain old corpus limits","Review connects digital history methods to current LLM work","Concept analysis faces same operationalization tradeoffs with LLMs","From early digital HPSS to LLMs: evaluation challenges persist"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"The three strands of earlier work adequately capture the main challenges and opportunities in computational conceptual history prior to LLMs.","fun_headline_variants_meta":{"raw":{"variants":["LLMs track scientific concepts but retain old corpus limits","Review connects digital history methods to current LLM work","Concept analysis faces same operationalization tradeoffs with LLMs","From early digital HPSS to LLMs: evaluation challenges persist"]},"model":"grok-4.3","cost_usd":0.004642,"raw_usage":{"total_tokens":2205,"prompt_tokens":643,"num_sources_used":0,"completion_tokens":63,"cost_in_usd_ticks":46415500,"prompt_tokens_details":{"text_tokens":643,"audio_tokens":0,"image_tokens":0,"cached_tokens":64},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":1499,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":643,"tokens_out":63,"duration_ms":11942,"temperature":1.0,"reasoning_tokens":1499,"cache_read_input_tokens":64,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-28T10:09:55.316146+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"An LLM-based study in HPSS that produces stable, reproducible results on conceptual change without requiring manual corpus curation or expert interpretation of outputs would undermine the claim that the longstanding problems are inherited.","supporting_citations":[],"review_version":1}