{"id":"99da31d3-01ef-4134-8802-673a4ad19ec8","arxiv_id":"2411.14877","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"This paper introduces Astro-HEP-BERT, a BERT model adapted to astrophysics and high-energy physics text, plus a large arXiv-based corpus, as a low-cost tool for studying conceptual change in science.","lead":"Astro-HEP-BERT is a language model built by continuing to train BERT on 21.84 million paragraphs from astronomy and particle physics papers on arXiv. The paper argues this cheap, laptop-scale approach can match much larger domain-specific models when tracing how scientific concepts change meaning.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The headline performance claim is not supported in this manuscript; all evaluation is deferred to a companion paper that reports a single-term ('Planck') case study.","rationale":"The reader's conditional verdict is appropriate: the model and corpus are described in enough detail to be useful resources, and the deferral to the companion paper is explicit and honest rather than hidden. My stress-test pass converges on the same bottom line, but through a different aperture. The reader's stated weakest assumption is that the 'full-paragraphs format' improves semantic coherence without ablation (Section 3). That is a real secondary risk, but it is not the most load-bearing concern for the paper's central claim: even if paragraph-level training is neutral or slightly harmful, the model could still perform comparably to from-scratch domain BERTs, so an ablation of that design choice would not by itself settle the abstract's comparative performance claim. The more decisive issue is that the central claim is an empirical claim about downstream CWE quality, and this manuscript contains no downstream evaluation whatsoever. Section 4 simply points to a separate paper by the same author, and the only in-manuscript evidence is a training-loss curve, which is not evidence of semantic quality. Furthermore, the cited companion study is explicitly framed as a single-term case study ('Planck'), so even accepting that study at face value, the abstract generalizes far beyond what a one-term analysis can support. This does not mean the model is bad or the author is being deceptive; it means the paper's headline claim is currently unverifiable from the submitted text. The reader's CONDITIONAL verdict already captures this by requiring evaluation to be included or the paper to be reframed as a resource announcement. I therefore keep the verdict unchanged rather than moving it, and I partially agree with the reader: the full-paragraphs concern is plausible but less load-bearing than the total absence of in-paper evaluation for the central performance claim.","tokens_in":6189,"tokens_out":4332,"duration_ms":44546,"concrete_test":"Read the companion paper (arXiv:2411.14073) and check whether it reports quantitative WSD/induction metrics (e.g., cluster purity, adjusted mutual information, F1, or semantic-change correlation scores) for Astro-HEP-BERT and each baseline (PhysBERT, astroBERT, SciBERT, BERT) on identical test material. If the claimed comparability rests on descriptive cluster inspection or on only the 'Planck' case, the central claim remains unestablished; if the companion paper contains such metrics and the comparison holds across multiple terms, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 4 (Evaluation) contains no quantitative results at all: it states only that 'In Simons (2024), I evaluate the performance of five BERT-based models ... using the polysemous physics term \"Planck\" as a test study.' The abstract's central claim — that Astro-HEP-BERT's CWEs 'perform comparably to domain-adapted BERT models trained from scratch on larger datasets for domain-specific word sense disambiguation and induction and related semantic change analyses' — therefore cannot be checked, reproduced, or falsified from this preprint. The only empirical evidence presented is the decreasing training-loss curve in Figure 4, which demonstrates that the model learns to mask-predict domain text, not that its contextualized embeddings are semantically superior or even comparable for the downstream HPSS tasks advertised. Moreover, the companion study is described as a case study of one term, 'Planck,' over a 30-year period; even a strong result on one term would not substantiate the general claim about 'word sense disambiguation and induction and related semantic change analyses.' This is a load-bearing evidential gap in the central performance claim, not merely a stylistic or formatting issue. The model and corpus release may be valuable as resources, but the headline comparative claim is currently unevidenced.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents Astro-HEP-BERT, a BERT-base model further pretrained for three epochs of masked language modeling on a newly curated corpus of 21.84 million paragraphs from more than 600,000 arXiv articles in astrophysics and high-energy physics. It describes the corpus construction pipeline, the model training configuration (whole-word masking, no NSP, paragraph-level sequences, dynamic batch sizing), and the feasibility of training on a single M2 MacBook over 48 days. The stated contribution is twofold: a reusable model and corpus for the history, philosophy, and sociology of science, and evidence that cost-effective domain adaptation can achieve performance comparable to domain-adapted BERT models trained from scratch. The abstract and conclusion make this comparative performance claim, but Section 4 defers all evaluation to a companion paper by the same author and contains no metrics, baselines, or error analysis.","tokens_in":6401,"tokens_out":8663,"duration_ms":80565,"significance":"If the comparative performance claim is borne out, this is a useful contribution: it would demonstrate that continued pretraining of a general BERT model on a modest compute budget can provide domain-appropriate contextualized embeddings for historical, philosophical, and sociological analyses of scientific concepts. The manuscript has concrete strengths: the model and corpus are publicly released, the corpus pipeline is described in unusual detail, filtering decisions are illustrated with distributions, and the training recipe is concrete enough to reproduce. The main significance is therefore conditional on evidence that is not present in this manuscript.","major_comments":[{"comment":"The load-bearing comparative claim is not supported by evidence in this manuscript. Section 4 states only that evaluation appears in Simons (2024) and reports no quantitative results for word sense disambiguation, induction, or semantic change detection. The abstract's claim that Astro-HEP-BERT's CWEs 'perform comparably to domain-adapted BERT models trained from scratch on larger datasets' and the conclusion's claim that Astro-HEP-BERT 'performs comparably with four leading BERT models' therefore cannot be checked, reproduced, or falsified from this preprint. Figure 4's decreasing training loss is evidence of mask-prediction optimization, not of semantic quality or downstream task performance. To make the paper self-contained, either include the evaluation data (task setups, metrics, baselines, and error analyses) or remove the comparative claim from the abstract and conclusion and re-scope the paper as a model and corpus resource description.","section":"Section 4 (Evaluation), Abstract, and Section 5 (Conclusion)"},{"comment":"The supporting evaluation is described as a single-term case study. Section 4 says the companion study uses 'the polysemous physics term \"Planck\" as a test study' over a 30-year period. Even if the companion results were included or summarized, one term is too narrow a basis for the abstract's general claim about 'domain-specific word sense disambiguation and induction and related semantic change analyses.' The manuscript should either report results on multiple concepts or explicitly qualify the claim as a single-case demonstration.","section":"Section 4 (Evaluation) and Abstract"},{"comment":"The claimed benefit of the full-paragraphs format is an untested design assumption. Section 3 states that this format 'recognizes the paragraph as the basic unit of meaning in academic writing' and that the author 'anticipate[s] even stronger semantic coherence,' but no ablation compares paragraph-level training against sentence-level or document-level training. Because this format is a distinctive feature of Astro-HEP-BERT relative to the comparators, the comparative evaluation in Simons (2024) cannot isolate its contribution. I would like to see a small ablation (even on a subset) or a reduced claim that does not attribute performance to paragraph-level coherence.","section":"Section 3 (The Astro-HEP-BERT Model)"}],"minor_comments":[{"comment":"The word 'fine-tuned' is used for continued masked-language-model pretraining (e.g., 'fine-tuned with the newly developed Astro-HEP Corpus'); this is task-specific fine-tuning terminology and should be replaced with 'domain-adapted' or 'further pretrained.'","section":"Section 3 and Section 5"},{"comment":"The chosen thresholds (250 characters; whitespace rates 0.1 and 0.2) are described as arising from frequency analysis and manual inspection, but the figures do not mark these cutoffs; please add reference lines and state how sensitive the corpus composition is to these thresholds.","section":"Section 2, Figures 2 and 3"},{"comment":"The sentence 'which could refer to weight in the particle is \"light\"' is ungrammatical and should be rewritten, for example as 'which could refer to the property of low mass in \"the particle is light\" or to the electromagnetic phenomenon in \"light is a particle.\"","section":"Section 1, Introduction"},{"comment":"The phrase 'test study .' contains a stray space before the period; please remove it.","section":"Section 4, first paragraph"},{"comment":"The corpus filtering removes paragraphs under 250 characters, yet the model section reports paragraph lengths ranging from 48 to 510 subwords; clarify the relationship between these numbers (characters versus subwords) to avoid an apparent inconsistency.","section":"Section 3, paragraph-length range"},{"comment":"The phrase 'the document-sentence input format proposed by Liu et al. (2019)' attributes a training-data format to RoBERTa; please cite the relevant analysis more precisely or rephrase to avoid implying that the RoBERTa paper proposed this format.","section":"Section 3, citation"},{"comment":"The term 'colexification' is used where 'polysemy' or 'homonymy' seems intended; colexification usually refers to a single word form covering multiple senses across languages, not to context-dependent meaning variation within one language.","section":"Section 1, Introduction"}],"recommendation":"major_revision","confidential_remarks":"The paper is honest about the location of the evaluation in Section 4, but as it stands the headline claim depends entirely on a companion paper by the same author. I recommend major revision: the authors should either integrate the evaluation into this manuscript or explicitly reframe it as a resource paper. The model and corpus are likely valuable to the HPSS community, and the training details are commendable, so I do not see a need for rejection if the claims are brought in line with the evidence."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things you should know about this one: the model and corpus are real, plausibly useful for HPSS work, and the headline performance claim is not supported by evidence in the paper itself. That is not a dirty secret—the paper says the evaluation lives in a companion preprint—but it shapes how you read it.\n\nWhat is genuinely new: the Astro-HEP Corpus, 21.84 million paragraphs from over 600,000 arXiv papers in astro and HEP (1986-2022), and a BERT model adapted to that text with three epochs of masked language modeling. The training pipeline—Pandoc-based LaTeX parsing, paragraph filtering with a 250-character cutoff and whitespace-rate thresholds, token-balanced batches of 8192 with padding capped at 20%—is described clearly enough to replicate. Doing it on a single M2 MacBook in 48 days is a useful feasibility result for small labs. Dropping NSP and using whole-word masking follows RoBERTa-style findings; the 'full-paragraphs' format is an intriguing idea but is asserted as beneficial without an ablation.\n\nThe soft spots are real. The abstract claims the model's CWEs 'perform comparably to domain-adapted BERT models trained from scratch.' In Section 4 you will find no numbers, no baselines, no error analysis. All evaluation is deferred to arXiv:2411.14073, and that companion is a single-term case study ('Planck' over 30 years). A decreasing loss curve shows the model learns mask prediction, not that its embeddings are better for word sense disambiguation or semantic change. The filtering thresholds are based on frequency analysis plus manual inspection; no validation shows they preserve corpus representativeness. And the paragraph-as-semantic-unit premise is untested. These are fixable: include the companion's key results, or reframe the paper as a resource announcement with evaluation marked forthcoming.\n\nThe reader's conditional verdict is fair. The paper is honest about the deferral, and the resource itself is valuable. I would send it to review—the corpus and model deserve referee time—but a referee should demand the abstract match the evidence and an ablation or qualifier on the paragraph format. If the companion delivers the promised comparison, this becomes a solid contribution. If not, the resource still stands, but the headline claim does not.","headline":"Useful new domain-adapted BERT and corpus for physics text, but the main performance claim is unevidenced in this preprint and rests on a single-term companion study.","tokens_in":6957,"tokens_out":3167,"would_cite":true,"duration_ms":29450,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that a general BERT model given three extra epochs of training on 21.84 million paragraphs of astrophysics and high-energy physics text produces contextualized word embeddings comparable to physics-specific models…","keywords":["contextualized word embeddings","domain adaptation","BERT","astrophysics","high energy physics","word sense disambiguation","lexical semantic change","history philosophy and sociology of science"],"falsifier":"A controlled ablation that trains Astro-HEP-BERT's identical setup on sentence-level sequences instead of full paragraphs, evaluated on the same word sense disambiguation and semantic change tasks, would settle whether the paragraph format actually helps; if the two versions perform equally, the assumed advantage of paragraphs over sentences is absent.","tokens_in":5945,"feed_emoji":"🔭","tokens_out":8419,"duration_ms":77670,"temperature":0.7,"pith_summary":"The paper presents Astro-HEP-BERT, a transformer language model made by taking a general-purpose BERT model and running three additional training passes over 21.84 million paragraphs from more than 600,000 astrophysics and high-energy physics articles. The author's central claim is that the resulting contextualized word embeddings perform about as well as domain-specific BERT models trained from scratch on larger physics corpora, at least for disambiguating and tracing word meanings such as the term \"Planck\". If that claim holds, researchers in the history, philosophy, and sociology of science could build custom tools for studying how scientific concepts shift in meaning using only open code, open data, and a single laptop, without the cost of training a model from scratch.","feed_headline":"Laptop-trained BERT matches physics models built from scratch","feed_subtitle":"A three-epoch tune-up on physics papers puts concept-meaning analysis within reach for humanities researchers.","key_machinery":"The machinery is continued pretraining of a general bidirectional transformer with three modifications to the original BERT protocol: masked language modeling without the next-sentence prediction objective, whole-word masking, and a \"full-paragraphs format\" in which each training sequence is a complete paragraph rather than a sentence or document. Training batches are organized to hold about 8,192 tokens with limited padding, so compute is spent on real text rather than placeholder tokens. The paragraph format is the distinctive design choice: the author argues that the paragraph, not the sentence, is the basic unit of meaning in academic writing, so embedding each paragraph as a unit should improve semantic coherence in the model's contextualized embeddings.","core_discovery":"Astro-HEP-BERT is a BERT model given three more epochs of masked-language-model training on the Astro-HEP Corpus, 21.84 million paragraphs drawn from more than 600,000 astrophysics and high-energy physics articles published between 1986 and 2022. The paper's central claim is that this modest continued pretraining, performed with freely available tools on one laptop, yields contextualized word embeddings that a companion evaluation finds comparable to physics-specific BERT models trained from scratch on larger corpora. The author frames this as evidence that domain adaptation, rather than from-scratch training, is a viable and affordable route for studying the meanings of scientific concepts.","pith_inferences":["A natural next test is to apply the same continued-pretraining recipe to other scientific literatures; if the laptop-scale result generalizes, domain-adapted transformers could become a standard tool for conceptual history across many fields.","The full-paragraphs format, if confirmed by ablation, would imply that academic paragraphs are a better semantic unit than sentences for language-model pretraining, a principle that could inform future model designs beyond this corpus.","Because the base model is uncased, case-only distinctions between terms (such as names that are also ordinary words) may be flattened; a cased or symbol-aware variant might improve fine-grained semantic analysis.","The comparability claim currently rests on the single test term \"Planck\"; a multi-term benchmark of homographs and polysemous words would show whether the result is a general property of continued pretraining or specific to that case."],"forward_implications":["Researchers in the history, philosophy, and sociology of science can build domain-adapted language models for new fields on a single laptop, using openly available code, weights, and text.","Astro-HEP-BERT can disambiguate and trace the meanings of terms such as \"Planck\" across the 1986–2022 corpus, with shifts tied to events like the Planck space mission.","The Astro-HEP Corpus provides a reusable dataset of 21.84 million paragraphs with article-level metadata for studying concept change in astrophysics and high-energy physics.","The decreasing masked-language-model loss over three epochs indicates that continued pretraining on physics text does capture domain-specific language, supporting the idea that from-scratch training is not required for useful domain embeddings."],"supporting_citations":[{"why":"Supplies the base BERT architecture, pretrained weights, and vocabulary that Astro-HEP-BERT starts from.","marker":"Devlin et al. (2018)"},{"why":"Source of the whole-word masking technique used in training.","marker":"Devlin (2019)"},{"why":"Justifies removing the next-sentence prediction objective and motivates the document-sentence input format that the full-paragraphs format refines.","marker":"Liu et al. (2019)"},{"why":"Cited as supporting the removal of the next-sentence prediction objective for distributional semantic quality.","marker":"Mickus et al. (2020)"},{"why":"Provides astroBERT, a from-scratch domain model used as the comparison baseline for evaluating Astro-HEP-BERT.","marker":"Grezes et al. (2021)"},{"why":"Provides PhysBERT, a from-scratch physics text embedding model used as the comparison baseline.","marker":"Hellert et al. (2024)"},{"why":"The companion study that actually evaluates Astro-HEP-BERT against four models for disambiguating and tracking the meaning of \"Planck\"; the comparability claim rests on this evaluation.","marker":"Simons (2024)"},{"why":"Frames lexical semantic change detection, the target task for the model's contextualized word embeddings.","marker":"Periti and Montanelli (2024)"},{"why":"Motivates the digital Begriffsgeschichte application in which the model's semantic change analyses are situated.","marker":"Wevers and Koolen (2020)"}],"fun_headline_variants":["Laptop-only training makes BERT competitive with from-scratch physics models","Three extra epochs on a laptop match costly physics BERTs","Cheap BERT tune-up rivals expensive physics models","A 3-epoch arXiv tune-up gives BERT physics-grade word senses"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The full-paragraphs format carries the argument: the paper assumes that a paragraph is a better unit of meaning than a sentence or document for academic writing, and offers no ablation to test that assumption.","fun_headline_variants_meta":{"raw":{"variants":["Laptop-only training makes BERT competitive with from-scratch physics models","Three extra epochs on a laptop match costly physics BERTs","Cheap BERT tune-up rivals expensive physics models","A 3-epoch arXiv tune-up gives BERT physics-grade word senses"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000178,"raw_usage":{"total_tokens":1281,"prompt_tokens":915,"completion_tokens":366,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":531,"completion_tokens_details":{"reasoning_tokens":291}},"tokens_in":531,"tokens_out":366,"duration_ms":4838,"temperature":1.0,"reasoning_tokens":291,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T14:45:20.138920+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A controlled ablation that trains Astro-HEP-BERT's identical setup on sentence-level sequences instead of full paragraphs, evaluated on the same word sense disambiguation and semantic change tasks, would settle whether the paragraph format actually helps; if the two versions perform equally, the assumed advantage of paragraphs over sentences is absent.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Source of the whole-word masking technique used in training."},{"cited_title":"What do you mean, BERT? Assessing BERT as a Distributional Semantics Model","cited_arxiv_id":"1911.05758","evidence_quote":"Cited as supporting the removal of the next-sentence prediction objective for distributional semantic quality."},{"cited_title":"Meaning at the Planck scale? Contextualized word embeddings for doing history, philosophy, and sociology of science","cited_arxiv_id":"2411.14073","evidence_quote":"The companion study that actually evaluates Astro-HEP-BERT against four models for disambiguating and tracking the meaning of \"Planck\"; the comparability claim rests on this evaluation."},{"cited_title":"and Montanelli, S","cited_arxiv_id":null,"evidence_quote":"Frames lexical semantic change detection, the target task for the model's contextualized word embeddings."},{"cited_title":"and Koolen, M","cited_arxiv_id":null,"evidence_quote":"Motivates the digital Begriffsgeschichte application in which the model's semantic change analyses are situated."}],"review_version":1}