{"id":"618ce57a-cb32-4f9f-81d7-8f99577d231d","arxiv_id":"2501.00391","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"A computational framework combining Kullback-Leibler divergence and embedding density estimation traces individual scholars' language against disciplinary knowledge evolution in 20th-century general relativity research.","lead":"This paper proposes two computational methods, relative entropy and embedding density estimation, to track how individual scientists' language aligns or diverges from their field over time. It applies them to Joseph Silk and Hans-Jürgen Treder in the post-war renaissance of general relativity, showing Silk tracked the field's future while Treder tracked its past.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Treder's divergence from the field may be an artifact of German-language indexing and translation rather than a real intellectual trajectory, and the paper's own citation-undercount evidence makes this corpus bias load-bearing for the complementarity claim.","rationale":"The reader's weakest assumption identified the NASA/ADS corpus as biased, with incomplete indexing of Treder's German-language journals as a key risk. My stress-test agrees and sharpens it: the paper itself provides quantitative evidence of underreporting for Treder, and the language imbalance between the two case-study authors makes the corpus artifact a direct alternative explanation for the central KLD/EDE correlation. This is not a mathematical flaw in the methods; it is an empirical premise about data quality that the authors explicitly leave unquantified. Because the central claim is a demonstration on two selected cases, the conditionality is appropriate. No change to the reader's CONDITIONAL verdict is needed; the concrete test would turn the condition into a checkable requirement. I found no independent reason to reject the methods or to accuse the authors of anything improper; the limitation is stated honestly in the paper, which is why the verdict remains CONDITIONAL rather than REJECT.","tokens_in":17111,"tokens_out":2595,"duration_ms":29448,"concrete_test":"Recompute the Section 4.3 KLD and Section 4.4 EDE analyses using only English-language publications for both Silk and Treder, with publication counts matched per two-year slice. Separately, take a random sample of Treder's German-language papers and compute KLD and EDE using the original German text with a German-language embedding model, comparing against a German-language field corpus. If Treder's English-only trajectory still shows rising KLD and falling EDE, the language/translation artifact is not the main driver; if the trajectory flattens or reverses, the paper's headline contrast is not robust to corpus construction.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that low KLD correlates with high EDE, using Treder as the inverse case: his rising KLD and falling EDE are read as showing that his work tracked the 'past' of GRG while Silk tracked its future. The most load-bearing concern is that this contrast may be produced by the corpus and preprocessing pipeline rather than by the authors' actual language use. Section 3.1 explicitly acknowledges 'incompleteness ..., bias (e.g., language/geographic, collection focus), and translation inaccuracies,' and the conclusion concedes that 'collection biases within the dataset, drawn from relatively small samples, remain mostly unquantified.' More concretely, Section 4.2 reports that a manual review of Treder's publications found references under-reported in the dataset (1.83 per publication versus 16.69 in manually verified data), demonstrating that ADS metadata for Treder is systematically incomplete. Treder published roughly half his works in German (175 German, 173 English) and primarily in journals such as Annalen der Physik and Astronomische Nachrichten, while Silk published in English in Nature and The Astrophysical Journal. If German-language abstracts are missing, translated inconsistently, or represented only by titles, both methods would be affected: KLD would inflate because of translationese and sparse text, and EDE would drop because the translated or truncated documents embed into different regions of the embedding space. Thus the observed rise in Treder's KLD and decline in his EDE could be an artifact of data quality, not a measure of his intellectual distance from the mainstream. Because the complementarity claim is supported by exactly this contrast, the unquantified corpus bias is load-bearing, not a peripheral caveat.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents two complementary quantitative methods for tracking the evolution of an individual researcher's language against a field-level corpus: (1) Kullback-Leibler divergence between unigram language models with Jelinek-Mercer smoothing and Welch t-test filtering, and (2) embedding density estimation (EDE) around individual documents. The methods are applied to a NASA/ADS-derived corpus of general relativity and gravitation (GRG) publications from 1911 to 2000, with case studies of Joseph Silk and Hans-Jürgen Treder over 1957-2000. The authors report that Silk's terminology increasingly aligns with the future mainstream of GRG while his EDE rises, whereas Treder's terminology aligns with the past and his EDE falls, and they interpret this as evidence that KLD and EDE are complementary indicators of individual versus system knowledge evolution.","tokens_in":1729,"tokens_out":1663,"duration_ms":54800,"significance":"The proposed combination of an information-theoretic corpus comparison and a contextual embedding density measure is a sensible and transferable contribution to quantitative history of science. The authors should be credited for making the corpus and code available, for using standard statistical machinery, and for explicitly listing dataset limitations. The two-author case study, however, cannot by itself validate the complementarity claim, and the empirical contrast is vulnerable to the acknowledged corpus biases, particularly for Treder. If the robustness issues are addressed, the framework would be a useful addition to the SEN toolkit; as it stands, the central finding is suggestive rather than demonstrated.","major_comments":[{"comment":"The central empirical contrast is not protected against the corpus-asymmetry problem that the paper itself documents. Section 4.2 reports that Treder's references are under-reported in ADS (1.83 per publication versus 16.69 in manually verified data), and Table 1 shows he published primarily in Annalen der Physik and Astronomische Nachrichten, with roughly half of his publications in German. Since both methods operate on titles and abstracts translated into English (Section 3.1), missing or poorly translated German abstracts would mechanically inflate KLD and depress EDE for Treder relative to Silk, producing exactly the 'past versus future' pattern reported in Sections 4.3 and 4.4. The conclusion in Section 5 concedes that 'collection biases ... remain mostly unquantified.' This is load-bearing: the authors should either quantify abstract and translation coverage per author per time slice, re-run the analysis on English-only subsets, or compare Treder's German-original and English-language outputs as a control. Without such a check, the claim that Treder tracks the past while Silk tracks the future is not established.","section":"Sections 3.1, 4.2, 5"},{"comment":"The EDE method's output is controlled by the KDE bandwidth h, but the manuscript never states the chosen value, the kernel implementation, or any sensitivity analysis. In Eq. (2), h is the only scale parameter; changing it can invert the relative density trajectories, and Figure 5's divergence between Silk and Treder after the mid-1970s is precisely the kind of qualitative pattern that a bandwidth choice can create or remove. Similarly, the choice of text-embedding-3-large is justified only as 'most accurate and meaningful' after experimentation, without reporting the compared models' scores or the embedding version. The authors should report h and the embedding configuration, and show that the main trends in Figure 5 are stable under reasonable variations of h and under an alternative embedding model.","section":"Section 3.2.2, Eq. (2), Fig. 5"},{"comment":"The asynchronous KLD result is interpreted as evidence that Silk's terminology aligns with the future of GRG while Treder's aligns with its past, but no statistical or effect-size measure is attached to the minima highlighted in Figure 4. Because KLD is asymmetric and the field distribution changes over time, a low KLD between an author's slice and a field slice can arise from a small number of high-frequency shared terms; the paper does not test whether the minima are distinguishable from a null model (e.g., random author-term distributions or cross-slice baselines). Without such a test, the 'future versus past' wording overstates what the figure shows. I request a significance or bootstrap-based assessment of the asynchronous minima.","section":"Section 4.3.3"},{"comment":"The paper's operating assumption at the start of Section 4, that low KLD and high EDE should correlate, is tested on exactly two authors. The conclusion appropriately notes this is not a general rule, but the complementarity claim is then used as the main takeaway. At minimum, the authors should test the correlation within authors over time slices (e.g., Spearman correlation between per-slice KLD and per-slice median EDE) and report the result; if the correlation is weak, the conclusion should be scaled back to a demonstration on two cases. As written, the 'complementary insights' claim is supported mainly by the same two cases that motivated the comparison.","section":"Section 4, first paragraph"}],"minor_comments":[{"comment":"'Kullback-Leiber' is misspelled; it should be 'Kullback-Leibler'.","section":"Section 3.2.1"},{"comment":"The numbered list uses 'a)' three times; the sub-steps should be labeled 2a, 2b, and 2c.","section":"Section 3.2.2"},{"comment":"The text says 'mean EDE starts to decline' while Figure 5's caption and the method description refer to median EDE; please make the statistic consistent.","section":"Section 4.4.1"},{"comment":"'Treder's use of German (173 publications in English, 175 in German)' sums to 348, not the 355 total reported for Treder; please reconcile the counts.","section":"Section 4.2"},{"comment":"The lemmatization step is not specified (tool, language resources, and version); since Method 1 depends on lemmas, this should be stated for reproducibility.","section":"Section 3.1"},{"comment":"The caption's description of the y-axis slices and time differences is confusing; consider simplifying it or adding a schematic to clarify which slice is compared with which.","section":"Figure 4 caption"}],"recommendation":"major_revision","confidential_remarks":"To the editor: the paper is honest about its limitations and makes data and code available, but the main empirical claim is currently dependent on unquantified corpus asymmetries, especially for Treder's German-language publications. The two-author design is appropriate for a methods illustration but not for validation of the complementarity claim. I recommend major revision with a specific request for robustness checks; if the authors cannot provide them, the paper should be reduced to a methodological proposal with clearly illustrative case studies. The manuscript fits venues in digital humanities and computational social science better than a core machine learning venue."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThis is a methods paper that combines two established techniques—KLD on unigram models and embedding density estimation—to compare an individual scientist's language to a field corpus. The new bit is the combination and the individual-level application, and it's a genuine contribution to digital history of science. The paper is clearly written, the statistics are standard (Jelinek-Mercer smoothing, Welch's t-test, KDE), and the authors are unusually transparent about limitations. They release the corpus, which is good practice.\n\nThe central claim is that low KLD correlates with high embedding density, using Silk (future-oriented) and Treder (past-oriented) as case studies. That claim is plausible but not fully proven. The biggest concern is the data. Treder published about half his output in German, in journals like Annalen der Physik, and the ADS corpus under-represents his citations (the paper reports 1.83 references per publication vs 16.69 in a manual check). If his German-language abstracts are missing or poorly translated, both methods would be affected: KLD would inflate from translationese and sparse text, and EDE would drop if the truncated documents land in odd regions of embedding space. The observed Treder trajectory could be partly an artifact. The authors acknowledge this bias but don't quantify it, and since the complementarity claim rests on exactly this contrast, it's load-bearing, not peripheral.\n\nThere are smaller issues: no link to the actual code (they mention a package but no repository URL), the embedding model choice is justified by 'experimentation' without a systematic comparison, and the two case studies are selected for convenience—they happen to differ in the desired ways. Error bars are absent throughout.\n\nThat said, the paper doesn't overclaim. It explicitly says the pattern needs broader testing, and the qualitative historical record supports the Treder-Silk divergence we see. So the stress-test concern is a call for more work, not a refutation.\n\nWho's this for? Historians of science and anyone doing diachronic corpus analysis at the individual level. It deserves a serious referee—send it out—but the revision should tackle corpus bias head-on, provide code, and add robustness checks (e.g., analyzing English-only Treder publications, or a multilingual embedding model).\n\nRecommendation: accept for review with heavy revision expected.","headline":"Useful methods demonstration for digital history, but the Silk-Treder contrast is vulnerable to the corpus bias the authors themselves acknowledge.","tokens_in":17974,"tokens_out":2980,"would_cite":true,"duration_ms":29949,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Word-frequency divergence and embedding density can track how an individual researcher aligns with the evolving mainstream of a scientific field, and the paper shows the two measures moving together for two physicists in opposite…","keywords":["knowledge evolution","socio-epistemic networks","Kullback-Leibler divergence","embedding density estimation","diachronic language change","history of general relativity","quantitative corpus analysis","document embeddings"],"falsifier":"A decisive check would be to rebuild the corpus with a manually verified and completed record of Treder's German-language publications and then recompute his KLD and EDE curves; if his late rise in divergence and drop in density vanish, the claimed past-oriented trajectory is a missing-data artifact rather than a knowledge evolution.","tokens_in":16867,"feed_emoji":"📈","tokens_out":10319,"duration_ms":95881,"temperature":0.7,"pith_summary":"This paper argues that the evolution of a scientific field and the position of an individual researcher within that evolution can be tracked quantitatively from the words and meanings of published titles and abstracts. It pairs two measures: the Kullback-Leibler divergence (KLD), which scores how far an author's word frequencies stray from the field's word frequencies in each two-year window, and embedding density estimation (EDE), which measures whether new papers cluster ever more densely around the semantic neighbourhood of the author's documents. Applied to a corpus of about 180,000 general-relativity and gravitation titles and abstracts from 1911 to 2000, analysed from 1957 onward, the two measures agree: the astrophysicist Joseph Silk's language stays close to the mainstream and his topic neighbourhood grows denser, while the East German physicist Hans-Jürgen Treder's language drifts toward the terminology of earlier decades and his neighbourhood thins out. The paper reads these signatures as Silk's work aligning with the future direction of general relativity and gravitation research and Treder's with its past.","feed_headline":"Silk's papers track the future of gravity research; Treder's, its past","feed_subtitle":"Pairing word-frequency divergence with topic density separates the researchers who lead a field from those who lag it.","key_machinery":"The machinery has two parts. First, the Kullback-Leibler divergence between two unigram language models, one built from an author's lemmatised titles and abstracts in a two-year slice and one from the rest of the field in the same slice, with Jelinek-Mercer smoothing and a Welch t-test filtering which terms count; summing it gives a divergence-over-time curve and pointwise values name the words doing the work. Second, embedding density estimation: documents are embedded with a large text-embedding model, dimension-reduced and clustered, and a Gaussian kernel density estimate is evaluated at each reference document across successive two-year slices, so a rising value means the field is publishing more documents semantically close to that paper. The paper also runs asynchronous KLD comparisons, aligning each author's slice against all past and future field slices, which is what lets it say an author's vocabulary resembles the field's past rather than its present.","core_discovery":"On the paper's own account, the discovery is that relative-entropy divergence and embedding density are two sides of one phenomenon: an author whose terminology matches the field's mainstream is an author whose topics other researchers continue to cluster around, and an author who drifts away from the mainstream is an author whose research neighbourhood disperses. The paper claims this correspondence holds across the two case studies, that KLD and EDE give complementary rather than redundant information—the first about the vocabulary that separates an individual from the collective, the second about how much ongoing work is semantically close—and that together they provide an operational way to connect an individual knowledge trajectory to a system-level field transformation. The specific case-study result is that Silk's trajectory tracks the future of general relativity and gravitation research while Treder's tracks its past, with the divergence visible in both synchronous and asynchronous comparisons.","pith_inferences":["A stress test the paper does not run is to apply the same pair of measures to an author whose career history is well documented but who lies between the two extremes; an intermediate case would show how much noise the KLD-EDE correlation tolerates.","Because the corpus contains only titles and abstracts translated into English, the measures may partly track changes in how scientists summarise their work, not only changes in the work itself; full-text analysis, which the paper names as future work, could separate these effects.","The paper's own manual comparison (1.83 versus 16.69 references per publication for Treder) gives a concrete robustness check: rebuild the corpus with a manually completed bibliography and recompute the trajectories.","The asynchronous comparison could be inverted into a 'field clock' that dates any document by the period in which its vocabulary would have been most typical, a tool for mapping when concepts become dominant and when they age."],"forward_implications":["In the two cases studied, the measures move together: Silk keeps a low divergence from mainstream terminology and a steady or rising embedding density, while Treder's divergence rises and his density falls from the mid-1970s onward.","The combined method supplies an operational way to place an individual micro-history inside a field-level macro-history, which the socio-epistemic network framework leaves as an open integration problem.","Because the measures work on any diachronic text corpus, the approach transfers to other sciences and to non-scientific settings such as migration-related knowledge dissemination or science-to-policy transitions.","Asynchronous KLD comparison adds a directional component: the field slice whose vocabulary is closest to an author's identifies which period of the field's history that author's work belongs to."],"supporting_citations":[{"why":"It supplies the relative-entropy method for detecting periods of diachronic linguistic change and the terms driving them, which Method 1 extends to the individual-vs-field comparison.","marker":"[24]"},{"why":"It demonstrates the KLD-based study of long-term scientific writing that the pointwise divergence workflow adapts.","marker":"[29]"},{"why":"It defines the GRG publication space and the network-analysis approach that the corpus selection and socio-epistemic framing build on.","marker":"[11]"},{"why":"It provides the quoted account of GRG's turn toward relativistic astrophysics that anchors the interpretation of the two authors' opposite trajectories.","marker":"[10]"},{"why":"It introduces the socio-epistemic network framework with its social, semiotic, and semantic layers that motivates the semantic-layer analysis.","marker":"[2]"},{"why":"It describes the embedding-to-cluster topic-modelling pipeline that Method 2's density estimation follows.","marker":"[55]"},{"why":"It sets out the Renaissance of General Relativity project whose selection criteria define the corpus.","marker":"[58]"}],"fun_headline_variants":["Entropy and density split gravity's future from its past","Two metrics expose which physicist shapes a field's future","Word shifts and topic clusters track a scientist's legacy","Silk's trajectory predicts, Treder's reflects—both via dual measures"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the large bibliographic corpus of titles and abstracts, translated into English, faithfully represents both the general-relativity field and the two authors' published work, so that word-frequency divergence and embedding density measure intellectual distance rather than gaps in indexing or translation artifacts.","fun_headline_variants_meta":{"raw":{"variants":["Entropy and density split gravity's future from its past","Two metrics expose which physicist shapes a field's future","Word shifts and topic clusters track a scientist's legacy","Silk's trajectory predicts, Treder's reflects—both via dual measures"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000205,"raw_usage":{"total_tokens":1363,"prompt_tokens":883,"completion_tokens":480,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":499,"completion_tokens_details":{"reasoning_tokens":410}},"tokens_in":499,"tokens_out":480,"duration_ms":5766,"temperature":1.0,"reasoning_tokens":410,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T22:50:30.436146+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A decisive check would be to rebuild the corpus with a manually verified and completed record of Treder's German-language publications and then recompute his KLD and EDE curves; if his late rise in divergence and drop in density vanish, the claimed past-oriented trajectory is a missing-data artifact rather than a knowledge evolution.","supporting_citations":[{"cited_title":"Degaetano-Ortlieb, E","cited_arxiv_id":null,"evidence_quote":"It supplies the relative-entropy method for detecting periods of diachronic linguistic change and the terms driving them, which Method 1 extends to the individual-vs-field comparison."},{"cited_title":"Bizzoni, S","cited_arxiv_id":null,"evidence_quote":"It demonstrates the KLD-based study of long-term scientific writing that the pointwise divergence workflow adapts."},{"cited_title":"Lalli, R","cited_arxiv_id":null,"evidence_quote":"It defines the GRG publication space and the network-analysis approach that the corpus selection and socio-epistemic framing build on."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"It provides the quoted account of GRG's turn toward relativistic astrophysics that anchors the interpretation of the two authors' opposite trajectories."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"It introduces the socio-epistemic network framework with its social, semiotic, and semantic layers that motivates the semantic-layer analysis."},{"cited_title":"Grootendorst, BERTopic: Neural topic modeling with a class-based TF-IDF procedure,","cited_arxiv_id":null,"evidence_quote":"It describes the embedding-to-cluster topic-modelling pipeline that Method 2's density estimation follows."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"It sets out the Renaissance of General Relativity project whose selection criteria define the corpus."}],"review_version":1}