{"id":"b773f02e-8f50-4489-a324-00fc1afce2a2","arxiv_id":"2411.15068","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"Precocity measured from the most forward-looking quarter of a text aligns better with citations and authorial youth than whole-text averages, and neural embeddings do not beat topic models.","lead":"This paper tests three ways of representing text, topic models, document embeddings, and perplexity, to see which best locates cultural change in literary studies, economics, and fiction. It finds that a text's most forward-looking passages predict citations and authorial youth better than its average, and that simple lexical topic models match neural methods.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The paper never verifies that the passages ranked most 'precocious' are the passages that actually attract citations or critical discussion; without that check, the top-quartile conclusion could be a document-level confound.","rationale":"Good faith: the paper is transparent, discloses preregistration deviations, provides code and data, reports a consistent pattern across three corpora and three representations, and explicitly addresses direct text-reuse circularity. None of that is in dispute. The remaining weak spot is between the measured quantity (document-level correlation) and the explanatory conclusion that a small forward-looking subset of a text drives social response. Because 'critical discussion' is operationalized as any mention in the literary-studies corpus (Section 2.2), it is especially vulnerable: mere mention tracks canonization and visibility, not necessarily engagement with innovative passages. The proposed citation-context test would settle whether the top-quartile passages are the ones actually responded to, and would also check whether the citation rows survive removing citing documents from the future comparison corpus. If the test fails, the mechanism claim should be withdrawn in favor of a descriptive claim about which aggregate representations correlate with social attention; if it passes, the CONDITIONAL verdict can be upgraded. The two apparent ties in Table 1 (perplexity, age variables) and the absence of confidence intervals further caution against reading 'in all tests' too strongly, but the passage-level evidence is the decisive gap.","tokens_in":9628,"tokens_out":10750,"duration_ms":125864,"concrete_test":"Use the released data and code to extract, for each cited article or critically discussed fiction work, the specific passages quoted, paraphrased, or substantively discussed in citing documents or literary-studies criticism (Semantic Scholar citation contexts and the full-text literary studies corpus). Compute the within-document precocity rank of each referenced passage and test whether referenced passages are concentrated in the top quartile (median rank < 0.75) and whether this concentration explains the document-level R2 advantage of top-quartile representation. Also rerun Table 1 excluding from the future comparison set any documents that cite the target; if the citation-row advantages collapse, the citation result is partly circular.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central inference in Section 4—that citations and critical references are 'motivated by innovations expressed in a relatively small part of a text'—is supported only by document-level R2 comparisons (Table 1). The authors never test whether the passages that citing or critical texts actually quote, paraphrase, or discuss are the same passages their metric labels as most precocious. This is load-bearing because the top-quartile advantage could arise from a document-level confound: prominent documents may contain a few future-sounding passages for reasons unrelated to the specific content that elicited the social response (prestige, visibility, length, author networks). Appendix C's exclusion of same-author and verbatim-quote chunks prevents one mechanical circularity, but it does not establish that the high-precocity chunks are the socially salient ones. The author-age results weaken a pure citation-circularity story, but they do not identify the locus of influence either. Without passage-level evidence, the headline claim that 'a text's impact may depend more on its most forward-looking moments' remains an interpretation of an aggregate correlation, not a demonstrated mechanism.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces a text-level measure called \"precocity,\" defined as the degree to which a text resembles future documents more than past documents, and compares three textual representations (topic models, fine-tuned document embeddings, and word-level perplexity) across three corpora (literary studies, economics, and fiction). The authors regress five social-outcome variables (citation counts in two fields, author age in two fields, and critical discussion in fiction) on precocity, comparing two document-level aggregation strategies: the mean over all passages and the mean over the top quartile of passages by precocity. They report that the top-quartile strategy aligns better with social evidence in every test they ran, while no single text representation consistently outperforms the others, and they interpret this as evidence that cultural influence may be concentrated in a small forward-looking fraction of a text.","tokens_in":9810,"tokens_out":3677,"duration_ms":39673,"significance":"If the top-quartile finding is robust, it would be a useful methodological result for computational humanities and cultural analytics, because most prior work averages novelty or prescience over entire documents. The paper is commendably broad in scope, covering three corpora and three representation families, and it makes data and code publicly available. It also shows unusual transparency in Appendix H, disclosing departures from the preregistration and reporting paths not taken, which strengthens confidence in the authors' good faith. However, the central claim currently rests on small differences in document-level R-squared values without uncertainty quantification, and the mechanism invoked to interpret those differences is not directly tested. The paper's practical significance depends on whether these gaps can be closed with additional analysis.","major_comments":[{"comment":"The claim that the top-quartile method \"aligned better with social evidence than a method that averaged all passages\" in all tests is not actually supported by Table 1. In the Perplexity rows for Age of literary scholars and Age of fiction writers, the reported R-squared values are identical under 0.25 and 1.0 aggregation (.024 vs .024 and .014 vs .014, respectively). Several other differences are extremely small, for example Topics on fiction-writer age (.051 vs .049) and Embeds on literary-scholar age (.035 vs .034). Because no confidence intervals, standard errors, or p-values are reported, the reader cannot tell whether these differences are anything beyond rounding or sampling noise. The claim of a consistent top-quartile advantage needs at least paired or clustered uncertainty estimates, and ideally a multiple-comparison correction across the 30 method-outcome combinations.","section":"Section 4, Table 1"},{"comment":"The entire method-comparison conclusion rests on R-squared values that are not accompanied by any measure of uncertainty. R-squared differences of 0.002 to 0.005 are reported as evidence for one aggregation strategy over another, but without error bars or significance tests these differences are not interpretable. This is load-bearing because the paper's headline conclusion is not that precocity has some association with social outcomes, but that the top-quartile representation is systematically better. Given that many modeling choices were explored and revised (Appendix H), the risk of overfitting to the reported outcome is real. I would like to see bootstrap confidence intervals for the R-squared differences, or an equivalent resampling procedure, applied to the central comparison in Section 4.","section":"Section 4, Table 1"},{"comment":"The inference that citations and critical references are \"motivated by innovations expressed in a relatively small part of a text\" is not directly tested. Table 1 contains only document-level regressions, which cannot establish that the specific passages ranked as most precocious are the passages that elicited the social response. A document-level confound is plausible: highly cited or prominent documents may contain a few future-sounding passages for reasons unrelated to the content that actually drew attention, such as visibility, author networks, or length. Appendix C excludes same-author and verbatim-quote chunks, which addresses one mechanical circularity, but it does not identify the locus of social influence. A passage-level validation would be needed: for example, locate passages that citing or critical texts explicitly quote, paraphrase, or discuss, and compare the precocity scores of those passages with the precocity scores of non-discussed passages, ideally conditioning on document-level effects.","section":"Section 4, paragraph beginning \"So what did we learn\""}],"minor_comments":[{"comment":"The statement that \"youth and citation frequency are negatively correlated in this corpus\" is reported without a correlation coefficient or sample-size caveat; since the age data cover only 2,646 of 40,407 articles, a brief summary of the strength and robustness of this correlation would help the reader assess the tension the paper relies on.","section":"Section 2.1"},{"comment":"The author-age inference procedure is described as achieving \"overall accuracy of greater than 90%,\" but no confidence interval or validation-set details are provided; because age is one of the two primary social-outcome variables, a sentence on the precision of the VIAF matching model would be useful.","section":"Appendix D"},{"comment":"The top-quartile strategy averages the 25% of chunks with the highest precocity within each document, but the number of chunks per document is not reported; for very short documents the quartile may be based on very few chunks, which could affect the comparison between the two aggregation strategies.","section":"Section 3.4"},{"comment":"The appendix honestly notes that the embedding strategy changed several times and that authorial age was added after preregistration; however, the main text does not flag how these departures affect the strength of the evidence for the claim that no representation is preferred, and a sentence acknowledging the potential for selective reporting would improve the paper's transparency.","section":"Appendix H"}],"recommendation":"major_revision","confidential_remarks":"The paper is honest, well-scoped, and accompanied by public data and code, which I regard as real strengths. The main issues are fixable within the manuscript's scope: reporting uncertainty for the central R-squared comparisons and adding a passage-level analysis that directly tests the locus-of-influence interpretation. I would not reject the paper, but the current evidence is not yet strong enough for the strength of the headline claim. The fit with the CHR audience is good, and the appendix-level disclosure of preregistration departures is a model of good practice."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know about arXiv:2411.15068. First, it's a genuinely useful empirical comparison. The authors test three text representations—topic models, document embeddings, and word-level perplexity—across three corpora, and ask which aligns best with external social evidence: citations, critical discussion, and author age. The consistent top-quartile result, where representing a document by its most forward-looking 25% of passages beats the full-document average, holds in every cell of Table 1. Second, the interpretation that impact is concentrated in the most forward-looking passages is an inference, not a demonstrated fact. Nothing in the data shows that the passages the metric flags are the ones that citing texts actually quote or engage.\n\nWhat's actually new is the scale and the systematic comparison. This is the first head-to-head of lexical topic models, fine-tuned transformers, and perplexity-based prescience on full-text literary studies, economics, and fiction. The practical takeaway—neural embeddings do not systematically beat a simple 250-topic LDA model—is worth the price of admission. The authors also handle the main circularity risk carefully by excluding same-author and verbatim-quote chunks, and they report that this exclusion barely changes the results. The public GitHub repo and preregistration record are a welcome change for the field.\n\nNow the soft spots, in proportion. The stress-test concern is real: the top-quartile advantage could be a document-level confound. Prominent documents may contain a few future-sounding passages for reasons unrelated to the specific content that drew responses. The author-age result weakens a pure prestige story, but it doesn't tell you whether the high-precocity passages are the socially salient ones. Second, Table 1 reports R-squared values without confidence intervals or any correction for multiple comparisons across 30 method-outcome cells. Several raw differences are small. Third, there is no control for document length or number of chunks, which could mechanically inflate the top-quartile mean for longer works. These are fixable with a confirmatory analysis.\n\nWho benefits: computational humanists, bibliometricians, and anyone measuring novelty or 'ahead-of-the-curve' signals in large text corpora. The paper deserves a serious referee. The data, code, and preregistration are all public, and the methodological guidance is useful. I would accept it for review with a request that the authors add significance testing or confidence bounds and, ideally, a passage-level validation using citation contexts. If the flagged passages survive that test, the claim is solid.","headline":"A solid empirical benchmark for textual precocity, but the headline causal interpretation needs passage-level evidence and significance tests before I'd take the top-quartile claim as established.","tokens_in":10383,"tokens_out":4270,"would_cite":true,"duration_ms":40875,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that a text's precocity — how much more it resembles the future than the past — aligns with citation counts and authorial youth across three corpora, and that this alignment is strongest when only the most forward-looking…","keywords":["precocity","cultural change","textual novelty","topic models","document embeddings","perplexity","citations","authorial age"],"falsifier":"Take a corpus in which later adoption can be traced passage by passage, for example by recording which specific passages later articles quote or which specific books later authors imitate. If later works are not disproportionately influenced by the high-precocity chunks identified by the method, then the top-quartile alignment with citations and youth would not show that precocity locates the leading edge of cultural change. The paper reports the aggregate alignment but does not perform this passage-level check.","tokens_in":9402,"feed_emoji":"📈","tokens_out":8198,"duration_ms":71551,"temperature":0.7,"pith_summary":"Researchers want to measure cultural change from text, but different similarity measures give different answers and no absolute ground truth exists. This paper tests three text representations — topic models, document embeddings, and word-level perplexity — against three social signals: citation counts, authorial youth, and critical discussion, across literary studies, economics, and fiction. Its central claim is that textual precocity, defined as looking more like later documents than earlier ones, aligns with these social signals in every corpus. The alignment is consistently strongest when a document is represented by the average precocity of its top quartile of passages rather than by all passages. That result suggests cultural impact is carried by a small forward-looking fraction of a text, not by sustained average novelty.","feed_headline":"A text's influence lives in its top quarter of passages","feed_subtitle":"Resembling the future more than the past, a text's top passages track citations and youth best.","key_machinery":"The load-bearing quantity is precocity, defined as novelty minus transience: novelty measures divergence from documents in the preceding twenty years, transience measures divergence from documents in the following twenty years, and a text with high precocity simply \"looks later than\" peers published in the same year. For perplexity, precocity is computed as $2\\cdot(\\mathrm{ppl}_{\\mathrm{past}}-\\mathrm{ppl}_{\\mathrm{future}})/(\\mathrm{ppl}_{\\mathrm{past}}+\\mathrm{ppl}_{\\mathrm{future}})$. Documents are cut into passages of about 512 tokens, each passage is characterized by its relation to the surrounding time window, and a document is represented either by the mean of all passages or by the mean of the top quartile of passages by precocity. The top-quartile operation is what carries the main conclusion: it concentrates the measure on a text's most forward-looking moments.","core_discovery":"On the paper's own terms, the discovery is that precocity is a measurable property of documents and that it tracks independent social evidence of cultural change. For every corpus and every representation, works by highly cited authors and by younger authors are textually ahead of the curve, even though citations and youth are anti-correlated in the data. The paper's most specific finding is that this alignment is stronger when a text is summarized by the mean precocity of its top 25% of passages than when all passages are averaged; the authors report that this held \"in all of the tests we ran.\" They also report no systematic advantage for neural embeddings over lexical topic models, and treat topic models as a strong baseline.","pith_inferences":["We infer that if influence concentrates in the most precocious passages, later works should quote or cite those passages disproportionately; a passage-level citation-tracing study could test this directly.","We infer the same top-quartile logic could be transferred to patents, scientific articles, or policy documents to ask whether breakthrough recognition is driven by a few forward-looking sections.","We infer that prestige-based evaluation may systematically underweight young authors whose innovation is concentrated in short anticipatory moments, since averaging hides the leading edge.","We infer a possible mechanism worth testing: the top-quartile passages may be structurally distinctive, such as introductions, conclusions, or theoretical framing, a pattern the paper did not analyze."],"forward_implications":["Averaging all passages dilutes the signal; later studies of textual change should consider measuring documents at their most precocious passages.","Lexical topic models, not just neural embeddings, can capture the leading edge; researchers should compare against such models before adopting embeddings.","Because precocity correlates with both citations and youth despite the two being anti-correlated, the measure is apparently not just tracking prestige.","Including both past and future information matters; the future-only prescience measure achieved less than half the explanatory power of the expanded version.","Journal-level and author-level plots of precocity can reveal editors or authors who anticipate later trends even when citation counts do not distinguish them."],"supporting_citations":[{"why":"Supplies the bibliometric finding that novelty predicts citation impact, motivating citations as a social signal.","marker":"[1]"},{"why":"Defines prescience by comparing sentence perplexity in present and future models; the paper extends this to include the past.","marker":"[2]"},{"why":"Introduces novelty, transience, and the composite resonance quantity on which precocity is modeled.","marker":"[3]"},{"why":"Explains the Matthew effect, the reason prominence alone cannot stand in for cultural influence.","marker":"[5]"},{"why":"Provides author birth years for the fiction corpus and the cohort-succession argument for treating youth as a signal of change.","marker":"[8]"},{"why":"Supplies the full-text journal corpora for literary studies and economics.","marker":"[9]"},{"why":"Supplies the citation counts used as the prominence variable.","marker":"[10]"},{"why":"Provides the topic-modeling algorithm used for the lexical representation of texts.","marker":"[13]"},{"why":"Provides the sentence-embedding architecture used to fine-tune document embeddings.","marker":"[14]"},{"why":"Provides the masked language model used to compute perplexity for past and future periods.","marker":"[20]"}],"fun_headline_variants":["Top quarter of passages marks cultural change","Text's forward-looking quarter predicts influence","Cultural change lies in a text's top passages","Leading edge of culture is in top passages","Precocity: top 25% of text tracks influence"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The argument depends on treating citation counts, authorial youth, and critical discussion as genuine signals of being on the leading edge of cultural change; if those variables mostly track prestige or visibility rather than influence, the alignment would not demonstrate what the paper claims.","fun_headline_variants_meta":{"raw":{"variants":["Top quarter of passages marks cultural change","Text's forward-looking quarter predicts influence","Cultural change lies in a text's top passages","Leading edge of culture is in top passages","Precocity: top 25% of text tracks influence"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000128,"raw_usage":{"total_tokens":1044,"prompt_tokens":794,"completion_tokens":250,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":410,"completion_tokens_details":{"reasoning_tokens":181}},"tokens_in":410,"tokens_out":250,"duration_ms":3056,"temperature":1.0,"reasoning_tokens":181,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T14:32:49.517609+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a corpus in which later adoption can be traced passage by passage, for example by recording which specific passages later articles quote or which specific books later authors imitate. If later works are not disproportionately influenced by the high-precocity chunks identified by the method, then the top-quartile alignment with citations and youth would not show that precocity locates the leading edge of cultural change. The paper reports the aggregate alignment but does not perform this passage-level check.","supporting_citations":[{"cited_title":"Zhang, Q","cited_arxiv_id":null,"evidence_quote":"Supplies the bibliometric finding that novelty predicts citation impact, motivating citations as a social signal."},{"cited_title":"Vicinanza, A","cited_arxiv_id":null,"evidence_quote":"Defines prescience by comparing sentence perplexity in present and future models; the paper extends this to include the past."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Explains the Matthew effect, the reason prominence alone cannot stand in for cultural influence."},{"cited_title":"Underwood, K","cited_arxiv_id":null,"evidence_quote":"Provides author birth years for the fiction corpus and the cohort-succession argument for treating youth as a signal of change."},{"cited_title":"Burns, A","cited_arxiv_id":null,"evidence_quote":"Supplies the full-text journal corpora for literary studies and economics."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the topic-modeling algorithm used for the lexical representation of texts."}],"review_version":1}