{"id":"4c1702e1-5588-4698-84ff-a2bbfdb11560","arxiv_id":"2506.22729","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"During the deep learning revolution, researchers who persisted in their prior topics suffered a citation-impact penalty, while moderate, selective pivoting was associated with the largest gains.","lead":"A study of 5,359 machine-learning researchers finds that after the 2012 deep learning breakthrough, those who stuck to their old research topics lost citation impact, while scientists who pivoted moderately gained the most. The paper argues that persistence, usually praised, became a liability during this paradigm shift.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The rigidity-penalty estimate may be an artifact of selective attrition from ICML/NeurIPS after 2013; the panel is unbalanced and the paper's own limitation admits only published papers were analyzed.","rationale":"The reader's weakest assumption is attrition and selective retention, and I agree that is the most load-bearing threat. The paper's headline effect is a conditional correlation within a non-randomly selected sample; unlike the magnitude typo (which is correctable), attrition can reverse the sign of the effect and is explicitly acknowledged in Section 6. I therefore focus on it. The proposed test is feasible because the authors already collected ICLR data for Appendix C. If the penalty persists when ICLR is included, the concern is mitigated; if not, the central claim fails. The scale error (0.1 × -0.177 does not equal a 1.77% decrease) is real and should be fixed, but it is a reporting error rather than a fundamental threat to the qualitative claim. The verdict remains CONDITIONAL, not REJECT, because the data and code could resolve the attrition issue with the suggested sensitivity analysis.","tokens_in":18933,"tokens_out":5395,"duration_ms":115856,"concrete_test":"Re-estimate the Table 1 model 6 specification on an expanded corpus that includes ICLR 2013-2022 publications for the same cohort, recomputing Research Persistence and Impact over ICML+NeurIPS+ICLR. If the post-2013 × persistence coefficient remains negative and significant (roughly -0.177), the penalty is not driven by venue-specific attrition; if it attenuates or flips, the penalty is an artifact of the restricted sample. As a secondary check, bound the missing years with Lee (2009) trimming or assign extreme counterfactual impact values (0 and 100) to authors with no publication in a given year; if the interaction is not robust to these bounds, the attrition channel cannot be ruled out.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The load-bearing condition for the central claim is that the two-way fixed-effects estimate of the post-2013 persistence penalty is identified from within-author variation in persistence that is uncorrelated with unobserved time-varying shocks, and that sample attrition is ignorable. This condition is least secure because both persistence and impact are observed only for scientist-years with at least one ICML/NeurIPS publication. After 2013, the venues shifted toward deep learning; authors who persisted in old topics were more likely to be rejected, to move to ICLR, or to leave the venue sample. Their observations drop out, so the negative coefficient on persistence×post in Table 1 is estimated only among authors who remained publishable in these venues. If remaining publishable is correlated with unobserved quality conditional on persistence, the penalty is biased; the sign could even be an artifact of selection. Section 6 acknowledges 'This research only considered published papers,' but no bound or sensitivity analysis is provided, so the central quantitative claim is not yet identified.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies how researchers' persistence -- measured as the bootstrapped text similarity between a scientist's current-year ICML/NeurIPS papers and their prior corpus -- relates to citation impact and productivity during the deep learning paradigm shift after 2013. Using a cohort of 5,359 scientists who published at ICML/NeurIPS between 2003 and 2012, the authors estimate two-way fixed effects panel regressions. They find a statistically significant negative interaction between persistence and the post-2013 period for impact (the 'rigidity penalty'), a positive association between persistence and productivity, and a negative association between prior impact and subsequent impact. They also report that previous productivity and large, older collaboration teams predict higher persistence. The paper concludes that persistence is context-dependent, that strategic partial adaptation is optimal, and that scientific disruptions redistribute influence.","tokens_in":19095,"tokens_out":7939,"duration_ms":89233,"significance":"If the central identification were credible, the paper would provide a valuable empirical contribution to the science-of-science literature: it would show that a clearly measured behavioral trait (textual persistence) has context-dependent returns during a paradigm shift, with implications for career strategy and research evaluation. The paper also offers a novel persistence measure that corrects for corpus-size differences via bootstrapping, and it combines macro-level venue-topic analysis with micro-level scientist panels. However, the main quantitative claim is threatened by sample-selection on the outcome and by a misreported effect size, so the significance as currently established is limited.","major_comments":[{"comment":"The interpretation of the interaction coefficient is incorrect. In Model 6, the coefficient on post2013×persistence is -0.177, with persistence measured on a 0-100 scale and impact measured as a 0-100 citation percentile. A 0.1-unit increase in persistence therefore corresponds to a change of -0.0177 percentile-rank points, not a 1.77% decrease. The current wording overstates the magnitude by a factor of 100 relative to the scale of the dependent variable. Please correct this sentence and any related claims in the abstract, main text, or discussion.","section":"Section 5.2.1, Table 1"},{"comment":"The panel regression is estimated only on scientist-year observations with at least one ICML/NeurIPS publication in that year, because both the persistence measure and the impact measure require a year-T+1 paper. After 2013, researchers who persisted in pre-deep-learning topics may have been more likely to be rejected at these venues, to move to ICLR, or to stop publishing in them; such years drop out of the sample. The negative interaction coefficient is therefore identified only among the selected group of scientists who continued to publish in ICML/NeurIPS, making the 'rigidity penalty' potentially an artifact of attrition. The paper's statement in Section 6 ('This research only considered published papers') acknowledges the issue but provides no sensitivity analysis. Please provide bounds, a selection correction, or at least a rigorous informal assessment of the direction and plausible magnitude of the resulting bias.","section":"Section 5.2.1 and Section 6"},{"comment":"The paper's policy-relevant conclusion that 'strategic adaptation' (moderate persistence, around 25-50% overlap) maximizes impact is not supported by the regression model. Table 1 contains only a linear persistence term and its interaction with the post-2013 indicator; it does not include a quadratic or spline term. Figure 4 is based on bivariate linear fits, not the multivariate specification. Please test for curvature explicitly in the main regression, or tone down the claim that the highest impact is achieved at an intermediate persistence level.","section":"Section 5.2.1 and Figure 4"},{"comment":"The construction of the 'bootstrapped text similarity' measure is not fully specified. The text does not state which text representation is used (e.g., TF-IDF vectors, word embeddings, or SPECTER2 embeddings) or which similarity metric (e.g., cosine, Jaccard, dot product) is applied. Appendix A describes SPECTER2 for the macro-level comparison, but it is not stated whether the same representation underlies the scientist-level persistence measure. Please provide a complete algorithmic definition, including the feature set, similarity function, and any preprocessing steps, so that the central independent variable is reproducible.","section":"Section 5.2.1, persistence measure"}],"minor_comments":[{"comment":"The R-squared for the impact model is 0.031, meaning the model explains roughly 3% of the variance in citation impact. While statistical significance is reported, the paper should acknowledge the small explanatory power when interpreting the economic significance of the persistence coefficient.","section":"Table 1"},{"comment":"The regression equation is not numbered, and the 'Post 2013' dummy is never explicitly defined. Please state whether it equals 1 for calendar years 2013 onward or for transitions into years after 2013, and clarify the timing of the persistence measure (from year T to T+1) relative to the outcome measured in year T+1.","section":"Section 5.2.1, regression equation"},{"comment":"The y-axis labels 'Impact Change' and 'Productivity Change' are undefined in the caption. The text mentions 'ratios of productivity and impact relative to the previous year,' but the units and the construction of these ratios should be stated explicitly in the figure caption.","section":"Figure 4"},{"comment":"The citation for SPECTER2 appears inconsistent: the text cites '[10, 47]' for the SPECTER2 model, but reference [10] is a Web Conference paper by Bao et al., not the Singh et al. SPECTER2 paper. Please verify all citations in Appendix A.","section":"Appendix A"},{"comment":"The limitations paragraph mentions that only published papers were considered, but it does not mention that persistence itself cannot be measured in years without publications. This is directly relevant to the attrition issue and should be stated explicitly.","section":"Section 6"}],"recommendation":"major_revision","confidential_remarks":"The paper has an interesting and timely research question, and the dataset is well suited to studying a concrete paradigm shift. However, the central claim is not yet identified because of the sample-selection issue, and the reported effect size is overstated. A major revision that adds selection sensitivity analyses and corrects the interpretation could make the contribution publishable. I would also recommend the editor ensure that the authors address the disconnect between the linear model and the 'strategic adaptation' inverted-U claim, since that overreach appears in the abstract and discussion."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know two things about this paper. First, it's a serious observational study of an interesting question: does persistence hurt researchers when a field undergoes a paradigm shift? The setting is clean (the AlexNet shock after 2012), the cohort is large (5,359 scientists who published in ICML/NeurIPS 2003–2012), and the persistence measure—bootstrapped text similarity between past and current papers—is reasonable. The main interaction (persistence × post-2013) is robustly negative across specifications, which indicates that within this cohort, staying close to old topics was associated with lower citation impact after deep learning took over. That is a real, new finding, distinct from the recent \"pivot penalty\" work, and worth taking seriously. The paper also does a nice job showing that high-status and high-productivity scientists were less likely to adapt, and that teams adapted through new collaborators. The authors are transparent about several limitations, which I respect.\n\nNow the soft spots, in proportion. The biggest one is the headline interpretation. The paper says a 0.1 increase in persistence explains a 1.77% decrease in impact. The coefficient in Table 1 (model 6) is −0.177 on a 0–100 impact scale. A 0.1 unit increase in persistence is therefore a 0.018 percentile-point decrease, not 1.77%. That is roughly a factor of 100 error, and it makes the effect look much larger than it is. With R² around 0.03, the model explains very little of the variance, so the substantive importance is modest, whatever the statistical robustness.\n\nThe second issue is attrition. Persistence and impact are observed only for scientist-years with at least one ICML/NeurIPS publication. After 2013, researchers who persisted in old topics were more likely to be rejected, move to ICLR, or leave these venues. If that attrition is correlated with unobserved quality or motivation, the interaction coefficient is identified only among survivors, and the sign could be an artifact. The paper's Section 6 admits \"only considered published papers,\" but there is no bound or sensitivity analysis. The stress-test note is right that this is the weakest identifying assumption. It is not fatal, but it is load-bearing.\n\nThird, the \"strategic adaptation\" inverted-U is asserted from a fitted figure, not formally tested. There is no quadratic term or breakpoint estimation, so that claim is more suggestive than established. And no code or data is provided, which makes it hard to check the persistence construction or reproduce the regressions.\n\nWho is this for? Science-of-science researchers and anyone studying careers under technological disruption. It's a useful hypothesis-generating study, not a definitive estimate. The interaction is real, but the headline magnitude is wrong, and the selection issue is unresolved. A serious editor should send this to peer review: the question is important, the design is mostly sound, and referees can request the corrections that would make it a solid contribution. I wouldn't cite the numerical claims in their current form, but I would absolutely bring this to a reading group for discussion.","headline":"The rigidity penalty is plausible and the study is honest, but the headline magnitude is off by 100x and the attrition channel is not bounded.","tokens_in":19670,"tokens_out":2319,"would_cite":false,"duration_ms":30454,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"During a paradigm shift like the deep learning revolution, persistence in one's established research direction measurably lowers citation impact, while partial pivoting that retains some thread to prior work maximizes it.","keywords":["persistence paradox","rigidity penalty","strategic adaptation","paradigm shift","deep learning revolution","scientific impact","career trajectories","citation percentile rank"],"falsifier":"The decisive check is to bring leavers back into the outcome: re-estimate the persistence-by-post-2013 interaction on a panel that adds this cohort's ICLR papers, their publications in other machine-learning and AI venues, and eventually their submitted-but-rejected papers. If the negative coefficient shrinks or vanishes once leavers are included, the rigidity penalty is an artifact of selective retention; if it grows, the conclusion is reinforced. A quicker placebo check on the existing data is to set the fake shift at, say, 2008 and confirm that persistence shows no comparable negative interaction in a period without a paradigm shift.","tokens_in":18664,"feed_emoji":"📉","tokens_out":16772,"duration_ms":144061,"temperature":0.7,"pith_summary":"Persistence is usually treated as a virtue in science, and this paper argues that its value is context-dependent: during a paradigm shift it becomes a measurable liability. Tracking more than 5,000 researchers who published in the top machine learning venues in the decade before AlexNet, the study shows that after 2013, the more a scientist's new output resembled their old output, the lower their citation percentile rank, while moderate pivoting brought the largest gains. Persistence kept raising productivity but began lowering impact, and the paper reads this as evidence that scientific breakthroughs act as mechanisms that reconfigure power structures within a field.","feed_headline":"A 0.1 rise in persistence explains a 1.77% drop in impact","feed_subtitle":"After deep learning's rise, staying close to old research cost citations at top ML venues; partial pivoting paid most.","key_machinery":"The workhorse is the research persistence index, a bootstrapped text-similarity score per scientist per year: the corpus of a scientist's ICML/NeurIPS papers up to year T is compared with their papers in year T+1, randomly sampling the larger corpus down to the smaller size 1,000 times and averaging the similarities, so a higher score means the scientist's latest work closely resembles their own past work. This index enters a two-way fixed-effects panel regression of impact and productivity on persistence, its interaction with the post-2013 period, prior performance, and coauthor-network features (count, familiarity, status). The persistence-by-post-2013 interaction coefficient is the evidence for the rigidity penalty, and the peak-in-the-middle pattern of impact against persistence identifies the strategic-adaptation zone.","core_discovery":"The paper's central claim is that in the era following the advent of deep learning, rigidity became costly: holding coauthor networks and prior performance fixed, a 0.1 increase in a scientist's year-over-year similarity to their own prior research explains a 1.77% marginal decrease in impact, measured by citation percentile rank at ICML and NeurIPS. Persistence remained positively associated with productivity, so the penalty is specific to influence rather than output. The paper also documents a redistribution of status: previous impact negatively predicts subsequent impact in this period, and scientists who were prolific or embedded in older, larger teams adapted more slowly. The maximum impact gain sits at moderate persistence - roughly 25-50% overlap with prior work - a zone the authors call strategic adaptation, which selectively adopts the new paradigm while keeping weak ties to old expertise.","pith_inferences":["Testable extension: apply the same within-author persistence design to later paradigm shifts - the large-language-model wave after 2018, or the CRISPR revolution in biology - to see whether the rigidity penalty is a general feature of scientific revolutions or specific to the deep-learning transition.","Implication the authors leave implicit: if rigidity reliably costs impact during revolutions, then evaluation and funding systems that reward uninterrupted productivity are systematically biased toward the strategy that loses influence in exactly those periods, and crediting adaptation during shifts would offset that bias.","Refinement with the same data: the persistence index is computed from published text, so a scientist who changes vocabulary without changing methods, or switches methods while keeping familiar vocabulary, is misclassified; separating topical similarity from methodological similarity would sharpen the 1.77% estimate."],"forward_implications":["A researcher who keeps their output highly similar to their own prior work after a paradigm shift will see citation percentile rank decline even when productivity and collaboration inputs are held constant.","Impact is redistributed in a revolution: high prior impact predicts lower subsequent impact in the same venues, so the scientists most established before 2013 are not the ones who benefit most after it.","The best-impact adaptation is partial: the largest gains occur at a 25-50% overlap with prior work, not at full continuity and not at a complete break.","Elite venues are not neutral: ICML and NeurIPS converged on the new deep-learning topics over time, so publishing in the same venue is not the same as staying in the same paradigm.","Older and larger collaboration teams lag in adaptation, and teams that take on new collaborators after the shock sustain their success."],"supporting_citations":[{"why":"The AlexNet breakthrough paper that defines the paradigm-shift event separating the pre- and post-deep-learning eras in the sample.","marker":"[32]"},{"why":"The Nature review that identifies AlexNet as the catalyst and fixes 2013 as the first year of the deep learning revolution.","marker":"[36]"},{"why":"Supplies the account of normal science and scientific revolutions that motivates why persistence turns costly when paradigms break.","marker":"[33]"},{"why":"The prior finding that topic-switching effects depend on career stage, which this paper extends by showing the payoff also depends on the paradigm environment.","marker":"[64]"},{"why":"The empirical baseline claiming adaptive researchers face a pivot penalty, the received view this paper's rigidity-penalty result contrasts with.","marker":"[27]"},{"why":"Evidence that removing a field's dominant leader opens room for new generations, supporting the redistribution-of-status mechanism.","marker":"[4]"},{"why":"The cumulative-advantage mechanism that explains why previously successful scientists and old teams persist rather than adapt.","marker":"[7]"},{"why":"The theory that disruptive events reallocate opportunities and status, which frames the paper's redistribution findings.","marker":"[65]"}],"fun_headline_variants":["In AI era, rigid scientists lose citation edge","Persistence costs citations after paradigm shifts","Stubbornness hurts: 0.1 persistence links to 1.77% citation loss","Sticking to old research costs ML scientists citations","Why persistence failed after deep learning took over"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise, which the paper itself flags in Section 6, is that the penalty is estimated only among researchers who kept getting published at ICML and NeurIPS - if the scientists who stopped appearing there (moving to ICLR, journals, or industry, or being rejected) are exactly the ones persistence hurt most, the measured penalty could reflect who stayed in the sample rather than what persistence does to impact.","fun_headline_variants_meta":{"raw":{"variants":["In AI era, rigid scientists lose citation edge","Persistence costs citations after paradigm shifts","Stubbornness hurts: 0.1 persistence links to 1.77% citation loss","Sticking to old research costs ML scientists citations","Why persistence failed after deep learning took over"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000556,"raw_usage":{"total_tokens":2630,"prompt_tokens":909,"completion_tokens":1721,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":525,"completion_tokens_details":{"reasoning_tokens":1643}},"tokens_in":525,"tokens_out":1721,"duration_ms":14345,"temperature":1.0,"reasoning_tokens":1643,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T22:00:14.273316+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"The decisive check is to bring leavers back into the outcome: re-estimate the persistence-by-post-2013 interaction on a panel that adds this cohort's ICLR papers, their publications in other machine-learning and AI venues, and eventually their submitted-but-rejected papers. If the negative coefficient shrinks or vanishes once leavers are included, the rigidity penalty is an artifact of selective retention; if it grows, the conclusion is reinforced. A quicker placebo check on the existing data is to set the fake shift at, say, 2008 and confirm that persistence shows no comparable negative interaction in a period without a paradigm shift.","supporting_citations":[{"cited_title":"Sutskever, and G","cited_arxiv_id":null,"evidence_quote":"The AlexNet breakthrough paper that defines the paradigm-shift event separating the pre- and post-deep-learning eras in the sample."},{"cited_title":"Bengio, and G","cited_arxiv_id":null,"evidence_quote":"The Nature review that identifies AlexNet as the catalyst and fixes 2013 as the first year of the deep learning revolution."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the account of normal science and scientific revolutions that motivates why persistence turns costly when paradigms break."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The prior finding that topic-switching effects depend on career stage, which this paper extends by showing the payoff also depends on the paradigm environment."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The empirical baseline claiming adaptive researchers face a pivot penalty, the received view this paper's rigidity-penalty result contrasts with."},{"cited_title":"Fons-Rosen, and J","cited_arxiv_id":null,"evidence_quote":"Evidence that removing a field's dominant leader opens room for new generations, supporting the redistribution-of-status mechanism."},{"cited_title":"Stuart, and Y","cited_arxiv_id":null,"evidence_quote":"The cumulative-advantage mechanism that explains why previously successful scientists and old teams persist rather than adapt."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The theory that disruptive events reallocate opportunities and status, which frames the paper's redistribution findings."}],"review_version":1}