{"id":"04102a6e-920a-4d1a-bf02-7184e4041b2d","arxiv_id":"1908.02951","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Using a Tobit gravity model on 484,903 biomedical papers, the authors report that leadership flow between Chinese institutions is shaped by all five proximity dimensions, with mixed evidence that these effects are weakening.","lead":"This paper measures 'research leadership' by the corresponding author's institution and models how it flows between Chinese universities using a gravity model with five kinds of proximity. It finds that geography, shared topics, same-province status, prior collaboration, and funding gaps all correlate with who leads whom, but the claim that these effects are declining is only partly supported.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Economic-proximity result flips sign in the paper's own robustness check, and the 'declining proximities' claim contradicts Table 7; the central all-five-proximities claim is not robust as stated.","rationale":"The reader's weakest assumption is construct validity of corresponding authorship. That is a genuine threat, but it depends on an external judgment about Chinese evaluation incentives. A more decisive problem is internal: the paper's own Table 10 reverses the sign of the economic-proximity coefficient from Table 6, and Table 7 contradicts the abstract's 'declining recently' generalization (cognitive and institutional proximity coefficients increase, economic proximity becomes significant only in the later period). These are not matters of interpretation; the reported numbers themselves conflict with the stated conclusions. The proposed re-estimation isolates the source of the economic sign flip and would settle whether the economic-proximity result is a measurement artifact or a specification artifact. The verdict should remain conditional: the core gravity-model idea and the descriptive network analysis are useful, but the headline claims need to be corrected and re-estimated before acceptance. If the sign flip is resolved and the sub-period claims are rewritten to match Tables 7 and 9, the paper could be accepted; as written it is not.","tokens_in":19148,"tokens_out":8122,"duration_ms":88782,"concrete_test":"Re-estimate Table 6 Model 5 using the full-count dependent variable from Section 5 with the same Tobit estimator, and re-estimate Table 10 Model 5 using the fractional dependent variable from Section 3 with a two-part model (logit hurdle plus OLS on positive values). If ln(Econprox) remains negative under Tobit on full counts, the economic result depends on the counting convention; if it turns positive, the sign flip is estimator-driven. The five-proximity claim is supportable only if the sign is stable across at least one of these crossed specifications.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central assertion that all five proximities are important determinants is not supported as stated because the paper's own robustness check reverses the economic-proximity result. In the main fractional-count Tobit model (Table 6, Model 5), ln(Econprox) is +0.277***, interpreted as center-periphery complementarity. In Section 5, switching to full counting and a zero-inflated negative binomial (Table 10, Model 5), the same variable is -0.030***. The text says this confirms the same conclusions, but for economic proximity the sign flips. Because the alternative specification changes both the dependent variable (fractional vs full count) and the estimator (Tobit vs ZINB), the two results cannot adjudicate each other and neither can be called a robustness check of the other. The 'in general, declining' headline is also contradicted by Table 7: cognitive proximity rises from 83.603 to 84.956, institutional proximity rises from 4.592 to 4.970, and economic proximity moves from 0.069 (n.s.) to 0.102***. At minimum, the qualitative claim should be restricted to geographical and social proximity and to the leadership-mass gap.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper defines 'research leadership' as the affiliation of the corresponding author, constructs a directed institution-level flow network from 484,903 life-sciences-and-biomedicine articles (with additional fields in Section 4.4), and estimates Tobit gravity models in which leadership masses and five proximity dimensions explain leadership flows among 244 Chinese institutions. It also reports sub-period estimates to study the evolution of proximity effects and presents a zero-inflated negative binomial model as a robustness check. The central claims are that leadership mass of both partners and geographical, cognitive, institutional, social, and economic proximity are all important determinants of research-leadership flow, and that the effects of these proximities have generally been declining.","tokens_in":19351,"tokens_out":5332,"duration_ms":53914,"significance":"If the estimates hold as stated, the paper would make a useful contribution by measuring directed research-leadership flows at the institutional level, jointly examining all five proximity dimensions, and providing a dynamic sub-period comparison. The strengths include a large bibliographic dataset, an explicit directed-network construction, the use of lagged regressors to mitigate endogeneity, and the inclusion of multiple research fields. The main empirical results for leadership mass and for geographical, cognitive, institutional, and social proximity are qualitatively plausible and internally consistent. However, the paper's central claims as written are not fully supported by its own tables: the economic-proximity result reverses sign under the stated robustness check, and the claim of generally declining proximity effects is contradicted by the sub-period coefficients for cognitive and institutional proximity. These issues are load-bearing because they concern the abstract's headline findings.","major_comments":[{"comment":"The robustness check reverses the sign of economic proximity: the main Tobit model in Table 6 Model 5 reports ln(Econprox) = +0.277***, while the zero-inflated negative binomial model in Table 10 Model 5 reports -0.030*** in the count equation (and -0.104*** in the inflation equation), yet the text says the alternative model 'leads to the same conclusions and confirms the robustness of the main model.' Because the alternative specification changes both the dependent variable (fractional vs. full count) and the estimator (Tobit vs. ZINB), it cannot adjudicate the direction of the economic-proximity effect. The paper should either provide a one-change-at-a-time robustness analysis, or explicitly acknowledge that the economic-proximity finding is sign-sensitive and restrict the corresponding conclusion.","section":"Section 5, Table 10 Model 5"},{"comment":"The abstract claims that 'in general, the effect of these proximities for research leadership flow has been declining recently,' and Section 6 repeats that 'the effects of these proximities have been declining recently.' This is contradicted by Table 7: cognitive proximity increases from 83.603 to 84.956, institutional proximity increases from 4.592 to 4.970, and economic proximity moves from 0.069 (not significant) to 0.102***. Section 4.3 itself acknowledges that cognitive and institutional coefficients increased and that economic proximity became significant. The qualitative claim should be restricted to geographical and social proximity (and to the narrowing leadership-mass gap), or the abstract and conclusions must be revised to reflect the mixed pattern.","section":"Abstract and Section 6 vs. Table 7"},{"comment":"The time-lag construction for the sub-period models is ambiguous. The main model in Section 3.3 uses independent variables measured over 2008-2012 to predict 2013-2017 flows, but Section 4.3 states that 'we take 2-year lagged independent variables' for the 2013-2014 and 2016-2017 sub-periods. No detail is given about whether the sub-period models use 2011-2012 and 2014-2015 measurements respectively, or whether they reuse the 2008-2012 variables. This matters because the sub-period comparison is the basis for the evolution claims, and a mismatch between the lag structure and the definitions in Table 2 would affect the interpretation of the Chow test results.","section":"Section 4.3"},{"comment":"The paper equates research leadership with corresponding authorship, and Section 2.2 justifies this by noting that corresponding authorship is the primary criterion in China's research evaluation system. Given those same incentives, corresponding authorship may be a strategic credit-allocation choice rather than a measure of intellectual leadership. The manuscript should either rename the construct 'corresponding-author-based leadership,' add an explicit validity discussion, or temper the policy conclusions in Section 6 that presuppose that the measured flows reflect leadership rather than administrative or incentive-driven attribution.","section":"Section 2.2 and Section 6"}],"minor_comments":[{"comment":"The cumulative percentage column appears internally inconsistent (for example, the Sun Yat Sen Univ row shows a cumulative value of 3.82% rather than the expected 11.24% after the previous rows). Please verify the cumulative calculations and the table formatting.","section":"Table 3"},{"comment":"The text says 'we examine the effect of all the four proximity dimensions (geographical, cognitive, institutional, and social proximity)' and then adds economic proximity; the wording should be corrected to say five dimensions, with economic proximity included in the list.","section":"Section 2.1"},{"comment":"The sentence on institutional proximity is garbled: 'The institutional proximity is positively significant, indicating that positive and significant, showing that...' should be rewritten for clarity.","section":"Section 4.2"},{"comment":"The variables ln(Cognprox) and ln(Econprox) are undefined if the cosine similarity is zero or if the absolute difference in NSFC project counts is zero. The paper should state whether such cases occur and how they are handled (for example, by adding a small constant before taking logs).","section":"Section 3.3 and Table 2"},{"comment":"The reference to 'Li T. (2002). Econometric Analysis of Cross Section and Panel Data' appears to be a misattribution; this is Wooldridge's book. Please correct the citation.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The paper is within the scope of a scientometrics journal and has a substantial dataset, but the abstract and conclusions currently overstate what the tables show. The sign reversal in the robustness check and the contradiction in the 'declining' claim are not merely presentational; they affect the headline findings. I would be willing to review a revised version that reconciles these issues. I would also suggest the authors verify the 'first quantitative research' novelty claim in the abstract against the cited and related literature before publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First, the load-bearing part: the paper's core contribution is the directed, institution-level 'research leadership flow' measure built from corresponding-author affiliations. That is a genuine extension over the undirected, static proximity literature (Fernandez et al. 2016; Plotnikova and Rake 2014), and the Tobit gravity model with lagged independent variables is a sensible way to test it. The main result—leadership mass and the five proximities are associated with RL flow—is mostly supported by Table 6, and the qualitative pattern across fields is coherent. Credit where due: the LDA-based cognitive proximity is not circular; social proximity is a lagged dummy, so it acts as a persistence control, not a definitional identity. The circularity burden is low.\n\nSoft spots, in proportion. The stress-test note is right on both counts. The abstract's claim that proximity effects are 'in general... declining' is contradicted by Table 7: cognitive and institutional proximity coefficients increase between 2013-2014 and 2016-2017, and economic proximity goes from insignificant to positive. The conclusion text is more careful, but the abstract and the opening of the conclusion repeat the blanket claim. That needs fixing.\n\nThe bigger issue is the robustness check. Section 5 says the ZINB with full counts 'leads to the same conclusions,' but Table 10 Model 5 shows ln(Econprox) at -0.030***, while the main Tobit has +0.277***. That is a sign flip, and because both the dependent variable scaling and the estimator change at once, the ZINB is not a clean robustness check for the economic-proximity result. The paper should either explain the discrepancy or drop the all-five strong claim.\n\nMinor: the 'first quantitative research' framing is a bit strong; the 'Chow test' comparing two Tobit models with different censoring fractions is not a standard Chow test, though it's not fatal. Data/code are not available, which limits reproducibility. And the leadership proxy—corresponding-author affiliation—is contestable under China's evaluation incentives, but the authors explicitly acknowledge that this is the institutional rule, so it's a defensible choice rather than a hidden flaw.\n\nVerdict: the core descriptive finding is probably reliable, but the headline claims need revision. This paper deserves serious peer review—it's a useful empirical contribution for scientometrics and science-policy audiences, especially those studying Chinese institutions. I would cite it once the robustness and declining-effects issues are addressed. Bring it to reading group as a case study in how an otherwise solid gravity model can be undermined by overstatement.","headline":"A real, useful new directed leadership-flow measure, but the abstract overclaims robustness and the 'declining proximity' headline is contradicted by the paper's own sub-period table.","tokens_in":19913,"tokens_out":3218,"would_cite":false,"duration_ms":29269,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A Tobit gravity model of 484,903 multi-institution Chinese papers shows that research leadership flows are directed, weighted, and governed by institutional leadership mass plus five proximity dimensions.","keywords":["research leadership","corresponding author","proximity dimensions","gravity model","Tobit regression","collaboration networks","Chinese institutions","scientometrics"],"falsifier":"Take a random sample of, say, 500 Chinese multi-institution life-sciences papers from 2013–2017, code leadership from explicit author-contribution statements or principal-investigator roles, and re-estimate the Tobit gravity model on that recoded dependent variable. If the five proximity coefficients do not keep their signs, the corresponding-author definition is doing the work, not leadership itself.","tokens_in":18907,"feed_emoji":"🧲","tokens_out":7550,"duration_ms":74389,"temperature":0.7,"pith_summary":"The paper is trying to establish that research leadership in multi-institution papers is a directed, measurable flow: the corresponding author's institution sends a share of leadership to every other participating institution. Using a Tobit regression-based gravity model on five years of Chinese publications, it argues that the leadership mass of both the leading and participating institutions, together with geographical, cognitive, institutional, social, and economic proximity, all shape how much leadership flows between institutions. If true, this turns the informal notion of who leads a project into a quantitative object that can be tracked across time and space. The authors further claim that most of these proximity effects weakened from 2013 to 2017, so Chinese collaboration networks are becoming flatter. That would matter for grant policy because it identifies which institutional pairings need support and which proximities are fading or strengthening.","feed_headline":"Gravity model on 484,903 papers maps who leads China's research","feed_subtitle":"Leadership flows follow distance, cognitive fit, province, prior ties, and funding gaps.","key_machinery":"The central object is 'research leadership mass,' built from a directed star network: each paper contributes one unit of leadership, and the corresponding author's institution—or several such institutions splitting the unit—distributes equal fractions to every participating institution. The gravity model then uses the log leadership masses of leader and participant as the two mass terms, the five proximity measures as distance-like terms, and Tobit regression to treat zero flows as left-censored observations. This combination converts authorship roles into a weighted directed network and lets the coefficients quantify which proximity pulls leadership hardest and how that pull changes between sub-periods.","core_discovery":"The central claim is that leadership flow obeys a gravity pattern. Direction matters: the leading institution's accumulated leadership mass has a larger coefficient than the participating institution's mass in every model, and the gap narrows over time. Geographical distance suppresses flow; cognitive similarity, same-province status, prior collaboration, and absolute gaps in national research funding increase it. The economic-gap effect is absent in 2013–2014 and positive in 2016–2017, which the authors read as consistent with center-periphery collaboration. The same signs hold in Technology, Physical Sciences, and HASS, and a zero-inflated negative binomial check on full-count flows reproduces the main results.","pith_inferences":["A natural test is to swap the dependent variable for one built from explicit author-contribution statements; if the proximity coefficients change sign or size, corresponding authorship is not a neutral proxy for leadership.","Because corresponding authorship is a reporting convention in China's evaluation system, part of the declining proximity effects could reflect shifts in authorship etiquette rather than real changes in collaboration; contribution-level data would separate those.","The same gravity specification could be applied to grant leadership (the applicant institution) or patent leadership; different 'leadership currencies' may show different proximity elasticities.","The model is estimated only on flows within China; running it on international corresponding-author flows would reveal whether the Chinese pattern is specific or general."],"forward_implications":["Institutions that lead many projects gain a compounding advantage, since current leadership mass predicts future leadership flows for both leader and participant.","Distance still blocks leadership, but increasingly less so, so large remote collaborations should keep becoming more common.","Cognitive similarity is a rising force, meaning funding bodies may get more leadership flow from pairing institutions with overlapping research profiles.","Funding-resource gaps began attracting leadership flow after 2014, so matching resource-rich with resource-poor institutions should stimulate collaboration.","Same-province and prior-collaboration effects remain positive but decline, so cross-provincial and first-time partnerships are where policy intervention has the most room."],"supporting_citations":[{"why":"Supplies the five-proximity taxonomy that structures the independent variables.","marker":"Boschma (2005)"},{"why":"Provides the Tobit gravity-model template for collaboration counts with many zeros.","marker":"Plotnikova and Rake (2014)"},{"why":"Establishes the institution-level proximity measures and the lagged-variable design that the paper adapts.","marker":"Fernandez, Ferrandiz et al. (2016)"},{"why":"Basis for geographical distance effects and the gravity framing of collaboration intensity.","marker":"Hoekman, Frenken et al. (2010)"},{"why":"Establishes the gravity-model approach for collaborative knowledge production in China.","marker":"Scherngell and Hu (2011)"},{"why":"Source of the fractional-count weighting used to measure collaboration intensity.","marker":"Jiang, Zhu et al. (2018)"},{"why":"Justifies treating the corresponding author as research leader and notes China's evaluation-system reliance on corresponding authorship.","marker":"Hu, Rousseau et al. (2010)"},{"why":"Provides the test used to show sub-period coefficients differ significantly.","marker":"Chow (1960)"}],"fun_headline_variants":["Gravity model explains China's research leadership flows","Proximity shapes research leadership across Chinese institutions","Leadership mass and five proximities drive research ties","Distance, cognition, institution, social, funding steer leadership","Who leads Chinese research? Gravity and proximity decide"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole network rests on treating the corresponding author's institution as the true leader of every multi-institution paper; if that label is a strategic or administrative choice rather than real intellectual leadership, then the flows, masses, and proximity effects describe a different thing than research leadership.","fun_headline_variants_meta":{"raw":{"variants":["Gravity model explains China's research leadership flows","Proximity shapes research leadership across Chinese institutions","Leadership mass and five proximities drive research ties","Distance, cognition, institution, social, funding steer leadership","Who leads Chinese research? Gravity and proximity decide"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000153,"raw_usage":{"total_tokens":1180,"prompt_tokens":891,"completion_tokens":289,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":507,"completion_tokens_details":{"reasoning_tokens":216}},"tokens_in":507,"tokens_out":289,"duration_ms":3583,"temperature":1.0,"reasoning_tokens":216,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T14:29:08.938313+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a random sample of, say, 500 Chinese multi-institution life-sciences papers from 2013–2017, code leadership from explicit author-contribution statements or principal-investigator roles, and re-estimate the Tobit gravity model on that recoded dependent variable. If the five proximity coefficients do not keep their signs, the corresponding-author definition is doing the work, not leadership itself.","supporting_citations":[],"review_version":1}