{"id":"cc1eb3c4-605c-4b9b-9c8f-d2b395934b72","arxiv_id":"1908.03813","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"high","formal_verification":"none","parameter_count":4,"one_line_summary":"Informal mentorship by highly cited senior coauthors is reported to raise junior researchers' later citation impact, with effects varying by number, age, and gender of mentors, though the paper's own tables cast doubt on the direction of the effect.","lead":"This study uses millions of coauthor pairs to estimate whether working with highly cited senior scientists improves a junior researcher's later citation impact. It reports a large mentorship benefit and gender-specific effects, but the paper's own supplementary tables may show the opposite direction.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The paper's own Supplementary Tables S3-S10 show the treatment group (higher big-shot mentors) has lower mean post-mentorship impact than controls in every comparison, contradicting the reported positive δ and the central claim.","rationale":"The reader's rejection is based on the same internal contradiction, and that is the most load-bearing concern. The stated weakest assumption, unconfoundedness, is secondary: causal identification cannot rescue a claim whose own matched means point in the opposite direction. My concern is therefore the sign of the effect in the paper's primary tables, not the sensitivity of the causal estimate. The reader's rationale identifies this sign inconsistency as critical, but their formal 'weakest_assumption' names unconfoundedness, so my agreement is partial. I recommend no change to the reader's verdict: REJECT remains appropriate because the central quantitative claim is contradicted by the supplementary evidence as printed. A revision would need to correct the sign error, clarify whether the table columns are mislabeled, and then address confounding before the causal claim could be considered.","tokens_in":18586,"tokens_out":3187,"duration_ms":35849,"concrete_test":"Recompute all δ values from the reported imp(C') and imp(T') columns in Supplementary Table S3 using the formula in Supplementary Note 3, δ = 100·(imp(T') − imp(C'))/imp(C'). For the first row, Q1 vs. Q2, verify that (11.66 − 14.86)/14.86 = −21.5%, not +27.4%. Repeat for all rows of Tables S3 and S4, and for one stratified table such as S5. If the recomputed δ values are negative wherever imp(T') < imp(C'), the positive big-shot effect and the magnitude 'up to 36%' are unsupported by the reported data.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim requires that higher mentorship quality increases post-mentorship impact. The paper's primary evidence, Supplementary Table S3, is internally inconsistent with that claim. The table defines C = Q_i and T = Q_{i+1}, reports imp(C') and imp(T'), and defines δ = 100·(imp(T') − imp(C'))/imp(C'). In every row of Table S3, imp(C') > imp(T'): for Q1 vs. Q2, imp(C')=14.86 and imp(T')=11.66, giving δ = −21.5%, yet the table reports δ = +27.4%. Similar reversals occur in every row: Q2 vs. Q3 gives −14.1% but is reported as +16.4%; Q3 vs. Q4 gives −17.7% but is reported as +21.5%; Q4 vs. Q5 gives −26.2% but is reported as +35.5%. The reported positive values appear to be computed with the inverted formula 100·(imp(C') − imp(T'))/imp(T'), which reverses the sign. The same pattern appears in the hub-effect table (Supplementary Table S4) and in the stratified big-shot tables S5–S10: in every comparison, the higher-exposure group has lower post-mentorship impact. This is not a minor numerical typo; it reverses the direction of the headline causal effect. Even if unconfoundedness were fully satisfied, the matched means as presented show a negative association between big-shot experience and post-mentorship impact. The causal language in the abstract and Results therefore is not supported by the paper's own evidence.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper investigates whether informal mentorship by senior coauthors affects junior scientists' later independent impact. Using Microsoft Academic Graph data, the authors define protégé-mentor pairs, measure mentorship quality by mentors' prior citation impact (\"big-shot experience\") and prior collaboration degree (\"hub experience\"), and use coarsened exact matching between adjacent quintiles of these measures to estimate effects on the protégé's average five-year citation impact after mentorship. The paper reports positive causal effects (up to 36% for high-status mentors), increasing with mentor count and with mentor age up to a threshold, and analyzes gender composition effects on both protégé outcomes and mentors' citation gains.","tokens_in":18858,"tokens_out":5899,"duration_ms":64490,"significance":"The research question is important, and the dataset is unusually large. I credit the authors for making their matching definitions and supplementary tables explicit; this transparency is what exposes the internal inconsistency discussed below. Nevertheless, the paper's primary evidence contradicts its headline claim: in the main matched comparison, higher mentorship-quality groups have lower post-mentorship impact than controls. Since this reversal affects the abstract, Results, and the policy-related gender conclusions, the significance of the claimed causal effect cannot be credited as presented.","major_comments":[{"comment":"The definition of δ in Supplementary Note 3 is δ = 100·(imp(T′) − imp(C′))/imp(C′), with C = Qi and T = Qi+1. For Q1 vs Q2, the table reports imp(C′) = 14.86 and imp(T′) = 11.66, which gives δ = −21.5%, not +27.4%. The same reversal occurs in every row: Q2 vs Q3 gives −14.1% (reported +16.4%), Q3 vs Q4 gives −17.7% (reported +21.5%), and Q4 vs Q5 gives −26.2% (reported +35.5%). The reported positive values are those that would be obtained from 100·(imp(C′) − imp(T′))/imp(T′), i.e., both the sign and the denominator are inverted. This is not a local typo: it reverses the direction of the headline causal effect.","section":"Supplementary Table S3"},{"comment":"The same pattern is systematic rather than isolated. In Table S4 (hub effect), each higher-quintile matched group has lower mean post-mentorship impact than the lower-quintile group (e.g., 18.61 vs 19.83 for Q1 vs Q2, reported as +6.6% instead of −6.2%), and the same inversion recurs in the large majority of stratified rows in Tables S5–S10 for time periods, mentor-age bins, mentor-count bins, university-rank bins, gender, and discipline. Some individual rows, such as Geology Q2 vs Q3 in Table S10, do report negative δ correctly, which shows that the issue is not a uniformly applied alternative convention. Because the matched control group consistently outperforms the treatment group, the abstract's central claim that higher mentorship quality increases post-mentorship impact is not supported by the paper's own evidence; the observed matched association is negative.","section":"Supplementary Tables S4–S10"},{"comment":"Even if the sign issue were corrected, the causal claim rests on an unconfoundedness assumption that is asserted rather than supported. The matching variables listed where CEM is introduced (number of mentors, first mentored-paper year, discipline, gender, affiliation rank, post-mentorship active years, average mentor age, and the other mentorship-quality measure) do not include the protégé's own prior productivity or prior citation impact, and no sensitivity analysis or placebo test is provided. Because both the treatment (mentors' prior citation impact) and the outcome (protégé's post-mentorship citation impact) are citation-based measures, selection of higher-ability protégés into mentors with higher citation impact is a concrete alternative explanation that the current design does not rule out.","section":"Results: Causal identification"}],"minor_comments":[{"comment":"The rows for two mentors and three mentors are numerically identical, which appears to be a copy-paste error and should be corrected.","section":"Supplementary Table S7"},{"comment":"The caption says to see Supplementary Tables S1 and S2 for details, but the actual CEM result tables for the big-shot and hub effects are Supplementary Tables S3 and S4.","section":"Figure 1 caption"},{"comment":"Reference [20] credits Fortunato et al. with the title \"Hot streaks in artistic, cultural, and scientific careers\"; the correct title of that review article is \"Science of science\". The hot streaks paper is already reference [17].","section":"Reference [20]"}],"recommendation":"reject","confidential_remarks":"The central problem is visible in the paper's own supplementary tables: the reported δ magnitudes and signs do not match the formula given in Supplementary Note 3. Because the matched treatment means are lower than controls in the primary comparison, the headline result and the gender-related policy conclusions reverse direction. A resubmission would need to recompute all estimates transparently or reframe the findings as negative associations; within the current scope, the error is load-bearing and cannot be resolved by local editing."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nShort version: this paper takes an important question—does informal mentorship from high-status senior coauthors move a junior scientist's later citation impact—and goes after it with an enormous MAG dataset and a sensible CEM design. The operationalization of informal mentorship via coauthorship is genuinely new, and the separate analysis of protege outcomes and mentor gains, split by gender, has clear policy relevance. The methods section is detailed and the reference list is appropriate.\n\nThe problem is that the headline result is contradicted by the paper's own supplementary tables. In Tables S3–S10, C′ and T′ are the matched lower- and higher-mentorship-quality groups, and imp(C′) and imp(T′) are their mean post-mentorship impacts. In every single row, imp(C′) > imp(T′). For example, Q1 vs. Q2 shows 14.86 vs. 11.66, and Q4 vs. Q5 shows 28.26 vs. 20.86. Yet the reported δ is positive throughout. The only way to produce those positive values is to use the inverted formula, 100·(imp(C′) − imp(T′))/imp(T′), instead of the stated 100·(imp(T′) − imp(C′))/imp(C′). The same inversion appears in the hub-effect table and in all of the stratified big-shot tables. So the paper's own matched means say that higher big-shot mentorship is associated with lower post-mentorship impact, not a 36% benefit.\n\nThis is not a minor numerical typo. It reverses the central causal claim in the abstract and in Figures 1–2. There are also the usual observational concerns—no sensitivity analysis, no control for the protege's own prior publication record, so the unconfoundedness assumption is doing heavy lifting. But even if unconfoundedness held, the sign is backwards.\n\nTo be fair, the data construction looks careful and the matching design is reasonable in principle. The gender-specific findings might survive re-analysis. But as submitted the paper cannot support its stated conclusions. I would not desk-reject-and-forget: the question and dataset are worth serious attention, and this looks fixable. But I would not accept it, and I would want the authors to correct the arithmetic and re-run before any referee spends time on it.\n\nBottom line: important question, real dataset, decisive internal inconsistency. Once corrected, it matters for science-of-science and gender-equity policy readers; don't cite the current version.","headline":"The paper's own supplementary tables show the mentorship effect running the wrong way, so the headline causal claim is not supported as written; fixable but not acceptable.","tokens_in":19430,"tokens_out":5280,"would_cite":false,"duration_ms":53995,"reading_group":"no","serious_thinker":"no","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Informal mentorship quality causally raises a junior scientist's later citation impact, by up to 36 percent, and the size depends on mentor count, mentor age, and the genders of both partners.","keywords":["informal mentorship","scientific impact","coarsened exact matching","big-shot experience","citation analysis","gender gaps in science","mentor-protege pairs","causal inference"],"falsifier":"A concrete check: re-run the matched comparison but add each protege's own citation count—or publication record—before the mentorship period as an additional exact-matching variable or regression control inside the matched sample. If the big-shot effect shrinks toward zero or reverses sign, the causal claim fails, because the apparent effect would then be explained by pre-existing differences among proteges rather than by the mentorship itself.","tokens_in":18324,"feed_emoji":"🎓","tokens_out":7090,"duration_ms":68155,"temperature":0.7,"pith_summary":"This paper tries to show that informal mentorship in academic collaborations—a junior scientist being supported by several senior coauthors without formal supervisory ties—has a causal effect on how much impact the junior scientist has later, after the mentorship ends. Mentorship quality is measured by the mentors' prior citation track record (the 'big-shot experience') and by their centrality in the collaboration network (the 'hub experience'). Comparing matched proteges with coarsened exact matching across 2.5 million mentor-protege pairs, the paper estimates that being mentored by higher-impact senior scientists raises a protege's post-mentorship citation impact by up to 36 percent, with the size growing with the number of mentors and peaking when mentors have about 30 years of experience. It also claims that a higher proportion of female mentors is associated with lower post-mentorship impact, and that female mentors lose about 18 percent of citations when mentoring female rather than male proteges, suggesting current female-female mentorship policies may trade retention for impact. A sympathetic reader would care because it locates an actionable lever—choice of informal mentors—on a measurable career outcome, and it complicates gender-based mentorship policies.","feed_headline":"Big-shot mentors raise junior scientists' impact by 36%","feed_subtitle":"Effect grows with mentor count, fades after 30 years, and shifts with mentor-protege gender pairing.","key_machinery":"The engine of the analysis is a pair of matched comparisons built with coarsened exact matching (CEM), a method that selects control proteges who resemble treated proteges on the number of mentors, first mentored-paper year, discipline, gender, affiliation rank, post-mentorship active years, and average mentor age. The treatment is 'big-shot experience'—the average annual citations, up to the first mentorship publication, of all a protege's informal mentors—or, in the secondary analysis, 'hub experience,' the average collaborator-network degree of those mentors. The outcome is 'c5,' the citations a protege's own post-mentorship papers (written without mentors) accumulate five years after publication. CEM is what lets the authors call the big-shot effect causal rather than merely correlational: it is the device that attempts to rule out differences in who gets high-impact mentors.","core_discovery":"The paper's central claim is that mentorship quality has a causal effect on the scientific impact of the papers a protege writes after the mentorship period ends. The strongest quantified form is the 'big-shot effect': moving a protege from one quintile of mentor prior-impact to the next increases post-mentorship impact by up to 36 percent, whereas the corresponding 'hub effect' from mentors' network centrality never exceeds 7 percent. The effect persists across disciplines, university ranks, and protege gender; it grows with the number of mentors and roughly doubles in recent decades, and it rises with mentors' academic age until around 30 years of experience, then declines. On gender, the paper claims that increasing the proportion of female mentors among a protege's informal mentors decreases the protege's later impact (by up to 35 percent in the matched comparisons), and that female mentors gain on average 18 percent fewer citations from mentoring female than male proteges, while male mentors' gains are unaffected. The paper reads these results as evidence that policies pushing female-female mentorships, however effective at retaining women in science, may reduce the later impact of the women who stay, and that opposite-gender mentorships should be encouraged instead.","pith_inferences":["Beyond the paper: a test restricting to same-subfield mentorships would discriminate knowledge transfer from a pure status halo, since the paper does not separate these channels.","Beyond the paper: adding each protege's own pre-mentorship citation trajectory as a matching variable would likely shrink the 36 percent figure, because the paper's match does not include prior output.","Beyond the paper: the gender findings imply that informal mentorship networks are a resource distributed unequally by gender, so opposite-gender mentorship policies could work only if enough high-impact male mentors are available and willing.","Beyond the paper: on the mentor side, the lower gain female mentors experience from female proteges could be driven by audience citation bias rather than mentorship quality, a channel the paper does not distinguish."],"forward_implications":["If mentorship quality is causal, then the impact of a young scientist's later independent work can be raised by changing who they collaborate with early, not just by changing their own effort or resources.","Because the big-shot effect grows with the number of mentors and is strongest with more than five mentors, supporting early-career scientists in building a broad set of senior collaborators is a plausible route to higher impact.","The peak at roughly 30 years of mentor experience means that the most productive informal mentoring may come from mid-to-late-career scientists, not the very newest or the most senior.","The gender results imply that policies designed to retain women by pairing them with female mentors may need to weigh a trade-off: retention gains against a measured reduction in later citation impact for those who stay, and against lower citation gains for female mentors themselves.","Because the hub effect is small relative to the big-shot effect, a mentor's standing in the collaboration network matters less than their demonstrated citation impact—so mentorship policy should prioritize research eminence over connectedness."],"supporting_citations":[{"why":"Supplies the method for classifying scientists into disciplines and defines the c5 citation-outcome measure used throughout.","marker":"[18]"},{"why":"Supplies the large bibliographic dataset from which the 2.5 million mentor-protege pairs are identified.","marker":"[36]"},{"why":"Supplies the coarsened exact matching technique used to build the causal comparisons.","marker":"[37]"},{"why":"Supplies the name-based gender classifier used to assign mentor and protege gender.","marker":"[39]"},{"why":"Supplies the university ranking used to control for affiliation rank.","marker":"[40]"}],"fun_headline_variants":["Informal mentorship raises protege impact by 36%","Female mentor ratio linked to 35% lower protege impact","Mentor quality, not centrality, drives later impact","Effect of mentorship peaks at 30 mentor years","Opposite-gender mentoring may boost women's research"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The estimate assumes that after matching on the listed covariates, proteges with high-impact mentors and proteges with lower-impact mentors differ only in their mentorship quality—that is, there is no unmeasured difference in the protege's own ability, prior publication record, or selection into high-status mentors that drives both who mentors them and their later impact.","fun_headline_variants_meta":{"raw":{"variants":["Informal mentorship raises protege impact by 36%","Female mentor ratio linked to 35% lower protege impact","Mentor quality, not centrality, drives later impact","Effect of mentorship peaks at 30 mentor years","Opposite-gender mentoring may boost women's research"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000591,"raw_usage":{"total_tokens":2815,"prompt_tokens":1033,"completion_tokens":1782,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":649,"completion_tokens_details":{"reasoning_tokens":1703}},"tokens_in":649,"tokens_out":1782,"duration_ms":15365,"temperature":1.0,"reasoning_tokens":1703,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T14:01:03.475825+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A concrete check: re-run the matched comparison but add each protege's own citation count—or publication record—before the mentorship period as an additional exact-matching variable or regression control inside the matched sample. If the big-shot effect shrinks toward zero or reverses sign, the causal claim fails, because the apparent effect would then be explained by pre-existing differences among proteges rather than by the mentorship itself.","supporting_citations":[{"cited_title":"K., Rahwan, T","cited_arxiv_id":null,"evidence_quote":"Supplies the method for classifying scientists into disciplines and defines the c5 citation-outcome measure used throughout."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the large bibliographic dataset from which the 2.5 million mentor-protege pairs are identified."},{"cited_title":"M., King, G","cited_arxiv_id":null,"evidence_quote":"Supplies the coarsened exact matching technique used to build the causal comparisons."},{"cited_title":"Gender prediction methods based on ﬁrst names with genderizer","cited_arxiv_id":null,"evidence_quote":"Supplies the name-based gender classifier used to assign mentor and protege gender."},{"cited_title":"The Impact of Informal Mentorship in Academic Collaborations","cited_arxiv_id":"1908.03813","evidence_quote":"Supplies the university ranking used to control for affiliation rank."}],"review_version":1}