{"id":"f37d633e-38c9-44ef-b5bf-48d46c3d5b77","arxiv_id":"2506.15237","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"Across 19,885 PLOS papers, women are significantly more likely than men to be acknowledged rather than listed as authors for investigation/analysis and writing, while higher citation status is associated with higher authorship rates regardless of gender.","lead":"This study analyzed about 20,000 open-access PLOS papers and found that women are acknowledged more often and listed as authors less often than men for investigation/analysis and writing roles. It may interest readers because it suggests that citation-based status, not gender alone, shapes who gets authorship credit.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Keyword mapping of acknowledgments to I&A is unvalidated by gender; generic 'help'/'assistance' could inflate women's non-author I&A roles and manufacture the headline gap.","rationale":"The reader identified the acknowledgment-to-CRiT role alignment as the weakest assumption, and my independent reading converges on the same point. The paper's most distinctive quantitative result is the role-specific gender gap in authorship rate for I&A, computed by classifying acknowledgment sentences with noun-based keywords. The taxonomy validation in Appendix A is thin: it reports manual checks of 100 random sentences per ambiguous word but gives no counts of misclassifications or agreement rates, and it does not address the most frequent and generic keywords ('help', 'assistance') that carry much of the I&A category. Because women could plausibly be thanked more often in generic support language and men in task-specific technical language, a gender-differential misclassification rate would directly create a spurious I&A gap. This is a concrete, testable threat rather than a merely philosophical objection. The paper's own Discussion acknowledges the role-alignment limitation, which supports treating the finding as conditional rather than rejectable: the raw direction of overall acknowledgment versus authorship may still hold, but the role-specific mechanism is not yet established. I therefore keep the reader's CONDITIONAL verdict and see no reason to move it.","tokens_in":12218,"tokens_out":5326,"duration_ms":56931,"concrete_test":"Sample 1,000 acknowledgment sentences containing 'help' or 'assistance' (and, for comparison, 'analysis' or 'code'), stratified by the predicted gender of the acknowledged person. Have two annotators blind to the study hypothesis classify the actual contribution as I&A, administrative/clerical, writing, or other, using CRediT definitions. Report per-gender label accuracy and inter-rater agreement. If accuracy differs by gender, or if generic 'help'/'assistance' is mostly non-I&A, recompute the paper-level authorship rate using corrected labels; if the I&A gap shrinks below statistical significance, the central claim fails.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The headline claim that women are more often acknowledged than authored specifically in Investigation and Analysis (I&A) rests on the assumption that the keyword taxonomy in Table 2 assigns acknowledgment sentences to the same contribution categories as CRediT roles. This assumption is weakest for the two most generic keywords, 'help' and 'assistance', which are placed in I&A with no ambiguity check; Appendix A reports manual checks only for words with multiple meanings in prior work (Paul-Hus and Desrochers 2019), and even those checks are described only as 'limited' with no counts or inter-rater agreement. Acknowledgment sentences such as 'We thank Jane for assistance' or '...for help with the manuscript' would be labeled I&A even when the contribution was clerical or writing-related. Since I&A is roughly 73-80% of women's acknowledgment roles, a small gender difference in how acknowledgments are phrased (e.g., women more often thanked for generic help/assistance, men for specific technical tasks like analysis or code) would directly produce a lower female authorship rate in I&A with no real difference in contribution level. The current validation cannot rule this out, and because the role-specific comparison is the central empirical claim, this is a load-bearing threat to the paper's main conclusion.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper analyzes 19,885 PLOS articles published 2016-2021, linking CRediT authorship roles with a keyword-based taxonomy of acknowledgment roles, and computes an Authorship Rate (AR) by gender and contribution role at both the paper level and the collaboration level. The central claim is that women are more likely to be acknowledged than listed as co-authors, especially in the Investigation and Analysis (I&A) role, where the paper-level AR is roughly 70% for men versus 65% for women. The paper also examines collaboration pairs by citation-based status, finding that highly cited scholars receive authorship more often regardless of gender, and that highly cited women paired with less-cited men are more likely to be authors than the men. The authors discuss limitations including the PLOS-only sample, the keyword-based acknowledgment taxonomy, and the use of gender-specific citation cutoffs.","tokens_in":12408,"tokens_out":5729,"duration_ms":55495,"significance":"If the central result is valid, the paper provides a large-scale, role-resolved quantitative description of gendered credit allocation, going beyond authorship-only analyses by systematically linking authors and acknowledgees. Its strengths include a large sample (122,072 authors and 63,343 acknowledgees), a rigorous linkage to scholar IDs via MAG, explicit research questions, and a transparent acknowledgment of several limitations. The finding that status dynamics may modulate gendered credit patterns is a potentially valuable contribution to the credit-allocation literature. However, the central claim depends on an acknowledgment-role taxonomy that is validated only by small, partially described manual checks, and the statistical tests ignore the non-independence of observations; these issues must be addressed before the conclusions can be regarded as established.","major_comments":[{"comment":"The acknowledgment-role classification is load-bearing for the central claim, but the keyword taxonomy in Table 2 is not validated by gender. The words 'help' and 'assistance' are assigned to Investigation and Analysis (I&A) with no ambiguity check, while Appendix A reports manual checks only for a handful of multi-meaning words, without counts, sampling fractions, or inter-rater agreement. Because I&A accounts for 73%-80% of women's acknowledgment roles (Fig. 2c), a differential tendency for women to be thanked for generic 'help' or 'assistance' rather than for specific technical tasks such as 'analysis' or 'code' could produce the observed lower female Authorship Rate in I&A without any true difference in contribution level. The authors should validate the mapping on a gender-stratified random sample of acknowledgment sentences with reported agreement statistics, or rerun the analysis using an unambiguous subset of keywords.","section":"§2.2.2, Table 2, Appendix A"},{"comment":"The t-tests used to compare Authorship Rates ignore the hierarchical structure of the data. At the paper level, each contributor is treated as an independent observation even though the same scholar appears in multiple papers, and at the collaboration level, the same scholar contributes to many pairs, so the 259,652 paired observations in the I&A comparison are not independent. The reported p-values therefore overstate precision, and with large N even substantively trivial differences (e.g., 88% vs 86% in collaboration-level I&A) become highly significant. I recommend hierarchical models or cluster-robust standard errors with clustering at the scholar or paper level, and reporting of effect sizes alongside p-values.","section":"§2.3, §3.2, Fig. 3"},{"comment":"The status analysis defines 'highly cited' and 'less cited' using gender-specific 10th and 90th percentiles of the citation distribution, so a 'high-woman' may have fewer citations than a 'high-man', and the absolute citation gap between a 'less-man' and a 'high-woman' is not controlled. The conclusion in §3.3 that 'power dynamics based on academic status or perceived success can override, or even reverse, gender-based disparities' is therefore weakened: the reversal in the less-man/high-woman pair could be an artifact of the relative-within-gender status labels rather than of an actual status effect. The paper acknowledges the definitional issue in §2.4, but the interpretation in §3.3 does not flag it as a caveat; the authors should either use gender-common citation cutoffs or, if they keep gender-specific cutoffs, explicitly recast the result as a relative-status comparison and test robustness to alternative cutoffs.","section":"§2.4, §3.3, Fig. 4"}],"minor_comments":[{"comment":"The abstract states 'over 20,000 authors and 60,000 acknowledged individuals', but §3.1 reports 122,072 authors and 63,343 acknowledged individuals; please reconcile the numbers.","section":"Abstract"},{"comment":"There is a typo: 'refered' should be 'referred'; also, the capitalization of 'CRediT' is inconsistent ('CrediT' appears in the same section).","section":"§2.1"},{"comment":"The sentence 'the number of remains relatively stable' appears to be missing a word; please revise for clarity.","section":"§3.1"},{"comment":"The keyword 'data' appears under both I&A and M&R; the disambiguation rule described in Appendix A (labeling 'data' as M&R only when accompanied by 'providing', 'provide', 'provided', or 'database') should be stated in a table note for clarity.","section":"Table 2"},{"comment":"The acknowledgment proportions for I&A differ noticeably between men (73.1%) and women (80.4%); the authors do not comment on this gap, which is directly relevant to the main analysis.","section":"Figure 2(c)"},{"comment":"The sentence 'The sample included as follows' is grammatically incomplete; please specify the sample being described.","section":"§2.4"},{"comment":"The paper does not state whether the analysis code or processed data are available; given that the underlying acknowledgment dataset is public (Kusumegi and Sano 2022), a data/code availability statement would strengthen reproducibility.","section":"Data and code availability"}],"recommendation":"major_revision","confidential_remarks":"The paper's main contribution is the role-specific comparison of authorship and acknowledgment, but the validity of the central claim hinges on the keyword-based acknowledgment taxonomy, which is only weakly validated. The authors should be asked to strengthen the validation (gender-stratified manual coding with agreement statistics) and to address the non-independence in the statistical tests. The novelty relative to Ross et al. (2022) and Paul-Hus et al. (2020) should be clarified in the revision, and the status analysis should be presented more cautiously given the gender-specific citation cutoffs."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe paper is worth a look if you work on scientific credit. It does something Ross et al. didn't: it matches authors and acknowledgees in the same PLOS papers, assigns contribution roles to both, and compares authorship rates by gender and citation status. The headline result — women are acknowledged more often than they are authored, especially in Investigation and Analysis — is in the same direction as prior work, but the matched design and the status interaction (highly cited scholars get authorship more, and highly cited women paired with less-cited men are especially likely to be authors) are genuinely new. The paper is also refreshingly upfront about its limitations.\n\nThe main soft spot is the acknowledgment taxonomy. It relies on keyword matching with only small manual checks, and the two most generic I&A keywords, 'help' and 'assistance', are placed in I&A with no ambiguity check at all. If those words are systematically used to thank women for clerical or writing-type support while men are thanked for analysis or code, the I&A gender gap could be inflated or even manufactured. That is a real validity threat, not a proven flaw — the paper gives no evidence of gendered phrasing, but the validation is too thin to rule it out. Add the t-tests ignoring repeated observations, and the absence of effect sizes or career-stage controls, and the precise magnitude of the gender gap is uncertain. The status analysis is interesting but built on gender-specific citation cutoffs, so 'highly cited' means different things for men and women; the authors acknowledge this, which is honest, but it limits the clean interpretation.\n\nOn balance, the direction of the central finding is probably right — it matches Ross et al. and other work, so I wouldn't bet against it. But the size of the I&A-specific gap, and the claim that role matters, rest on that keyword mapping. I'd want a hand-coded validation set, ideally blinded to gender, before treating the role-specific numbers as solid. I'd also want confidence intervals and models with random effects for individuals and papers.\n\nWho it's for: science-of-science and bibliometrics people, plus anyone thinking about credit equity in publishing. It deserves a serious referee — enough newness and careful description to warrant review, not desk rejection. I'd send it out, expecting major revision on the taxonomy and the statistics.","headline":"A solid, honest empirical extension of prior credit-gap work, but the role-specific claim rests on a weakly validated keyword taxonomy and statistics that ignore dependence.","tokens_in":12980,"tokens_out":2463,"would_cite":true,"duration_ms":26438,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that women are more likely to be acknowledged than listed as co-authors, especially for investigation and analysis roles, and that citation status can override gender in who receives authorship credit.","keywords":["gender bias","authorship","acknowledgment","credit allocation","scientometrics","CRediT taxonomy","citation status","collaboration"],"falsifier":"Take a random sample of the acknowledgment sentences labeled by the keyword taxonomy, have annotators assign contribution roles by hand, and recompute the gender-specific authorship rates; if the gender gap in investigation and analysis disappears or reverses, the reported disparity is an artifact of the keyword mapping rather than a real difference in credit.","tokens_in":11982,"feed_emoji":"⚖️","tokens_out":12487,"duration_ms":113646,"temperature":0.7,"pith_summary":"This paper asks whether women receive less formal credit than men for the same scientific contributions. It compares authorship records with acknowledgment statements for more than 20,000 authors and 60,000 acknowledged individuals in open-access journals, classifying both sides into shared contribution roles. The central finding is a persistent gender gap: for investigation and analysis, women are more often named in acknowledgments than in author lists, while men's equivalent contributions are more likely to be credited as authorship. The gap shrinks or reverses when citation status is taken into account: highly cited scholars receive authorship at higher rates regardless of gender, and a highly cited woman collaborating with a less-cited man is more likely to be the author than he is. Because authorship shapes hiring, funding, and awards, the results matter for equity policies.","feed_headline":"Women are more often acknowledged than credited as authors","feed_subtitle":"For investigation and analysis, credit follows status: highly cited researchers get authorship, and women overall lose out.","key_machinery":"The central object is the Authorship Rate, the share of contributors in a given role who are listed as authors rather than only acknowledged. The paper computes this rate separately for men and women, at the paper level and within man-woman collaborator pairs, for the three contribution types that appear on both the authorship and acknowledgment sides: Investigation and Analysis, Materials and Resources, and Writing. Authorship roles come from the CRediT taxonomy, a standardized contributor-role system; acknowledgment roles are inferred from a keyword-based taxonomy built on an existing codebook of acknowledgment wording. The status analysis splits scholars into the top and bottom 10% by citation count, within gender, and classifies collaborator pairs as high/high, high/low, low/high, or low/low.","core_discovery":"The paper's central claim is that being listed as an author rather than only acknowledged is gendered. For contributions classified as Investigation and Analysis, women's authorship rate is lower than men's at the paper level, about 65% versus 70%, and at the collaboration level, about 86% versus 88%; a smaller but significant gap appears for Writing, while Materials and Resources shows no significant difference. The paper further claims that citation-based status interacts with these disparities: when a highly cited scholar collaborates with a less-cited scholar, the highly cited scholar is more likely to be listed as author irrespective of gender, and in the specific case of a highly cited woman paired with a less-cited man, the woman's authorship rate exceeds the man's. The authors read this as evidence that credit allocation tracks perceived success and power more than gender alone, while still leaving women disadvantaged overall because women are underrepresented among highly cited scholars.","pith_inferences":["The authors leave implicit that the gender authorship gap may be partly a downstream effect of citation inequality; if so, one testable prediction is that the gap narrows or disappears when men and women with equal citation counts collaborate in matched pairs.","Because the dataset excludes individuals who are acknowledged but never appear as authors, the true gender gap in credit may be larger than measured, since women are overrepresented among acknowledgees.","A direct extension would compare the keyword-based authorship rates with rates obtained by human annotation of the same acknowledgment sentences; if the two diverge by gender, the taxonomy itself needs correction."],"forward_implications":["Women's contributions to investigation and analysis are systematically under-credited as authorship relative to men's, so auditing role-specific credit could reduce the gap.","Because citation status strongly predicts authorship, raising women's visibility and citation counts could improve their credit share even without changing authorship norms.","The finding that highly cited women paired with less-cited men receive authorship more often suggests that status can override gender bias, making status hierarchies a relevant intervention point.","The absence of a gender gap in Materials and Resources indicates the inequity is tied to particular intellectual roles rather than to all contributions.","The U-shaped relationship between number of authors and number of acknowledgees implies that large-team research follows different credit-allocation norms that merit separate study."],"supporting_citations":[{"why":"Supplies the prior finding that women are credited less than men in science, which the paper's comparison extends.","marker":"Ross et al. 2022"},{"why":"Provides the codebook and word-category frequencies from which the acknowledgment taxonomy is adapted.","marker":"Paul-Hus and Desrochers 2019"},{"why":"Provides the dataset of identified acknowledged scholars that supplies all acknowledgment names and identifiers.","marker":"Kusumegi and Sano 2022"},{"why":"Supplies the CRediT-based division of labor analysis used for the authorship role mapping.","marker":"Larivière, Pontille, and Sugimoto 2021"},{"why":"Reports earlier gender differences in contributor roles that motivate the authorship-rate comparison.","marker":"Macaluso et al. 2016"},{"why":"Documents CRediT contribution disclosure and authorship norms in the journals under study.","marker":"Sauermann and Haeussler 2017"},{"why":"Supports the accuracy of the name-to-gender inference service used to assign gender.","marker":"Sebo 2021"}],"fun_headline_variants":["In science, women get thanked, men get bylines","Science credit: status trumps gender, but women still lose","Authorship credit favors the highly cited, study finds","Acknowledgment vs authorship: the gender gap in science credit"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole comparison depends on the assumption that the words used in acknowledgment sentences, such as 'help,' 'data,' and 'discussion,' identify the same contribution roles that CRediT assigns to authors, so that a gender difference in role classification reflects a difference in credit rather than a difference in wording.","fun_headline_variants_meta":{"raw":{"variants":["In science, women get thanked, men get bylines","Science credit: status trumps gender, but women still lose","Authorship credit favors the highly cited, study finds","Acknowledgment vs authorship: the gender gap in science credit"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000913,"raw_usage":{"total_tokens":3915,"prompt_tokens":933,"completion_tokens":2982,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":549,"completion_tokens_details":{"reasoning_tokens":2914}},"tokens_in":549,"tokens_out":2982,"duration_ms":25873,"temperature":1.0,"reasoning_tokens":2914,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T19:39:24.749367+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a random sample of the acknowledgment sentences labeled by the keyword taxonomy, have annotators assign contribution roles by hand, and recompute the gender-specific authorship rates; if the gender gap in investigation and analysis disappears or reverses, the reported disparity is an artifact of the keyword mapping rather than a real difference in credit.","supporting_citations":[],"review_version":2}