Pith. sign in

REVIEW 3 major objections 4 minor 12 references

Do male leading authors retract more articles than female leading authors?

T0 review · 3 major / 4 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read Male first authors have a 23% higher retraction rate than female first authors after accounting for how much each gender publishes.

desk verdict A genuinely large and useful descriptive study that states a plausible global claim, but the missing gender data is serious enough that the headline MFRR needs a sensitivity bound before it should be read as a fact. read the letter →

arxiv 2507.17127 v1 pith:6R2HMPGP submitted 2025-07-23 cs.DL

classification cs.DL
keywords retractionrategenderdisparityscientificmisconductfirstauthorscorrespondingresearchintegrityscientometrics
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper tries to establish that gender differences in retraction are real once publication volume is taken into account. Across all fields of science, it finds male first authors are retracted at a rate of 6.38 per 10,000 articles versus 5.19 for female first authors, a male/female retraction ratio of 1.23 (95% CI 1.18 to 1.28). The excess is concentrated in misconduct categories such as plagiarism, authorship disputes, duplication, fabrication/falsification, and ethical issues; for honest mistakes, no significant gender difference appears. The same pattern holds for corresponding authors, and the direction reverses in mathematics and computer science, where female first authors have higher retraction rates. This matters because a productivity-adjusted gap points to gendered patterns in misconduct or its detection that raw counts cannot reveal.

What carries the argument

The central object is the male/female retraction ratio (MFRR), defined as the retraction rate for male first authors divided by the retraction rate for female first authors, where each gender's retraction rate is the number of retracted articles first-authored by that gender divided by all articles first-authored by that gender. MFRR converts raw retraction counts into productivity-adjusted risk, so a value above 1 means men are retracted disproportionately to their output. The paper layers retraction reasons, a five-field subject classification, country, and a corresponding-author analysis onto this ratio, with confidence intervals used to determine which differences are significant.

What would settle it

Recompute the MFRR after assigning gender to the unidentified first authors, especially the Chinese-affiliated majority, using a validated name-gender dictionary, and check whether the ratio stays clearly above 1; if plausible assignments bring it to roughly 1.0 or below, the central claim would not survive.

Watch

Extended reading notes

Core claim

The central discovery is that male leading authors have a higher retraction rate than female leading authors relative to their publication output. For first authors, the retraction rate is 6.38 per 10,000 for men and 5.19 for women, giving an MFRR of 1.23 (95% CI 1.18 to 1.28); for corresponding authors, the ratio is 1.20 (95% CI 1.15 to 1.25). The authors read the mistake category as the key contrast: because retractions attributed to mistakes show no significant gender difference, they conclude the male excess is driven by scientific misconduct rather than a general tendency to err. The largest male excesses appear for plagiarism (MFRR 1.99) and authorship issues (MFRR 1.73). The gender gap is not uniform: men retract more in biomedical and health sciences and in life and earth sciences, while women retract more in mathematics and computer science.

Load-bearing premise

The central ratio is unbiased only if the roughly half of retracted articles whose first-author gender could not be inferred would show the same male/female pattern as the identified half, and the paper notes that most of the missing authors are at Chinese institutions, where name-based gender inference is least reliable.

Editorial extensions

If this is right

  • Retraction-risk comparisons between genders should be expressed per publication rather than as raw counts, since male overrepresentation in raw retractions is partly a reflection of higher output.
  • Integrity and editorial policies aimed at misconduct categories with the largest male excess, such as plagiarism and authorship issues, would target where the productivity-adjusted gap is strongest.
  • A single all-science gender gap does not describe the full picture: in mathematics and computer science, female first authors have significantly higher retraction rates, so field-specific monitoring is warranted.
  • Because mistakes show no significant gender difference, attempts to reduce retractions through better error-checking may not address the gendered part of the problem.
  • The corresponding-author results reproduce the first-author pattern, indicating the finding is not an artifact of choosing one author position as the leading author.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The paper covers only about 53% of retracted articles for first-author gender, and more than 80% of the missing authors are at Chinese institutions; a sensitivity analysis that assigns plausible gender distributions to the missing group could push the MFRR above or below 1.23.
  • A retraction requires detection and a decision, not just a mistake or misconduct, so the misconduct-driven gap invites a companion study of whether reporting and investigation practices differ by gender.
  • Because 13.7% of retracted articles carry multiple reasons, the per-reason MFRRs are not independent; reweighting by unique articles could identify which misconduct type actually drives the male excess.
  • The country-level result that female first authors have higher retraction rates in China is flagged by the authors as fragile because Chinese names resist name-based gender inference; a name-dictionary validation would determine whether that reversal is real.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper investigates gender differences in retraction rates among first and corresponding authors, using a dataset of 11,622 retracted and 19,475,437 non-retracted articles from Web of Science and Retraction Watch (2008–2023). Gender is inferred from author first names via the CWTS pipeline with a reported accuracy threshold of 90%. The central descriptive claim is that male first authors have a higher retraction rate (6.38 vs 5.19 per 10,000; MFRR = 1.23, 95% CI 1.18–1.28), with the male excess concentrated in misconduct categories and varying by field and country. The corresponding-author analysis gives a similar MFRR of 1.20 (95% CI 1.15–1.25).

Significance. If the finding holds, it is a valuable contribution because it measures retraction risk relative to publication volume with a much larger sample than earlier studies, and it separates misconduct-related retractions from honest-error retractions. The analysis is transparent in its definitions of RR and MFRR, uses standard confidence intervals, and explicitly acknowledges the non-binary gender limitation and the difficulty of gender inference for Chinese names. The consistency between first-author and corresponding-author results is a useful internal check. However, the paper's global claim rests on a complete-case analysis in which gender is unknown for 47.1% of retracted articles and 23.3% of non-retracted articles, and the missingness is concentrated in a country (China) for which the paper itself reports a female-favoring retraction-rate pattern. Because the paper provides no sensitivity analysis bounding the effect of this missingness, the quantitative global MFRR and the misconduct-versus-mistake interpretation are not yet fully supported.

major comments (3)
  1. [§2.2, Table 1, and §4.4] The analysis is a complete-case comparison: only articles with inferred first-author gender are retained, giving 52.9% of retracted articles and 76.7% of non-retracted articles (Table 1). Missingness is not ignorable: §4.4 reports that over 80% of unidentified authors are affiliated with Chinese institutions, and Figure 5 shows that among the ten countries with the most retractions, China is one where female first authors have significantly higher retraction rates. Excluding these articles can therefore shift the headline MFRR of 1.23 (Section 3) in either direction, and the direction of the bias is unknown without additional assumptions. The paper should provide a sensitivity analysis, such as deterministic bounds under worst-case assumptions, inverse probability weighting, or multiple imputation, before the global claim can be read without this qualification.
  2. [§3.2 and Figure 2] The absence of a significant gender difference for 'mistakes' is central to the conclusion in §4.1 that the male excess is driven by misconduct rather than honest error. However, the reason-specific analysis is also restricted to the 52.9% of retracted articles with inferred gender, and the excluded Chinese-affiliated retractions are likely non-random with respect to retraction reason (for example, paper-mill retractions often appear under fabrication or multiple-reason labels). If unclassified retractions are disproportionately of a particular reason and gender, the null result for 'mistakes' and the elevated MFRRs for misconduct categories could be artifacts of missingness. At minimum, the authors should report the gender-inference rate by retraction reason and ideally provide bounds for the reason-specific MFRRs.
  3. [§3.6 and Supplementary Table S1] The corresponding-author analysis shares the same gender-inference pipeline and has similar missingness (55.1% of retracted and 79.0% of non-retracted articles with known gender). It is therefore not an independent robustness check: any systematic bias in the CWTS inference or in the missingness mechanism affects both analyses in similar ways. The statement in §3.6 that consistency 'provides further support for the robustness of our findings' is stronger than the evidence supports unless a sensitivity analysis is added. I would frame the corresponding-author result as a replication of the first-author analysis under the same missingness assumptions, not as an independent validation.
minor comments (4)
  1. [Abstract and Section 3] The headline RR and MFRR values should be explicitly qualified as computed on the complete-case sample of articles with inferred gender, e.g., 'among articles with inferred gender of the first author'.
  2. [Figure 6] Figure 6 is very dense and the small multiples are difficult to read, especially for the cross-analysis by field and reason. A supplementary table reporting the point estimates and 95% confidence intervals for each cell would substantially improve readability.
  3. [Abstract and Section 1] The abstract states that 'no comprehensive study has explored this issue across all fields of science,' but earlier work such as Fanelli et al. (2015) covers multiple fields. The wording should be softened to reflect the scale and scope of this study rather than claiming absolute novelty.
  4. [Section 4.4] The statement that 'over 80% of the authors with unidentified gender were affiliated with institutions in China' would be more informative if broken down separately for retracted and non-retracted articles, since the differential missingness between the two groups is the main methodological concern.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the headline MFRR is computed directly from observed counts with no fitted parameter that re-enters the derivation.

full rationale

The paper's central claim is an empirical rate comparison. The retraction rate RR is defined as the ratio of retracted articles to all articles for each gender, and the MFRR is the ratio of the male RR to the female RR (Section 2.5). The reported values, such as MFRR = 1.23 (95% CI 1.18, 1.28), are direct arithmetic functions of the observed counts in Table 1; no parameter is fitted to a subset and then relabeled as a prediction, and no result is assumed in the definitions. The gender-inference pipeline and missing-data limitations are acknowledged in Sections 2.2 and 4.4, but missingness is a validity threat or correctness risk, not a circularity: it does not make the derived ratio equivalent to an input by construction. The robustness check using corresponding authors reuses the same gender-inference methodology, but reusing a measurement tool is not the same as importing the target conclusion. Citations to prior classification work (e.g., Tang et al. 2020, Zhang et al. 2020) support an empirical taxonomy and do not function as a self-citation chain that forces the outcome. Overall, the derivation chain is self-contained with respect to the claims it makes, so the circularity burden is negligible.

Assumptions & free parameters 1 free parameters · 4 assumptions · 0 invented entities

The analysis introduces no invented entities and no fitted parameter beyond the gender inference threshold. It rests on the validity of external databases, name-based gender inference, and the attribution of retraction responsibility to leading authors.

free parameters (1)
  • Gender inference accuracy threshold = 90%
    Authors only assign gender when the inferred accuracy is at least 90%. This hand-chosen threshold affects which articles are included and could influence the male/female ratio.
assumptions (4)
  • domain assumption Name-based gender inference with at least 90% reported accuracy approximates true author gender.
    Used for all gender assignment; no validation against self-reported gender is possible in this dataset.
  • domain assumption Web of Science and Retraction Watch provide complete coverage of retractions and retraction reasons.
    The entire analysis depends on the completeness and correctness of these external databases.
  • domain assumption The first author or corresponding author is the appropriate person to attribute a retraction to.
    The authors acknowledge other authors may be responsible; this assumption is central to the research question.
  • domain assumption Binary gender classification is sufficient for the analysis.
    Non-binary authors are excluded, as the authors acknowledge in the limitations.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Do male leading authors retract more articles than female leading authors?." pith.science (2026). https://pith.science/paper/6R2HMPGP

@misc{pith2026250717127,
  author       = {Pith},
  title        = {Pith review of: Do male leading authors retract more articles than female leading authors?},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/6R2HMPGP}},
  note         = {Machine review of arXiv:2507.17127}
}
read the original abstract

Scientific retractions reflect issues within the scientific record, arising from human error or misconduct. Although gender differences in retraction rates have been previously observed in various contexts, no comprehensive study has explored this issue across all fields of science. This study examines gender disparities in scientific misconduct or errors, specifically focusing on differences in retraction rates between male and female first authors in relation to their research productivity. Using a dataset comprising 11,622 retracted articles and 19,475,437 non-retracted articles from the Web of Science and Retraction Watch, we investigate gender differences in retraction rates from the perspectives of retraction reasons, subject fields, and countries. Our findings indicate that male first authors have higher retraction rates, particularly for scientific misconduct such as plagiarism, authorship disputes, ethical issues, duplication, and fabrication/falsification. No significant gender differences were found in retractions attributed to mistakes. Furthermore, male first authors experience significantly higher retraction rates in biomedical and health sciences, as well as in life and earth sciences, whereas female first authors have higher retraction rates in mathematics and computer science. Similar patterns are observed for corresponding authors. Understanding these gendered patterns of retraction may contribute to strategies aimed at reducing their prevalence.

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

12 extracted references · 11 canonical work pages

  1. [1]

    A., & Caprasecca, A

    Abramo, G., D’Angelo, C. A., & Caprasecca, A. (2009). Gender differences in research productivity: A bibliometric analysis of the Italian academic system. Scientometrics, 79(3), 51 7–539. https://doi.org/10.1007/s11192-007-2046-8 Aksnes, D. W., Rorstad, K., Piro, F., & Sivertsen, G. (2011). Are female researchers less cited? A large‐ scale study of Norweg...

  2. [3]

    https://doi.org/10.3390/publications1030087 Brainard, J. (2018). Rethinking retracti ons. Science, 362(6413), 390 –393. https://doi.org/10.1126/science.362.6413.390 Budd, J. M., Sievert, M., & Schultz, T. R. (1998). Phenomena of retraction: Reasons for retraction and citations to the publications. JAMA, 280(3), 296–297. https://doi.org/10.1001/jama.280.3....

  3. [4]

    B., & Konecni, D

    https://doi.org/10.1177/2378023117738903 Konecni, V ., Ebbeson, E. B., & Konecni, D. K. (1976). Decision processes and risk taking in traffic: Driver response to the onset of yellow light. Journal of Applied Psychology , 61(3), 359–367. https://doi.org/10.1037/0021-9010.61.3.359 Kozlowski, D., Larivière, V ., Sugimoto, C. R., & Monroe-White, T. (2022). In...

  4. [5]

    Knowledge-Enriched Distributional Model Inversion Attacks

    https://doi.org/10.1371/journal.pone.0248625 Shema, H., Hahn, O., Mazarakis, A., & Peters, I. (2019). Retractions from altmetric and bibliometric 30 perspectives. Information - Wissenschaft & Praxis , 70(2–3), 98 –110. https://doi.org/10.1515/iwp-2019-2006 Shuai, X., Rollins, J., Moulinier, I., Custis, T., Edmunds, M., & Schilder, F. (2017). A multidimens...

  5. [8]

    https://doi.org/10.3389/frma.2023.1064230 Schumann, K., & Ross, M. (2010). Why women apologize more than men: Gender differences in thresholds for perceiving offensive behavior. Psychological Science , 21(11), 1649 –1655. https://doi.org/10.1177/0956797610384150 Serghiou, S., Marton, R. M., & Ioannidis, J. P. A. (2021). Media and social media attention to...

  6. [10]

    The Editor -in-Chief has retracted this article because several images in this article appear to overlap with those of a previously published article by different authors

    https://doi.org/10.29024/joa.7 Zhang, L., Gou, Z., Fang, Z., Sivertsen, G., & Huang, Y . (2023). Who tweets scientific publications? A large-scale study of tweeting audiences in all areas o f research. Journal of the Association for Information Science and Technology, 74(13), 1485–1497. https://doi.org/10.1002/asi.24830 Zhang, Q., Abraham, J., & Fu, H. -Z...

  7. [149]

    https://doi.org/10.1038/494149a Fanelli, D. (2013b). Why growing retractions are (mostly) a good sign. PLOS Medicine , 10(12), e1001563. https://doi.org/10.1371/journal.pmed.1001563 Fanelli, D., Costas, R., Fang, F. C., Casadevall, A., & Bik, E. M. (2019). Testing hypotheses on risk factors for scientific misconduct via matched-control analysis of papers ...

  8. [457]

    F., Jin, G

    https://doi.org/10.1038/541455a Lu, S. F., Jin, G. Z., Uzzi, B., & Jones, B. (2013). The retraction p enalty: Evidence from the web of science. Scientific Reports, 3(1),

Show all 12 references
  1. [2015]

    mistakes

    35 Supplementary Materials We conducted the same analysis for corresponding authors as we did for first authors to assess whether the results would differ from the first-author analysis and would therefore depend on who was believed to be the leading author. As detailed in Tab...

  2. [2023]

    Error bars indicate 95% confidence intervals. While gender differences in retraction rates for most retraction reasons are not statistically significant in most years, there are more years in which male corresponding authors have higher retraction rates (Figure S3). Consequent...

  3. [3146]

    B., Nielsen, M

    https://doi.org/10.1038/srep03146 Madsen, E. B., Nielsen, M. W., Bjørnholm, J., Jagsi, R., & Andersen, J. P. (2022). Meta-research: Author- level data confirm the widening gender gap in publishing rates during COVID -19. eLife, 11, e76559. https://doi.org/10.7554/eLife.76559 M...

  4. [5233]

    1038/s41598-019- 31 41695-z Trivers, R

    https://doi.org/10. 1038/s41598-019- 31 41695-z Trivers, R. (1972). Parental investment and sexual selection. In Sexual Selection and the Descent of Man (Bernard Campbell, pp. 136–179). https://roberttrivers.com/publications/ Van Noorden, R. (2011). Science publishing: The tro...

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.