Pith. sign in

REVIEW 3 major objections 4 minor 18 references

Delineating Feminist Studies through bibliometric analysis

T0 review · 3 major / 4 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read A two-stage method turns 289 Gender Studies journals into a 1.9-million-document map of feminist research.

desk verdict Useful two-stage core-and-keyword workflow for delineating gender studies, but the precision claim is unvalidated and the only internal check is circular; needs external validation and artifact release before it can be trusted as a corpus. read the letter →

arxiv 2411.18306 v1 pith:QUONGNQD submitted 2024-11-27 cs.DL cs.IR

classification cs.DLcs.IR
keywords feministstudiesgenderbibliometricstopicmodelingBERTopicDimensionsdatabasegender/sexrelatedcore-and-extensionmethod
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper tries to establish that Feminist and Gender Studies can be delimited bibliometrically by a two-stage hybrid: start from a manually curated core of 289 specialized journals, let BERTopic topic modeling extract the vocabulary of that core, and then search that vocabulary in article titles across the whole Dimensions database. The authors claim this approach beats a basic keyword search because the keywords are anchored in the field's own literature rather than in a researcher's manual enumeration, while manual review of the topics corrects the model's gaps and biases. The resulting dataset of 1,967,302 documents in English, Spanish, French, and Portuguese from 1668 to 2023 would provide the empirical basis for mapping the topics, citation flows, collaboration patterns, and institutional and regional participation of gender/sex related research inside and outside the core. If the method transfers, the same recipe could delineate other research areas that escape disciplinary classifications, such as Black Studies, Human Rights, or Agroecology.

What carries the argument

The machinery is the core-to-periphery vocabulary transfer. A set of 289 manually curated Gender Studies journals forms the core; BERTopic topic modeling of that core's titles and abstracts produces 330 topics with their characteristic words; manual refinement turns those words into the final list of 229 keywords (259 regular expressions) that is then run against article titles in every journal. The central object is this keyword list, since it is the instrument that decides whether a document enters the 1.9-million corpus. Its constituent words do the work of carrying the field's vocabulary from the social sciences and humanities into biomedical and clinical journals, and the paper's results are all downstream of that transfer.

What would settle it

Randomly sample 500 documents from the Not Core segment that were retrieved by generic tokens such as 'sex', 'gender', or 'woman' in biomedical journals, and have domain experts judge whether each is substantively about gender/sex. If most are routine clinical or biological studies that merely mention sex as a variable, then the title-keyword assumption overcounts and the 1.9-million corpus is not a faithful delineation.

Watch

Extended reading notes

Core claim

On its own terms, the paper's central claim is that a hybrid core-and-extension pipeline delineates gender/sex related studies more faithfully than any existing classification. The authors build a core of 289 Social Sciences and Humanities journals specializing in Gender Studies, apply BERTopic to their titles and abstracts to obtain 330 topics, manually condense the characteristic words into 229 keywords (259 regular expressions) in four languages, and search article titles across all Dimensions journals. This yields 1,967,302 documents from 1668 to 2023, with 91.9% of them coming from outside the core. The asymmetry between the Core and the Not Core segments is read as evidence that the method captures both gender studies as a discipline and gender/sex as a transversal perspective across biomedicine, health, and other fields; for example, 'gender' dominates the Core while 'sex' dominates the Not Core, and 'feminis' terms are notably concentrated in the Core. The authors contend that this two-stage design reflects the dynamic interaction between Gender Studies and the disciplines it influences.

Load-bearing premise

The load-bearing premise is that an article title containing any of the 229 keywords is a reliable sign that the article belongs to gender/sex related studies, including in biomedical journals, even though the keywords were derived from a Social Sciences and Humanities core.

Editorial extensions

If this is right

  • Researchers can map gender/sex related research by topic, citation, collaboration, institution, country, and language, using a corpus that does not depend on journal or paper-level disciplinary labels.
  • The core/not-core split gives a quantitative read on how feminist vocabulary diffuses: 'gender' and 'feminis' terms stay concentrated in the core while 'sex', 'abortion', and 'menstrual' dominate the biomedical periphery.
  • The method tracks conceptual change in titles, such as the rise of 'gender' over 'sex' since the 1980s and the growing co-occurrence of the two terms.
  • The same two-stage workflow can be applied to other 'more-than-disciplinary' conversations whose boundaries are not captured by existing classifications.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Because the search is restricted to titles, the dataset probably misses gender/sex research that signals its topic only in abstracts or full text; an abstract-based variant would trade precision for recall and could be tested against the current corpus.
  • The precision of the 1.9-million count is an empirical question about generic terms: many biomedical hits on 'sex' or 'woman' may be routine variable mentions rather than gender scholarship, so a precision study on a random sample would determine how much of the corpus is substantively gender/sex related.
  • A natural extension is to use the same pipeline with abstracts rather than titles, or with full texts, and compare topic distributions across languages and regions to see whether the concept of 'gender/sex' spreads unevenly.
  • The method's portability is not automatic: the curated core is the hard-won part, and other social-movement fields will need their own core, their own topic-modeling step, and their own manual review before the keyword list is trustworthy.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper proposes a two-stage bibliometric method to delineate gender/sex-related studies using the Dimensions database. Stage one builds a manually curated core of 289 Gender Studies journals. Stage two applies BERTopic topic modeling to the core's titles and abstracts, manually refines the resulting topic terms into a list of 229 keywords (259 regular expressions), and performs a title-only keyword search over all Dimensions journals in English, Spanish, French, and Portuguese. The resulting corpus contains 1,967,302 documents, of which 1,807,272 come from outside the core. The authors claim that this hybrid system 'surpasses basic keyword search' and produces a dataset suitable for bibliometric analyses of Feminist/Gender Studies across disciplines.

Significance. If validated, this work would provide a substantial, openly reusable corpus and a transferable methodological template for delineating 'more-than-disciplinary' fields. The authors give a detailed, transparent account of their workflow, combine expert curation with NLP in a way that is explicitly designed to reduce manual-keyword bias, and commit to publishing the core journal list, topics, keywords, and document identifiers. These are real strengths. However, the central claim of superiority over simple keyword search and the overall validity of the corpus rest on precision and recall evidence that the paper does not provide. The only internal check (the core overlap shown in Figure 1) is expected by construction and cannot establish that the extended set is not dominated by false positives, especially in biomedical fields. As such, the significance is conditional on the addition of a credible external evaluation.

major comments (3)
  1. [Data and methods; Figure 1] The only validation reported for the keyword retrieval is that 'more than half of the Core overlaps with the documents identified through keyword retrieval' (Figure 1). This overlap is not informative about precision or recall because the keyword list was derived from that same Core via BERTopic and manual refinement. The authors need an external gold standard, such as a random sample of Not Core documents manually annotated for relevance (with agreement statistics), and should report precision, recall, and F-score for the whole corpus and for major disciplinary subsets. Without this, the claim that the hybrid system 'surpasses basic keyword search' is unsupported.
  2. [Data and methods; Figure 3; Figure 4A] The keyword vocabulary was extracted from a core restricted to Social Sciences and Humanities journals, but it is then applied to all journals in Dimensions, including Biomedical and Clinical Sciences (37.2% of Not Core), Health Sciences (12%), and Biological Sciences (6.1%) (Figure 3). In those fields, generic title terms such as 'sex', 'women', and 'gender' are often incidental (e.g., 'sex differences in sepsis outcomes', 'women with breast cancer'), and Figure 4A shows that these are indeed the top three keywords in the Not Core segment. Since the search uses titles only and ignores abstracts, the risk of false positives is substantial. The authors should either restrict the expansion to fields where the vocabulary is likely to be indicative, or quantify precision within the biomedical and health sciences subsets. A simple manual audit of a random sample from those disciplines would address this concern.
  3. [Data and methods; Discussion] The authors justify rejecting Dimensions' Gender Studies group because only about 63% of its documents align with their definition, yet they provide no equivalent precision measurement for their own method. A direct comparison is needed: on a labeled sample, estimate the precision and recall of the proposed hybrid method, of the Dimensions group, and of a baseline keyword-only query (e.g., titles containing 'gender', 'sex', 'woman', or 'feminist'). Such a comparison is the minimal evidence required to substantiate the statement that this approach 'surpasses basic keyword search.'
minor comments (4)
  1. [Introduction] The phrase 'inter-, intra-, inter-, post-disciplinary' appears to contain a duplicated or erroneous prefix; likely 'inter-, intra-, post-disciplinary' was intended.
  2. [Results, Figure 4B caption and text] The sentence 'more inclusive acronyms such as LGBT gained popularity in the 20th century' should probably read 'in the 21st century,' since the rise of the inclusive acronym is a recent phenomenon.
  3. [Table 2] Table 2 lists 282 journals in the Core segment, while the Data and methods section states that the core comprises 289 scientific journals. This discrepancy should be reconciled (e.g., if 7 journals had no indexed articles).
  4. [Declarations, Data transparency] The data transparency statement says the lists and DOIs 'will be published in a public repository' once the article is accepted. For reproducibility, it would be preferable to provide them as supplementary material or in a repository at the time of submission, even in preliminary form.

Circularity Check

1 steps flagged · score 6.0 of 10

Keyword validation is circular: the Core overlap cited as evidence of precision is guaranteed by deriving the keywords from that same Core.

  1. fitted input called prediction [Data and methods (BERTopic keyword generation) and Results (Figure 1 paragraph)]
    "Applying BERTopic (Grootendorst, 2022) to this corpus, we generated a list of 330 topics... The resulting keyword list is a combination of terms that were already used for the journal search... As shown in Figure 1, more than half of the Core overlaps with the documents identified through keyword retrieval. This shows that the keywords used were indeed characteristic of the discipline."

    The keyword list is fitted to the Core corpus: it is generated by BERTopic over the Core's titles and abstracts, then manually refined from those topics. Figure 1's overlap between Core and keyword-retrieved documents is therefore a property of the construction, not an independent confirmation that the keywords delineate gender/sex related studies. The authors offer no external gold standard or manual precision sample outside the Core, and they reject Dimensions' Gender Studies group for having only about 63% precision without measuring their own method's precision. Consequently, the claim that this 'hybrid system surpasses basic keyword search' rests on a validation that reduces to the fitting procedure.

full rationale

The paper's central validation step is circular. The authors derive the 229-keyword query from the 289-journal Core via BERTopic and manual refinement, then cite the fact that more than half of the Core overlaps with keyword-retrieved documents (Figure 1) as evidence that 'the keywords used were indeed characteristic of the discipline' and that the hybrid system 'surpasses basic keyword search.' That overlap is expected by construction: a query built from a corpus will tend to retrieve documents from that same corpus. No independent benchmark, human-annotated sample outside the Core, or comparison against a standard keyword baseline is provided to support the precision of the 1.8 million Not Core documents, which include large Biomedical and Health Sciences segments where title terms such as 'sex' and 'women' are often incidental. The circularity is localized to the validation and to the superiority claim; the dataset itself is a reproducible operationalization, so the paper is not wholly circular. Score 6 reflects that the only reported evidence for the central claim reduces by construction, while the rest of the descriptive analysis is independent of that validation.

Assumptions & free parameters 4 free parameters · 3 assumptions · 0 invented entities

The paper introduces no new theoretical entities, particles, forces, or conserved quantities. It contributes a dataset and a workflow, both built from hand-chosen lists and assumptions about title-based retrieval. The main free parameters are the journal and keyword lists, which are manual and project-specific.

free parameters (4)
  • Core journal list = 289 journals
    Manually curated selection (WoS Women's Studies, DOAJ, LATINDEX, regex, and manual review) defines the seed corpus; all downstream topics and keywords depend on it.
  • Keyword list = 229 keywords / 259 regular expressions
    Manually curated from BERTopic topics; directly determines the extended corpus via title search and is load-bearing for corpus composition.
  • BERTopic min_topic_size = 50
    Threshold for keeping topics in the model; affects the set of topics and thus the vocabulary used for retrieval.
  • Language scope = English, Spanish, French, Portuguese
    Chosen based on the authors' language proficiency; excludes other languages and shapes the geographic and linguistic coverage of the dataset.
assumptions (3)
  • domain assumption Presence of a keyword in the title is a sufficient proxy for belonging to gender/sex related studies.
    The keyword search is run on titles only, with no validation that title matches correspond to true positives. Entered in Data and methods: 'we conducted keyword searches within article titles, disregarding their abstracts.'
  • domain assumption The vocabulary extracted from the Social Sciences and Humanities core generalizes to gender/sex related work in medical, biological, and other disciplines.
    The method assumes that BERTopic on the core corpus produces terms that also identify gender-related work outside those disciplines, where terminology may differ.
  • domain assumption Dimensions provides sufficient coverage of non-English and non-Western literature for a global delineation.
    Stated in Data and methods with citations; coverage limitations for Africa and Asia are acknowledged in the Discussion.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Delineating Feminist Studies through bibliometric analysis." pith.science (2026). https://pith.science/paper/QUONGNQD

@misc{pith2026241118306,
  author       = {Pith},
  title        = {Pith review of: Delineating Feminist Studies through bibliometric analysis},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/QUONGNQD}},
  note         = {Machine review of arXiv:2411.18306}
}
read the original abstract

The multidisciplinary and socially anchored nature of Feminist Studies presents unique challenges for bibliometric analysis, as this research area transcends traditional disciplinary boundaries and reflects discussions from feminist and LGBTQIA+ social movements. This paper proposes a novel approach for identifying gender/sex related publications scattered across diverse scientific disciplines. Using the Dimensions database, we employ bibliometric techniques, natural language processing (NLP) and manual curation to compile a dataset of scientific publications that allows for the analysis of Gender Studies and its influence across different disciplines. This is achieved through a methodology that combines a core of specialized journals with a comprehensive keyword search over titles. These keywords are obtained by applying Topic Modeling (BERTopic) to the corpus of titles and abstracts from the core. This methodological strategy, divided into two stages, reflects the dynamic interaction between Gender Studies and its dialogue with different disciplines. This hybrid system surpasses basic keyword search by mitigating potential biases introduced through manual keyword enumeration. The resulting dataset comprises over 1.9 million scientific documents published between 1668 and 2023, spanning four languages. This dataset enables a characterization of Gender Studies in terms of addressed topics, citation and collaboration dynamics, and institutional and regional participation. By addressing the methodological challenges of studying "more-than-disciplinary" research areas, this approach could also be adapted to delineate other conversations where disciplinary boundaries are difficult to disentangle.

Figures

Figures reproduced from arXiv: 2411.18306 by the authors.

Figure 1
Figure 1. Overlapping between the core and the rest of the dataset. The earliest publications in journals specifically focused on Gender Studies that appear in Dimensions date back to 1970. The publication volume in the Core experienced a rapid growth until 1975, after which it stabilized at a growth rate similar to the rest of the dataset. This trend aligns with literature on the institutionalization of Feminist Studies afte… view at source ↗
Figure 2
Figure 2. Evolution of the number of articles in each part of the dataset (Core: 1976-2022, Not Core: 1950-2022) Regarding the disciplines associated with each article ( [PITH_FULL_IMAGE:figures/full_fig_p009_2.png] view at source ↗
Figure 3
Figure 3. Distribution of disciplines in each part of the dataset (<2% for both were grouped in "Others"). Figure 4A illustrates the most frequently used keywords in each segment of the dataset. The keywords used for the queries were consolidated into 63 representative terms. That is, keywords referring to the same topic, expressed in different languages, were grouped under a single term. Notably, in both segments, the top th… view at source ↗
Figures from the paper (2 more)
Figure 4
Figure 4. Figure 4: A: Relative frequency of keywords in each part of the dataset; B: Evolution of the relative frequency of a selection of keywords. Figure 5A compares the evolution of the relative frequencies of the terms "sex" and "gender" across each segment of the dataset. The popula…
Figure 5
Figure 5. Figure 5: A: Keywords "sex" and "gender", evolution in each part of the dataset; B: Evolution of the usage of sex and gender in titles, simultaneously. Discussion This paper presents a methodology for creating a bibliometric corpus on gender/sex related studies. As a result, we …

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

18 extracted references · 18 canonical work pages

  1. [1]

    more-than-disciplinary

    1 Delineating Feminist Studies through bibliometric analysis Author information Natsumi S. Shokida*, Diego Kozlowski** and Vincent Larivière*** *natsumi.solange.shokida@umontreal.ca ORCID 0009-0008-9198-3888 University of Montreal, EBSI, Montreal, Canada (corresponding author) ** diego.kozlowski@umontreal.ca ORCID 0000-0002-5396-3471 University of Montrea...

  2. [2]

    Table 2 provides absolute and relative sizes of these segments

    the rest (Not Core), comprises articles identified through keyword searches, excluding those overlapped with the Core. Table 2 provides absolute and relative sizes of these segments. As expected, the Core features fewer journals but more articles per journal. The Not Core segment is much larger, containing mostly articles from Biomedical and Clinical Scie...

  3. [3]

    women",

    Distribution of disciplines in each part of the dataset (<2% for both were grouped in "Others"). Figure 4A illustrates the most frequently used keywords in each segment of the dataset. The keywords used for the queries were consolidated into 63 representative terms. That is, keywords referring to the same topic, expressed in different languages, were grou...

  4. [4]

    sex" and

    A: Relative frequency of keywords in each part of the dataset; B: Evolution of the relative frequency of a selection of keywords. Figure 5A compares the evolution of the relative frequencies of the terms "sex" and "gender" across each segment of the dataset. The popularization of the term "gender" is evident, particularly during the 1980s. In the Core seg...

  5. [5]

    sex" and

    or a particular scientific domain (Dehdarirad et al., 2015; Lähdesmäki & Vlase, 2023). In other cases, they resort to using a limited list of keywords (Fox et al., 2022; Majumder et al., 2021). There is a need for a more comprehensive approach to bibliometrically delineate gender/sex related research. The historical evolution of terms like "sex" and "gend...

  6. [8]

    feminist

    Topics with their specific words, and number of documents. Topic Count 0_queer_lesbian_gay_lgbt 2,091 27_masculinity_masculinities_man_masculine 263 31_trafficking_human_prostitution_sexualidad 248 82_pregnancy_maternal_motherhood_mother 139 93_feminismo_feministas_feminismos_brasil 133 94_black_race_justice_freedom 133 138_labor_workers_migrant_care 101 ...

  7. [11]

    Conversely, the Not Core segment is more focused on Biomedical and Clinical Sciences (37.2%), Health Sciences (12%), and Biological Sciences (6.1%)

    Evolution of the number of articles in each part of the dataset (Core: 1976-2022, Not Core: 1950-2022) Regarding the disciplines associated with each article (Figure 3), it is noteworthy that the Core features a significant proportion from Human Society (69%) and History, Heritage, and Archaeology (6.8%). Conversely, the Not Core segment is more focused o...

  8. [15]

    http://sedici.unlp.edu.ar/handle/10915/61652 Cislak, A., Formanowicz, M., & Saguy, T. (2018). Bias against research on gender bias. Scientometrics, 115(1), 189–200. https://doi.org/10.1007/s11192-018-2667-0 Dehdarirad, T., Villarroya, A., & Barrios, M. (2015). Research on women in science and higher education: A bibliometric analysis. Scientometrics, 103(...

Show all 18 references
  1. [16]

    https://doi.org/10.3389/frma.2020.593494 Haig, D. (2004). The Inexorable Rise of Gender and the Decline of Sex: Social Change in Academic Titles, 1945–2001. Archives of Sexual Behavior, 33(2), 87–96. https://doi.org/10.1023/B:ASEB.0000014323.56281.0d Harding, S. (1991a). How t...

  2. [18]

    https://doi.org/10.3998/ptpbio.2096 17 Rojas, F. (2010). From Black Power to Black Studies: How a Radical Social Movement Became an Academic Discipline. JHU Press. Schiebinger, L. (2000). Has Feminism Changed Science? Signs, 25(4), 1171–1175. https://www.jstor.org/stable/31755...

  3. [741]

    S., & Larivière, V

    https://doi.org/10.1007/s00429-023-02750-8 Pradier, C., Kozlowski, D., Shokida, N. S., & Larivière, V. (2024, July 26). Science for whom? The influence of the regional academic circuit on gender inequalities in Latin America. arXiv.Org. https://arxiv.org/abs/2407.18783v1 Prum,...

  4. [1970]

    a multidisciplinary stream of feminist scholarship on gender and science

    The publication volume in the Core experienced a rapid growth until 1975, after which it stabilized at a growth rate similar to the rest of the dataset. This trend aligns with literature on the institutionalization of Feminist Studies after the 'second wave' of US feminism. Th...

  5. [2015]

    thematic research

    and addresses topics of interest to the feminisms, or as "thematic research" and "thematic feminist studies" (Lykke, 2011). We will refer to this object as "gender/sex related studies" to emphasize their thematic focus on the category of gender/sex, or as "feminist studies" to...

  6. [2017]

    overlapping circles

    or South Africa (Hassim & Walker, 1993)—, there may be gender/sex related science without a strong institutionalization into Women's, Gender, or Feminist Studies. These characteristics imply that feminist discussions within the scientific sphere exceeds the definition of a tra...

  7. [2019]

    the combined biological and cultural elements of human sex and social behavior

    to refer to "the combined biological and cultural elements of human sex and social behavior" (Prum, 2023). 3 the particularities of gender/sex related studies and reflect on the challenges they imply for their bibliometric operationalization. Characteristics of Feminist Studie...

  8. [2021]

    4405" corresponds to

    and metadata quality. In particular, it includes journals from non-English speaking countries and developing countries (Basson et al., 2022; Guerrero-Bote et al., 2021). In addition, indexing is not based on restrictive selection criteria (such as citations or reputation), but...

  9. [2022]

    This list of journals was then manually curated to ensure they belong to what is traditionally considered as Gender Studies

    6 across the four chosen languages: gender, sex, woman, women, feminism, feminist, masculinities, lgbt, lesbian, gay, homosexual, bisexual, queer, girl. This list of journals was then manually curated to ensure they belong to what is traditionally considered as Gender Studies....

  10. [2023]

    This is accomplished by combining bibliometrics and NLP techniques, with manual curation

    The corpus is divided into core Gender Studies publications, and articles from different disciplines. This is accomplished by combining bibliometrics and NLP techniques, with manual curation. These techniques complement each other to avoid biases or gaps, both in the model's o...

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.