{"id":"dd152bd1-b1b2-4c89-b931-6cd87953b6cb","arxiv_id":"2411.18306","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A hybrid core-journal plus topic-modeling keyword workflow delineates 1.97 million gender/sex related publications in four languages from the Dimensions database, without relying on standard disciplinary classifications.","lead":"Researchers built a 1.97 million document dataset meant to map feminist and gender studies across 60,000+ journals by combining a curated core of 289 specialized journals with automated topic modeling and a manually refined keyword list. The value is a method for tracing research areas that, like gender studies, cross traditional disciplinary boundaries and reflect social movements.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The corpus's precision is unvalidated: the only internal check (core overlap, Fig. 1) is circular, so the claim that this hybrid system surpasses basic keyword search is not supported, especially in biomedical fields where 'sex' and 'women' are often incidental title terms.","rationale":"The reader's weakest_assumption is correct, and I agree with it. The central claim of the paper is not purely technical but a measurement claim: the hybrid method yields a valid corpus of gender/sex related studies. For that claim to hold, the keyword query must have acceptable precision across all disciplines, not just the SSH core from which the vocabulary was drawn. The paper provides no external validation: no gold standard, no manual precision sample, no comparison to a baseline keyword list. The only evidence, the core-overlap percentage in Figure 1, is circular because the keywords were generated from the core itself. This is not an internal logical flaw in the workflow; the two-stage design is coherent and the manual curation is transparently described. The risk is that the method's output is dominated by over-inclusion in biomedical literature, where terms such as 'sex' and 'women' are commonly used as variables or descriptors rather than as indicators of gender scholarship. The paper actually shows the Not Core segment is about 55% biomedical/health/biological fields (Figure 3), so this is not a marginal issue. The fact that the authors rejected Dimensions' own Gender Studies group because only 63% of its documents matched their definition underscores that precision matters; they do not measure their own. Given this, the appropriate outcome is the reader's CONDITIONAL verdict: the contribution is a promising, clearly described method, but the core claim of superior delineation requires external validation and the release of the keyword list and dataset for reproducibility. My read does not change that verdict.","tokens_in":12356,"tokens_out":5331,"duration_ms":48899,"concrete_test":"Sample 400 documents from the Not Core segment in Biomedical and Clinical Sciences, stratified by keyword frequency (e.g., top generic terms 'sex', 'women', 'gender' plus a random sample of other keywords). Two independent annotators, blind to the retrieval category, each classify the title and abstract as either (a) substantively about gender/sex as an object of study or a feminist/gender perspective, or (b) using sex/gender only as a demographic variable or incidentally. Compute percent agreement and Cohen's kappa, then estimate precision. If the lower bound of the 95% confidence interval for precision falls below 0.70, the extended corpus is over-inclusive for the stated object and the 'surpasses keyword search' claim must be revised. A parallel sample from the Core (e.g., 100 documents) can serve as a sanity check on annotation reliability.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The most load-bearing assumption is that a title containing any of the 229 keywords indicates that the document belongs to gender/sex related studies in every one of the 60,919 journals in the extended set. The vocabulary was extracted from a 289-journal Social Sciences and Humanities core (Data and methods) and then applied without restriction to Biomedical and Clinical Sciences, Health Sciences, and Biological Sciences, which make up 37.2%, 12%, and 6.1% of the Not Core segment (Figure 3). In these fields, 'sex' frequently denotes a biological covariate ('sex differences in sepsis outcomes') and 'women' often identifies a study population ('women with breast cancer'), not an engagement with gender studies. The paper's own Figure 4A shows 'women', 'sex', and 'gender' are the top three keywords in the Not Core segment, so the risk is substantial. The only reported validation is that more than half of the Core also matches the keyword query (Figure 1). That overlap is expected, because the keywords were derived from the Core via BERTopic; it cannot demonstrate precision. The authors reject Dimensions' Gender Studies group because only about 63% of its documents align with their definition, yet they never measure their own method's precision. Thus the central claim that this hybrid system surpasses basic keyword search is not established.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a two-stage bibliometric method to delineate gender/sex-related studies using the Dimensions database. Stage one builds a manually curated core of 289 Gender Studies journals. Stage two applies BERTopic topic modeling to the core's titles and abstracts, manually refines the resulting topic terms into a list of 229 keywords (259 regular expressions), and performs a title-only keyword search over all Dimensions journals in English, Spanish, French, and Portuguese. The resulting corpus contains 1,967,302 documents, of which 1,807,272 come from outside the core. The authors claim that this hybrid system 'surpasses basic keyword search' and produces a dataset suitable for bibliometric analyses of Feminist/Gender Studies across disciplines.","tokens_in":12577,"tokens_out":2600,"duration_ms":25204,"significance":"If validated, this work would provide a substantial, openly reusable corpus and a transferable methodological template for delineating 'more-than-disciplinary' fields. The authors give a detailed, transparent account of their workflow, combine expert curation with NLP in a way that is explicitly designed to reduce manual-keyword bias, and commit to publishing the core journal list, topics, keywords, and document identifiers. These are real strengths. However, the central claim of superiority over simple keyword search and the overall validity of the corpus rest on precision and recall evidence that the paper does not provide. The only internal check (the core overlap shown in Figure 1) is expected by construction and cannot establish that the extended set is not dominated by false positives, especially in biomedical fields. As such, the significance is conditional on the addition of a credible external evaluation.","major_comments":[{"comment":"The only validation reported for the keyword retrieval is that 'more than half of the Core overlaps with the documents identified through keyword retrieval' (Figure 1). This overlap is not informative about precision or recall because the keyword list was derived from that same Core via BERTopic and manual refinement. The authors need an external gold standard, such as a random sample of Not Core documents manually annotated for relevance (with agreement statistics), and should report precision, recall, and F-score for the whole corpus and for major disciplinary subsets. Without this, the claim that the hybrid system 'surpasses basic keyword search' is unsupported.","section":"Data and methods; Figure 1"},{"comment":"The keyword vocabulary was extracted from a core restricted to Social Sciences and Humanities journals, but it is then applied to all journals in Dimensions, including Biomedical and Clinical Sciences (37.2% of Not Core), Health Sciences (12%), and Biological Sciences (6.1%) (Figure 3). In those fields, generic title terms such as 'sex', 'women', and 'gender' are often incidental (e.g., 'sex differences in sepsis outcomes', 'women with breast cancer'), and Figure 4A shows that these are indeed the top three keywords in the Not Core segment. Since the search uses titles only and ignores abstracts, the risk of false positives is substantial. The authors should either restrict the expansion to fields where the vocabulary is likely to be indicative, or quantify precision within the biomedical and health sciences subsets. A simple manual audit of a random sample from those disciplines would address this concern.","section":"Data and methods; Figure 3; Figure 4A"},{"comment":"The authors justify rejecting Dimensions' Gender Studies group because only about 63% of its documents align with their definition, yet they provide no equivalent precision measurement for their own method. A direct comparison is needed: on a labeled sample, estimate the precision and recall of the proposed hybrid method, of the Dimensions group, and of a baseline keyword-only query (e.g., titles containing 'gender', 'sex', 'woman', or 'feminist'). Such a comparison is the minimal evidence required to substantiate the statement that this approach 'surpasses basic keyword search.'","section":"Data and methods; Discussion"}],"minor_comments":[{"comment":"The phrase 'inter-, intra-, inter-, post-disciplinary' appears to contain a duplicated or erroneous prefix; likely 'inter-, intra-, post-disciplinary' was intended.","section":"Introduction"},{"comment":"The sentence 'more inclusive acronyms such as LGBT gained popularity in the 20th century' should probably read 'in the 21st century,' since the rise of the inclusive acronym is a recent phenomenon.","section":"Results, Figure 4B caption and text"},{"comment":"Table 2 lists 282 journals in the Core segment, while the Data and methods section states that the core comprises 289 scientific journals. This discrepancy should be reconciled (e.g., if 7 journals had no indexed articles).","section":"Table 2"},{"comment":"The data transparency statement says the lists and DOIs 'will be published in a public repository' once the article is accepted. For reproducibility, it would be preferable to provide them as supplementary material or in a repository at the time of submission, even in preliminary form.","section":"Declarations, Data transparency"}],"recommendation":"major_revision","confidential_remarks":"The core issue is validation: the circularity of the only reported check and the absence of any precision/recall measurement against external judgments. This is fixable within the scope of the paper by adding a manual annotation study of the Not Core set, particularly in biomedical/health fields, and a comparison with existing baselines. If such an evaluation is added and the results are reasonable, the paper would make a solid contribution to bibliometric methodology and to gender studies research infrastructure. If the authors are unable to provide external validation, the central claims should be substantially softened."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear colleague,\n\nThe short version: this paper is a clearly written, well-motivated attempt to delineate gender/sex related research across all of Dimensions, and the 1.9M-document dataset it produces could be useful to the field. But the paper's central claim — that the hybrid system surpasses basic keyword search — is not backed by any external validation, and the only internal check they report (Figure 1) is circular because the keywords were extracted from the same core used to measure the overlap. The precision problem is concrete: in the Not Core segment, which is mostly biomedical, 'sex' and 'women' are the top terms, and in those fields those words often indicate a covariate or study population, not engagement with feminist scholarship. So I'd treat the dataset as a promising starting point rather than a validated corpus.\n\nWhat's genuinely new is the specific combination: a manually curated 289-journal core, BERTopic topic modeling over titles and abstracts, manual refinement into 229 keywords (259 regexes) in four languages, and a title-only search over Dimensions. That workflow is described in enough detail that it could be reproduced or adapted to other 'more-than-disciplinary' fields, and the authors are appropriately careful about issues like regional coverage and the role of books in feminist theory. The literature review is honest about the limitations of previous keyword-based approaches.\n\nThe soft spots, in order of seriousness:\n- No gold-standard evaluation. They reject Dimensions' Gender Studies group because only 63% of its documents match their definition, but they never measure their own precision or recall. That asymmetry is glaring.\n- The overlap check is expected by construction. The authors say Figure 1 'shows that the keywords used were indeed characteristic of the discipline,' but a keyword list derived from the core will match the core. It says nothing about false positives outside it.\n- Title-only search with generic terms. Terms like 'sex', 'women', 'gender' in a biomedical title are often incidental. The paper offers no evidence that these terms transfer from the SSH core to medicine and biology, where most of the extended corpus sits (Figure 3 shows Biomedical + Health + Biological make up over 55% of Not Core).\n- Artifacts are promised but not yet released; the keyword list, core journal list, and DOIs are essential for anyone to assess the result.\n\nIs this fatal? Not for the method itself — the two-stage core-and-extension idea is sound and the workflow is a real contribution. But the empirical claim that this approach is better than keyword search is not established by the evidence in the paper. A revision that includes an external validation (e.g., manual annotation of a random sample from the Not Core, stratified by discipline, and comparison with an existing classification or key-author approach) plus the release of artifacts could bring it to a convincing state.\n\nWho this is for: scientometricians working on field delineation, especially for fields tied to social movements; also anyone building corpora for gender studies research. I'd send it to a serious referee — the problem is important and the method is worth engaging with — but I'd expect heavy revision before it's publishable as a validated method.\n\nBest,\n[Your name]","headline":"Useful two-stage core-and-keyword workflow for delineating gender studies, but the precision claim is unvalidated and the only internal check is circular; needs external validation and artifact release before it can be trusted as a corpus.","tokens_in":13152,"tokens_out":3095,"would_cite":true,"duration_ms":27787,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A two-stage method turns 289 Gender Studies journals into a 1.9-million-document map of feminist research.","keywords":["feminist studies","gender studies","bibliometrics","topic modeling","BERTopic","Dimensions database","gender/sex related studies","core-and-extension method"],"falsifier":"Randomly sample 500 documents from the Not Core segment that were retrieved by generic tokens such as 'sex', 'gender', or 'woman' in biomedical journals, and have domain experts judge whether each is substantively about gender/sex. If most are routine clinical or biological studies that merely mention sex as a variable, then the title-keyword assumption overcounts and the 1.9-million corpus is not a faithful delineation.","tokens_in":12115,"feed_emoji":"📚","tokens_out":8442,"duration_ms":67423,"temperature":0.7,"pith_summary":"The paper tries to establish that Feminist and Gender Studies can be delimited bibliometrically by a two-stage hybrid: start from a manually curated core of 289 specialized journals, let BERTopic topic modeling extract the vocabulary of that core, and then search that vocabulary in article titles across the whole Dimensions database. The authors claim this approach beats a basic keyword search because the keywords are anchored in the field's own literature rather than in a researcher's manual enumeration, while manual review of the topics corrects the model's gaps and biases. The resulting dataset of 1,967,302 documents in English, Spanish, French, and Portuguese from 1668 to 2023 would provide the empirical basis for mapping the topics, citation flows, collaboration patterns, and institutional and regional participation of gender/sex related research inside and outside the core. If the method transfers, the same recipe could delineate other research areas that escape disciplinary classifications, such as Black Studies, Human Rights, or Agroecology.","feed_headline":"New method maps 1.9 million gender-studies papers","feed_subtitle":"Combining 289 core journals with topic-modeled keywords traces feminist research across disciplines and languages.","key_machinery":"The machinery is the core-to-periphery vocabulary transfer. A set of 289 manually curated Gender Studies journals forms the core; BERTopic topic modeling of that core's titles and abstracts produces 330 topics with their characteristic words; manual refinement turns those words into the final list of 229 keywords (259 regular expressions) that is then run against article titles in every journal. The central object is this keyword list, since it is the instrument that decides whether a document enters the 1.9-million corpus. Its constituent words do the work of carrying the field's vocabulary from the social sciences and humanities into biomedical and clinical journals, and the paper's results are all downstream of that transfer.","core_discovery":"On its own terms, the paper's central claim is that a hybrid core-and-extension pipeline delineates gender/sex related studies more faithfully than any existing classification. The authors build a core of 289 Social Sciences and Humanities journals specializing in Gender Studies, apply BERTopic to their titles and abstracts to obtain 330 topics, manually condense the characteristic words into 229 keywords (259 regular expressions) in four languages, and search article titles across all Dimensions journals. This yields 1,967,302 documents from 1668 to 2023, with 91.9% of them coming from outside the core. The asymmetry between the Core and the Not Core segments is read as evidence that the method captures both gender studies as a discipline and gender/sex as a transversal perspective across biomedicine, health, and other fields; for example, 'gender' dominates the Core while 'sex' dominates the Not Core, and 'feminis' terms are notably concentrated in the Core. The authors contend that this two-stage design reflects the dynamic interaction between Gender Studies and the disciplines it influences.","pith_inferences":["Because the search is restricted to titles, the dataset probably misses gender/sex research that signals its topic only in abstracts or full text; an abstract-based variant would trade precision for recall and could be tested against the current corpus.","The precision of the 1.9-million count is an empirical question about generic terms: many biomedical hits on 'sex' or 'woman' may be routine variable mentions rather than gender scholarship, so a precision study on a random sample would determine how much of the corpus is substantively gender/sex related.","A natural extension is to use the same pipeline with abstracts rather than titles, or with full texts, and compare topic distributions across languages and regions to see whether the concept of 'gender/sex' spreads unevenly.","The method's portability is not automatic: the curated core is the hard-won part, and other social-movement fields will need their own core, their own topic-modeling step, and their own manual review before the keyword list is trustworthy."],"forward_implications":["Researchers can map gender/sex related research by topic, citation, collaboration, institution, country, and language, using a corpus that does not depend on journal or paper-level disciplinary labels.","The core/not-core split gives a quantitative read on how feminist vocabulary diffuses: 'gender' and 'feminis' terms stay concentrated in the core while 'sex', 'abortion', and 'menstrual' dominate the biomedical periphery.","The method tracks conceptual change in titles, such as the rise of 'gender' over 'sex' since the 1980s and the growing co-occurrence of the two terms.","The same two-stage workflow can be applied to other 'more-than-disciplinary' conversations whose boundaries are not captured by existing classifications."],"supporting_citations":[{"why":"Supplies BERTopic, the neural topic-modeling procedure used to generate the 330 topics and characteristic words from the core corpus.","marker":"Grootendorst, 2022"},{"why":"Introduces the Dimensions database, the sole data source for journal selection and keyword-based document retrieval.","marker":"Herzog et al., 2020"},{"why":"Provides the large-scale comparison of bibliographic sources that justifies Dimensions' balance of coverage and metadata quality.","marker":"Visser et al., 2021"},{"why":"Establishes the core-and-extension logic and warns that relying on 'sex' or 'gender' terms alone cannot delimit Gender Studies.","marker":"Lundgren et al., 2015"},{"why":"Supplies the 'more-than-discipline' and bottom-up frameworks that motivate the two-stage, vocabulary-led delineation.","marker":"Sugimoto & Weingart, 2015"},{"why":"Characterizes Women's Studies as a field of overlapping circles, supporting the idea of a specialized core plus a dispersed periphery.","marker":"Buker, 2003"}],"fun_headline_variants":["Method pinpoints 1.9M gender studies docs","NLP traces feminist research across 4 languages","Hybrid core-and-extension maps gender studies","From 1668 to 2023: 1.9M gender papers","Bibliometric mix reveals gender studies' reach"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that an article title containing any of the 229 keywords is a reliable sign that the article belongs to gender/sex related studies, including in biomedical journals, even though the keywords were derived from a Social Sciences and Humanities core.","fun_headline_variants_meta":{"raw":{"variants":["Method pinpoints 1.9M gender studies docs","NLP traces feminist research across 4 languages","Hybrid core-and-extension maps gender studies","From 1668 to 2023: 1.9M gender papers","Bibliometric mix reveals gender studies' reach"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000234,"raw_usage":{"total_tokens":1524,"prompt_tokens":997,"completion_tokens":527,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":613,"completion_tokens_details":{"reasoning_tokens":448}},"tokens_in":613,"tokens_out":527,"duration_ms":5151,"temperature":1.0,"reasoning_tokens":448,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T11:18:22.257763+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Randomly sample 500 documents from the Not Core segment that were retrieved by generic tokens such as 'sex', 'gender', or 'woman' in biomedical journals, and have domain experts judge whether each is substantively about gender/sex. If most are routine clinical or biological studies that merely mention sex as a variable, then the title-keyword assumption overcounts and the 1.9-million corpus is not a faithful delineation.","supporting_citations":[],"review_version":1}