REVIEW 3 major objections 5 minor 15 references
Causal Language in Observational Studies: Sociocultural Backgrounds and Team Composition
T0 review · 3 major / 5 minor · reviewed 2026-08-09 · deepseek-v4-flash
Pith's one-line read Who writes an observational study shapes how causal its abstract sounds.
desk verdict A large, carefully controlled study of causal language in observational abstracts; the experience, team-size, and gender findings hold up, but the cross-cultural UAI claim rests on unadjusted country means and needs reframing or reanalysis. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The machinery is a two-stage measurement-and-model pipeline. First, a BioBERT-based sentence classifier, fine-tuned on 3,061 manually annotated conclusion sentences, assigns each abstract-conclusion sentence a label of causal, correlational, or neither (macro-F1 0.89); a conclusion is counted as causal if at least one sentence is causal, with conditional statements such as 'may cause' folded into the correlational category. Second, a logistic linear mixed-effects regression predicts that binary outcome from first- and last-author experience, author gender, author country, team size, with controls including journal rank, publication-year and journal random effects, conclusion length, and over 1,400 MeSH terms. The regression coefficients on the demographics and the country-level correlation with the uncertainty avoidance index are the quantities carrying the argument.
What would settle it
Take a stratified random sample of several thousand abstracts, re-label each conclusion by hand with annotators blind to author identity, and rerun the regression; if the demographic coefficients vanish or reverse under human labels, the transformer's systematic misclassification is the real driver. A cheaper check is to compute the classifier's error rate as a function of author country, gender, and experience and test whether errors correlate with the predictors.
Extended reading notes
Core claim
On the paper's own terms, the central discovery is a set of systematic demographic and cultural correlates of causal language in observational-study conclusions. More experienced authorship predicts more conservative wording: each doubling of a first author's publication count lowers the odds of a causal claim by about 8%, and each doubling of a last author's publication count by about 11%. Larger teams also write more cautiously, with about 9% lower odds per doubling of coauthor count, and male last authors use causal claims at higher rates than female last authors. At the country level, the proportion of conclusions with causal claims correlates with the uncertainty avoidance index (r=0.45 across all 40 countries; r=0.70 among 22 Western countries), so cultures that are less tolerant of ambiguity tend to produce more definitive causal statements.
Load-bearing premise
The result stands on the assumption that the automated classifier's causal/correlational labels are equally valid for authors of different genders, levels of experience, and countries; if the model misreads certain groups' sentence patterns, the measured associations could be artifacts rather than real sociolinguistic signals.
Editorial extensions
If this is right
- Journals and peer reviewers could treat unhedged causal phrasing in observational abstracts as a stylistic signal to check rather than as a neutral description of evidence.
- If seniority and team size are genuine moderators, training and co-authorship practices could be used to reduce overstated causal claims.
- The country-level correlation implies that editorial policies about causal language may have different effects across cultures, and that international review boards might standardize language expectations.
- The same measurement pipeline can be applied to other datasets, such as preprints or non-biomedical observational literatures, to see whether the demographic and cultural correlates generalize.
- A quantitative measure of causal strength, once available, would let researchers separate the evidence-driven portion of causal language from the social portion.
Reading between the lines
- The paper's own admitted gap—no direct measure of causal strength—means the headline conclusion 'not solely driven by evidence' could be weakened if high-UAI countries also happen to publish studies with stronger designs; adding a strength rating would settle that.
- The ecological country-level correlation cannot distinguish individual author attitudes from national writing conventions; a within-country, multilingual extension would test whether the effect is cultural or linguistic.
- Because conditional causal statements were merged into the correlational category, the findings describe unhedged causal claims specifically; different social predictors might emerge for hedging behavior.
- The gender last-author effect could partly be an experience or seniority pathway; a mediation analysis would tell whether male last authors' higher causal-language rates are independent of publication history.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper analyzes 91,933 observational-study abstracts from PubMed and asks whether the use of causal language in abstract conclusions is associated with author experience, team size, author gender, and national culture, after controlling for journal, publication year, study design, and more than 1,400 MeSH terms. The dependent variable is a BioBERT-based classifier label indicating whether a conclusion contains at least one causal sentence. Using logistic linear mixed-effects regression, the authors report that causal language is more common among less experienced first and last authors, smaller teams, male last authors, and authors from countries with higher uncertainty-avoidance index (UAI) scores, and they interpret this as evidence that sociocultural backgrounds and team composition shape scientific communication beyond the strength of the evidence.
Significance. If the results hold, the paper makes a useful contribution to the sociology of science and to debates about overstatement in observational research. The main regression is carefully specified: it includes random effects for journal and year, controls for study design, journal rank, and an extensive MeSH-term set, and uses a large, publicly documented corpus with open code and data. The experience and team-size effects are robust and substantively interesting, and the gender finding for last authors is a meaningful addition to the literature on gender and scientific language. The UAI result, however, is not supported by the analysis as currently presented, and because that result is one of the four headline findings, the paper as a whole requires substantial revision. The ecological, partially post-hoc nature of the UAI analysis is the main load-bearing weakness; the measurement-error issue in the dependent variable is a secondary but important concern.
major comments (3)
- [Section 2, 'Country and uncertainty avoidance culture'; Figure 2; Table 1] The abstract and Discussion claim that authors from higher-UAI countries use causal language more after controlling for journal, design, year, and MeSH terms, but no model with a UAI term is estimated. Table 1 and SI Table S3 include country fixed effects but no UAI covariate, and Figure 2 plots raw country means against UAI. Raw country means can be confounded by country-specific field mixes, journal submission patterns, and study-design distributions, and with only 40 unweighted countries the correlation is fragile. The 'Western' subset (r = 0.70) appears to have been selected after inspecting the data; the main text does not state this was a pre-specified hypothesis. Because the UAI finding is a headline conclusion, please re-run the analysis with UAI included as a country-level covariate in the mixed-effects model (or using the adjusted country effects from Table 1), report the coefficient and confidence interval for all 40 countries, and justify or pre-specify any subgroup analysis.
- [Section 4 and SI B.1 (dependent variable)] The regression treats the BioBERT classifier's predicted label as the true outcome without propagating measurement error. The classifier has macro-F1 0.89, which is good, but if misclassification is correlated with the predictors (e.g., if certain author groups or countries use sentence constructions the classifier misreads), the reported associations could be biased. Given that the entire analysis rests on this outcome, please add a sensitivity analysis: for example, use the classifier's probability or confidence scores as a continuous outcome, re-estimate on a manually validated subset, or at least discuss the likely direction and magnitude of bias from known error patterns. This concern is distinct from the UAI issue and applies to all four headline findings.
- [Section 4, 'Author Country' and SI B.2] Papers are assigned to a country only when all author affiliations belong to that country with average confidence above 0.8; 18,399 multinational papers are pooled into 'Others'. This means the country-level analysis covers only single-country papers, and the UAI correlation is computed on those means. The paper should state explicitly that the UAI result pertains to single-country teams only and discuss whether the exclusion of multinational collaborations could affect the cultural interpretation. Adding UAI as a country-level predictor in the full mixed model would at least use the model's country effects rather than raw means.
minor comments (5)
- [Abstract and throughout] The phrase 'causal language are more common' is grammatically incorrect; it should be 'causal language is more common' or 'causal expressions are more common.'
- [Table 1] The country 'Columbia' should be spelled 'Colombia'.
- [SI Table S3] There are typographical errors in the random-effects labels: 'Astract_conclusion_length' should be 'Abstract_conclusion_length', and 'absense' should be 'absence'.
- [Figure 2 caption and main text] The '22 Western cultural countries' are not defined; please provide a list or a criterion in the supplementary material so the subset is reproducible.
- [SI B.1] The sentence 'our specifically trained model still have advantages over ChatGPT' has subject-verb agreement errors; also, the comparison to ChatGPT is reported only by citation to prior work, so please state whether those comparisons used the same data and evaluation protocol as Table S1.
Circularity Check
No significant circularity; the regression associations are empirical and do not reduce to fitted inputs or self-citation.
full rationale
The paper's central claim is an empirical association between author/team/country variables and the prevalence of causal language in observational-study abstracts. The dependent variable is produced by a BioBERT classifier trained in the authors' prior work (Yu et al., 2019). This is a measurement instrument, not a derived prediction: the regression coefficients for experience, team size, gender, and country are not constructed from the classifier's labels in any way that would force the reported associations. The classifier is validated by cross-validation and by comparisons to ChatGPT-family models, so the self-citation is supporting evidence rather than a load-bearing circular premise. The UAI analysis in Figure 2 is an unadjusted ecological correlation of country means against Hofstede's UAI, while Table 1 contains country fixed effects but no UAI term; this means the abstract's 'after controlling for' claim is not actually tested for UAI, and the raw means may be confounded. That is a correctness/evidential gap, not a circular derivation, because the correlation is computed directly from observed country means and the UAI index rather than from the paper's own fitted parameters. No equation, fitted coefficient, or self-cited theorem is reused as its own output. The paper is self-contained against external benchmarks (Cofield et al.'s 31% comparison, external gender/name benchmarks, S2AND disambiguation), so no circularity score is warranted.
Assumptions & free parameters
free parameters (3)
- Gender confidence thresholds =
0.82 male / 0.78 female
- Country assignment confidence threshold =
0.8
- Journal rank imputation =
0.49
assumptions (5)
- domain assumption The BioBERT classifier's causal/correlational labels accurately operationalize 'causal language' as used in the study.
- domain assumption Hofstede's Uncertainty Avoidance Index validly represents national cultural attitudes toward uncertainty.
- domain assumption PubMed's 'Observational Study' MeSH publication type correctly identifies observational studies.
- domain assumption Semantic Scholar author disambiguation (S2AND) provides acceptable author identity resolution.
- standard math Logistic linear mixed-effects model assumptions hold (no perfect separation, random effects adequately specified).
Cite this review
Pith. "Pith review of Causal Language in Observational Studies: Sociocultural Backgrounds and Team Composition." pith.science (2026). https://pith.science/paper/PTZ65V6P
@misc{pith2026250212159,
author = {Pith},
title = {Pith review of: Causal Language in Observational Studies: Sociocultural Backgrounds and Team Composition},
year = {2026},
howpublished = {\url{https://pith.science/paper/PTZ65V6P}},
note = {Machine review of arXiv:2502.12159}
}
read the original abstract
The use of causal language in observational studies has raised concerns about overstatement in scientific communication. While some argue that such language should be reserved for randomized controlled trials, others contend that rigorous causal inference methods can justify causal claims in observational research. Ideally, causal language should align with the strength of the underlying evidence. However, through the analysis of over 90,000 abstracts from observational studies using computational linguistic and regression methods, we found that causal language are more common in work by less experienced authors, smaller research teams, male last authors, and researchers from countries with higher uncertainty avoidance indices. Our findings suggest that the use of causal language is not solely driven by the strength of evidence, but also by the sociocultural backgrounds of authors and their team composition. This work provides a new perspective for understanding systematic variations in scientific communication and emphasizes the importance of recognizing these human factors when evaluating scientific claims.
Figures
Reference graph
Works this paper leans on
-
[2]
Division of Infectious Diseases
For journals in our dataset that could not be matched to the 2024 data via their ISSN, we attempted to retrieve their information from previous years. If a journal remained unmatched (which occurred for only 0.35% of the papers), we assigned it an SJR score of 0.49, corresponding to the bottom quartile in our list of 2https://www.scimagojr.com/journalrank...
-
[4]
doi:10.1177/0361684310392728. URL https://journals. sagepub.com/doi/10.1177/0361684310392728. Jinhyuk Lee, Wonjin Yoon, Sungdong Kim, Donghyeon Kim, Sunkyu Kim, Chan Ho So, and Jae- woo Kang. BioBERT: a pre-trained biomedical language representation model for biomedical text mining. Bioinformatics, 36(4):1234–1240,
-
[8]
Judea Pearl and Dana Mackenzie
Accessed: 2025-02-13. Judea Pearl and Dana Mackenzie. The Book of Why. Basic Books, New York,
work page 2025
-
[9]
S2AND: A benchmark and evaluation system for author name disambiguation
Shivashankar Subramanian, Daniel King, Doug Downey, and Sergey Feldman. S2AND: A benchmark and evaluation system for author name disambiguation. 2021 ACM/IEEE Joint Conference on Dig- ital Libraries (JCDL) , pages 170–179,
work page 2021
-
[13]
URL https://www.aclweb.org/anthology/D19-1473.pdf. SI Supplemental Material A Data Collection A.1 Obtaining Observational Studies This study adopts the definition of an observational study provided by the NIH National Library of Medicine1, which states: “[...] a work that reports on the results of a clinical study in which participants may receive diagnos...
work page 2014
-
[14]
We then excluded publications that were also labeled as Randomized Con- trolled Trialor Clinical Trial, leaving us with 176,336 exclusive observational studies. Next, we removed 1https://www.ncbi.nlm.nih.gov/mesh/68064888 12 papers that lacked either a structured abstract with a conclusion subsection in the abstract or an English full text, resulting in 1...
-
[16]
on 91,933 observations. In our model, the dependent variable indicates whether the conclusion subsection of an article’s structured abstract contains at least one sentence with a causal claim. The primary explanatory variables include: (1) Author writing experience (measured as the number of observational studies published by the first/last authors, log-t...
work page 2013
-
[1988]
doi:10.2307/3178066. Miguel A Hernán. The c-word: scientific euphemisms do not improve causal inference from observa- tional data. American journal of public health, 108(5):616–619,
Show all 15 references
-
[1995]
Eva Thue V old
URL https://hbr.org/1995/09/ the-power-of-talk-who-gets-heard-and-why . Eva Thue V old. Epistemic modality markers in research articles: a cross-linguistic and cross-disciplinary study. International Journal of Applied Linguistics , 16(1):61–87,
1995
-
[2006]
Jun Wang
doi:10.1111/j.1473- 4192.2006.00106.x. Jun Wang. A lightgbm-based approach to predicting gender likelihood from personal names. https: //github.com/junwang4/name-to-gender-inference ,
2006
-
[2011]
Campbell Leaper and Rachael D
doi:10.1515/text.2011.004. Campbell Leaper and Rachael D. Robnett. Women are more likely than men to use tentative language, aren’t they? a meta-analysis testing for gender differences and moderators. Psychology of Women Quarterly, 35(1):129–142,
2011 doi
-
[2014]
Charles Bazerman
doi:10.18637/jss.v067.i01. Charles Bazerman. Scientific writing as a social act: A review of the literature of the sociology of science. In Paul V . Anderson, R. John Brockmann, and Carolyn R. Miller, editors, New Essays in Technical and Scientific Communication: Research, The...
-
[2019]
An NLP Analysis of Exaggerated Claims in Science News
Yingya Li, Jieke Zhang, and Bei Yu. An NLP Analysis of Exaggerated Claims in Science News. In Proceedings of the 2017 EMNLP Workshop: Natural Language Processing meets Journalism, pages 106–111,
2017
-
[2020]
URL https://academic.oup.com/bioinformatics/article/36/4/1234/5566506
doi:10.1093/bioinformatics/btz682. URL https://academic.oup.com/bioinformatics/article/36/4/1234/5566506. Marc J Lerchenmueller, Olav Sorenson, and Anupam B Jena. Gender differences in how scientists present the importance of their research: observational study. bmj, 367,
-
[2024]
org/abs/2411.09675
URL https://arxiv. org/abs/2411.09675. Bei Yu, Yingya Li, and Jun Wang. Detecting causal language use in science findings. In EMNLP’2019, pages 4656–4666,
2019 arXiv
Reviewed August 9, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.