Pith. sign in

REVIEW 3 major objections 5 minor 15 references

Causal Language in Observational Studies: Sociocultural Backgrounds and Team Composition

T0 review · 3 major / 5 minor · reviewed 2026-08-09 · deepseek-v4-flash

Pith's one-line read Who writes an observational study shapes how causal its abstract sounds.

desk verdict A large, carefully controlled study of causal language in observational abstracts; the experience, team-size, and gender findings hold up, but the cross-cultural UAI claim rests on unadjusted country means and needs reframing or reanalysis. read the letter →

arxiv 2502.12159 v2 pith:PTZ65V6P submitted 2025-02-04 physics.soc-ph cs.CL

classification physics.soc-phcs.CL
keywords causallanguageobservationalstudiesscientificcommunicationuncertaintyavoidanceauthorexperiencegenderdifferencesteamsizeBioBERT
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper asks whether the use of causal wording in observational-study abstracts is shaped only by the quality of the evidence or also by who is doing the writing. Analyzing 91,933 structured abstracts with a transformer-based sentence classifier and logistic mixed-effects regression, it finds that causal claims appear more often when the first or last author has published fewer observational studies, when the team is smaller, when the last author is male, and when authors come from countries with higher uncertainty-avoidance scores. These associations survive controls for journal rank, publication year, study design, and more than 1,400 topic terms. If the pattern is real, it means the rhetorical caution of scientific prose is partly a social and cultural product, not merely an epistemic one.

What carries the argument

The machinery is a two-stage measurement-and-model pipeline. First, a BioBERT-based sentence classifier, fine-tuned on 3,061 manually annotated conclusion sentences, assigns each abstract-conclusion sentence a label of causal, correlational, or neither (macro-F1 0.89); a conclusion is counted as causal if at least one sentence is causal, with conditional statements such as 'may cause' folded into the correlational category. Second, a logistic linear mixed-effects regression predicts that binary outcome from first- and last-author experience, author gender, author country, team size, with controls including journal rank, publication-year and journal random effects, conclusion length, and over 1,400 MeSH terms. The regression coefficients on the demographics and the country-level correlation with the uncertainty avoidance index are the quantities carrying the argument.

What would settle it

Take a stratified random sample of several thousand abstracts, re-label each conclusion by hand with annotators blind to author identity, and rerun the regression; if the demographic coefficients vanish or reverse under human labels, the transformer's systematic misclassification is the real driver. A cheaper check is to compute the classifier's error rate as a function of author country, gender, and experience and test whether errors correlate with the predictors.

Watch

Extended reading notes

Core claim

On the paper's own terms, the central discovery is a set of systematic demographic and cultural correlates of causal language in observational-study conclusions. More experienced authorship predicts more conservative wording: each doubling of a first author's publication count lowers the odds of a causal claim by about 8%, and each doubling of a last author's publication count by about 11%. Larger teams also write more cautiously, with about 9% lower odds per doubling of coauthor count, and male last authors use causal claims at higher rates than female last authors. At the country level, the proportion of conclusions with causal claims correlates with the uncertainty avoidance index (r=0.45 across all 40 countries; r=0.70 among 22 Western countries), so cultures that are less tolerant of ambiguity tend to produce more definitive causal statements.

Load-bearing premise

The result stands on the assumption that the automated classifier's causal/correlational labels are equally valid for authors of different genders, levels of experience, and countries; if the model misreads certain groups' sentence patterns, the measured associations could be artifacts rather than real sociolinguistic signals.

Editorial extensions

If this is right

  • Journals and peer reviewers could treat unhedged causal phrasing in observational abstracts as a stylistic signal to check rather than as a neutral description of evidence.
  • If seniority and team size are genuine moderators, training and co-authorship practices could be used to reduce overstated causal claims.
  • The country-level correlation implies that editorial policies about causal language may have different effects across cultures, and that international review boards might standardize language expectations.
  • The same measurement pipeline can be applied to other datasets, such as preprints or non-biomedical observational literatures, to see whether the demographic and cultural correlates generalize.
  • A quantitative measure of causal strength, once available, would let researchers separate the evidence-driven portion of causal language from the social portion.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The paper's own admitted gap—no direct measure of causal strength—means the headline conclusion 'not solely driven by evidence' could be weakened if high-UAI countries also happen to publish studies with stronger designs; adding a strength rating would settle that.
  • The ecological country-level correlation cannot distinguish individual author attitudes from national writing conventions; a within-country, multilingual extension would test whether the effect is cultural or linguistic.
  • Because conditional causal statements were merged into the correlational category, the findings describe unhedged causal claims specifically; different social predictors might emerge for hedging behavior.
  • The gender last-author effect could partly be an experience or seniority pathway; a mediation analysis would tell whether male last authors' higher causal-language rates are independent of publication history.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper analyzes 91,933 observational-study abstracts from PubMed and asks whether the use of causal language in abstract conclusions is associated with author experience, team size, author gender, and national culture, after controlling for journal, publication year, study design, and more than 1,400 MeSH terms. The dependent variable is a BioBERT-based classifier label indicating whether a conclusion contains at least one causal sentence. Using logistic linear mixed-effects regression, the authors report that causal language is more common among less experienced first and last authors, smaller teams, male last authors, and authors from countries with higher uncertainty-avoidance index (UAI) scores, and they interpret this as evidence that sociocultural backgrounds and team composition shape scientific communication beyond the strength of the evidence.

Significance. If the results hold, the paper makes a useful contribution to the sociology of science and to debates about overstatement in observational research. The main regression is carefully specified: it includes random effects for journal and year, controls for study design, journal rank, and an extensive MeSH-term set, and uses a large, publicly documented corpus with open code and data. The experience and team-size effects are robust and substantively interesting, and the gender finding for last authors is a meaningful addition to the literature on gender and scientific language. The UAI result, however, is not supported by the analysis as currently presented, and because that result is one of the four headline findings, the paper as a whole requires substantial revision. The ecological, partially post-hoc nature of the UAI analysis is the main load-bearing weakness; the measurement-error issue in the dependent variable is a secondary but important concern.

major comments (3)
  1. [Section 2, 'Country and uncertainty avoidance culture'; Figure 2; Table 1] The abstract and Discussion claim that authors from higher-UAI countries use causal language more after controlling for journal, design, year, and MeSH terms, but no model with a UAI term is estimated. Table 1 and SI Table S3 include country fixed effects but no UAI covariate, and Figure 2 plots raw country means against UAI. Raw country means can be confounded by country-specific field mixes, journal submission patterns, and study-design distributions, and with only 40 unweighted countries the correlation is fragile. The 'Western' subset (r = 0.70) appears to have been selected after inspecting the data; the main text does not state this was a pre-specified hypothesis. Because the UAI finding is a headline conclusion, please re-run the analysis with UAI included as a country-level covariate in the mixed-effects model (or using the adjusted country effects from Table 1), report the coefficient and confidence interval for all 40 countries, and justify or pre-specify any subgroup analysis.
  2. [Section 4 and SI B.1 (dependent variable)] The regression treats the BioBERT classifier's predicted label as the true outcome without propagating measurement error. The classifier has macro-F1 0.89, which is good, but if misclassification is correlated with the predictors (e.g., if certain author groups or countries use sentence constructions the classifier misreads), the reported associations could be biased. Given that the entire analysis rests on this outcome, please add a sensitivity analysis: for example, use the classifier's probability or confidence scores as a continuous outcome, re-estimate on a manually validated subset, or at least discuss the likely direction and magnitude of bias from known error patterns. This concern is distinct from the UAI issue and applies to all four headline findings.
  3. [Section 4, 'Author Country' and SI B.2] Papers are assigned to a country only when all author affiliations belong to that country with average confidence above 0.8; 18,399 multinational papers are pooled into 'Others'. This means the country-level analysis covers only single-country papers, and the UAI correlation is computed on those means. The paper should state explicitly that the UAI result pertains to single-country teams only and discuss whether the exclusion of multinational collaborations could affect the cultural interpretation. Adding UAI as a country-level predictor in the full mixed model would at least use the model's country effects rather than raw means.
minor comments (5)
  1. [Abstract and throughout] The phrase 'causal language are more common' is grammatically incorrect; it should be 'causal language is more common' or 'causal expressions are more common.'
  2. [Table 1] The country 'Columbia' should be spelled 'Colombia'.
  3. [SI Table S3] There are typographical errors in the random-effects labels: 'Astract_conclusion_length' should be 'Abstract_conclusion_length', and 'absense' should be 'absence'.
  4. [Figure 2 caption and main text] The '22 Western cultural countries' are not defined; please provide a list or a criterion in the supplementary material so the subset is reproducible.
  5. [SI B.1] The sentence 'our specifically trained model still have advantages over ChatGPT' has subject-verb agreement errors; also, the comparison to ChatGPT is reported only by citation to prior work, so please state whether those comparisons used the same data and evaluation protocol as Table S1.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity; the regression associations are empirical and do not reduce to fitted inputs or self-citation.

full rationale

The paper's central claim is an empirical association between author/team/country variables and the prevalence of causal language in observational-study abstracts. The dependent variable is produced by a BioBERT classifier trained in the authors' prior work (Yu et al., 2019). This is a measurement instrument, not a derived prediction: the regression coefficients for experience, team size, gender, and country are not constructed from the classifier's labels in any way that would force the reported associations. The classifier is validated by cross-validation and by comparisons to ChatGPT-family models, so the self-citation is supporting evidence rather than a load-bearing circular premise. The UAI analysis in Figure 2 is an unadjusted ecological correlation of country means against Hofstede's UAI, while Table 1 contains country fixed effects but no UAI term; this means the abstract's 'after controlling for' claim is not actually tested for UAI, and the raw means may be confounded. That is a correctness/evidential gap, not a circular derivation, because the correlation is computed directly from observed country means and the UAI index rather than from the paper's own fitted parameters. No equation, fitted coefficient, or self-cited theorem is reused as its own output. The paper is self-contained against external benchmarks (Cofield et al.'s 31% comparison, external gender/name benchmarks, S2AND disambiguation), so no circularity score is warranted.

Assumptions & free parameters 3 free parameters · 5 assumptions · 0 invented entities

The paper introduces no new theoretical entities. The main numerical inputs that shape the results are the classifier thresholds and the country/gender confidence cutoffs, all fitted to benchmark or training data. The Hofstede UAI score and the MeSH study-design labels are taken from prior literature and external databases.

free parameters (3)
  • Gender confidence thresholds = 0.82 male / 0.78 female
    Tuned on a 5,779-name benchmark to maximize F1 under a constraint that at most 12.5% of names are unknown (SI B.3). Affects the gender variable used in the regression.
  • Country assignment confidence threshold = 0.8
    A paper is assigned to a country only if the average confidence of affiliation classification exceeds 0.8 and all affiliations agree (SI B.2). Defines the 'others' reference group.
  • Journal rank imputation = 0.49
    Unmatched journals (0.35% of papers) are assigned the bottom-quartile SJR score (SI A.2). Assumed to be a reasonable proxy for unlisted journals.
assumptions (5)
  • domain assumption The BioBERT classifier's causal/correlational labels accurately operationalize 'causal language' as used in the study.
    The dependent variable is a predicted label with macro-F1 0.89; measurement error is not propagated into the regression (SI B.1).
  • domain assumption Hofstede's Uncertainty Avoidance Index validly represents national cultural attitudes toward uncertainty.
    Used in Figure 2 to interpret country-level differences; the index is a widely used but contested cultural metric (Hofstede et al., 2010).
  • domain assumption PubMed's 'Observational Study' MeSH publication type correctly identifies observational studies.
    Inclusion criterion for the corpus (SI A.1). Misclassification would introduce selection bias.
  • domain assumption Semantic Scholar author disambiguation (S2AND) provides acceptable author identity resolution.
    Experience counts rely on author IDs for 99.8% of authors; errors in disambiguation could bias the experience variable (SI A.3).
  • standard math Logistic linear mixed-effects model assumptions hold (no perfect separation, random effects adequately specified).
    Standard statistical model with lme4; no diagnostics such as variance components or convergence checks are reported (SI C).

how reviews work

0 comments
Cite this review

Pith. "Pith review of Causal Language in Observational Studies: Sociocultural Backgrounds and Team Composition." pith.science (2026). https://pith.science/paper/PTZ65V6P

@misc{pith2026250212159,
  author       = {Pith},
  title        = {Pith review of: Causal Language in Observational Studies: Sociocultural Backgrounds and Team Composition},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/PTZ65V6P}},
  note         = {Machine review of arXiv:2502.12159}
}
read the original abstract

The use of causal language in observational studies has raised concerns about overstatement in scientific communication. While some argue that such language should be reserved for randomized controlled trials, others contend that rigorous causal inference methods can justify causal claims in observational research. Ideally, causal language should align with the strength of the underlying evidence. However, through the analysis of over 90,000 abstracts from observational studies using computational linguistic and regression methods, we found that causal language are more common in work by less experienced authors, smaller research teams, male last authors, and researchers from countries with higher uncertainty avoidance indices. Our findings suggest that the use of causal language is not solely driven by the strength of evidence, but also by the sociocultural backgrounds of authors and their team composition. This work provides a new perspective for understanding systematic variations in scientific communication and emphasizes the importance of recognizing these human factors when evaluating scientific claims.

Figures

Figures reproduced from arXiv: 2502.12159 by the authors.

Figure 1
Figure 1. Distribution of the dependent variable (causal vs correlational) along with the key variables of [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. Authors from countries with higher uncertainty avoidance index scores—reflecting a greater cul [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

15 extracted references · 12 canonical work pages

  1. [2]

    Division of Infectious Diseases

    For journals in our dataset that could not be matched to the 2024 data via their ISSN, we attempted to retrieve their information from previous years. If a journal remained unmatched (which occurred for only 0.35% of the papers), we assigned it an SJR score of 0.49, corresponding to the bottom quartile in our list of 2https://www.scimagojr.com/journalrank...

  2. [4]

    URL https://journals

    doi:10.1177/0361684310392728. URL https://journals. sagepub.com/doi/10.1177/0361684310392728. Jinhyuk Lee, Wonjin Yoon, Sungdong Kim, Donghyeon Kim, Sunkyu Kim, Chan Ho So, and Jae- woo Kang. BioBERT: a pre-trained biomedical language representation model for biomedical text mining. Bioinformatics, 36(4):1234–1240,

  3. [8]

    Judea Pearl and Dana Mackenzie

    Accessed: 2025-02-13. Judea Pearl and Dana Mackenzie. The Book of Why. Basic Books, New York,

  4. [9]

    S2AND: A benchmark and evaluation system for author name disambiguation

    Shivashankar Subramanian, Daniel King, Doug Downey, and Sergey Feldman. S2AND: A benchmark and evaluation system for author name disambiguation. 2021 ACM/IEEE Joint Conference on Dig- ital Libraries (JCDL) , pages 170–179,

  5. [13]

    URL https://www.aclweb.org/anthology/D19-1473.pdf. SI Supplemental Material A Data Collection A.1 Obtaining Observational Studies This study adopts the definition of an observational study provided by the NIH National Library of Medicine1, which states: “[...] a work that reports on the results of a clinical study in which participants may receive diagnos...

  6. [14]

    We then excluded publications that were also labeled as Randomized Con- trolled Trialor Clinical Trial, leaving us with 176,336 exclusive observational studies. Next, we removed 1https://www.ncbi.nlm.nih.gov/mesh/68064888 12 papers that lacked either a structured abstract with a conclusion subsection in the abstract or an English full text, resulting in 1...

  7. [16]

    In our model, the dependent variable indicates whether the conclusion subsection of an article’s structured abstract contains at least one sentence with a causal claim

    on 91,933 observations. In our model, the dependent variable indicates whether the conclusion subsection of an article’s structured abstract contains at least one sentence with a causal claim. The primary explanatory variables include: (1) Author writing experience (measured as the number of observational studies published by the first/last authors, log-t...

  8. [1988]

    Miguel A Hernán

    doi:10.2307/3178066. Miguel A Hernán. The c-word: scientific euphemisms do not improve causal inference from observa- tional data. American journal of public health, 108(5):616–619,

Show all 15 references
  1. [1995]

    Eva Thue V old

    URL https://hbr.org/1995/09/ the-power-of-talk-who-gets-heard-and-why . Eva Thue V old. Epistemic modality markers in research articles: a cross-linguistic and cross-disciplinary study. International Journal of Applied Linguistics , 16(1):61–87,

  2. [2006]

    Jun Wang

    doi:10.1111/j.1473- 4192.2006.00106.x. Jun Wang. A lightgbm-based approach to predicting gender likelihood from personal names. https: //github.com/junwang4/name-to-gender-inference ,

  3. [2011]

    Campbell Leaper and Rachael D

    doi:10.1515/text.2011.004. Campbell Leaper and Rachael D. Robnett. Women are more likely than men to use tentative language, aren’t they? a meta-analysis testing for gender differences and moderators. Psychology of Women Quarterly, 35(1):129–142,

  4. [2014]

    Charles Bazerman

    doi:10.18637/jss.v067.i01. Charles Bazerman. Scientific writing as a social act: A review of the literature of the sociology of science. In Paul V . Anderson, R. John Brockmann, and Carolyn R. Miller, editors, New Essays in Technical and Scientific Communication: Research, The...

  5. [2019]

    An NLP Analysis of Exaggerated Claims in Science News

    Yingya Li, Jieke Zhang, and Bei Yu. An NLP Analysis of Exaggerated Claims in Science News. In Proceedings of the 2017 EMNLP Workshop: Natural Language Processing meets Journalism, pages 106–111,

  6. [2020]

    URL https://academic.oup.com/bioinformatics/article/36/4/1234/5566506

    doi:10.1093/bioinformatics/btz682. URL https://academic.oup.com/bioinformatics/article/36/4/1234/5566506. Marc J Lerchenmueller, Olav Sorenson, and Anupam B Jena. Gender differences in how scientists present the importance of their research: observational study. bmj, 367,

  7. [2024]

    org/abs/2411.09675

    URL https://arxiv. org/abs/2411.09675. Bei Yu, Yingya Li, and Jun Wang. Detecting causal language use in science findings. In EMNLP’2019, pages 4656–4666,

Pith tools

Reviewed August 9, 2026 · model on record in the stance chip above.