{"id":"86c74bdd-3fd6-4ac9-a19a-eae105c9bbd3","arxiv_id":"2502.12159","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Causal language in observational-study abstracts is more common among less experienced authors, smaller teams, male last authors, and authors from high-uncertainty-avoidance countries.","lead":"This paper analyzed over 90,000 published abstracts of observational studies and found that causal wording appears more often when the first or last author is less experienced, when the team is smaller, when the last author is male, and when the authors come from countries that score high on cultural uncertainty avoidance. It suggests that how strongly scientists phrase their conclusions is shaped not only by the evidence but also by who is writing.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The UAI finding is computed from unadjusted country means and a post-hoc Western subset; the paper's claim that the cultural effect survives controls is not actually tested.","rationale":"The reader's weakest assumption is classifier measurement error. I agree that predicted labels are treated as truth, but the model is validated (macro-F1=0.89) and misclassification would need to be systematically correlated with author country, gender, or experience to generate the reported pattern; that is possible but not demonstrated. The UAI correlation is a clearer vulnerability: the paper's own model formula (SI Table S3) has no UAI term, and the Figure 2 analysis is at country level using averages that the text never says are covariate-adjusted. Since the abstract explicitly claims the UAI association survives journal/design/MeSH/year controls, a raw aggregate correlation is insufficient. The Western-subset r=0.70 is especially fragile: 22 unweighted points, subset chosen after seeing the full 40-country plot, with no robustness check. This does not invalidate the experience/team-size/gender findings, which are estimated in the adjusted individual-level regression, so the conditional verdict stands; the condition should be that the authors supply an adjusted UAI analysis.","tokens_in":12753,"tokens_out":5320,"duration_ms":55845,"concrete_test":"Refit the model with UAI as the country-level predictor: replace or augment the country dummies in SI Table S3 with Hofstede UAI (or add UAI as a second-level variable in a country random-intercept logistic model), keeping all individual-level controls and random effects. Test the UAI coefficient; then repeat using precision-weighted regression of the estimated country effects from Table 1 on UAI, with leave-one-country-out and separate Western/non-Western estimates. If the adjusted UAI effect is not positive and significant, the cultural component of the central claim is unsupported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The most load-bearing part of the central claim is the uncertainty-avoidance (UAI) result. Section 2 ('Country and uncertainty avoidance culture') and Figure 2 report Pearson r = 0.45 across 40 countries and r = 0.70 for 22 'Western' countries, but the y-axis is described only as 'countries' average use of causal language.' The regression in Table 1/SI Table S3 includes country dummies and controls for journal rank, study design, year, and 1,400+ MeSH terms, but it contains no UAI term, and the main text does not state that the Figure 2 averages are adjusted using those controls. Raw country means can be confounded by country-specific field mixes, journal submission patterns, or study-design distributions; with only 40 (or 22) unweighted countries, the correlation is fragile, and the Western subset was selected after inspecting the data. Therefore the abstract's claim that authors from high-UAI countries use causal language more 'after controlling for' confounders is not actually supported by the analysis as presented.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper analyzes 91,933 observational-study abstracts from PubMed and asks whether the use of causal language in abstract conclusions is associated with author experience, team size, author gender, and national culture, after controlling for journal, publication year, study design, and more than 1,400 MeSH terms. The dependent variable is a BioBERT-based classifier label indicating whether a conclusion contains at least one causal sentence. Using logistic linear mixed-effects regression, the authors report that causal language is more common among less experienced first and last authors, smaller teams, male last authors, and authors from countries with higher uncertainty-avoidance index (UAI) scores, and they interpret this as evidence that sociocultural backgrounds and team composition shape scientific communication beyond the strength of the evidence.","tokens_in":12968,"tokens_out":3225,"duration_ms":36840,"significance":"If the results hold, the paper makes a useful contribution to the sociology of science and to debates about overstatement in observational research. The main regression is carefully specified: it includes random effects for journal and year, controls for study design, journal rank, and an extensive MeSH-term set, and uses a large, publicly documented corpus with open code and data. The experience and team-size effects are robust and substantively interesting, and the gender finding for last authors is a meaningful addition to the literature on gender and scientific language. The UAI result, however, is not supported by the analysis as currently presented, and because that result is one of the four headline findings, the paper as a whole requires substantial revision. The ecological, partially post-hoc nature of the UAI analysis is the main load-bearing weakness; the measurement-error issue in the dependent variable is a secondary but important concern.","major_comments":[{"comment":"The abstract and Discussion claim that authors from higher-UAI countries use causal language more after controlling for journal, design, year, and MeSH terms, but no model with a UAI term is estimated. Table 1 and SI Table S3 include country fixed effects but no UAI covariate, and Figure 2 plots raw country means against UAI. Raw country means can be confounded by country-specific field mixes, journal submission patterns, and study-design distributions, and with only 40 unweighted countries the correlation is fragile. The 'Western' subset (r = 0.70) appears to have been selected after inspecting the data; the main text does not state this was a pre-specified hypothesis. Because the UAI finding is a headline conclusion, please re-run the analysis with UAI included as a country-level covariate in the mixed-effects model (or using the adjusted country effects from Table 1), report the coefficient and confidence interval for all 40 countries, and justify or pre-specify any subgroup analysis.","section":"Section 2, 'Country and uncertainty avoidance culture'; Figure 2; Table 1"},{"comment":"The regression treats the BioBERT classifier's predicted label as the true outcome without propagating measurement error. The classifier has macro-F1 0.89, which is good, but if misclassification is correlated with the predictors (e.g., if certain author groups or countries use sentence constructions the classifier misreads), the reported associations could be biased. Given that the entire analysis rests on this outcome, please add a sensitivity analysis: for example, use the classifier's probability or confidence scores as a continuous outcome, re-estimate on a manually validated subset, or at least discuss the likely direction and magnitude of bias from known error patterns. This concern is distinct from the UAI issue and applies to all four headline findings.","section":"Section 4 and SI B.1 (dependent variable)"},{"comment":"Papers are assigned to a country only when all author affiliations belong to that country with average confidence above 0.8; 18,399 multinational papers are pooled into 'Others'. This means the country-level analysis covers only single-country papers, and the UAI correlation is computed on those means. The paper should state explicitly that the UAI result pertains to single-country teams only and discuss whether the exclusion of multinational collaborations could affect the cultural interpretation. Adding UAI as a country-level predictor in the full mixed model would at least use the model's country effects rather than raw means.","section":"Section 4, 'Author Country' and SI B.2"}],"minor_comments":[{"comment":"The phrase 'causal language are more common' is grammatically incorrect; it should be 'causal language is more common' or 'causal expressions are more common.'","section":"Abstract and throughout"},{"comment":"The country 'Columbia' should be spelled 'Colombia'.","section":"Table 1"},{"comment":"There are typographical errors in the random-effects labels: 'Astract_conclusion_length' should be 'Abstract_conclusion_length', and 'absense' should be 'absence'.","section":"SI Table S3"},{"comment":"The '22 Western cultural countries' are not defined; please provide a list or a criterion in the supplementary material so the subset is reproducible.","section":"Figure 2 caption and main text"},{"comment":"The sentence 'our specifically trained model still have advantages over ChatGPT' has subject-verb agreement errors; also, the comparison to ChatGPT is reported only by citation to prior work, so please state whether those comparisons used the same data and evaluation protocol as Table S1.","section":"SI B.1"}],"recommendation":"major_revision","confidential_remarks":"The paper's core regression findings on experience, team size, and last-author gender appear solid, and the open data/code is a real strength. The UAI result, however, is currently presented in a way that overstates the evidence: no adjusted model with UAI is shown, and the Western-subset correlation looks post hoc. If the authors can add a proper multilevel or country-level covariate analysis and address the classifier measurement-error sensitivity, the paper would be much stronger. The use of the authors' own BioBERT classifier is legitimate, but the paper should acknowledge more directly that the dependent variable is a model prediction."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper is worth a serious read, but the cultural result is the soft spot. What's actually new: they scale causal-language detection to 91,933 observational-study abstracts using a BioBERT classifier they trained (macro-F1 0.89), then run a mixed-effects logistic regression with journal, year, conclusion length, study design, and 1,400+ MeSH terms as controls. The associations for author experience, team size, and last-author gender are strong, highly significant, and survive a demanding control set. The data and code are on GitHub. That is reproducible, careful work, and the scale is a real step beyond the manual analyses in prior literature.\n\nThe problem is the uncertainty-avoidance (UAI) finding. Figure 2 plots raw country means against Hofstede's UAI; there is no UAI term in the regression, and the text gives no indication that the means are adjusted for anything. The Western-subset correlation (r = 0.70 vs. r = 0.45 for all 40 countries) looks like exactly the kind of post-hoc split that inflates apparent effects. So the abstract's phrasing that authors from high-UAI countries use causal language more 'after controlling for' confounders is not supported by the analysis as presented. That claim would require a model with UAI as a country-level predictor (or at least a formal interaction test), not a separate plot of unadjusted averages. The regression does include country dummies, but dummies don't test UAI as a continuous cultural dimension.\n\nA secondary issue is that the dependent variable comes from predicted labels without propagating classifier uncertainty. The authors validated the classifier, but a sensitivity analysis—treating borderline predictions as uncertain, or excluding low-confidence cases—would strengthen confidence that the associations aren't artifacts of misclassification correlated with author characteristics.\n\nAll that said, the core findings—less experienced authors, smaller teams, and male last authors using more causal language—are plausible, consistent with prior work on hedging and gender, and delivered with honest limitations (the authors note they did not assess underlying causal strength). The cultural claim needs to be fixed or explicitly downgraded to an exploratory ecological correlation.\n\nThis paper deserves a serious referee. The resource is valuable, the methods are mostly sound, and the questions matter for metascience and science communication. I'd send it to review with a clear request: either fit UAI properly or reframe the cultural discussion as hypothesis-generating.","headline":"A large, carefully controlled study of causal language in observational abstracts; the experience, team-size, and gender findings hold up, but the cross-cultural UAI claim rests on unadjusted country means and needs reframing or reanalysis.","tokens_in":13457,"tokens_out":2080,"would_cite":true,"duration_ms":22202,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Who writes an observational study shapes how causal its abstract sounds.","keywords":["causal language","observational studies","scientific communication","uncertainty avoidance","author experience","gender differences","team size","BioBERT"],"falsifier":"Take a stratified random sample of several thousand abstracts, re-label each conclusion by hand with annotators blind to author identity, and rerun the regression; if the demographic coefficients vanish or reverse under human labels, the transformer's systematic misclassification is the real driver. A cheaper check is to compute the classifier's error rate as a function of author country, gender, and experience and test whether errors correlate with the predictors.","tokens_in":12567,"feed_emoji":"🔬","tokens_out":8346,"duration_ms":73829,"temperature":0.7,"pith_summary":"The paper asks whether the use of causal wording in observational-study abstracts is shaped only by the quality of the evidence or also by who is doing the writing. Analyzing 91,933 structured abstracts with a transformer-based sentence classifier and logistic mixed-effects regression, it finds that causal claims appear more often when the first or last author has published fewer observational studies, when the team is smaller, when the last author is male, and when authors come from countries with higher uncertainty-avoidance scores. These associations survive controls for journal rank, publication year, study design, and more than 1,400 topic terms. If the pattern is real, it means the rhetorical caution of scientific prose is partly a social and cultural product, not merely an epistemic one.","feed_headline":"Who writes a study shapes how causal its abstract sounds","feed_subtitle":"Over 90,000 abstracts reveal experience, team size, gender, and culture predict causal wording.","key_machinery":"The machinery is a two-stage measurement-and-model pipeline. First, a BioBERT-based sentence classifier, fine-tuned on 3,061 manually annotated conclusion sentences, assigns each abstract-conclusion sentence a label of causal, correlational, or neither (macro-F1 0.89); a conclusion is counted as causal if at least one sentence is causal, with conditional statements such as 'may cause' folded into the correlational category. Second, a logistic linear mixed-effects regression predicts that binary outcome from first- and last-author experience, author gender, author country, team size, with controls including journal rank, publication-year and journal random effects, conclusion length, and over 1,400 MeSH terms. The regression coefficients on the demographics and the country-level correlation with the uncertainty avoidance index are the quantities carrying the argument.","core_discovery":"On the paper's own terms, the central discovery is a set of systematic demographic and cultural correlates of causal language in observational-study conclusions. More experienced authorship predicts more conservative wording: each doubling of a first author's publication count lowers the odds of a causal claim by about 8%, and each doubling of a last author's publication count by about 11%. Larger teams also write more cautiously, with about 9% lower odds per doubling of coauthor count, and male last authors use causal claims at higher rates than female last authors. At the country level, the proportion of conclusions with causal claims correlates with the uncertainty avoidance index (r=0.45 across all 40 countries; r=0.70 among 22 Western countries), so cultures that are less tolerant of ambiguity tend to produce more definitive causal statements.","pith_inferences":["The paper's own admitted gap—no direct measure of causal strength—means the headline conclusion 'not solely driven by evidence' could be weakened if high-UAI countries also happen to publish studies with stronger designs; adding a strength rating would settle that.","The ecological country-level correlation cannot distinguish individual author attitudes from national writing conventions; a within-country, multilingual extension would test whether the effect is cultural or linguistic.","Because conditional causal statements were merged into the correlational category, the findings describe unhedged causal claims specifically; different social predictors might emerge for hedging behavior.","The gender last-author effect could partly be an experience or seniority pathway; a mediation analysis would tell whether male last authors' higher causal-language rates are independent of publication history."],"forward_implications":["Journals and peer reviewers could treat unhedged causal phrasing in observational abstracts as a stylistic signal to check rather than as a neutral description of evidence.","If seniority and team size are genuine moderators, training and co-authorship practices could be used to reduce overstated causal claims.","The country-level correlation implies that editorial policies about causal language may have different effects across cultures, and that international review boards might standardize language expectations.","The same measurement pipeline can be applied to other datasets, such as preprints or non-biomedical observational literatures, to see whether the demographic and cultural correlates generalize.","A quantitative measure of causal strength, once available, would let researchers separate the evidence-driven portion of causal language from the social portion."],"supporting_citations":[{"why":"Supplies the annotated corpus of 3,061 sentences and the BioBERT-based classification task that defines the paper's dependent variable.","marker":"Yu et al., 2019"},{"why":"Provides the pretrained BioBERT model that the classifier fine-tunes.","marker":"Lee et al., 2020"},{"why":"Defines the uncertainty avoidance index used for the country-level cultural analysis.","marker":"Hofstede et al., 2010"},{"why":"Provides the previously published 31% prevalence of causal language in a hand-coded sample, used as a validation point.","marker":"Cofield et al., 2010"},{"why":"Provides the mixed-effects logistic regression implementation used in the main analysis.","marker":"Bates et al., 2014"},{"why":"Justifies focusing on first and last authors and on gender in biomedical authorship.","marker":"Lerchenmueller et al., 2019"},{"why":"Supplies the benchmark dataset used to validate the name-to-gender inference tool.","marker":"Santamaría and Mihaljevic, 2018"},{"why":"Supplies the author-name disambiguation method used to measure publication experience.","marker":"Subramanian et al., 2021"}],"fun_headline_variants":["Causal claims in abstracts vary with author experience and team size","Male last authors and uncertainty-avoiding cultures overstate causality","Less experienced authors and smaller teams use more causal language","Team makeup and culture predict causal overstatement in observational studies"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The result stands on the assumption that the automated classifier's causal/correlational labels are equally valid for authors of different genders, levels of experience, and countries; if the model misreads certain groups' sentence patterns, the measured associations could be artifacts rather than real sociolinguistic signals.","fun_headline_variants_meta":{"raw":{"variants":["Causal claims in abstracts vary with author experience and team size","Male last authors and uncertainty-avoiding cultures overstate causality","Less experienced authors and smaller teams use more causal language","Team makeup and culture predict causal overstatement in observational studies"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00022,"raw_usage":{"total_tokens":1398,"prompt_tokens":849,"completion_tokens":549,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":465,"completion_tokens_details":{"reasoning_tokens":481}},"tokens_in":465,"tokens_out":549,"duration_ms":6812,"temperature":1.0,"reasoning_tokens":481,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-09T13:56:58.396813+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a stratified random sample of several thousand abstracts, re-label each conclusion by hand with annotators blind to author identity, and rerun the regression; if the demographic coefficients vanish or reverse under human labels, the transformer's systematic misclassification is the real driver. A cheaper check is to compute the classifier's error rate as a function of author country, gender, and experience and test whether errors correlate with the predictors.","supporting_citations":[{"cited_title":"S2AND: A benchmark and evaluation system for author name disambiguation","cited_arxiv_id":null,"evidence_quote":"Supplies the author-name disambiguation method used to measure publication experience."}],"review_version":1}