REVIEW 3 major objections 6 minor 43 references
Statistical realism is not evidence that LLMs can estimate treatment effects in social science experiments
T0 review · 3 major / 6 minor · reviewed 2026-08-02 · deepseek-v4-flash
Pith's one-line read This paper challenges the assumption that descriptive fit qualifies LLMs to estimate treatment effects; it argues that statistical realism and causal accuracy are separate targets in behavioral simulation.
desk verdict A large-scale, careful demonstration that descriptive fit of LLM survey simulators does not certify treatment-effect accuracy; the qualitative claim holds, but the headline numbers need a noise-floor analysis before being taken at face value. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing machinery is a paired measurement design: the same simulated responses are scored twice, once for statistical realism (individual-level mean absolute error against observed responses) and once for causal fidelity (absolute error of predicted average treatment effects against human experimental ATEs). Variance decomposition, intraclass correlations, permutation tests, and text-perturbation probes are used to show that the two errors are structurally distinct. The key mechanism behind the divergence is the models' tendency to infer behavior from attitudes more strongly than humans do, which the authors trace to narrative text in training data and the next-token prediction obje
What would settle it
A re-analysis of the same data that weights countries by sample representativeness and finds a strong country-level correlation (e.g., r > 0.7) between descriptive mean absolute error and ATE error—or a prompting strategy that consistently lowers both on held-out countries—would contradict the paper's central claim.
Extended reading notes
Core claim
The central discovery is a quantified separation between descriptive fit and causal fidelity in LLM-based behavioral simulation. When the models predicted responses to 11 interventions, they reproduced country-level means and individual-level attitudinal patterns reasonably well, but their average treatment effect (ATE) estimates deviated from human experimental benchmarks by roughly 9 percentage points on a 0–100 scale—about double the error of simple supervised baselines. The two accuracy dimensions were governed by different variance structures: outcome type dominated descriptive error, while method, country, and their interactions dominated causal error. The same prompting strategy that
Load-bearing premise
The human experimental ATEs used as ground truth are assumed to be unbiased estimates of each country's true intervention effect; if national samples are unrepresentative or noisy, part of what the paper calls LLM causal error is actually benchmark error.
Editorial extensions
If this is right
- Descriptive validation alone is insufficient for any causal use of LLM simulators; deploying them without direct causal benchmarks risks propagating unseen errors into decisions.
- Prompting strategies chosen for descriptive fit can worsen ATE accuracy, so model and prompt selection should be based on causal criteria, not realism.
- Causal errors vary systematically by intervention logic: interventions that rely on evoking internal experience (situational simulation) show the largest overestimation, while cultural and group norm interventions show the smallest bias.
- Behavioral outcomes are more distorted than attitudinal ones, so an intervention that looks effective at the attitude level in simulation may not translate into behavior.
- Descriptive accuracy at the country level does not identify countries with causal accuracy; fairness assessments based only on descriptive fit can miss the populations whose intervention effects are most misestimated.
Reading between the lines
- The attitude-behavior over-coupling is likely not unique to climate psychology; because narrative text generally compresses the intention-behavior gap, similar descriptive-causal divergence may appear in health, finance, or civic behavior simulations.
- The paper's paired-evaluation design could be turned into a routine diagnostic for LLM-generated synthetic data: always report ATE error alongside realism, and treat realism as insufficient for causal claims.
- If the weak descriptive-causal correlation holds broadly, the only reliable way to select prompts and models for causal tasks is to build small targeted experimental benchmarks per domain rather than relying on distributional fit.
- A stronger testable extension would be to check whether fine-tuning an LLM on observed responses (rather than prompt engineering) reduces the descriptive-causal gap or merely improves realism while leaving causal error unchanged.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper evaluates three LLMs (GPT-4o-mini, Gemini 2.5 Flash Lite, Claude 3 Haiku) as behavioral simulators of climate-psychology interventions, using the ICCP experiment (59,508 participants, 62 countries) and replicating key analyses in two additional cross-national experiments (Spampatti et al., Većkalov et al.). It contrasts descriptive fit (country-level means, individual MAE, distributional metrics) with causal fidelity (absolute error of predicted country-level ATEs against observed experimental ATEs). The main claims are that prompting refinements improve descriptive fit but not causal fidelity, that descriptive and causal errors have different variance structures, that LLMs impose stronger attitude-behavior coupling than human data, and that country-level descriptive quality is weakly correlated with causal accuracy. The paper concludes that statistical realism is insufficient evidence for treatment-effect accuracy.
Significance. If the results are robust, this is a valuable and timely cautionary contribution to the growing literature on LLM-based behavioral simulation. The study uses an unusually large pre-registered cross-national experimental dataset, tests multiple LLMs and prompting methods, replicates the central analyses in two independent datasets, and employs permutation tests and multilevel decompositions. It also avoids the most obvious circularity by selecting prompting methods on development-set descriptive MAE rather than on ATE error. However, the causal benchmark is itself noisy, and the paper does not quantify how that noise affects the headline correlations, variance decompositions, and ICCs. The central conceptual claim is plausible and important, but several load-bearing quantitative claims need additional support.
major comments (3)
- [Materials and Methods, 'Causal fidelity'; Figs. 3 and 5a] The country-level ATEs from the ICCP experiment are treated as noise-free ground truth. With roughly 59,508/(62×12)≈80 participants per arm per country, the sampling SE of a country-level ATE is about 4–6 p.p. for 0–100 outcomes with SD≈25, and many intervention ATEs have 95% CIs crossing zero (Fig. 4a). Under the metric |predicted ATE – observed ATE|, this measurement error attenuates the descriptive–causal correlations in Fig. 5a, injects variance into the method and method×outcome terms interpreted in Fig. 3a, and—because interventions within a country share a control group—adds a country-level component to the causal ICCs in Fig. 3c. The Discussion lists non-representative national samples as a limitation, but no noise-floor quantification or robustness check is reported. Please add a bootstrap/simulation analysis that propagates the uncertainty in the observed ATEs, report confidenc
- [Materials and Methods, 'Descriptive-causal correlation'; Fig. 5a] The paper reports r=0.22, 0.15, and 0.11 without confidence intervals, and the unit of analysis is ambiguous. The Methods say the correlation is computed 'within each outcome-by-method combination' across 62 countries, but Fig. 5a plots baseline, few-shot, and VBN-CoT points together. It is unclear whether the reported r is for baseline only with n=62, pooled across methods with n=186, or some other aggregation. Please state the exact specification and report CIs; the headline claim that the association is 'weak' depends on this choice.
- [Results, 'Descriptive and causal errors are governed by different structures'; Fig. 3a] The variance decomposition showing that outcome type dominates descriptive error (92.8–94.5%) is partly mechanical: action MAE (~40 p.p.) is much larger than belief/policy MAE (~16 p.p.), so outcome necessarily accounts for most descriptive variance, whereas ATE errors are more comparable across outcomes. The claim that causal error has a 'different structure' would be more convincing if the decomposition were also shown on standardized or within-outcome metrics, or if the scale difference were explicitly modeled. Without this, part of the structural shift may reflect the different scaling of the two error metrics rather than a substantive difference in error sources.
minor comments (6)
- [Abstract and Results] The abstract uses 'statistical realism' while the main text uses 'descriptive fit.' Define this equivalence at first use to avoid ambiguity.
- [Results, Fig. 2j paragraph] The sentence 'VBN-CoT reduced mean absolute ATE error by 33.9%–42.8%... with 3.9 p.p. for belief, 3.5 p.p. for policy, and 8.7 p.p. for action' is ambiguous: it should state explicitly that these values refer to GPT (or whichever model).
- [Materials and Methods, Statistical analysis] The Results mention 'precision-weighted analyses,' but no precision-weighting procedure is described in the Methods. Add a description or remove the phrase.
- [Supplementary Table S1] Table S1 compares zero-shot LLMs with OLS/LASSO trained on 80% of the observed responses. This is not an apples-to-apples comparison; the text should not use it to infer that LLM causal error is not 'solely a consequence of task difficulty.' Soften the claim or provide a genuinely comparable baseline.
- [Fig. 4a] Please define 'direction flipped' and 'CI crosses 0' explicitly in the legend. It is currently unclear whether the direction-flipped marker refers to the predicted-vs-observed sign or only to the observed CI including zero.
- [Materials and Methods, Perturbation analysis] The text says perturbed texts were generated using 'GPT-5.4 thinking.' Specify the model version and date, since future readers may need to reproduce this step.
Circularity Check
No significant circularity: the central comparison is a benchmark against external human experiments, and prompting refinements were selected on descriptive MAE, not ATE error.
full rationale
The paper's central claim—that statistical realism does not reliably imply causal fidelity—rests on an empirical comparison between LLM-simulated responses and externally published experimental data (ICCP: Vlasceanu et al.; Spampatti et al.; Veckalov et al.). The two key metrics are defined independently: country-level MAE is the mean of person-level absolute differences between predicted and observed construct scores, while causal fidelity is defined as |predicted ATE – true ATE|, with true ATEs taken from the randomized experiments (Materials and Methods, Evaluation: 'Causal fidelity'). Neither metric is defined in terms of the other, and the weak country-level correlations (r = 0.11–0.22, Fig. 5a) are an empirical result, not an algebraic identity. Critically, the prompting refinements (few-shot and VBN-CoT) were selected using individual-level MAE on a control-group development sample (n = 5,093), not on ATE error, so the reported divergence—and even the worsening of ATE error under few-shot in some cells—is not a fitted artifact of the target quantity. The manuscript contains no load-bearing self-citation: the authors (Li, Ji) cite external benchmark datasets and prior external studies on descriptive–causal divergence, but no uniqueness theorem or prior result of their own is invoked to force the conclusion. The one limitation passage that deserves flagging is in the Discussion: 'the national samples may not be uniformly representative of each country’s full population, so the human data used for causal fidelity may be more reliable in some settings than in others.' This is a benchmark-validity caveat (sampling noise in ground-truth ATEs could attenuate the reported MAE–ATE correlations), but it is a measurement-error problem affecting external validity, not a circular step: the model outputs are not fit to those ATEs. Thus the derivation is self-contained against external benchmarks; the appropriate circularity verdict is near zero, with the benchmark-noise caveat noted under correctness risk rather than circularity.
Assumptions & free parameters
assumptions (3)
- domain assumption Human experimental data from ICCP, Spampatti, and Veckalov provide unbiased ground-truth ATEs for the simulated targets.
- domain assumption Three lightweight LLMs queried at temperature=0 constitute a valid test of the general 'LLMs' claim.
- ad hoc to paper The a priori categorization of 11 interventions into deductive reasoning, situational simulation, and cultural/group norms reflects underlying psychological mechanisms.
Cite this review
Pith. "Pith review of Statistical realism is not evidence that LLMs can estimate treatment effects in social science experiments." pith.science (2026). https://pith.science/paper/MDTV3Z5T
@misc{pith2026260402458,
author = {Pith},
title = {Pith review of: Statistical realism is not evidence that LLMs can estimate treatment effects in social science experiments},
year = {2026},
howpublished = {\url{https://pith.science/paper/MDTV3Z5T}},
note = {Machine review of arXiv:2604.02458}
}
read the original abstract
Large language models (LLMs) are increasingly used to simulate human responses and estimate treatment effect of interventions when real-world experiments are costly or infeasible. The treatment-effect estimates are often evaluated using statistical realism, the degree to which simulated responses reproduce properties of observed human responses, although whether realism predicts treatment-effect accuracy remains unknown. Here we test this proxy relationship by jointly measuring statistical realism and treatment-effect accuracy on the same simulated responses in a cross-national experiment with 59,508 participants from 62 countries using three LLMs. The correlation between statistical realism and treatment-effect accuracy is weak, and optimizing for statistical realism can even worsen treatment-effect accuracy when selecting models, prompts, and target populations. The pattern replicates in two additional cross-national experiments spanning 12 and 27 countries with 20,785 participants. The divergence between the two reflects distinct error structures and is larger for behavioral outcomes, where models appear to extrapolate behavioral effects from attitudinal patterns. Because this divergence may remain hidden in deployment, errors can propagate into simulation-informed decisions. We introduce a diagnostic framework for LLM-generated synthetic data and discuss how treatment-effect validation should proceed under varying availability of experimental benchmarks. Simulated responses and simulated treatment effects are distinct estimation targets, and evidence for one does not certify the other.
Reference graph
Works this paper leans on
-
[1]
E. P. Fenichel, C. Castillo-Chavez, M. G. Ceddia, et al. Adaptive human behavior in epidemiological models.Proceedings of the National Academy of Sciences, 108(15):6306–11, 2011
2011
-
[2]
Alexander Haslam, et al
Kai Ruggeri, Friederike Stock, S. Alexander Haslam, et al. A synthesis of evidence for policy from behavioural science during covid-19.Nature, 625(7993):134–147, 2024
2024
-
[3]
Validating vignette and conjoint survey experiments against real- world behavior.Proceedings of the National Academy of Sciences, 112(8):2395–2400, 2015
Jens Hainmueller, Dominik Hangartner, and Teppei Yamamoto. Validating vignette and conjoint survey experiments against real- world behavior.Proceedings of the National Academy of Sciences, 112(8):2395–2400, 2015
2015
-
[4]
Patrik Michaelsen, Aksel Sundström, and Sverker C. Jagers. Mass support for conserving 30% of the earth by 2030: Experimen- tal evidence from five continents.Proceedings of the National Academy of Sciences, 122(35):e2503355122, 2025
-
[5]
Bruch and J
E. Bruch and J. Atwell. Agent-based models in empirical so- cial research.Sociological Methods&Research, 44(2):186–221, 2015
2015
-
[6]
Maria del Rio-Chanona, et al
Marco Pangallo, Alberto Aleta, R. Maria del Rio-Chanona, et al. The unequal effects of the health–economy trade-offduring the covid-19 pandemic.Nature Human Behaviour, 8(2):264–275, 2024
2024
-
[7]
Sorgente, R
A. Sorgente, R. Caliciuri, M. Robba, M. Lanz, and B. D. Zumbo. A systematic review of latent class analysis in psychology: Exam- ining the gap between guidelines and research practice.Behavior Research Methods, 57(11):301, 2025. When simulations look right but causal effects go wrong: Large language models as beha vioral simulators12
2025
-
[8]
Argyle, Ethan C
Lisa P. Argyle, Ethan C. Busby, Nancy Fulda, et al. Out of one, many: Using language models to simulate human samples. Political Analysis, 31(3):337–351, 2023
2023
Show all 43 references
-
[9]
Clinton, Cassy Dorff, Brenton Kenkel, and Jennifer M
James Bisbee, Joshua D. Clinton, Cassy Dorff, Brenton Kenkel, and Jennifer M. Larson. Synthetic replacements for human survey data? the perils of large language models.Political Analysis, 32(4):401–416, 2024
2024
-
[10]
Specializing large language models to simulate survey response distributions for global populations
Yong Cao, Haijiang Liu, Arnav Arora, et al. Specializing large language models to simulate survey response distributions for global populations. InProceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Compu- tational Linguistics: Huma...
2025
-
[11]
Finetuning llms for human behavior prediction in so- cial science experiments
Akaash Kolluri, Shengguang Wu, Joon Sung Park, and Michael S Bernstein. Finetuning llms for human behavior prediction in so- cial science experiments. InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, page 30084–30099, 2025
2025
-
[12]
Christopher A. Bail. Can generative ai improve social sci- ence?Proceedings of the National Academy of Sciences, 121(21):e2314021121, 2024
2024
-
[13]
Beyond weird: Can synthetic survey participants substitute for humans in global policy research?Behavioral Science&Policy, 10(2):26–45, 2024
Pujen Shrestha, Dario Krpan, Fatima Koaik, et al. Beyond weird: Can synthetic survey participants substitute for humans in global policy research?Behavioral Science&Policy, 10(2):26–45, 2024
2024
-
[14]
Large lan- guage models surpass human experts in predicting neuroscience results.Nature Human Behaviour, 9(2):305–315, 2025
Xiaoliang Luo, Akilles Rechardt, Guangzhi Sun, et al. Large lan- guage models surpass human experts in predicting neuroscience results.Nature Human Behaviour, 9(2):305–315, 2025
2025
-
[15]
Argyle, Ethan C
Lisa P. Argyle, Ethan C. Busby, Joshua R. Gubler, et al. Test- ing theories of political persuasion using ai.Proceedings of the National Academy of Sciences, 122(18):e2412815122, 2025
2025
-
[16]
Large language models empowered agent-based modeling and simulation: a survey and perspectives.Humanities and Social Sciences Communications, 11(1):1259, 2024
Chen Gao, Xiaochong Lan, Nian Li, et al. Large language models empowered agent-based modeling and simulation: a survey and perspectives.Humanities and Social Sciences Communications, 11(1):1259, 2024
2024
-
[17]
Simulating human opinions with large language models: Opportunities and challenges for personalized survey data modeling, 2025
Carolin Kaiser, Jakob Kaiser, Vladimir Manewitsch, Lea Rau, and Rene Schallner. Simulating human opinions with large language models: Opportunities and challenges for personalized survey data modeling, 2025
2025
-
[18]
Pat Pataranutaporn, Nattavudh Powdthavee, Chayapatr Archi- waranguprok, and Pattie Maes. Simulating human well-being with large language models: Systematic validation and misestima- tion across 64,000 individuals from 64 countries.Proceedings of the National Academy of Science...
2025
-
[19]
A founda- tion model to predict and capture human cognition.Nature, 644(8078):1002–1009, 2025
Marcel Binz, Elif Akata, Matthias Bethge, et al. A founda- tion model to predict and capture human cognition.Nature, 644(8078):1002–1009, 2025
2025
-
[20]
Using large language models to simulate multiple humans and replicate human subject studies
Gati V Aher, Rosa I Arriaga, and Adam Tauman Kalai. Using large language models to simulate multiple humans and replicate human subject studies. InInternational conference on machine learning, page 337–371. PMLR, 2023
2023
-
[21]
A large-scale replication of scenario-based experiments in psychology and management using large language models.Nature Computational Science, 5(8):627–634, 2025
Ziyan Cui, Ning Li, and Huaikang Zhou. A large-scale replication of scenario-based experiments in psychology and management using large language models.Nature Computational Science, 5(8):627–634, 2025
2025
-
[22]
Marcelo Sartori Locatelli, Pedro Dutenhefner, Arthur Buzelin, et al. Ai and climate change discourse: What opinions do large language models present? InProceedings of the 2nd Workshop on Natural Language Processing Meets Climate Change (Cli- mateNLP 2025), page 113–125. Associ...
2025
-
[23]
Agentsociety: Large- scale simulation of llm-driven generative agents advances un- derstanding of human behaviors and society.arXiv preprint arXiv:2502.08691, 2025
Jinghua Piao, Yuwei Yan, Jun Zhang, et al. Agentsociety: Large- scale simulation of llm-driven generative agents advances un- derstanding of human behaviors and society.arXiv preprint arXiv:2502.08691, 2025
2025 arXiv
-
[24]
Predicting results of social science experiments using large language models.Preprint, 2024
Luke Hewitt, Ashwini Ashokkumar, Isaias Ghezae, and Robb Willer. Predicting results of social science experiments using large language models.Preprint, 2024
2024
-
[25]
Take caution in using llms as human surrogates.Proceedings of the National Academy of Sciences, 122(24):e2501660122, 2025
Yuan Gao, Dokyun Lee, Gordon Burtch, and Sina Fazelpour. Take caution in using llms as human surrogates.Proceedings of the National Academy of Sciences, 122(24):e2501660122, 2025
2025
-
[26]
Gui, Tianyi Peng, et al
Olivier Toubia, George Z. Gui, Tianyi Peng, et al. Database report: Twin-2k-500: A data set for building digital twins of over 2,000 people based on their answers to over 500 questions.Marketing Science, 44(6):1446–1455, 2025
2025
-
[27]
This human study did not involve human subjects: Validat- ing llm simulations as behavioral evidence.arXiv preprint arXiv:2602.15785, 2026
Jessica Hullman, David Broska, Huaman Sun, and Aaron Shaw. This human study did not involve human subjects: Validat- ing llm simulations as behavioral evidence.arXiv preprint arXiv:2602.15785, 2026
2026
-
[28]
Doell, Boryana Todorova, Madalina Vlasceanu, et al
Kimberly C. Doell, Boryana Todorova, Madalina Vlasceanu, et al. The international climate psychology collaboration: Climate change-related data collected from 63 countries.Scientific Data, 11(1):1066, 2024
2024
-
[29]
Doell, Joseph B
Madalina Vlasceanu, Kimberly C. Doell, Joseph B. Bak-Coleman, et al. Addressing climate change with behavioral science: A global intervention tournament in 63 countries.Science Advances, 10(6):eadj5778, 2024
2024
-
[30]
Tobia Spampatti, Ulf J. J. Hahnel, Evelina Trutnevyte, and Tobias Brosch. Psychological inoculation strategies to fight climate disinformation across 12 countries.Nature Human Behaviour, 8(2):380–398, 2024
2024
-
[31]
Geiger, František Bartoš, et al
Bojana Ve´ckalov, Sandra J. Geiger, František Bartoš, et al. A 27- country test of communicating the scientific consensus on climate change.Nature Human Behaviour, 8(10):1892–1905, 2024
1905
-
[32]
The challenge of using llms to simulate human behavior: A causal inference perspective.arXiv preprint arXiv:2312.15524, 2023
George Gui and Olivier Toubia. The challenge of using llms to simulate human behavior: A causal inference perspective.arXiv preprint arXiv:2312.15524, 2023
2023
-
[33]
V . A. Shaffer, E. S. Focella, A. Hathaway, L. D. Scherer, and B. J. Zikmund-Fisher. On the usefulness of narratives: An in- terdisciplinary review and theoretical model.Ann Behav Med, 52(5):429–442, 2018
2018
-
[34]
Vieira, S
J. Vieira, S. L. Castro, and A. S. Souza. Psychological barriers moderate the attitude-behavior gap for climate change.PLoS One, 18(7):e0287404, 2023
2023
-
[35]
Generative language models exhibit social identity biases.Nature Computa- tional Science, 5(1):65–75, 2025
Tiancheng Hu, Yara Kyrychenko, Steve Rathje, et al. Generative language models exhibit social identity biases.Nature Computa- tional Science, 5(1):65–75, 2025
2025
-
[36]
Gallegos, Ryan A
Isabel O. Gallegos, Ryan A. Rossi, Joe Barrow, et al. Bias and fairness in large language models: A survey.Computational Linguistics, 50(3):1097–1179, 2024
2024
-
[37]
Moore, Daniel M
Suhaib Abdurahman, Alireza Salkhordeh Ziabari, Alexander K. Moore, Daniel M. Bartels, and Morteza Dehghani. A primer for evaluating large language models in social-science research. Advances in Methods and Practices in Psychological Science, 8(2):25152459251325174, 2025
2025
-
[38]
The work for environmental protection task: A consequential web-based procedure for study- ing pro-environmental behavior.Behavior Research Methods, 54(1):133–145, 2022
Florian Lange and Siegfried Dewitte. The work for environmental protection task: A consequential web-based procedure for study- ing pro-environmental behavior.Behavior Research Methods, 54(1):133–145, 2022
2022
-
[39]
99% of expert climate scientists agree that Earth is warming and climate change is happening, mainly because of human activity
Paul C. Stern, Thomas Dietz, Troy Abel, Gregory A. Guagnano, and Linda Kalof. A value-belief-norm theory of support for social movements: The case of environmentalism.Human Ecology Review, 6(2):81–97, 1999. When simulations look right but causal effects go wrong: Large languag...
1999
-
[40]
evidence: list 2–4 most relevant explicit profile items (must be explicit; no guessing)
-
[41]
values: Based on this person’s political orientation, social position, and education, infer their dominant value orientation on the Schwartz self-enhancement↔self-transcendence dimension. – self-enhancement: prioritizes personal status, wealth, and authority – self-transcenden...
-
[42]
How likely does the profile indicate they have strong perceived severity of climate impacts? – AR: Low|Medium|High
VBN labels (derived from the profile and inferred values; interpreted for belief/accuracy judgments): – AC: Low|Medium|High. How likely does the profile indicate they have strong perceived severity of climate impacts? – AR: Low|Medium|High. How likely does the profile suggest ...
-
[43]
evidence
synthesis: 2–3 sentences (max 50 words) citing which evidence supports the labels, and 2–3 sentences inferring how these VBN factors drive the final answer. Output JSON ONLY: { "evidence": ["...", "..."], "values": "self-enhancement | mixed | self-transcendence", "VBN": {"AC":...
Reviewed August 2, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.