REVIEW 4 major objections 5 minor 8 references
Are we facing a reproducibility crises in materials synthesis? A systematic review of Turkevich AuNP synthesis and CVD MoS2 growth
T0 review · 4 major / 5 minor · reviewed 2026-08-02 · deepseek-v4-flash
Pith's one-line read Systematic review finds most gold nanoparticle and MoS2 synthesis papers omit key parameters needed to replicate the work.
desk verdict Useful reporting audit of two synthesis communities, but it overclaims a meta-analysis that isn't there. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The analysis uses a structured checklist, adapted from the STROBE reporting framework, to score each paper's methodological transparency. The central object is the 'methodological score' defined, for each synthesis type, as a weighted sum of binary answers to whether a paper reported parameters such as precursor concentration, temperature, pH, stirring rate, heating ramp, tube geometry, etc. A weighted threshold of 75% is used to classify a study as methodologically sound. This score operationalizes reproducibility as 'reported completeness', and the paper's conclusions are derived from the distribution of these scores across the literature.
What would settle it
A concrete interlaboratory replication study would settle the question: take a defined set of Turkevich AuNP and CVD MoS2 protocols, run them in at least three independent laboratories or in one lab with deliberately varied conditions (e.g., different stirring rates, pH, or boat distances), and measure the distribution of particle sizes or film morphologies. If outcomes vary widely even when all checklist parameters are matched, the paper's claim that reproducibility is primarily a matter of reporting would be falsified. Conversely, if outcomes match when parameters are controlled, the paper's
Extended reading notes
Core claim
Using an adapted PRISMA/SPIDER systematic-review framework, the authors evaluated 1,299 AuNP papers and 1,573 MoS2 CVD papers, scoring them against weighted checklists of essential experimental parameters. Only 19 AuNP studies (4% of the screened set) and 115 MoS2 studies (24% of that screened set) met a 75% reporting-quality threshold. The least-reported AuNP parameters were solution pH (7%), stirring rate (7%), and number of replicates (2%); for MoS2, the least-reported parameters were boat dimensions (2%), precursor-substrate distance (16%), and tube dimensions (27%). The paper's central claim is that the reproducibility crisis in materials synthesis is largely a crisis of incomplete meth
Load-bearing premise
The paper assumes that a 75% score on its weighted reporting checklist is a valid measure of a study's reproducibility; if a paper can be irreproducible while reporting every checklist item (or reproducible while omitting some), then the paper's conclusions describe reporting habits rather than actual reproducibility.
Editorial extensions
If this is right
- If the paper is right, the perception that Turkevich AuNP synthesis is standard and reproducible needs to be tempered, because the published record lacks the data needed to confirm or refute batch-to-batch consistency.
- For CVD MoS2 growth, the omission of reactor-geometry parameters (tube dimensions, boat distances, precursor position) means that readers cannot reliably transfer growth recipes between laboratories, directly affecting the scalability of 2D materials production.
- The paper's framework provides a transferable method for auditing other nanomaterial synthesis routes, such as quantum dots or metal oxides, to identify which parameters are most commonly underreported.
- Scientific publishers and the community could use the identified underreported parameters (pH, stirring, replicates, precursor distances) as the basis for synthesis-specific reporting checklists and supplementary-information templates.
- If the observed reporting gaps are acknowledged, future studies may begin to include the missing variables, improving the statistical foundation for metaanalyses that compare synthesis outcomes across labs.
Reading between the lines
- A natural extension of the paper's logic is that 'methodologically sound' as defined by a 75% reporting threshold is not the same as 'empirically reproducible' — a paper could report every item on their checklist yet still fail when independently replicated, because checklists cannot capture tacit knowledge and lab-specific details. Conversely, a paper that omits a few items might still produce hi
- If the reporting-based interpretation is accepted, the remedy implied is a shift toward structured, per-synthesis reporting templates (like a 'synthesis recipe card') that journals could enforce. This would be an inexpensive intervention compared to full factorial replication studies.
- The authors note that only one MoS2 publication performed a comprehensive statistical reproducibility assessment (using design-of-experiments). A testable extension is to run a designed interlaboratory study that intentionally varies the least-reported parameters (pH, stirring rate, boat distance) and measure how much of the outcome variance they explain.
- The paper's findings suggest that the meta-analysis stage of such systematic reviews is limited by the very underreporting it identifies: missing data introduce bias. Future reviews might need to impute or handle missing parameters explicitly to make quantitative cross-study comparisons meaningful.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents a systematic review of two widely used synthesis routes—Turkevich gold nanoparticle synthesis and CVD growth of MoS2—claiming to evaluate the reproducibility of these methods via an adapted PRISMA/SPIDER/STROBE framework. The authors retrieved more than 1,300 records per case study, screened abstracts with a Python-based classifier validated against human reviewers, and evaluated full texts with weighted reporting checklists. Their central finding is that only a small fraction of studies (about 2% of AuNP papers and 7% of MoS2 papers from the initial corpora) reached the authors' threshold for methodological soundness, with critical parameters such as pH, stirring rate, and reactor geometry frequently underreported. The abstract and title frame the contribution as a 'systematic review and meta-analysis' of synthesis reproducibility, and Methodology §6 promises a quantitative meta-analysis using the METAFOR environment. However, no meta-analytic results—no pooled effect sizes, heterogeneity statistics, moderator analyses, or forest plots—appear anywhere in the Results or Conclusion.
Significance. If the central claim were supported, this would be an important contribution to the reproducibility discussion in materials synthesis. The paper has genuine strengths: a large, transparent literature retrieval protocol; explicit inclusion/exclusion criteria; a reproducibility check on the abstract-screening step (kappa-based agreement); public code and data on OSF; and a robustness check that relaxes the abstract threshold. These elements are valuable and reproducible. However, the evidence actually presented is a reporting-compliance audit, not an evaluation of reproducibility itself. A fully reported synthesis can be irreproducible across laboratories, and an underreported one can be robust; the paper's own data cannot distinguish these possibilities. The missing meta-analysis and the construct-validity gap in the 'methodologically sound' threshold are load-bearing. The study can be made publishable as a reporting audit with appropriately narrowed claims, but the current framing overreaches.
major comments (4)
- [Methodology §6 and Results (AuNP and MoS2 sections)] The stated meta-analysis is absent. Methodology §6 says a quantitative meta-analysis was conducted using METAFOR to evaluate whether similar experimental conditions yield statistically consistent results, but the Results contain only abstract-screening counts, methodological-score distributions, and item-level reporting frequencies. There are no pooled effect sizes, heterogeneity statistics (e.g., I²), moderator analyses, or forest plots for AuNP size or MoS2 domain size/layer number. The abstract's claim of a 'systematic review and meta-analysis evaluating reproducibility' is therefore unsupported. Either the meta-analysis must be added, or the title/abstract/conclusions must be scaled back to describe a systematic review of methodological reporting.
- [Methodology §5 and Results] The operationalization of 'methodologically sound' as scoring ≥75% on an author-designed weighted checklist makes the low pass rates at least partly tautological. A study is labeled 'rigorously addressed synthesis reproducibility' only if it satisfies this checklist, and then the low pass rate is used to conclude that the literature does not rigorously address reproducibility. No external validation is provided—e.g., whether papers passing the checklist are in fact more reproducible in interlaboratory replication, or whether the checklist items and weights predict reproducibility outcomes. The absence of a sensitivity analysis of the 75% threshold and the arbitrary weights is especially problematic given that the central claim depends entirely on this threshold. Please justify the weights and threshold with reference to known reproducible/irreproducible cases, or soften the causal claims
- [Results — AuNP Abstract Screening and Methodological Evaluation] The internal numbers conflict. The text says 'A total of 466 articles were approved in this step, representing a total of 36% (Figure 3B)' and later '466 articles (31%) were retained'; these fractions are inconsistent for the same denominator (1,299). In the methodological evaluation, the text states that after excluding 9 paywalled articles the set was 457, but then refers to 'None of the 458 articles analyzed' and '351 out of 458 articles.' Similarly, the text reports 'only 19 studies reached the threshold,' while Figure 3D's caption says 'only 20 met the methodological quality threshold.' These inconsistencies may stem from a mid-analysis decision or a figure typo, but as written they undermine confidence in the audit's accuracy and must be corrected systematically across the text, figures, and supplementary tables.
- [Methodology §4] The manuscript reports that inter-rater reliability was quantified with Cohen's kappa and that the Python classifier was trained 'until achieving agreement with human evaluations higher than 70%, calculated by the κ coefficient,' but no κ values are reported anywhere in the text or figures. Without actual κ statistics, the reader cannot judge the reliability of the abstract classifications. Furthermore, the methodological (full-text) evaluation was 'carried only by human evaluators'—but no inter-rater reliability is reported for that stage either. This is not an optional addition: the entire quantitative takeaway depends on the reliability of these binary scores. Please report κ (or an equivalent) for both screening stages, and for the full-text checklist if it was double-coded.
minor comments (5)
- [Title and Abstract] The title uses 'crises' where 'crisis' is the appropriate singular form. The keyword 'Meta-analisys' is misspelled ('Meta-analysis').
- [Figure 3 caption vs. text] The Figure 3 caption assigns panel (E) to 'Reporting frequency of key experimental parameters' and panel (F) to 'Distribution of methodological scores,' but the text describes Figure 3E as the score distribution and Figure 3F as the reporting profile. The panels and callouts need to be reconciled.
- [AuNP Abstract Screening relaxation] The sentence 'increased the approval rate during the abstract screaming from 36% to 75%, resulting in a total of 978 articles' is arithmetically odd: 75% of 1,299 is 974.25, not 978. Please verify the denominator and the exact counts.
- [Introduction vs. Results (MoS2)] The Introduction states that 'only one publication reported a comprehensive assessment of the reproducibility of the CVD growth process' (ref. 66), but the Results later report 115 articles approved in the methodological evaluation. The two statements are not necessarily contradictory (one 'comprehensive assessment' vs. many adequately reported papers), but the distinction is not explained and will confuse readers.
- [References and language] There are several typos ('aprowed', 'cheklist', 'Pedratory', 'the the', 'Emial', 'working principal'), and some references are duplicated (e.g., refs 21 and 32 appear identical; refs 22 and 109; refs 23 and 110). A careful proofreading pass is needed.
Circularity Check
Central 'reproducibility gap' is partly defined by the authors' own ≥75% reporting checklist; the advertised meta-analysis that would tie reporting to actual consistency is absent, so the headline conclusion overreaches.
-
self definitional
[Methodology §5; AuNP Results (Fig. 3E); Abstract]
"Studies achieving ≥ 75% compliance were considered methodologically sound and included in the subsequent meta-analysis. ... None of the 458 articles analyzed achieved a perfect score, and only 19 studies reached the threshold of 75% of the points ( Figure 3E), corresponding to 4% at this stage and 2% of the initial database. ... only a small fraction of studies rigorously addressed synthesis reproducibility."
The abstract's 'small fraction rigorously addressed synthesis reproducibility' is a restatement of the pass rate under a checklist whose item weights and ≥75% cutoff were chosen by the authors. 'Methodologically sound' is defined as scoring ≥75% on that checklist, so the low pass rate is not independent evidence about reproducibility: changing the cutoff or weights changes the 'small fraction' without any change in actual synthesis outcomes. The paper never validates the checklist against interlaboratory replication, so the inference 'underreporting → limited reproducibility' is built into the scoring instrument rather than tested.
full rationale
Most of the screening architecture is self-contained: the search strategy, sentinel-article retrieval checks, human-reviewer κ validation, and the raw reporting frequencies (e.g., AuNP pH reporting in 7% and replicate reporting in 2%) are direct observations that would stand regardless of the threshold. The self-citations (refs. 17 and 20) are used as sentinel articles and as background on pH/stirring effects; they are not the sole or load-bearing evidence. However, the headline inference does contain a definitional component: 'methodologically sound' is defined as ≥75% compliance on a weighted checklist chosen by the authors, and the Results then translate the 19/458 pass count into 'only a small fraction rigorously addressed synthesis reproducibility.' This makes the central 'small fraction' claim partly a property of the chosen cutoff and weights, not of externally measured interlaboratory consistency. In addition, Methodology §6 announces a quantitative METAFOR meta-analysis to evaluate whether similar reported conditions yield statistically consistent results, but no pooled effect sizes, heterogeneity statistics, or moderator analyses appear in the Results; the conclusion that underreporting 'limits inter-laboratory comparability and reproducibility' is therefore an inference from a reporting audit rather than from the advertised consistency analysis. These are partly overreach/missing-evidence issues as much as circularity; the raw reporting-frequency findings themselves are not circular. Score of 5 reflects one central self-definitional step while acknowledging the independent reporting data.
Assumptions & free parameters
free parameters (5)
- Methodological approval threshold =
75%
- Checklist item weights =
e.g., Au precursor concentration weight 3; other AuNP items weight 1; MoS2 high/low weights per Supplementary Note 8
- Classifier validation threshold =
Cohen's kappa ≥70–75% (exact values not reported)
- Search period cutoffs =
AuNP 1999–2024; MoS2 2012–2024
- Sentinel article sets =
15 AuNP + 18/26 MoS2 references
assumptions (5)
- domain assumption Reporting completeness of a fixed checklist is a valid operational measure of synthesis reproducibility.
- domain assumption Two case studies (Turkevich AuNP and APCVD MoS2) are representative of materials synthesis reproducibility.
- domain assumption Scopus/Web of Science English-language corpus with exclusions (reviews, non-primary, 'predatory' journals) is an unbiased sample of the relevant literature.
- ad hoc to paper Adapted PRISMA/SPIDER/STROBE frameworks remain valid when converted to binary weighted checklists with arbitrary thresholds.
- domain assumption Inter-rater agreement κ ≥75% validates the classifier; no κ values are actually reported.
Cite this review
Pith. "Pith review of Are we facing a reproducibility crises in materials synthesis? A systematic review of Turkevich AuNP synthesis and CVD MoS2 growth." pith.science (2026). https://pith.science/paper/FS2MB733
@misc{pith2026260714849,
author = {Pith},
title = {Pith review of: Are we facing a reproducibility crises in materials synthesis? A systematic review of Turkevich AuNP synthesis and CVD MoS2 growth},
year = {2026},
howpublished = {\url{https://pith.science/paper/FS2MB733}},
note = {Machine review of arXiv:2607.14849}
}
read the original abstract
Reproducibility remains a major challenge in materials synthesis, particularly for nanomaterials whose properties are highly sensitive to experimental conditions. Here, we present a systematic review and meta-analysis evaluating the reproducibility of two widely used synthesis routes: the Turkevich method for gold nanoparticles (AuNPs) and the chemical vapor deposition (CVD) growth of MoS2. An adapted PRISMA-based protocol combined with a modified SPIDER framework was applied to assess methodological transparency, parameter reporting, and experimental consistency across the literature. More than 1,300 articles for each case study were retrieved from Scopus and Web of Science and systematically screened using structured checklists and a Python-based text classification algorithm validated against independent human reviewers. Despite the extensive literature and the widespread perception of these methods as reproducible, only a small fraction of studies rigorously addressed synthesis reproducibility. Critical experimental parameters were frequently underreported, and statistical analyses were rarely included, limiting inter-laboratory comparability and reproducibility. These findings demonstrate the value of systematic reviews and meta-analyses as tools for identifying reproducibility gaps and guiding the development of more transparent and reliable synthesis protocols in materials science.
Reference graph
Works this paper leans on
-
[29]
https://doi.org/10.1038/s41699-020-00162-4. (116) Yang, R.; Fan, Y.; Zhang, Y.; Mei, L.; Zhu, R.; Qin, J.; Hu, J.; Chen, Z.; Hau Ng, Y.; Voiry, D.; Li, S.; Lu, Q.; Wang, Q.; Yu, J. C.; Zeng, Z. 2D Transition Metal Dichalcogenides for Photocatalysis. Angew Chem Int Ed 2023, 62 (13), e202218016. https://doi.org/10.1002/anie.202218016. (117) Xu, Y.; Ge, R.; ...
-
[472]
https://doi.org/10.1021/nl4033704. (39) Jeon, J.; Jang, S. K.; Jeon, S. M.; Yoo, G.; Jang, Y. H.; Park, J.-H.; Lee, S. Layer-Controlled CVD Growth of Large-Area Two-Dimensional MoS 2 Films. Nanoscale 2015, 7 (5), 1688–
-
[480]
(109) Ojea-Jiménez, I.; Bastús, N
https://doi.org/10.1039/C0CC02075C. (109) Ojea-Jiménez, I.; Bastús, N. G.; Puntes, V. Influence of the Sequence of the Reagents Addition in the Citrate-Mediated Synthesis of Gold Nanoparticles. The Journal of Physical Chemistry C 2011, 115 (32), 15752–15757. https://doi.org/10.1021/jp2017242. (110) Bastús, N. G.; Comenge, J.; Puntes, V. Kinetically Contro...
-
[905]
https://doi.org/10.1007/s10876-022-02261-2. (61) Grasseschi, D.; De O. Pereira, M. L.; Shinohara, J. S.; Toma, H. E. Facile Synthesis of Labile Gold Nanodiscs by the Turkevich Method. J Nanopart Res 2018, 20 (2), 35. https://doi.org/10.1007/s11051-018-4149-y. (62) Oluwatosin Kudirat, S.; Tawakalitu, A.; A. Saka, A.; O. Kamaldeen, A.; Mercy T, B.; Jimoh Ol...
arXiv 2018
-
[1695]
(40) Kang, K.; Xie, S.; Huang, L.; Han, Y.; Huang, P
https://doi.org/10.1039/C4NR04532G. (40) Kang, K.; Xie, S.; Huang, L.; Han, Y.; Huang, P. Y.; Mak, K. F.; Kim, C.-J.; Muller, D.; Park, J. High-Mobility Three-Atom-Thick Semiconducting Films with Wafer-Scale Homogeneity. Nature 2015, 520 (7549), 656–660. https://doi.org/10.1038/nature14417. (41) Wang, W.; Zeng, X.; Wu, S.; Zeng, Y.; Hu, Y.; Ding, J.; Xu, ...
-
[1882]
https://doi.org/10.1039/c2an16108g. (54) Stein, R.; Friedrich, B.; Mühlberger, M.; Cebulla, N.; Schreiber, E.; Tietze, R.; Cicha, I.; Alexiou, C.; Dutz, S.; Boccaccini, A. R.; Unterweger, H. Synthesis and Characterization of Citrate-Stabilized Gold-Coated Superparamagnetic Iron Oxide Nanoparticles for Biomedical Applications. Molecules 2020, 25 (19), 4425...
-
[2562]
(44) Li, X.; Kahn, E.; Chen, G.; Sang, X.; Lei, J.; Passarello, D.; Oyedele, A
https://doi.org/10.3390/ma11122562. (44) Li, X.; Kahn, E.; Chen, G.; Sang, X.; Lei, J.; Passarello, D.; Oyedele, A. D.; Zakhidov, D.; Chen, K.-W.; Chen, Y.-X.; Hsieh, S.-H.; Fujisawa, K.; Unocic, R. R.; Xiao, K.; Salleo, A.; Toney, M. F.; Chen, C.-H.; Kaxiras, E.; Terrones, M.; Yakobson, B. I.; Harutyunyan, A. R. Surfactant- Mediated Growth and Patterning...
arXiv 2020
-
[5118]
https://doi.org/10.1021/acs.jpcc.7b10536. (57) Ali, M. M.; A. Rajab, N.; A. Abdulrasool, A. Preparation, Characterization and Optimization of Etoposide-Loaded Gold Nanoparticles Based on Chemical Reduction Method. IJPS 2020, 29 (2), 107–121. https://doi.org/10.31351/vol29iss2pp107-121. (58) Díaz-García, V.; Haensgen, A.; Inostroza, L.; Contreras-Trigo, B....
Reviewed August 2, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.