{"id":"c43eb05b-e734-4565-be18-3ef2d225d331","arxiv_id":"2412.10477","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"A new human-graded benchmark of 70 open-ended ALD questions shows GPT-4o passing overall but producing vague answers, hallucinations, and lower scores on harder and more specific questions.","lead":"The authors built ALDbench, a set of 70 expert-written open-ended questions about atomic layer deposition, and scored GPT-4o's answers by human domain experts. The model passed overall, but struggled with precise quantitative details, and harder or more specific questions tended to get lower scores.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The reported Fisher exact p-values are not trustworthy: Table VII's first contingency table sums to 265 reviews rather than 236, and the tests ignore clustering of reviews within questions and experts.","rationale":"The reader's weakest assumption correctly identified clustered reviews as a threat to the Fisher exact p-values. I agree that this is a serious concern, but I also found a more immediately verifiable problem: the printed contingency table for quality vs. difficulty does not sum to the stated number of reviews, and no table sums to 236. This means the p-values cannot be independently reproduced from the manuscript as written. The benchmark itself is a genuine contribution: 70 expert-written open-ended ALD questions, a documented GPT-4o baseline, explicit grading rubrics, and a transparent hallucination discussion. However, the paper's headline statistical correlations are load-bearing for the claim that question difficulty and specificity predict response quality, and those correlations are not currently supported by a statistically valid analysis. A clustered reanalysis and table correction could easily resolve this, so the appropriate verdict remains conditional rather than accept or reject. If the raw data were released and the results survived the reanalysis, the paper would be a solid benchmark contribution; if not, the quantitative conclusions would need to be substantially weakened to a descriptive level. The hallucination analysis, while qualitative, is plausible and appropriately hedged, so it does not drive my concern.","tokens_in":10186,"tokens_out":3994,"duration_ms":44975,"concrete_test":"Obtain the raw 236 expert-question ratings from the Supporting Information, rebuild all eight 2x2 tables, and reconcile the row sums with 236. Then recompute the three claimed correlations with a clustered permutation test that shuffles question-level labels while keeping all reviews from a given expert/question together, and apply a Bonferroni or FDR correction across the eight tests. If p-values for quality-difficulty, relevance-difficulty, and accuracy-specificity remain below the corrected threshold, the correlations are robust; if they do not, the headline quantitative claim should be downgraded.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central quantitative claim — three significant correlations — rests on Fisher exact tests in Table VI applied to contingency tables in the Appendix. Two independent problems undermine those p-values. First, the first contingency table in Table VII (quality vs. difficulty) sums to 265 reviews, while §III B states there were 236 independent reviews; every other table sums to 235. Since no table equals 236, the printed tables are internally inconsistent and the reported p=0.033 cannot be reconstructed from them. Second, even with corrected tables, Fisher's exact test treats all reviews as independent. But the 236 reviews are nested: seven experts each graded many of the same 70 questions, and experts were free to choose subsets, so ratings from the same expert or same question are positively correlated. This clustering reduces the effective sample size and makes the p-values 0.033, 0.016, and 0.007 anticonservative. With eight tests and no multiple-comparison adjustment, the quality-difficulty p=0.033 would not survive even a Bonferroni threshold. The correlations may be real, but the paper as written does not demonstrate them.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces ALDbench, a benchmark of 70 open-ended questions in atomic layer deposition (ALD), graded by seven domain experts who scored each question's difficulty and specificity and each GPT-4o response's overall quality, specificity, relevance, and accuracy on 1–5 Likert scales. The authors report a composite quality score of 3.7, with 36% of questions receiving at least one below-average score, at least five suspected hallucinations, and three statistically significant correlations from Fisher exact tests (quality–difficulty, relevance–difficulty, accuracy–specificity). The stated goal is to provide a domain-specific evaluation tool for LLMs in materials synthesis.","tokens_in":10377,"tokens_out":2664,"duration_ms":28011,"significance":"If the descriptive results stand, ALDbench is a valuable community resource: it is one of the few expert-reviewed, open-ended benchmarks in materials synthesis, and it deliberately includes questions at the frontier of expert knowledge. The multi-criteria rubric (quality, specificity, relevance, accuracy) and the qualitative hallucination analysis are useful contributions. The paper also ships the question set and responses in the Supporting Information, which supports reuse and independent verification. However, the paper's headline statistical claims are currently under-supported because of internal inconsistencies in the reported contingency tables and the use of a statistical test that assumes independence of clustered observations.","major_comments":[{"comment":"The first contingency table in Table VII (response quality vs. question difficulty) sums to 265 reviews (60+38+132+35), whereas §III B states that 236 independent reviews were gathered and every other contingency table in Tables VII–VIII sums to 235. This internal inconsistency means the reported p=0.033 for the quality–difficulty correlation cannot be reconstructed from the printed data. The authors should correct the table and recompute the p-value, and should also verify the marginal totals of all tables.","section":"§III B and Table VII"},{"comment":"The Fisher exact tests treat each of the 236 reviews as an independent observation, but reviews are nested within questions (70 questions) and within the seven expert raters, who were free to review different subsets of questions. Ratings from the same expert or the same question are likely positively correlated, so the effective sample size is smaller than 236 and the reported p-values (0.033, 0.016, 0.007) are anticonservative. The authors should use an analysis that accounts for this clustering—for example, mixed-effects logistic regression with random intercepts for expert and question, or cluster-robust standard errors—or should aggregate to the question level and perform an appropriate test.","section":"§III B and Tables VI–VIII"},{"comment":"The paper performs eight Fisher exact tests without any adjustment for multiple comparisons. Even if the first contingency-table issue is fixed, the quality–difficulty p=0.033 would not survive a Bonferroni correction (threshold 0.00625), and the relevance–difficulty p=0.016 would also fail. Only the accuracy–specificity p=0.007 would potentially remain significant, but its reliability is still subject to the clustering problem. The authors should either pre-specify a limited set of hypotheses, apply a correction, or present the results as exploratory rather than confirmatory.","section":"§III B and Table VI"}],"minor_comments":[{"comment":"The dichotomization of scores into above-average (4–5) and at-or-below-average (1–3) is arbitrary and discards ordinal information. The authors should justify this split or show that the conclusions are robust to alternative thresholds (e.g., median split per expert or the full ordinal scale).","section":"§III B"},{"comment":"No inter-rater reliability statistic (e.g., Cohen's kappa, weighted kappa, or intraclass correlation) is reported for the expert graders. Given the observed dispersion in Figures 1–3, such a measure would help the reader calibrate the reliability of the ratings.","section":"General"},{"comment":"Reference 11, 'Atomic limits ald database', is incomplete; the authors should provide a full citation or URL.","section":"References"},{"comment":"There are several typographical errors, including 'In Figure 1 We show' (capital W) in §III A, and 'How does the temperature effect the growth' in Question 51 (should be 'affect').","section":"Text"},{"comment":"The color-map representations of expert scores are difficult to read in grayscale print; a dot-plot or heatmap with numeric labels might be clearer.","section":"Figures 1–2"}],"recommendation":"major_revision","confidential_remarks":"The manuscript describes a useful benchmark and a mostly sensible evaluation protocol, but the statistical analysis is not yet publication-ready. The contingency-table inconsistency and the clustering issue are fixable; I would encourage the editor to request a revision rather than reject, because the underlying dataset and qualitative findings are valuable. The authors should also be asked to supply the raw per-expert/per-question data in a machine-readable format so that the corrected analyses are verifiable."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a real resource—70 open-ended ALD questions graded by seven domain experts on four response criteria—but the statistical claims in Section III.B are not yet trustworthy. The first contingency table in Table VII sums to 265 reviews while the text says 236; every other table sums to 235. That is not rounding. So the reported Fisher exact p=0.033 cannot be reconstructed from the printed numbers, and the other p-values sit on tables that are also internally short. On top of that, the reviews are clustered: the same seven experts graded many of the same 70 questions, and experts chose subsets. Treating each review as independent inflates the effective sample size and makes the three “significant” correlations anticonservative. With eight tests and no multiple-comparison correction, the weakest p=0.033 would not survive Bonferroni anyway. The correlations may be real—the direction is plausible, and the accuracy–specificity anti-correlation matches their qualitative observation that GPT-4o gives vague ranges for quantitative questions—but this paper does not demonstrate them as written.\n\nWhat earns credit: the benchmark itself is a genuine contribution, not just ChemBench rehashed. The questions span graduate to state-of-the-art expert level, the rubric in Table I is clear, and the authors report GPT-4o’s composite 3.7 with an honest per-question breakdown showing 36% of questions got at least one below-average grade. The hallucination analysis, though qualitative, gives concrete examples (NF3/SF6 as MgF2 co-reactants, cobalt nitrate for Co3O4) and reads them sensibly as context mis-association. The near-zero correlation between difficulty and specificity (Pearson 0.12) is worth keeping.\n\nSoft spots besides the stats: the same group wrote the questions and graded the responses; there is no inter-rater reliability measure; the 4–5 vs 1–3 split is arbitrary. Those are normal for an expert-curated benchmark but deserve explicit acknowledgment. The hallucination count of “at least five” is not audited—no full list, no independent adjudication.\n\nBottom line: this deserves a serious referee, but the revision must fix the contingency tables and redo the significance analysis with clustering in mind (e.g., a mixed model or a permutation test resampling by question/expert). As a benchmark dataset, ALDbench is worth having; as a paper about LLM performance correlations, it needs reanalysis before the numbers are cited. Send it to review with a request for major revision.","headline":"A genuinely useful 70-question ALD benchmark with careful expert grading, but the headline correlations rest on internally inconsistent tables and clustered reviews; the stats need a redo before the main claims are reliable.","tokens_in":10895,"tokens_out":2285,"would_cite":false,"duration_ms":24710,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"GPT-4o earns a passing but fragile grade on expert atomic layer deposition questions, with 36% of questions drawing at least one below-average score and five suspected hallucinations.","keywords":["atomic layer deposition","large language models","benchmark","materials synthesis","hallucination detection","expert evaluation","open-ended questions","GPT-4o"],"falsifier":"Reanalyze the contingency tables with expert and question as random effects, for example using a mixed-effects logistic regression or a cluster-robust version of the Fisher exact test; if the associations between difficulty and quality, difficulty and relevance, and specificity and accuracy lose statistical significance, the paper's headline correlation claims would not survive.","tokens_in":10001,"feed_emoji":"🧪","tokens_out":4817,"duration_ms":44495,"temperature":0.7,"pith_summary":"This paper introduces ALDbench, a benchmark of 70 open-ended questions about atomic layer deposition (ALD), a thin-film growth technique, with questions ranging from graduate level to state-of-the-art expert knowledge. Six ALD experts wrote the questions and seven experts graded the GPT-4o answers on four criteria: overall quality, specificity, relevance, and accuracy. The model earned a composite quality score of 3.7 out of 5, a passing grade, but 36% of questions received at least one below-average score and the authors identified at least five suspected hallucinations, mainly invented precursor chemistries. The paper also finds statistically significant correlations: harder questions get lower quality and relevance scores, and more specific questions get lower accuracy scores, while question difficulty and specificity are themselves nearly uncorrelated. The authors argue that multi-criteria, open-ended expert grading reveals failure modes that multiple-choice and NLP-style benchmarks miss.","feed_headline":"GPT-4o passes ALD expert quiz but stumbles on hard specifics","feed_subtitle":"New 70-question benchmark: average scores pass, but 36% of questions flunk and five answers invent chemistry.","key_machinery":"The central object is ALDbench, a hand-curated set of 70 open-ended questions about atomic layer deposition, each graded by multiple human ALD experts on question difficulty and specificity, and on the model's response quality, specificity, relevance, and accuracy using 1-5 Likert rubrics. The statistical engine is the 2x2 contingency table split at 'above average' (scores 4-5) versus 'at or below average' (scores 1-3), tested with the Fisher exact test to detect correlations between question attributes and response scores. The in-depth qualitative analysis of individual responses, looking for hallucinations and precision failures, completes the machinery.","core_discovery":"On its own terms, the paper's central claim is that a state-of-the-art general LLM, GPT-4o, performs at a passing but uneven level on expert-level ALD knowledge: aggregate scores are above average across all four criteria, yet a substantial minority of questions draw low scores, and the model fabricates plausible-sounding but unreferenced chemistry, such as NF3 or SF6 as fluorine sources for MgF2 ALD and cobalt(II) nitrate as a Co3O4 precursor. The paper further claims that response quality and relevance drop as question difficulty rises, and accuracy drops as question specificity rises, based on Fisher exact tests of 2x2 contingency tables built from 236 expert reviews. It presents ALDbench itself as a reusable open-ended benchmark for materials synthesis that captures dimensions—relevance and specificity—that standard multiple-choice or NLP benchmarks do not.","pith_inferences":["If the clustering of expert reviews is taken into account, the reported Fisher-exact p-values (0.033, 0.016, 0.007) are likely too small; a mixed-effects reanalysis could weaken or erase the claimed correlations.","A natural next experiment is to run ALDbench on the same model with access to the Atomic Limits ALD database or a retrieval tool; the paper's own hallucination analysis suggests that specificity and accuracy scores would rise.","The 'at least five' hallucination figure is probably a floor, since it depends on the reviewing experts' personal knowledge; automated checking of every proposed precursor pair against a database would give a reproducible rate.","The specificity-accuracy anticorrelation, if it generalizes, implies that LLM fluency in quantitative domains is partly a trade-off: precise numeric answers are sacrificed for plausible-sounding ranges, a testable hypothesis for other synthesis fields."],"forward_implications":["GPT-4o's aggregate passing scores across all four criteria show that general-purpose LLMs already hold substantial declarative knowledge about a specialized synthesis field like ALD.","The statistically significant correlations mean that a user asking hard or highly specific ALD questions should expect lower-quality, less relevant, or less accurate answers than a user asking general ones.","Because question difficulty and question specificity are nearly uncorrelated (Pearson r = 0.12), benchmarks must grade both dimensions separately to reveal where LLMs fail.","Open-ended, expert-graded benchmarks can expose failure modes—such as invented precursor chemistries and imprecise quantitative ranges—that multiple-choice or NLP-style benchmarks would pass over."],"supporting_citations":[{"why":"Supplies the contrasting ChemBench results and the 'superhuman chemists' claim that ALDbench qualifies with open-ended expert grading.","marker":"[1]"},{"why":"Defines atomic layer deposition and its self-limited surface chemistry, which the benchmark questions test.","marker":"[6]"},{"why":"Represents the NLP-style multi-task materials benchmark that ALDbench is designed to complement.","marker":"[8]"},{"why":"Provides the correct precedent for HF-pyridine as a fluorine source in metal fluoride ALD, used to identify hallucinated alternatives.","marker":"[9]"},{"why":"Documents a valid MgF2 ALD process, the baseline against which the model's invented NF3/SF6 chemistry is judged.","marker":"[10]"},{"why":"Serves as the reference database for checking whether the proposed NF3 or SF6 ALD processes exist, supporting the hallucination finding.","marker":"[11]"},{"why":"Explains why the model plausibly but wrongly suggested SF6 plasma, by showing SF6 is used in related fluoride ALD processes.","marker":"[13]"},{"why":"Cited as the untested strategy of augmenting LLMs with chemistry tools that could improve performance on specific questions.","marker":"[16]"}],"fun_headline_variants":["GPT-4o flunks 36% of ALD expert questions","LLM passes ALD quiz but invents chemistry","GPT-4o uneven on expert ALD knowledge","Benchmark: LLM weak on hard ALD specifics","GPT-4o fabricates ALD precursors in expert test"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The statistical correlations rest on treating all 236 expert reviews as independent observations, but the reviews are clustered: the same seven experts graded many questions each, so the effective sample size is smaller and the reported p-values are likely too optimistic.","fun_headline_variants_meta":{"raw":{"variants":["GPT-4o flunks 36% of ALD expert questions","LLM passes ALD quiz but invents chemistry","GPT-4o uneven on expert ALD knowledge","Benchmark: LLM weak on hard ALD specifics","GPT-4o fabricates ALD precursors in expert test"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000214,"raw_usage":{"total_tokens":1431,"prompt_tokens":958,"completion_tokens":473,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":574,"completion_tokens_details":{"reasoning_tokens":390}},"tokens_in":574,"tokens_out":473,"duration_ms":4393,"temperature":1.0,"reasoning_tokens":390,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T16:37:25.609761+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Reanalyze the contingency tables with expert and question as random effects, for example using a mixed-effects logistic regression or a cluster-robust version of the Fisher exact test; if the associations between difficulty and quality, difficulty and relevance, and specificity and accuracy lose statistical significance, the paper's headline correlation claims would not survive.","supporting_citations":[{"cited_title":"Alvaro \\ and\\ author A","cited_arxiv_id":null,"evidence_quote":"Documents a valid MgF2 ALD process, the baseline against which the model's invented NF3/SF6 chemistry is judged."},{"cited_title":"Pilvi , author T","cited_arxiv_id":null,"evidence_quote":"Explains why the model plausibly but wrongly suggested SF6 plasma, by showing SF6 is used in related fluoride ALD processes."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Cited as the untested strategy of augmenting LLMs with chemistry tools that could improve performance on specific questions."}],"review_version":1}