{"id":"3bdfffe9-0411-4f45-a71a-a1b4b599d412","arxiv_id":"2607.22573","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"low","formal_verification":"none","parameter_count":5,"one_line_summary":"A public dataset of 46,899 Materials Project-derived crystals pairing phonopy spectra with deterministic stable/unstable labels from a fixed MatterSim-phonopy workflow.","lead":"This paper releases PhononBench-MP40: 46,899 crystals from the Materials Project, each with a computed phonon spectrum and a stable/unstable label from one fixed simulation workflow. It gives machine-learning researchers a large, auditable dataset for screening crystals for dynamic stability.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"No executed audit demonstrates that the released YAML files reproduce the completed labels; pairing and threshold consistency remain unverified.","rationale":"The reader's weakest assumption is precisely that the 46,899 completed records are correctly paired triples with no audit of token collisions or agreement between the raw workflow_label and the threshold-derived completed_label. My stress test converges on the same point: the paper's strongest claim is that labels are recomputable from the released YAML spectra, but the authors provide no executed validation of this reproducibility. The cohort arithmetic and workflow description are internally consistent, and the stated limitations are transparent. The missing audit is the single most load-bearing gap because it sits between the release artifacts and every downstream use: if the pairing or the label derivation is wrong, the counts of 16,683 Stable and 30,216 unstable records are unreliable, and the benchmark cannot serve as an auditable reference. The reader already assigned CONDITIONAL, so my concern does not change the verdict; it reinforces the condition. The proposed full-programmatic audit is concrete and cheap, and would either confirm the claim or expose specific mismatches. No other concern — DOI inactivity, missing baselines, or path-sampling boundaries — is as central to the dataset's integrity as verifying that the released YAML files actually yield the published labels.","tokens_in":13758,"tokens_out":5181,"duration_ms":52908,"concrete_test":"Download the archived release and run a full programmatic audit: (1) verify that every yaml_relative_path in the completed table resolves to a file; (2) parse each YAML, extract all band frequencies, compute ωmin, apply the threshold ω < −1e-3 THz, and compare the resulting label to completed_label; (3) check that structure_name tokens are unique across the completed table and match exactly one YAML filename; (4) build the confusion matrix between workflow_label and completed_label and report the count of disagreements. If any completed_label differs from the threshold-derived label, or if any token collision or unresolved path occurs, the central claim fails. A 1,000-record random sample would be a minimal check, but a full 46,899-record audit is computationally inexpensive and would be definitive.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim — that every completed record's label is deterministically derivable from its released YAML spectrum — rests on two unverified assumptions. First, each YAML file must actually contain the high-symmetry-path frequencies from which ωmin and the −1e-3 THz threshold are supposed to be computed. Second, the raw workflow_label and the derived completed_label must be mutually consistent, and the structure_name token must uniquely pair each completed table row with the correct YAML file. The paper reports intersection counts but no executed validation: S10 lists checks (path existence, threshold relabeling, split leakage) without presenting results for any of them. Without a demonstration that parsing the YAML files and applying the stated rule reproduces the 16,683/30,216 split, the dataset's headline numbers could be wrong due to token collisions, mismatched files, or a threshold mismatch between the raw PhononBench labels and the derived labels. This is load-bearing because any such error directly changes the published counts and invalidates the benchmark's auditable-reuse promise, even though the workflow description itself is internally consistent.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"PhononBench-MP40 is a data-descriptor paper releasing 46,899 completed records for workflow-defined phonon stability, with paired stability labels and phonopy YAML spectra, derived from an MP40-derived task cohort of 47,969 tasks. The authors report 16,683 Stable and 30,216 completed-phonon unstable records, plus 1,067 relaxation failures reported outside the completed denominator. The central design is that the stability label is derived from each released YAML spectrum by a deterministic threshold rule (unstable iff any sampled band frequency is below −1e-3 THz, S3), making the labels auditable and threshold-refinable. The paper also provides a formula-group reference split, coverage statistics, benchmark tasks, and a stated boundary that the labels are workflow-defined rather than claims of absolute physical stability.","tokens_in":13943,"tokens_out":3242,"duration_ms":37273,"significance":"If the dataset is correct and accessible, it would be a valuable resource for materials informatics: it is substantially larger than most open phonon-stability datasets, it pairs labels with the underlying spectral object, it explicitly separates relaxation failures from completed phonon instabilities, and it provides a deterministic label rule and split. The paper also benefits from a consistent internal accounting of task, label, YAML, and completed-record counts, and from explicit discussion of interpretation boundaries. The main scientific value depends on verifiability of the label–spectrum pairing, which is currently asserted rather than demonstrated.","major_comments":[{"comment":"The central claim of the paper is that every completed label is deterministically derived from the released YAML spectrum, so that users can re-extract ωmin and reproduce the label. However, no executed validation is reported. Section S10 lists a 'Threshold relabeling' check, a 'YAML path existence' check, a 'Split leakage check', and other integrity checks, but the paper gives no results for any of them. This is load-bearing: the 16,683/30,216 split and the auditable-reuse promise stand or fall on whether parsing the released YAMLs and applying the S3 rule reproduces the completed labels. I ask that the authors provide, in the release or supplement, a runnable audit script and a summary table reporting, over all 46,899 completed records: (a) the number of YAML paths that resolve, (b) the number of YAML files containing parseable high-symmetry-path frequencies, (c) the number of records","section":"§6 and S10, Table S5"},{"comment":"The relationship between the raw 'workflow_label' and the derived 'completed_label' is not made explicit. Section S2 defines workflow_label as the raw PhononBench label and completed_label as the label 'after enforcing the label+YAML completion rule', while S3 states that the completed stability label is derived from high-symmetry-path frequencies using the −1e-3 THz threshold. If the raw PhononBench workflow_label was produced by a different threshold, path convention, or post-processing step, then the completed_label might not equal the threshold-derived label for some records. The paper never reports the agreement between workflow_label and threshold-derived completed_label. I request a cross-tabulation (or, at minimum, a statement that completed_label was recomputed from the YAML frequencies using S3 and that the raw labels were not used except as an audit cross-check).","section":"S2 and S3"},{"comment":"The data availability statement says the dataset 'is being archived' and that the DOI 'should be cited once registration is active'. This means the central artifact is not available for reviewers or readers to verify any of the claims in the paper. For a data-descriptor paper, an active and accessible archival record is not a minor detail; it is the primary evidence. I recommend that acceptance be conditioned on an active Science Data Bank DOI with the completed table, split table, failure table, metadata, checksum file, and the 46,899 YAML spectra available under the claimed paths.","section":"Data availability statement"}],"minor_comments":[{"comment":"The accounting is internally consistent, but the three YAML outputs that do not have a matching label (46,902 YAMLs vs 46,899 completed records) are not discussed. A sentence explaining what those three records are would help users understand the exact matching rule.","section":"Table 1 and §2"},{"comment":"The matching key structure_name is defined as a 'workflow key' and source_id as the first token. Given that labels, YAML paths, and table rows are joined on this token, a uniqueness audit of structure_name in the master, completed, and failure tables should be included in the S10 validation results. This is a minor presentation point only if the proposed audit passes; it is part of the major concern if it has not been run.","section":"S2, structure_name"},{"comment":"The discussion of limitations is clear and appropriately cautious. However, the phrase 'within the adopted tolerance' for MgB2 (ωmin = 0.0000 THz) may mislead readers into thinking there is a nonzero tolerance on the stable side. The rule is a one-sided threshold (ω < −1e-3 THz), so a zero minimum is consistent with Stable. Consider stating this explicitly to avoid confusion.","section":"§9 and S8"},{"comment":"The SHA-256 split rule should specify the exact byte representation of the formula key (e.g., UTF-8 encoding of the parsed formula token) so that users can independently reproduce the split. This is a reproducibility detail, not a correctness issue.","section":"S5"},{"comment":"The recommended reporting checklist is useful. A small addition would be to require reporting the exact PhononBench-MP40 version and Git commit of the parsing utilities, as already suggested in S3, so that any relabeling or threshold experiments can be traced.","section":"§7, baseline protocol"}],"recommendation":"major_revision","confidential_remarks":"This is a data paper whose central claims rest on the availability and verifiability of the released artifact. The lack of executed validation results and the inactive DOI are the deciding factors; both are fixable within the scope of a revision. I would not reject, but I would not accept before the authors provide a reproducible label-rederivation audit and an accessible archival record."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe short version: if phonon stability is in your orbit, this is a dataset to know about. The authors release 46,899 MP-derived crystals with phonopy YAML spectra and deterministic stability labels, plus 1,067 relaxation failures kept separate. The core design is sound and the claims are carefully bounded.\n\nWhat's actually new: the spectral object. Most phonon datasets give you a label table; here the YAML files are the primary release, so any user can re-derive omega_min and the Stable/unstable decision from the raw frequencies. That is a real step forward. The label rule is simple and stated: unstable iff any sampled band frequency is below -1e-3 THz. The cohort arithmetic checks out (46,899 + 1,067 = 47,966; 16,683 + 30,216 = 46,899), the split is formula-disjoint, and the failure cohort is correctly excluded from the completed denominator. The authors are also honest about what this is not: it is workflow-defined stability under MatterSim/phonopy, not a universal physical statement.\n\nWhere it's soft: the paper lists a validation checklist (S10) but reports no results. For a dataset whose whole pitch is auditable reuse, the one load-bearing check—whether parsing the released YAMLs and applying the threshold reproduces the published counts—is absent. The stress-test note has this exactly right. I don't think it's hiding anything; the internal accounting is consistent, but it is an omission that needs fixing before the dataset can be trusted as shipped. The DOI isn't active yet, and no commit hash or checksums are pinned, so the deliverable is not verifiable right now. And there are zero baseline scores; the paper defers them to future work. That's acceptable for a data paper, but it makes the 'benchmark' label aspirational rather than demonstrated.\n\nNone of these sink the core design. The pairing risk is real but likely minor given the simple token scheme; the threshold is arbitrary but disclosed; the workflow is what it is. The paper shows clear thinking and honest engagement with the literature.\n\nRecommendation: send it to peer review. Ask for an executed validation pass (at least a script that re-derives labels from the released YAMLs and reports how many match), an active DOI with checksums, and at least one simple baseline. Once those land, it is a genuinely useful community resource. I would cite it in that form.","headline":"A useful, honestly scoped phonon-stability dataset that needs an executed audit and a live DOI before it fully delivers on its auditable-reuse promise.","tokens_in":14524,"tokens_out":2922,"would_cite":true,"duration_ms":32165,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper presents a benchmark dataset of 46,899 crystal phonon spectra whose stability labels are deterministic functions of the released spectra, so every stable/unstable tag can be reproduced and re-thresholded.","keywords":["phonon stability","phonon spectra","benchmark dataset","materials informatics","minimum frequency","imaginary modes","workflow-defined labeling","data release"],"falsifier":"Re-run the release's own audit: take the master task table and the released YAML paths, count how many structure-name tokens fail to resolve or resolve to multiple YAML files, then recompute the completed label from every YAML by applying the −10⁻³ THz rule and compare with the completed table's labels. If the pairwise counts differ from 46,899/16,683/30,216, or if any released YAML gives a different label than the table, the central audit claim is falsified.","tokens_in":13543,"feed_emoji":"🔬","tokens_out":7685,"duration_ms":65171,"temperature":0.7,"pith_summary":"This paper presents a spectrum-resolved benchmark dataset in which workflow-defined phonon stability is made auditable. The central claim is that 46,899 crystal records — 16,683 stable and 30,216 unstable — can be released together with the local band-frequency spectra from which those labels were derived, under one fixed rule: any sampled band frequency below −10⁻³ THz means unstable. It also separates 1,067 relaxation failures that produced no spectrum from the completed denominator, so failed calculations are not mistaken for physical instabilities. A sympathetic reader should care because large phonon datasets typically offer only binary flags, hiding the difference between a shallow soft mode, a deep imaginary branch, and a missing result; here the spectrum behind each label is the release object, making classification reproducible, threshold studies possible, and failure-aware triage a distinct task. If the data hold up, the field gains a shared target for stability classification, minimum-frequency regression, and workflow-robustness studies.","feed_headline":"46,899 phonon spectra get stability labels anyone can reproduce","feed_subtitle":"Every tag is derived from the released spectrum by one fixed rule, so users can recheck it or move the threshold","key_machinery":"The central object is the local YAML spectrum: the file storing the sampled high-symmetry-path band frequencies for each completed calculation. It carries the evidence behind every label — the dispersion, the minimum frequency ωmin, and the threshold margin — and it is the object through which all derived uses flow: reconstructing dispersions, extracting ωmin, and re-deriving labels under new thresholds. The two supporting mechanisms are the completion rule (a completed record is the intersection of a recovered workflow label and a matching YAML output, which yields 46,899 records and sends 1,067 relaxation failures to a separate table) and the deterministic label threshold (ω < −10⁻³ THz ⇒","core_discovery":"The core discovery is a dataset design: each completed phonon-stability record consists of a recovered workflow label joined to a matching local YAML spectrum, and the manuscript label is not an independent model output but a derived view of that spectrum. The label rule is deterministic and simple — unstable iff any sampled high-symmetry-path band frequency is below −10⁻³ THz; otherwise stable — so users can re-extract the minimum sampled frequency and reproduce the binary label from the YAML file. The completed cohort totals 46,899 records (16,683 stable, 30,216 completed-phonon unstable), with 1,067 relaxation failures listed separately. The release also supplies a formula-group split int","pith_inferences":["The paper stops short of saying that the 64.4% unstable fraction is partly a threshold artifact, but the released spectra make that testable: reporting what fraction of 'unstable' records have ωmin between −10⁻³ and, say, −0.1 THz would show how many borderline cases drive the high instability rate.","A natural three-class extension not in the paper is to split completed-phonon unstable into soft-mode (near-threshold) and deep-imaginary records; this would give downstream validation a more graded screening target.","Because all labels inherit a single machine-learned potential, a spot-check of a few hundred near-threshold YAML spectra against higher-fidelity calculations would bound the workflow's systematic error — a test the authors flag implicitly but do not run.","The failure table could be modeled as a missingness signal; if relaxation failures correlate with chemistry or symmetry, ignoring them would bias any classifier trained on the completed cohort, and weighting or joint modeling would be needed."],"forward_implications":["Anyone can verify a label by parsing the released YAML and applying the −10⁻³ THz rule; no label table has to be trusted on its own.","The same released objects support minimum-frequency regression, so models can distinguish near-threshold soft modes from deep imaginary branches instead of collapsing them into one 'unstable' class.","Because the formula-group split keeps every formula in exactly one partition, model comparisons on the completed cohort avoid exact-formula leakage.","The 1,067 relaxation failures form a separate failure-aware triage target: predicting whether a workflow produces an auditable spectrum, which should not be mixed with predicting instability.","Threshold studies become measurable: moving the −10⁻³ THz cut and relabeling gives a quantitative map of borderline records, useful for calibration in screening pipelines."],"fun_headline_variants":["Every phonon label is re-derivable from its own spectrum","46,899 spectra, one rule: unstable if any frequency below -1e-3 THz","PhononBench: labels derived, not predicted — recheck any spectrum","Auditable stability: 46,899 phonon spectra with one deterministic rule"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The load-bearing premise is that the 46,899 completed records are correctly paired label-plus-YAML triples joined on the structure-name token, so that the released counts and threshold-derived labels are internally consistent; if token collisions occurred or the raw labels used a different threshold, the numbers would not mean what they appear to mean.","fun_headline_variants_meta":{"raw":{"variants":["Every phonon label is re-derivable from its own spectrum","46,899 spectra, one rule: unstable if any frequency below -1e-3 THz","PhononBench: labels derived, not predicted — recheck any spectrum","Auditable stability: 46,899 phonon spectra with one deterministic rule"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000487,"raw_usage":{"total_tokens":2241,"prompt_tokens":750,"completion_tokens":1491,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":494,"completion_tokens_details":{"reasoning_tokens":1415}},"tokens_in":494,"tokens_out":1491,"duration_ms":12307,"temperature":1.0,"reasoning_tokens":1415,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T12:18:22.307659+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the release's own audit: take the master task table and the released YAML paths, count how many structure-name tokens fail to resolve or resolve to multiple YAML files, then recompute the completed label from every YAML by applying the −10⁻³ THz rule and compare with the completed table's labels. If the pairwise counts differ from 46,899/16,683/30,216, or if any released YAML gives a different label than the table, the central audit claim is falsified.","supporting_citations":[],"review_version":1}