{"id":"599b2062-3c23-448f-b260-dd03cde2c088","arxiv_id":"2508.20143","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":7,"one_line_summary":"CrystalICL is a few-shot crystal generation model combining space-group tokenization with multi-task instruction tuning.","lead":"This paper presents a method for generating crystal structures using language models that learn from a few examples. It introduces a symmetry-based tokenization and hybrid instruction tuning, reporting gains over existing zero-shot methods.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Unconditional-generation superiority claim contradicted by Table 2: DiffCSP/CDVAE beat CrystalICL on key property-distribution metrics, so the abstract overstates results.","rationale":"The reader's weakest assumption (SGS tokenization harming chemical-formula generation) is a real limitation but is explicitly acknowledged in Sec. 4.2 and does not falsify the core method: CrystalICL with XYZ still outperforms CrystalLLM. The more load-bearing problem is that the paper's own Table 2 contradicts the abstract's claim of superiority over leading baselines on unconditional generation. CDVAE and DiffCSP are named baselines, and they achieve lower Wasserstein distances on density and formation energy than CrystalICL(SGS) on MP20 and C24. This is not a matter of interpretation; it is a direct internal inconsistency in the claimed contribution. A conditional acceptance remains appropriate because the ICL mechanism and the gains over CrystalLLM are plausible and could be salvaged by rewriting the claims and adding a more comprehensive comparison. The concrete test of reproducing Table 2 would settle whether the contradiction is due to an error in the table or in the claim, but either way the current text overstates the results. I therefore agree with the reader that the paper needs revision, but the most load-bearing concern is the unconditional superiority over DiffCSP/CDVAE, not the SGS tradeoff.","tokens_in":19248,"tokens_out":10219,"duration_ms":103037,"concrete_test":"Reproduce Table 2's property-distribution metrics (wdist(ρ), wdist(E), wdist(Nel)) from the generated structures using the standard evaluation script for CDVAE/DiffCSP/CrystalICL on MP20 and C24. If the numbers match the paper—particularly C24 wdist(E)≈1.31 for CrystalICL(SGS) vs. ≈0.04 for DiffCSP—then the abstract's unconditional superiority claim must be revised to a comparison against CrystalLLM only, and the conclusion should not claim superiority over leading baselines.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim (Abstract, Sec. 1) is that CrystalICL outperforms 'the leading baseline methods' on conditional and unconditional generation. For unconditional generation, Sec. 4.1 explicitly names CDVAE, DiffCSP, and CrystalLLM as baselines. Table 2 contradicts this: on C24, DiffCSP achieves wdist(E)=0.0415 and CDVAE 0.2206, while CrystalICL(SGS) achieves 1.3061; on MP20, DiffCSP wdist(ρ)=0.1907 and wdist(E)=0.1394 vs. CrystalICL(SGS) 0.6039 and 0.2568. Even CrystalICL(XYZ) does not beat DiffCSP on these metrics. Thus the unconditional superiority claim is false as stated; at best CrystalICL is superior to CrystalLLM. This is an internal contradiction, not a matter of external consensus, and it directly weakens the paper's headline claim. Conditional comparisons also only include CrystalLLM, so 'leading baseline methods' is unsupported there as well.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes CrystalICL, an LLM-based crystal generation method. It introduces a space-group based tokenization (SGS) that represents crystals by space group, Wyckoff positions, and representative atoms; a condition-structure aware instruction-tuning framework with three example-selection strategies; and a multi-task auxiliary property-prediction objective. The method is evaluated on MP20, MP30, P5, and C24 for conditional and unconditional generation, with CrystalLLM, CDVAE, and DiffCSP as baselines. The paper claims that CrystalICL is the first method to leverage LLM in-context learning for crystal generation and that it outperforms leading baselines on conditional and unconditional tasks.","tokens_in":19541,"tokens_out":6498,"duration_ms":68640,"significance":"The SGS representation is a sensible attempt to reduce the burden of learning crystallographic symmetry in autoregressive text generation, and the retrieval-augmented prompt construction is a plausible way to inject few-shot structure–property priors. The ablation study and the preprocessing-cost appendix are useful. However, the central broad claim of superiority over leading baselines is not supported by the reported numbers: unconditional property-distribution metrics favor DiffCSP on several datasets, and the conditional comparison is limited to CrystalLLM. The contribution can still be valuable if the claims are re-scoped and the evaluation protocol is made fairer, but the current framing overstates the evidence.","major_comments":[{"comment":"The abstract claims 'superiority of CrystalICL over the leading baseline methods on conditional and unconditional generation tasks.' Table 2 does not support this for unconditional generation. On MP20, DiffCSP has wdist(rho)=0.1907 and wdist(E)=0.1394, while CrystalICL(SGS) has 0.6039 and 0.2568; on C24, DiffCSP has wdist(E)=0.0415 versus CrystalICL(SGS)=1.3061, and CDVAE also beats CrystalICL on wdist(E). CrystalICL is consistently better than CrystalLLM, but not than the diffusion/VAE baselines. The abstract and conclusion should be restricted to the comparisons actually supported.","section":"Abstract; §4.4 (Table 2)"},{"comment":"For conditional generation the only baseline is CrystalLLM (§4.1), so 'leading baseline methods' is unsupported. Moreover, the reported advantage is format-dependent: on MP20 Pretty Formula, CrystalICL(SGS) 3-Shot reaches 0.8868 versus CrystalLLM(XYZ) 0.9394; on MP30 the corresponding numbers are 0.9641 versus 0.9699. The text in §4.2 calls this 'slightly reduced,' but the drop exceeds 5 percentage points on MP20. Since SGS is a main contribution, the paper should explicitly frame the tradeoff between symmetry-conditioning gains and composition-generation losses, and compare within the same representation.","section":"§4.1–4.2 (Table 1)"},{"comment":"The paper states that when sampling for unconditional generation, 'If a sampled string cannot be parsed as a valid CIF, the sample is rejected and re-sampled.' This makes the reported validity and coverage numbers conditional on a retry loop that diffusion/VAE baselines do not have. As a result, the validity comparisons in Table 2 are not apples-to-apples. The authors should report the raw parse/validity rate before re-sampling and either apply the same rejection criterion to all baselines or drop the validity comparison.","section":"§4.4, unconditional generation protocol"},{"comment":"The conclusion that 'hybrid instruction tuning effectively enhances the capabilities of CrystalICL across various scenarios' is not consistently supported by Table 3. In the 3-Shot setting on MP20, the full CrystalICL achieves Pretty Formula 0.9214, while C-only achieves 0.9340 and noAux achieves 0.9347; for Space Group the full model scores 0.9948 versus 0.9954 (C) and 0.9959 (noAux). The full hybrid set is thus not clearly superior in few-shot inference. This needs a more careful analysis or a softened conclusion.","section":"§4.5 (Table 3)"}],"minor_comments":[{"comment":"The bar chart is difficult to interpret: the percentages above the bars are not clearly tied to the four metric categories, and the 0%/100% labels are unexplained. Also, GPT-3.5 Turbo is used as motivation but is not included as a baseline in the main experiments.","section":"Figure 1"},{"comment":"The sentence 'demonstrating that randomly example selection strategy leads to failing to derive task-relevant information from the demonstrations, thus losing ICL capability' is grammatically unclear and does not follow from the preceding comparison of F, CF, and C. Please rewrite.","section":"§4.5"},{"comment":"The table header mixes 'Validity Check' with separate 'Composition' and 'Structural' columns. Please clarify which numbers correspond to the definitions in Appendix C and whether 'Validity Check' is a combined score.","section":"Table 2 / Appendix C"},{"comment":"The physical-realism metrics (atomic overlap, symmetry adherence, energy feasibility) are reported only in the appendix and are not referenced in the main evaluation. This evidence is relevant to the core claims and should be summarized in the main text.","section":"Appendix H"},{"comment":"No code or data splits are provided. Given the many hyperparameters in Appendix E (LoRA rank, alpha, dropout, learning rate, temperature, top-p, number of shots), a public implementation would substantially improve reproducibility.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The reader's concern about Table 2 is valid: the abstract overstates unconditional superiority. The conditional comparison is also too narrow. I do not see evidence of methodological circularity; the issue is overclaiming and an unfair retry protocol for LLM parsing. The paper is likely salvageable by re-scoping claims, adding within-representation comparisons, and reporting raw parse rates before re-sampling."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know. CrystalICL is a real step: it adapts an LLM to few-shot crystal generation using space-group/Wyckoff tokenization plus hybrid instruction tuning, and the few-shot gains over CrystalLLM look genuine. But the headline claim — superiority over leading baselines on both conditional and unconditional generation — is not supported by the paper's own numbers.\n\nWhat's actually new: the SGS tokenization is clever. Merging atoms by Wyckoff position cuts coordinate complexity and gives the model explicit symmetry guidance, which measurably helps space-group-conditioned generation. The condition-structure aware demonstration selection, especially condition-based selection, is well motivated, and the ablations (Tables 3–5) show it matters. Adding the property-prediction auxiliary task also helps zero-shot performance. The physical realism appendix (atomic overlap, symmetry adherence) is a nice touch. The careful comparison against CrystalLLM under identical formats is useful.\n\nWhere it goes soft: the stress-test concern lands. Table 2 shows that on MP20, DiffCSP gets wdist(ρ)=0.19 and wdist(E)=0.14; CrystalICL-SGS gets 0.60 and 0.26. On C24, DiffCSP wdist(E)=0.04 vs CrystalICL-SGS 1.31. CDVAE also beats CrystalICL on several metrics. So the unconditional superiority claim is false as stated; at best CrystalICL beats CrystalLLM, not the diffusion baselines. Conditional generation is compared only against CrystalLLM, so the phrase \"leading baseline methods\" is overreach there too. The paper also admits SGS trades off chemical-formula success rate — on MP20, CrystalICL-SGS 3-shot gets 0.887 vs CrystalLLM-XYZ 0.939 — yet still concludes broad superiority. That's a real tension. On top of that, P5 and C24 conditional results are only shown as figures, and no code is released, which makes verification hard. The property-conditioned evaluation uses MEGNet with sign-matching and 0.5 eV tolerances; that's lenient, though common.\n\nOverall: the core ICL idea and SGS tokenization are worth taking seriously, and the ablations show honest engineering. The paper overstates what it demonstrates and needs a revised set of claims, full numerical tables, and ideally code. It deserves a serious referee — a good editor should send it out with those expectations.","headline":"A genuinely new few-shot ICL + Wyckoff-tokenization framework for crystal generation, but the central superiority claim is undercut by its own Table 2; worth refereeing after reframing and fuller data.","tokens_in":19997,"tokens_out":1977,"would_cite":true,"duration_ms":21831,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"CrystalICL claims that giving LLMs a compact space-group/Wyckoff representation plus few-shot, condition-aware prompts lets them carry over in-context learning to crystal generation, outperforming zero-shot baselines on four benchmarks.","keywords":["crystal generation","in-context learning","large language models","space group","Wyckoff positions","instruction tuning","few-shot learning","materials discovery"],"falsifier":"Run chemical-formula-conditioned generation on MP20 with SGS at several demonstration counts (e.g., 8 or 16) and compare with XYZ-format CrystalLLM/CrystalICL; if SGS cannot reach the 0.9394 pretty-formula success of the XYZ baseline (or its own 3-shot XYZ 0.9906), the compact representation is discarding composition information that examples cannot restore.","tokens_in":19177,"feed_emoji":"💎","tokens_out":7504,"duration_ms":73213,"temperature":0.7,"pith_summary":"The paper is trying to show that LLMs can design crystals the way human experts do: take a few known structures with similar target properties and modify them. It proposes CrystalICL, a model that keeps the LLM's few-shot in-context learning ability by representing crystals compactly, mixing zero-shot and few-shot instruction tuning, and adding a property-prediction auxiliary task. On four datasets, MP20, MP30, P5, and C24, it reports that three-shot prompts outperform zero-shot prompts and that the resulting crystals better match property distributions than the leading LLM baseline. If the claim holds, materials researchers could steer generation by selecting a handful of relevant reference crystals rather than retraining or imposing symmetry constraints from outside.","feed_headline":"Few-shot prompts let LLMs design crystals from examples","feed_subtitle":"A compact space-group tokenization pushes 3-shot crystal generation past zero-shot baselines on four benchmarks.","key_machinery":"The load-bearing object is the space-group-based crystal tokenization (SGS), which uses Wyckoff positions: points whose site-symmetry groups are conjugate, so atoms sharing a Wyckoff position can be collapsed into one representative atom. This cuts the number of coordinates to generate and removes the need for the model to enforce symmetry. The second mechanism is the condition-structure aware hybrid instruction set, which supplies few-shot demonstrations chosen by property matching, CrystalNN fingerprint distance, or both; the third is a multi-task property-prediction instruction set that masks a property in the text and trains the model to fill it in. Together they are meant to make a fine","core_discovery":"The paper claims that in-context learning is not lost when an LLM is specialized to crystal generation, provided the crystal is written in a form the LLM can reason about. Its space-group-based tokenization (SGS) replaces all atoms in equivalent Wyckoff positions with one representative atom, so the model predicts space-group identity, lattice parameters, and Wyckoff sites rather than every fractional coordinate. A hybrid instruction set mixes zero-shot prompts with few-shot prompts whose examples are chosen by property similarity, structural similarity via CrystalNN fingerprints, or both; a multi-task instruction set also asks the model to predict masked properties from structure text. The","pith_inferences":["The SGS tokenization trades chemical-composition fidelity for symmetry fidelity: on MP20 the 3-shot SGS chemical-formula success is 0.8868 versus 0.9394 for XYZ, so a hybrid representation or a second decoding path for composition might recover both strengths; this is an extension the paper does not test.","The finding that shot count barely matters while example ordering does suggests the demonstrations may act more as format anchors than as a source of transferable chemical knowledge; testing with intentionally misleading but well-formatted examples would separate the two mechanisms.","CrystalNN-based structure retrieval is a natural candidate for replacing hand-designed property filters; a learned retriever or embedding-based scorer could make the few-shot selection cheaper and more effective on large databases, but that is outside the paper's experiments.","Because the model is trained and evaluated on DFT-relaxed structures, few-shot ICL could plausibly extend to other crystal-property mappings, such as synthesizability or ionic conductivity, if structure text for those properties is available in the prompt; the paper leaves this unexplored."],"forward_implications":["If CrystalICL holds, few-shot prompting becomes a viable interface for materials generation: a user supplies a target property plus a handful of similar crystals and gets candidate structures without task-specific retraining.","Space-group-conditioned generation becomes practical in LLMs; the paper reports a jump in space-group success rate from XYZ text (around 6-11%) to SGS text (above 98% on MP20/MP30), with symmetry adherence above 96%.","The auxiliary property-prediction task can serve as a general recipe: training an LLM to invert structure-property mappings improves its generative fidelity even when generation is the primary task.","Unconditional generation over narrow domains, such as the all-carbon C24 set, is made tractable for LLMs: SGS reduces Wasserstein distances for density and formation energy versus the XYZ-format baseline.","Condition-based demonstration selection matters more than demonstration count; the paper finds shuffling the order of examples hurts, while varying the number of shots does not significantly change success."],"supporting_citations":[{"why":"Baseline CrystalLLM; it fine-tunes Llama-2 on XYZ-format crystal text in zero-shot settings and serves as the main comparative baseline and the motivation for inheriting ICL.","marker":"Gruver et al., 2024"},{"why":"Supplies the XYZ text format that CrystalICL compares against and that CrystalLLM is built on.","marker":"Flam-Shepherd and Aspuru-Guzik, 2023"},{"why":"Defines the CIF format that SGS simplifies, framing the tokenization problem.","marker":"Hall et al., 1991"},{"why":"Supplies the Wyckoff-position concept on which the SGS tokenization is based.","marker":"LIPSON, 1949"},{"why":"Defines CrystalNN fingerprints used for structure-based and condition-structure example selection.","marker":"Zimmermann and Jain, 2020"},{"why":"Provides the Materials Project database behind MP20 and MP30 and the reference data used for property validation.","marker":"Jain et al., 2013"},{"why":"Supplies the P5 perovskite dataset used as a domain-specific benchmark.","marker":"Castelli et al., 2012"},{"why":"Supplies the C24 carbon dataset generated by AIRSS and used as a narrow-domain benchmark.","marker":"Pickard, 2020"},{"why":"Provides Llama-2, the pretrained base model that CrystalICL fine-tunes with LoRA.","marker":"Touvron et al., 2023"},{"why":"Supplies LoRA, the low-rank adaptation method used for efficient fine-tuning of the LLM.","marker":"Hu et al., 2022"}],"fun_headline_variants":["CrystalICL: Few-shot crystal generation with space-group tokens","Space-group tokens unlock few-shot crystal design in LLMs","CrystalICL harnesses in-context learning for crystal generation","Fewer examples, better crystals: CrystalICL's few-shot edge"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The compact space-group/Wyckoff representation must preserve the compositionally relevant details of a crystal; the paper's own Table 1 shows a clear drop in chemical-formula success under SGS, so if that drop is not recoverable by demonstrations, the method's generality is only partial.","fun_headline_variants_meta":{"raw":{"variants":["CrystalICL: Few-shot crystal generation with space-group tokens","Space-group tokens unlock few-shot crystal design in LLMs","CrystalICL harnesses in-context learning for crystal generation","Fewer examples, better crystals: CrystalICL's few-shot edge"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00022,"raw_usage":{"total_tokens":1252,"prompt_tokens":682,"completion_tokens":570,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":426,"completion_tokens_details":{"reasoning_tokens":496}},"tokens_in":426,"tokens_out":570,"duration_ms":6381,"temperature":1.0,"reasoning_tokens":496,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T15:35:25.758160+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run chemical-formula-conditioned generation on MP20 with SGS at several demonstration counts (e.g., 8 or 16) and compare with XYZ-format CrystalLLM/CrystalICL; if SGS cannot reach the 0.9394 pretty-formula success of the XYZ baseline (or its own 3-shot XYZ 0.9906), the compact representation is discarding composition information that examples cannot restore.","supporting_citations":[{"cited_title":"Lawrence Zitnick, and Zachary W","cited_arxiv_id":null,"evidence_quote":"Baseline CrystalLLM; it fine-tunes Llama-2 on XYZ-format crystal text in zero-shot settings and serves as the main comparative baseline and the motivation for inheriting ICL."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the CIF format that SGS simplifies, framing the tokenization problem."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines CrystalNN fingerprints used for structure-based and condition-structure example selection."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the P5 perovskite dataset used as a domain-specific benchmark."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the C24 carbon dataset generated by AIRSS and used as a narrow-domain benchmark."}],"review_version":1}