{"id":"3b9fe69e-c300-4c84-82ce-95312a39354b","arxiv_id":"2509.02240","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Style-translated simulated AFM images improve machine-learning structure discovery on experimental AFM data, as judged by agreement with simulation-derived structural distributions.","lead":"This paper trains a style-translation model to turn simulated atomic force microscopy images into experimental-looking ones, and shows this improves machine-learning predictions of water structure on a gold surface. It matters because it offers a way to train structure-discovery models when no labelled experimental data exists.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Evaluation is circular: Fig. 7 reference distributions come from the same simulated configuration set M used to train every model, so improved distributional agreement may reflect conformity to the training prior, not better structure recovery on experimental data.","rationale":"The reader's weakest assumption correctly identifies the load-bearing issue: the evaluation reference is derived from the same simulated configuration set that generated all training data, so the reported improvements conflate true structure recovery with conformity to the training prior. My stress test finds this concern is real and specific. The paper's own SI strengthens it by showing that style-translated models lose sensitivity to lower-layer atoms and that forward PPM reconstructions do not discriminate between models. The style-gap reduction itself is well supported by the machine-expert and FID/WD results, so the paper's engineering contribution stands. However, the headline claim about improved performance on real experimental inputs is not independently validated. Since the reader already returned CONDITIONAL, my read does not change that verdict; it reinforces the need for an independent structure benchmark before accepting the claim as established.","tokens_in":23163,"tokens_out":2905,"duration_ms":35341,"concrete_test":"Obtain an independent benchmark with known atomic structure—either a DFT-resolved experimental system or a held-out simulation configuration set M_holdout not used in training or reference construction—and compute atom-level precision/recall/F1 and per-atom RMSD for FU versus F˜V. If F˜V does not outperform FU on atom-level recovery, or does so only for top-layer atoms, the distributional improvement in Fig. 7 is explained by training-prior conformity rather than improved structure discovery.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that style-translated training data improve structure discovery on real experimental AFM images—rests entirely on the distributional evaluation in Fig. 7. The reference target distributions are computed from the top-layer water molecules of the same configuration set M that provided every training label (Section V: 'The reference theoretical target distributions in Fig. 7 are obtained from the top-layer water molecules in the configuration set M as shown in Fig. 6'). A model trained on images whose labels come from M, then evaluated by distance to M-derived distributions, is partially measuring how well it reproduces the training prior, not whether it recovered the unknown experimental structure. This is not merely hypothetical: the SI admits style-translated models are more conservative and systematically miss lower-layer molecules (Figs. 9–11 and the θZOH discussion around Fig. 15), and the forward PPM checks 'appear similar' across models, so they do not discriminate. Because the claimed improvement over the baseline is measured only on metrics that reward agreement with M, the headline conclusion 'significantly better performance on real experimental inputs' is not yet substantiated.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper addresses the simulation-to-experiment gap in machine-learning-based structure discovery from AFM images. The authors train a CycleGAN to translate simulated 3D AFM images (represented as 2D slices) into an experimental style, then train structure-prediction models on the translated images. They evaluate these models on real experimental AFM images using distributional comparisons of local structural properties (dOO, dOH, θHOH, θZOH, hydrogen-bond geometry, tetrahedral order) against reference distributions derived from the same simulated configuration set M used to generate all training labels. The paper reports that style-translated training data yield improved structural-property distributions relative to a baseline trained on pure simulation, and concludes that style translation improves structure discovery on real experimental inputs.","tokens_in":23406,"tokens_out":3357,"duration_ms":43202,"significance":"If the central claim is valid, the work offers a practical route to leveraging unlabeled experimental AFM images for training structure discovery models, which would be valuable for the microscopy and ML communities. The paper has several strengths: the style-gap reduction itself is convincingly demonstrated through authenticity-score shifts, Wasserstein distance, and FID (Fig. 3); the study includes multiple distributional metrics, 10 independent replicas per configuration, and detailed supporting information; and the authors commit to releasing code and data. The approach of evaluating without ground-truth atomic structures is thoughtful and clearly motivated. However, the central claim of improved structure discovery rests on an evaluation that is partially circular with respect to the training data, and the paper's own supporting information documents a known sensitivity loss for lower-layer molecules. These issues must be addressed before the conclusion can be considered established.","major_comments":[{"comment":"The evaluation of structure-discovery performance is circular with respect to the training data. The reference theoretical target distributions in Fig. 7 are explicitly obtained from the top-layer water molecules of configuration set M, which is the same set used to generate every training image and label (Fig. 4, Section IV). Thus the metrics measure how well the predicted structures reproduce the training-set statistics, not how accurately the model recovers the unknown experimental atomic structure. A model that outputs structures near the M prior will score well regardless of image content. The forward PPM comparisons in the SI (Figs. 8–11) are also described as 'appear similar' across models, so they do not discriminate. To substantiate the headline claim, an independent evaluation is needed: e.g., a held-out simulation set with different physics, a known experimental crystal struct","section":"§V, Fig. 7"},{"comment":"The paper's own supporting information documents that style-translated models systematically miss lower-layer water molecules, creating a spurious peak in θZOH near 170°. Yet the main-text evaluation (Fig. 7) aggregates across properties with min-max normalization and does not prominently report this failure. Since the missed lower-layer molecules are a known limitation directly tied to the style-translation pipeline, the claim of 'significantly better performance on real experimental inputs' is overstated. The authors should either restrict the claim to top-layer properties explicitly, or provide evidence that the distributional agreement correlates with true structure recovery despite the missed molecules.","section":"SI, Fig. 15 and §V"},{"comment":"The style translator is applied independently to each 2D slice of the simulated AFM image, and the translated slices are then stacked to form a 3D experimental-style image. The paper does not validate that slice-wise independent translation preserves the vertical consistency of the 3D AFM signal. Since the structure-discovery model consumes 3D images, any slice-independent artifacts or noise correlations introduced by the generator could alter the apparent height-dependent features and thus affect predictions. This assumption is load-bearing for the method; an ablation or validation (e.g., comparing vertical profiles before and after translation) would strengthen the paper.","section":"§IV, Fig. 4E"},{"comment":"The attribution of improved performance to the learned style translation is not fully isolated from generic augmentation. Handcrafted perturbations also improve some metrics (Fig. 7A), and the hybrid model performs best, suggesting that added stochasticity or regularization may account for part of the gain. The paper compares against handcrafted perturbations but does not control for the intensity or amount of added noise. An ablation that matches the perturbation strength between handcrafted and style-translated images, or a test with noise level as a hyperparameter, would clarify whether the improvement is due to 'style' or simply to data augmentation.","section":"§IV and §V"},{"comment":"The performance scores in Fig. 7 use min-max normalization across all models in the computational experiments, and the reported error bars are standard errors over replicas. No statistical significance tests or effect sizes are provided. For some properties (e.g., tetrahedral order parameters), the authors state that gains are not evident. Without significance testing, it is difficult to assess whether the reported improvements are robust or within replica noise. Reporting confidence intervals or pairwise significance tests for each property would make the evaluation more rigorous.","section":"Fig. 7 and Materials and Methods"}],"minor_comments":[{"comment":"'Handcrafted permutations' should be 'handcrafted perturbations' (also in the text near Fig. 3C).","section":"Fig. 3 caption and §III"},{"comment":"The variables m and n are used for both domain sizes and batch sizes; this is confusing. Use distinct notation (e.g., M_batch, N_batch).","section":"Eq. (1)"},{"comment":"References [33] and [47] are the same arXiv paper; [34] and [47] also overlap. These duplication issues should be cleaned up.","section":"References"},{"comment":"The radar charts are visually dense and the normalization procedure is not intuitive from the figure alone. Adding a caption note or a table with the actual normalized distances would improve readability.","section":"§V, Fig. 7"}],"recommendation":"major_revision","confidential_remarks":"The circularity issue is the core problem: the evaluation reference comes from the same configuration set M that generated all training labels, so the reported improvement may be largely a measure of how well the model reproduces the training prior. I would advise the editor that the paper needs an independent structural reference—either a separate simulation ensemble not used in training, or an experimentally known structure—to support the central claim. The SI's admission of missed lower-layer molecules further weakens the 'significantly better performance' statement. The style-gap reduction part is solid, so a major revision is appropriate rather than rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know before reading: this is a solid engineering paper, not a breakthrough. The CycleGAN-based style translation works—the image-level gap between simulated and experimental AFM images is convincingly reduced, as shown by authenticity scores, Wasserstein distance, and FID. Handcrafted perturbations don't achieve the same effect, which is a fair and useful comparison. The idea of evaluating structure predictions through physically meaningful structural distributions (dOO, dOH, angles, hydrogen-bond geometry, tetrahedral order) is genuinely useful for the no-ground-truth setting, and the paper is honest about its own limitations.\n\nThe soft spot is exactly where the stress-test note points. The reference distributions in Fig. 7 come from the top-layer water molecules of configuration set M, the same simulated set that generated every training label. So when style-translated models show better agreement with those distributions, part of that improvement may simply reflect conformity to the training prior. The baseline FU shares that prior, which softens the circularity—both models are trained on M-derived labels—but it doesn't remove the problem: without an independent experimental ground truth, closeness to M distributions is not proof of better structure recovery. The paper itself acknowledges that style-translated models are more conservative and miss lower-layer molecules (Figs. 9–11 and the θZOH discussion), and that the forward PPM checks \"appear similar\" across models. So the central claim of \"significantly better performance on real experimental inputs\" is plausible but not conclusively established. That's the key thing a referee should push on.\n\nMinor points: code and data are promised but not yet public; the per-slice 2D-to-3D translation is not tested for 3D consistency; and the evaluation has several free parameters (λc, λi, cutoffs, MMD bandwidth) that could be tuned to favor a particular model. None of these are fatal. The citation pattern is fine—prior work from the same group is appropriately cited.\n\nThis paper is for researchers working on ML-based interpretation of AFM/SPM images, especially those facing the sim-to-exp gap. It deserves a serious referee and, if the circularity is addressed with an independent benchmark or a system with known structure, it would be a useful contribution. I'd bring it to a reading group and would cite it in related work.\n\nRecommendation: send it to peer review, but with a clear request for an evaluation that does not rely solely on distributions derived from the training set.","headline":"A careful CycleGAN style-transfer study for AFM that convincingly closes the image-level sim-to-exp gap, but the headline claim of better structure discovery rests on an evaluation that is partly circular because the reference distributions come from the same simulated configuration set used for training.","tokens_in":23894,"tokens_out":1621,"would_cite":true,"duration_ms":21952,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":["68.37.Ps"],"model":"deepseek-v4-flash","headline":"By translating simulated AFM images into experimental style with a CycleGAN, the paper shows that structure discovery models trained on the translated images predict local water structure on real experimental AFM images significantly better","keywords":["atomic force microscopy","structure discovery","style translation","CycleGAN","domain gap","water on Au(111)","probe particle model","machine learning"],"falsifier":"An experimental AFM dataset on which the atomic structure is independently determined (for example, by another high-resolution technique or by a system with a known surface registry) would settle the claim: the style-translated model should recover more of the independently known atoms—especially lower-layer species—than the pure-simulation baseline. The paper itself notes that style-translated models suppress low-lying atoms, so a case where independently confirmed lower-layer molecules are systematically missed would falsify the claim that style translation improves structure discovery rathe","tokens_in":23038,"feed_emoji":"🔬","tokens_out":6222,"duration_ms":58464,"temperature":0.7,"pith_summary":"Machine learning models that reconstruct atomic structures from atomic force microscopy (AFM) images are trained on simulations, because experimental images lack ground-truth structures—but simulations look suspiciously clean, and performance drops on real data. This paper tries to close that gap by first translating simulated AFM images into the style of experimental ones using an unpaired image-to-image translation model (CycleGAN), then training the structure discovery model on the translated images. Applied to bilayer water on Au(111), the approach makes predicted configurations match reference distributions of local structural properties—oxygen distances, hydrogen-bond geometry, tetrahedral order—more closely than training on pure simulation or on handcrafted noise perturbations. The point of the work is a practical one: when labelled experimental data is unavailable, style translation can substitute for it.","feed_headline":"Sim-to-experiment style transfer boosts AFM structure discovery","feed_subtitle":"Training on CycleGAN-style simulated water images beats pure simulation on real AFM inputs.","key_machinery":"The load-bearing machinery is a cycle-consistent generative adversarial network (CycleGAN) acting as an unpaired image-to-image translator. Two generators, GU (simulation-to-experiment) and GV (experiment-to-simulation), are trained with adversarial losses, a cycle-consistency loss (translating and returning should recover the original image), and an identity loss (inputs already in the target style should pass through unchanged). The forward generator GU is applied slice-by-slice to 3D simulated AFM images to produce experimental-style training volumes. The paper also builds a separate evaluation machinery: a 'machine expert' binary classifier scores image authenticity, Wasserstein distance","core_discovery":"The paper's central claim is that the simulation-to-experiment 'style gap'—the noise, artefacts, and subtle distortions present in real AFM images but absent in particle-probe-model simulations—degrades the structure discovery model trained on simulation only, and that reducing this gap in the training data recovers much of the lost performance. The authors train a CycleGAN on unpaired sets of 729 simulated and 728 experimental 2D AFM slices; the forward generator GU maps simulated slices to experimental style, and the stacked 3D volumes are used to train structure discovery models. On six real experimental AFM images of water on Au(111), models trained on style-translated (and hybrid style-","pith_inferences":["The evaluation's reference distributions come from the same simulated configuration set that generated the training images, so part of the reported 'improvement' may reflect better conformity to the training prior rather than truer recovery of the unknown experimental structure; a direct test would require experimental images with independently known atomic structures.","The paper notes a generalisation-versus-sensitivity trade-off: style-translated models ignore weak signals from lower-layer molecules, so future work could attach per-atom confidence scores to predictions, a direction the authors flag.","The same unpaired style-translation recipe should transfer to other scanning probe modalities (STM) and other adsorbate/substrate systems, wherever simulation-trained models meet unlabelled experimental images.","Since the structural-property distributions are used only for evaluation here, turning them into training constraints is a natural extension that the paper itself raises as an open question."],"forward_implications":["Training structure discovery models on style-translated simulated AFM images yields better agreement with reference structural distributions on real experimental inputs than training on pure simulation.","Handcrafted perturbations (Gaussian noise, cutout, gradient background) barely shift the authenticity distribution and produce only narrow performance gains, whereas the learned style translation reduces both Wasserstein and FID distances to the experimental domain.","The reverse generator GV acts as an image denoiser, removing noise and artefacts from experimental images, which the paper suggests as a separate practical application.","A hybrid dataset combining style translation with handcrafted perturbations gives the most balanced improvement across the six structural metrics.","The distribution-based evaluation scheme offers a way to assess structure discovery performance on experimental data without ground-truth atomic structures."],"supporting_citations":[{"why":"CycleGAN; supplies the unpaired image-to-image translation method that is the core of the style translation.","marker":"[26]"},{"why":"Supplies the structure discovery model architecture and the bilayer water on Au(111) simulated dataset used for training and reference distributions.","marker":"[18]"},{"why":"The probe particle model (PPM) that generates simulated AFM images from atomic configurations, defining the forward problem.","marker":"[8]"},{"why":"The earlier automated structure discovery pipeline that establishes the task of predicting atomic structure from AFM images.","marker":"[13]"},{"why":"FID, the metric used to measure the residual style gap between translated and experimental image distributions.","marker":"[36]"},{"why":"Maximum mean discrepancy, one of the three distributional distance metrics used to score predicted versus reference structural distributions.","marker":"[50]"},{"why":"Energy distance, the second distributional metric used in the performance evaluation.","marker":"[49]"},{"why":"Molecule graph reconstruction from AFM images; the line of structure discovery models this work augments.","marker":"[17]"}],"fun_headline_variants":["Style transfer closes AFM's sim-to-real gap","Sim-to-real style transfer sharpens AFM structure models","Style-translated AFM training beats pure simulation on real data","AFM style gap bridged to boost real structure discovery","Style translation unlocks AFM structure discovery from real images"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The load-bearing premise is that closeness to the simulation-derived reference distributions of local structural properties measures structure discovery accuracy; if those distributions do not track the true (unknown) experimental atomic structure—they are computed from the same simulated configuration set that made the training images—the reported gains could partly reflect conformity to the training prior rather than better structure recovery.","fun_headline_variants_meta":{"raw":{"variants":["Style transfer closes AFM's sim-to-real gap","Sim-to-real style transfer sharpens AFM structure models","Style-translated AFM training beats pure simulation on real data","AFM style gap bridged to boost real structure discovery","Style translation unlocks AFM structure discovery from real images"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001191,"raw_usage":{"total_tokens":4722,"prompt_tokens":688,"completion_tokens":4034,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":432,"completion_tokens_details":{"reasoning_tokens":3954}},"tokens_in":432,"tokens_out":4034,"duration_ms":28262,"temperature":1.0,"reasoning_tokens":3954,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T11:44:39.166972+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"An experimental AFM dataset on which the atomic structure is independently determined (for example, by another high-resolution technique or by a system with a known surface registry) would settle the claim: the style-translated model should recover more of the independently known atoms—especially lower-layer species—than the pure-simulation baseline. The paper itself notes that style-translated models suppress low-lying atoms, so a case where independently confirmed lower-layer molecules are systematically missed would falsify the claim that style translation improves structure discovery rathe","supporting_citations":[{"cited_title":"Carracedo-Cosme, C","cited_arxiv_id":null,"evidence_quote":"CycleGAN; supplies the unpaired image-to-image translation method that is the core of the style translation."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the structure discovery model architecture and the bilayer water on Au(111) simulated dataset used for training and reference distributions."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The probe particle model (PPM) that generates simulated AFM images from atomic configurations, defining the forward problem."},{"cited_title":"Albrecht, N","cited_arxiv_id":null,"evidence_quote":"The earlier automated structure discovery pipeline that establishes the task of predicting atomic structure from AFM images."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"FID, the metric used to measure the residual style gap between translated and experimental image distributions."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Maximum mean discrepancy, one of the three distributional distance metrics used to score predicted versus reference structural distributions."},{"cited_title":"Offei-Danso, A","cited_arxiv_id":null,"evidence_quote":"Energy distance, the second distributional metric used in the performance evaluation."},{"cited_title":"Heggemann, Y","cited_arxiv_id":null,"evidence_quote":"Molecule graph reconstruction from AFM images; the line of structure discovery models this work augments."}],"review_version":1}