{"id":"360d0808-b2e2-4b79-b856-3e5b4d072913","arxiv_id":"2608.09100","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Fine-tuning optical structure recognizers with labeled real depictions, even at 5 to 10 percent of the training mixture, is the most effective intervention for closing the synthetic-to-real gap.","lead":"This paper shows that adding a small amount of real labeled chemical structure images to synthetic training data is the main lever that improves recognition on real patents and journal figures. It also shows that the best base model and whether to adapt the vision part depend strongly on how much real data is available.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Real-data effect conflates depiction realism with molecular-domain match; without synthetic renders of the real pool, the central claim is underdetermined.","rationale":"The reader identified single-seed training as the weakest assumption. That is a genuine limitation and is properly flagged by the paper itself (Secs. 9, C.1); however, it affects the precision of the reported magnitudes and rankings, not the direction of the real-data effect, which is large and consistent across bases. The missing 2×2 control is more load-bearing because it bears on what the effect actually is: 'real labeled training images' vs 'real molecular structures.' The paper's own data show that synthetic scale alone does not help, but that experiment never adds real-world molecules in synthetic form, so it cannot rule out molecular-domain match as the active ingredient. The abstract and contribution (1) make a causal claim about depiction realism; the design conflates two variables. I therefore partially agree with the reader: single-seed replication is worth doing, but the synthetic-rendered real-molecule arm is the experiment most likely to change the interpretation of the central claim. The verdict stays CONDITIONAL: the headline finding is plausible and well supported in direction, but its attribution to image realism, and hence the practical recommendation, needs the missing control.","tokens_in":21265,"tokens_out":7354,"duration_ms":82793,"concrete_test":"Train a matched 20%-real-fraction cell in which the same real-pool molecules used in the existing real arm are rendered with the same synthetic toolkits (RDKit, CoordGen, Indigo) used for the synthetic arm, keeping total mixture size, steps, and LoRA surface identical to the controlled sweep; evaluate on ACS and the other real sets. Run at least 3 seeds per arm. If the synthetic-rendered real-molecule cell reproduces most of the gain of the real-depiction cell, the operative factor is molecular distribution rather than depiction realism, and the claim should be re-scoped; if it does not, the realism interpretation is confirmed.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The most load-bearing gap is that the 'real-data effect' is never separated from molecular provenance. In the controlled 3×6 sweep (Sec. 4.3) and the Qwen dose series (Sec. 4.2), every real training example is simultaneously a real document depiction and a molecule drawn from the real-world chemical distribution (USPTO patents, journal figures, hand-drawn sources; Sec. A.2), while every synthetic example is both a synthetic depiction and a molecule from the curated one-million-molecule set. Raising the real fraction therefore replaces both depiction style and molecular distribution at once. The paper's own limitation statement (Sec. 9) admits that 'raising the real fraction adds images and molecules together,' and the InChIKey blocklist removes exact identity overlap but not scaffold or source-style overlap. The synthetic scaling experiments (Sec. 4.1) do not control for this: they add more molecules from the curated collection, not real-world molecules rendered synthetically, and show no monotonic gain on external PubChem. Consequently, the observed ACS/CLEF-IP/USPTO gains could be driven by training on molecules close to the test distribution rather than by the realism of the training depictions. Contribution (1) — 'real labeled training images are the most effective way' — requires a 2×2 control: real vs synthetic depictions crossed with real vs curated molecular provenance, or at least synthetic renders of the real-pool molecules. Without that arm, the practical recommendation to prioritize collecting real depictions over structure databases is not fully supported.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies optical chemical structure recognition (OCSR) with vision-language models, comparing fine-tuning on synthetic renders versus mixtures that include labeled real document depictions. Across three VLM bases and six real-data fractions in a controlled sweep, plus an independent Qwen dose series, the authors report that increasing the real-data fraction improves exact match on real document benchmarks (ACS, CLEF-IP, UOB, USPTO) while slightly lowering rendered-clean accuracy. They also report that vision-tower LoRA helps InternVL3-8B strongly, helps GLM-4.1V-9B modestly, and does not help Qwen2.5-VL-7B, and that base-model rankings change with the real-data fraction. The paper's main contribution is the claim that real labeled training images are the most effective way to close the synthetic-to-real gap, supported by a controlled experimental design and a common evaluation harness.","tokens_in":21503,"tokens_out":3183,"duration_ms":37115,"significance":"If the central claim is taken at face value, the practical implication is important: OCSR practitioners should prioritize collecting labeled real depictions over adding more synthetic renders or renderer diversity. The paper is commendable for its controlled 3x6 design, fixed mixture size and training budget, paired statistical tests, shared evaluation protocol across specialist recognizers and chemical VLMs, and unusually frank limitations section. The main weakness is that the real-data effect is measured under a confound between depiction realism and molecular provenance: every real training example is also a molecule from the real-world distribution, while every synthetic example is a molecule from the curated one-million set. Because the paper itself acknowledges this in Section 9, the core contribution needs additional experimental separation before the strong practical recommendation is fully supported.","major_comments":[{"comment":"The central claim that real labeled images close the synthetic-to-real gap is confounded with molecular provenance. In the controlled 3×6 sweep and the Qwen dose series, raising the real-data fraction simultaneously replaces the depiction style (real documents vs. RDKit/CoordGen/Indigo renders) and the molecular distribution (real-pool molecules such as USPTO patents and journal figures vs. the curated one-million-molecule set). Section 9 explicitly states that 'raising the real fraction adds images and molecules together,' and Section A.2 notes that the InChIKey blocklist removes exact identity overlap but not scaffold or source-style overlap. The observed ACS/CLEF-IP/USPTO gains could therefore be driven by training on molecules closer to the test distribution rather than by the realism of the training depictions. To support Contribution (1), the paper needs a control that separates depiction style from molecular provenance, for example synthetic renders of the real-pool molecules or at least a molecular-provenance-matched arm. Without such an arm, the claim that real depictions are the effective ingredient is underdetermined.","section":"§4.3, §9, §A.2"},{"comment":"The paper does not establish that real data is 'the most effective way' to close the gap because the alternative interventions—synthetic scale, renderer diversity, degradation, and the verifiable RL objective—were not evaluated on the real-document benchmark suite. Table 2 reports these interventions only on rendered conditions, and the text itself says 'no conclusion about real-document transfer can be obtained from these runs alone.' The only other intervention measured on real documents is the vision-LoRA contrast, which shows no effect for Qwen. Thus the empirical content is that real data is the only tested intervention that improves real-document accuracy, not that it is more effective than alternatives that could have been evaluated on the same real benchmarks. At minimum, one strong synthetic-only recipe should be evaluated on ACS, CLEF-IP, UOB, and USPTO under the same protocol to give the comparative claim substance.","section":"§4.1, §5.2"},{"comment":"The ranking and between-base spread claims rest on single-seed cells. The controlled 18-cell sweep trains exactly one model per cell, and Section 9 acknowledges that the reported confidence intervals capture evaluation-set uncertainty, not training-seed variation. Section C.1 further states that the 0.060 between-base spread at 70% real data 'should not be interpreted as a stable ranking without multi-seed replication.' This directly undermines Contribution (4) and the associated discussion of base-model reordering, since those conclusions depend on the assumption that a single fine-tuning run represents each condition. The monotonic real-data trend is less vulnerable because it is replicated across three bases and an independent descriptive series, but the ranking-change and spread-shrinkage claims need multi-seed replication for at least the 0%, 5%, and 70% cells, or they should be explicitly downgraded to single-run observations.","section":"§9, §C.1"}],"minor_comments":[{"comment":"The sentence 'These paiors are not cells of a full base×real fraction×adaptation surface design' contains a typo: 'paiors' should be 'pairs'.","section":"§3.1"},{"comment":"The acronym is introduced as 'Optical Chemical Structure Recognition (OSCR)' but the standard and consistently used abbreviation in the paper is OCSR; please correct the introduction to read 'Optical Chemical Structure Recognition (OCSR)'.","section":"§1"},{"comment":"The phrase 'In computation chemistry' should be 'In computational chemistry'.","section":"§1"},{"comment":"The table row 'b/c—11/11 16/27 25/43 22/22' is cryptic; the b/c values should be labeled explicitly as the discordant-pair counts for the McNemar test, and the formatting should be aligned with other tables.","section":"§5.2, Table 6"},{"comment":"The description of the effective epoch count is clear, but the relationship between the 496k-sample mixture and the exact 10,000-step budget would benefit from stating the effective batch size explicitly in the main text rather than only in the appendix.","section":"§A.1"}],"recommendation":"major_revision","confidential_remarks":"The paper is honest and well-executed on many fronts, and the reviewer agrees with the stress-test concern: the real-data effect is genuinely confounded with molecular provenance, and the paper's own limitation statement concedes this. The fix is not merely editorial; it requires either a new experimental arm (synthetic renders of the real-pool molecules) or a substantial reframing of the central claim to the narrower, defensible statement that adding real labeled images in a one-image-per-molecule regime improves real-document accuracy. The single-seed issue is also load-bearing for the ranking claims. I would not reject the paper, but I would not accept it without these points being addressed."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know. First, this is a genuinely useful empirical paper: 21 fine-tuned recognizers, a controlled 3×6 base-by-real-fraction sweep, and a common evaluation harness. The headline result—labeled real depictions are worth more than more synthetic renders—is supported by consistent gains across three bases and an independent dose series. Second, the paper is more careful than most about its own limitations, and one of those limitations is load-bearing: the real-data effect is never separated from molecular provenance. Every real training image is a real depiction and a real-world molecule; every synthetic image is synthetic and from a curated pool. The authors admit this in Sec. 9 (\"raising the real fraction adds images and molecules together\"), but it means Contribution (1) as stated—\"real labeled training images are the most effective way\"—is only partly established. A 2×2 crossing depiction realism with molecular provenance, or at least synthetic renders of the real-pool molecules, would settle whether the gains come from the images or the molecules.\n\nWhat is genuinely new: the controlled design itself, the base-dependent effect of vision LoRA (nothing for Qwen, large for InternVL), and the careful paired McNemar analyses on identical images. The paper also earns credit for transparency: single-seed cells, weak consensus labels, non-grid checkpoints, and the failed full-unfreeze arms are all disclosed, and the planned release of predictions, code, and construction metadata is concrete.\n\nSoft spots, in proportion. The single-seed cells are real but secondary; the effect sizes are large enough that the central trend probably holds, though the 0.060 between-base spread at 70% real data should not be read as a stable ranking. The synthetic scaling experiments show no benefit from adding more curated molecules, which argues against a pure molecular-diversity story, but they do not test real-world molecular distribution. The interpretation that depiction realism drives the gains is therefore underdetermined. That is the main revision request, not a rejection.\n\nWho this is for: anyone working on OCSR or on fine-tuning VLMs for structured prediction. It deserves serious peer review; the central confound should be addressed by adding the 2×2 arm or softening the claim. I would bring it to my reading group—the discussion of what counts as a controlled comparison in VLM fine-tuning is worth the time.","headline":"A well-designed empirical study showing real training images beat synthetic scaling for OCSR, though the reason—depiction realism vs molecular provenance—is underdetermined; deserves peer review.","tokens_in":22044,"tokens_out":3732,"would_cite":true,"duration_ms":42893,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Labeled real training images, not more synthetic renders, close the synthetic-to-real gap in optical chemical structure recognition, with exact match on a held-out journal-figure benchmark rising from 0.154 to 0.459 as the real-data…","keywords":["optical chemical structure recognition","synthetic-to-real gap","vision-language models","LoRA fine-tuning","real labeled data","SMILES generation","document image understanding","exact match evaluation"],"falsifier":"Re-run the controlled sweep with multiple independent seeds per cell at 0%, 20%, and 70% real data using the same fixed evaluation sets. If between-seed spread at 70% equals or exceeds the observed 0.060 between-base spread, or if base orderings flip across seeds, the paper's ranking and contraction conclusions would not survive.","tokens_in":21050,"feed_emoji":"🧪","tokens_out":8669,"duration_ms":84904,"temperature":0.7,"pith_summary":"The paper takes on a mismatch that matters for chemistry: optical chemical structure recognition (OCSR) looks solved on synthetic renderings but collapses on real documents, where most structures actually live. It tries to pin down what actually closes that gap by fine-tuning 21 recognizers on controlled mixtures of synthetic and labeled real depictions, varying the base model, the fraction of real training data, and where LoRA adaptation is applied. The central finding is that labeled real training images are the most effective lever: on the primary held-out journal-figure benchmark, exact match rises from 0.154 with no real data to 0.372 at 9.5% real data and 0.459 at 50.2%, while synthetic scaling, renderer diversity, degradation, and a verifiable reward produce small or inconsistent gains. A controlled sweep across three base models reproduces the trend. If the paper is right, OCSR should reallocate effort from synthetic render scaling to collecting and labeling real depictions.","feed_headline":"Real data beats synthetic renders in chemical structure recognition","feed_subtitle":"A controlled sweep across three base models shows labeled real depictions are the biggest lever.","key_machinery":"The key machinery is the controlled 3-by-6 factorial sweep: three pretrained vision-language model bases crossed with six real-data fractions (0%, 5%, 10%, 20%, 40%, 70%) under a fixed approximately 496k-sample mixture, a fixed 10k-step budget, and a fixed light vision-LoRA adaptation surface. This design isolates the real-data fraction from base choice and training recipe, and it is paired with a canonical exact-match scoring metric that parses predicted SMILES with RDKit and compares molecular identity through InChIKey; the same scoring protocol is applied to every configuration and every external comparison.","core_discovery":"On the paper's own terms, the central discovery is that representative real supervision is the main thing that closes the synthetic-to-real gap in OCSR, and that this effect is separable from the choice of base model or adaptation strategy. The controlled 3-by-6 sweep crosses three pretrained vision-language model bases with six real-data fractions, holding mixture size, training budget, objective, and adaptation surface fixed, and shows mean real-document exact match rising monotonically from as low as 0.027 at 0% real data to between 0.674 and 0.734 at 70%, with the first 5% real data delivering about two-thirds of the total gain. The paper also shows that base-model rankings change with the real-data fraction, that the between-base spread contracts from 0.212 at 0% to 0.060 at 70%, and that extending LoRA to the vision tower helps one base by 22.8 to 34.6 points, helps another modestly, and leaves a third unchanged.","pith_inferences":["A testable extension: if the observed real-data dose response is a general phenomenon, then even a small labeled real set (around 5% of the mixture) may be the most cost-effective next step, since roughly two-thirds of the total gain arrives there.","The paper's consensus-reconciliation labeling rule implies a bootstrapping loop in which the recognizer itself helps generate the real-data resource needed to improve it; that loop is described but not fully evaluated as a training-data acquisition strategy.","The cross-task probes suggest the base-dependence pattern may generalize to other image-to-structure tasks, but single-run results make that an open question rather than a conclusion."],"forward_implications":["OCSR development should prioritize collecting labeled real depictions from patents, journals, and hand-drawn collections over further synthetic render scaling.","The real-data fraction is an operating point, not something to maximize: real-document accuracy rises while rendered-clean accuracy falls by roughly 10 to 14 points across bases.","Base-model choice must be revisited at the intended real-data fraction, since the ranking of the three bases changes as real data is added.","Vision-tower adaptation should be justified by a matched within-base comparison on the target task, because its measured effect ranges from about 23 to 35 points for one base to zero for another.","Synthetic-only accuracy is not a reliable predictor of real-document accuracy for OCSR."],"supporting_citations":[{"why":"Supplies the 7B vision-language model base whose synthetic-to-real gap is characterized and whose real-data dose series is run.","marker":"[Bai et al., 2025]"},{"why":"Supplies the LoRA adaptation method used for all fine-tuned recognizers in the study.","marker":"[Hu et al., 2022]"},{"why":"Supplies RDKit for parsing and canonicalizing predicted SMILES and for computing molecular identity in exact-match scoring.","marker":"[Landrum et al., 2024]"},{"why":"Supplies InChIKey, the molecular identity key used for exact-match scoring and evaluation-set deduplication.","marker":"[Heller et al., 2015]"},{"why":"Supplies MolScribe, a specialist baseline, and the ACS journal-figure benchmark used as the primary held-out real-document test.","marker":"[Qian et al., 2023]"},{"why":"Supplies DECIMER, the direct image-to-SMILES baseline, and the synthetic-rendering approach that defines the synthetic training condition.","marker":"[Rajan et al., 2020]"},{"why":"Supplies MolNexTR, a specialist graph-decoder baseline used in the common evaluation.","marker":"[Chen et al., 2024a]"},{"why":"Supplies OCSRGlyph and MarkushGlyph, the strongest patent-document baselines against which the real-data checkpoints are compared.","marker":"[Andonian et al., 2026]"},{"why":"Supplies the DECIMER-HDM hand-drawn image collection used as one source of labeled real depictions in the training pool.","marker":"[Rajan et al., 2023b]"}],"fun_headline_variants":["Real data, not synthetic, closes chemical structure recognition gap","Real images are the biggest lever for reading chemical drawings","Synthetic renders mislead; real data fixes chemical OCR","Real supervision outperforms synthetic in chemical structure reading","Real data shrinks the synthetic-to-real gap in OCSR"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that each training cell in the controlled 18-run sweep is represented by one seed, while the paper's uncertainty intervals measure evaluation-set sampling only; if seed-to-seed training variation is large, the base ranking and the dose-response shape could change.","fun_headline_variants_meta":{"raw":{"variants":["Real data, not synthetic, closes chemical structure recognition gap","Real images are the biggest lever for reading chemical drawings","Synthetic renders mislead; real data fixes chemical OCR","Real supervision outperforms synthetic in chemical structure reading","Real data shrinks the synthetic-to-real gap in OCSR"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000727,"raw_usage":{"total_tokens":3364,"prompt_tokens":1156,"completion_tokens":2208,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":772,"completion_tokens_details":{"reasoning_tokens":2128}},"tokens_in":772,"tokens_out":2208,"duration_ms":15645,"temperature":1.0,"reasoning_tokens":2128,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T23:31:02.525836+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the controlled sweep with multiple independent seeds per cell at 0%, 20%, and 70% real data using the same fixed evaluation sets. If between-seed spread at 70% equals or exceeds the observed 0.060 between-base spread, or if base orderings flip across seeds, the paper's ranking and contraction conclusions would not survive.","supporting_citations":[{"cited_title":"RDKit : Open-source cheminformatics","cited_arxiv_id":null,"evidence_quote":"Supplies RDKit for parsing and canonicalizing predicted SMILES and for computing molecular identity in exact-match scoring."},{"cited_title":"InChI , the IUPAC international chemical identifier","cited_arxiv_id":null,"evidence_quote":"Supplies InChIKey, the molecular identity key used for exact-match scoring and evaluation-set deduplication."},{"cited_title":"2023 , publisher=","cited_arxiv_id":null,"evidence_quote":"Supplies MolScribe, a specialist baseline, and the ACS journal-figure benchmark used as the primary held-out real-document test."},{"cited_title":"DECIMER : towards deep learning for chemical structure recognition","cited_arxiv_id":null,"evidence_quote":"Supplies DECIMER, the direct image-to-SMILES baseline, and the synthetic-rendering approach that defines the synthetic training condition."},{"cited_title":"MarkushGlyph and OCSRGlyph: Improved Chemical Structure Recognition","cited_arxiv_id":"2607.28532","evidence_quote":"Supplies OCSRGlyph and MarkushGlyph, the strongest patent-document baselines against which the real-data checkpoints are compared."}],"review_version":1}