{"id":"9c49b475-c17e-4c28-8201-136ea362474e","arxiv_id":"2608.00393","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"MolecularCanvas, an interactive system using structure-level annotations and constraints to steer an LLM's molecular edit plans, received higher user ratings and expert-assessed molecule quality than a baseline workflow in a 12-participant study.","lead":"A new interactive system, MolecularCanvas, lets chemists annotate specific parts of a molecule as keep, avoid, or modify, then uses that structure-level context to guide an LLM in generating candidate drug molecules. A user study with 12 participants reports that this integrated workflow improves self-rated and expert-rated outcomes compared to using separate general-purpose tools.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Paired user-study comparisons are confounded by task: Section 5.1.4 assigns the two systems to two different tasks, so questionnaire and expert-rating differences may reflect task difficulty rather than system quality.","rationale":"I read the study design in good faith. The reader correctly flags the uncontrolled baseline as a threat to internal validity, but the more load-bearing problem is the task-system pairing. Section 5.1.4 says participants completed the study tasks using two systems across two tasks, which in a 1.5-hour session with a warm-up and two 20-minute tasks almost certainly means each participant did one task per system. The two formal tasks are not matched: they start from different molecules (nalidixic acid vs. cefazolin) and have different optimization goals and constraints. Therefore every paired statistical test reported in Section 5.2.1, including the expert rating of final molecules, compares different tasks as well as different systems. The paper provides no per-task breakdown, so the reader cannot determine whether MolecularCanvas would still appear superior when task difficulty is held constant. This concern is concrete and directly undermines the strongest claim the paper makes. The qualitative themes and formative study still provide useful design insights, but the quantitative demonstration of 'usefulness and effectiveness' is currently unverified. A reanalysis with task as a covariate, or a re-run with matched tasks, could settle it; until then the central comparative claim should not be accepted as established. For that reason I would move the verdict from CONDITIONAL to UNVERDICTED, because the condition is not merely missing ablation or a better baseline, but the basic comparability of the two study conditions.","tokens_in":20236,"tokens_out":5935,"duration_ms":63845,"concrete_test":"Request or reconstruct the per-task, per-condition data and the participant-task-system assignment matrix. Fit a mixed-effects model with system, task, and their interaction as fixed effects and participant as a random effect. If the task effect is significant, or if the system effect reverses or attenuates within either task, the reported Wilcoxon p-values do not support the central claim. As a simpler check, recompute each paired comparison separately for the two subgroups (participants who used MolecularCanvas on Task A vs. those who used it on Task B); if the direction or magnitude of the difference differs across subgroups, the headline comparison is confounded. If the design was actually fully crossed (each participant used both systems on both tasks), the paper must state this explicitly and provide per-task paired analyses; with n=12, that design would also be underpowered but at","verdict_should_be":"UNVERDICTED","load_bearing_attack":"The central claim that MolecularCanvas improves molecular optimization outcomes rests on the user study in Section 5.2.1. The study design described in Section 5.1.4 has participants complete the two formal tasks (Task A: nalidixic acid to quinolone; Task B: cefazolin to cephalosporin) using the two systems, one task per system, in counterbalanced order. The two tasks have different starting molecules, different goals, and different constraint complexity. Thus every paired questionnaire comparison (Q1-Q13), the SUS comparison, and the expert rating of final molecules is a within-subject comparison across two different tasks, not across two systems on the same task. Counterbalancing the order of systems does not remove the task confound: if one task is intrinsically easier, more satisfying, or produces higher-quality outputs, the observed superiority of MolecularCanvas could be produced by task assignment alone. The paper does not report per-task means or include a task term in any analysis, so the reader cannot separate system effects from task effects. This is more fundamental than the reader's concern about the uncontrolled baseline toolset: even a perfectly matched alternative interface would not fix the task confound.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"MolecularCanvas is an interactive system for LLM-assisted small-molecule optimization. It lets users specify high-level goals, structure-level annotations (anchor/avoid/free-form), property targets, and reference-based preferences; these are compiled into a structured intent representation, used to generate an explicit RDKit-executable edit plan, and candidates are validated, ranked, and presented with supporting evidence from ChEMBL/PubChem. The paper reports a formative study with six experts that yields five design requirements, the system design, and a user study with 12 participants in which MolecularCanvas is claimed to outperform a baseline on all 13 perceived-effectiveness items, SUS (72.5 vs. 60.8), and expert-rated final molecule quality (0.73 vs. 0.58). The authors also report pipeline reliability logs (95% executable edit plans, 57% valid candidates after filtering, 91% constraint satisfaction).","tokens_in":20533,"tokens_out":5019,"duration_ms":59816,"significance":"If the effectiveness claim were well supported, the paper would make a useful contribution: it addresses a real gap in generative molecular design tools by making structure-level intent explicit, providing evidence provenance, and integrating evaluation tools into one environment. The structured intent formalism and the two-stage edit-plan generation are sensible and reusable ideas, and the pipeline reliability logs are a valuable empirical addition. The formative study is appropriate for deriving design requirements. The main weakness is that the user-study evidence, which is load-bearing for the central claim, is confounded by the assignment of different tasks to the two conditions and by an uncontrolled baseline; the current data do not cleanly isolate the system effect.","major_comments":[{"comment":"The central effectiveness claim rests on a comparison that confounds system with task. Each participant used one system on Task A (nalidixic acid → quinolone) and the other system on Task B (cefazolin → cephalosporin), with only the order counterbalanced. The two tasks differ in starting molecule, optimization goal, and constraint complexity. Thus each paired difference on Q1–Q13, SUS, and expert ratings is a difference across two different tasks, not a difference between systems on the same task. If one task is intrinsically easier or more satisfying, the observed superiority of MolecularCanvas could arise from task assignment even with no true system benefit. The paper does not report per-task means or include task as a factor in any analysis. To support the claim, the authors must provide per-task descriptive statistics and either (a) compare systems within each task using the appropr","section":"§5.1.4 / §5.2.1"},{"comment":"The baseline condition is not a fixed comparator: 'participants were allowed to use any tools they were familiar with in the baseline conditions.' These tools vary per participant (different conversational AI, property calculators, editors), and the paper reports no measure of prior familiarity or fluency with the baseline tools or with MolecularCanvas after the tutorial. A participant who is less skilled with their self-selected baseline may rate it lower for reasons unrelated to MolecularCanvas, and a novelty or tutorial effect could inflate the test condition. The authors should either use a matched alternative interface or, at minimum, measure and statistically control for tool familiarity and prior experience, and report the robustness of the results under such controls.","section":"§5.1.2"},{"comment":"The quantitative analysis reports p-values for 13 questionnaire items and the SUS without multiple-comparison correction and without effect sizes or confidence intervals. With n=12 and a large number of significance tests, the reported 'all p<0.05' pattern is difficult to interpret. The authors should report exact test statistics (e.g., Wilcoxon W or t-value), effect sizes (e.g., rank-biserial or Cliff's delta), and either apply a correction or explicitly justify why it is unnecessary. This is also needed to help readers evaluate the magnitude of the perceived differences, which appear substantial on several items but are not quantified beyond means and standard deviations.","section":"§5.2.1"}],"minor_comments":[{"comment":"Near-duplicate removal uses Tanimoto similarity >0.85 and evidence retrieval uses ≥0.8, but these thresholds are not justified or varied in sensitivity analysis. Since they are free parameters of the pipeline, a brief rationale or robustness check would strengthen the presentation.","section":"§4.3.2/§4.3.3"},{"comment":"The paper states 'Wilcoxon signed-rank test, p<0.001' for expert ratings and 'Wilcoxon p<0.05' for SUS, but the questionnaire items do not state which test was used. Please specify the test and whether it was paired or otherwise, and clarify how ties in Likert ratings were handled.","section":"§5.2.1"},{"comment":"Figure 3 shows paired distributions for Q1–Q13, but individual-level paired data would be more informative, especially for judging the consistency and direction of the differences. Consider adding a paired-by-participant visualization or data table in the supplementary.","section":"§5.1.4 / Figure 3"},{"comment":"The generation reliability numbers (95% executable plans, 57% valid candidates, 91% constraint satisfaction) are useful, but the 57% is the fraction passing validity and diversity filtering; the paper should state the base number of attempts and the number of candidates actually presented to users, so readers can gauge the practical user-facing yield.","section":"§5.2.1"},{"comment":"The limitations section is honest about missing ablations, but a sentence in Section 5.2.1 or Section 6 could note that the user study does not establish which components (e.g., anchor/avoid annotations, evidence panels, integrated property views) drive the observed differences. This would help set expectations for future work.","section":"§6.2"}],"recommendation":"major_revision","confidential_remarks":"The system and study are potentially valuable for UIST, but the evaluation confound is real. I would not reject outright because the per-task comparison may be recoverable from the existing data if the authors have recorded task assignment and can reanalyze appropriately, or they could add a short follow-up study. I would ask the authors to share the anonymized participant-level data and analysis scripts for the per-task analysis as part of the revision. The lack of a fixed baseline and the absence of effect sizes are secondary but should also be addressed."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The strongest thing about this paper is the system concept itself: structure-grounded anchor/avoid annotations on a molecular canvas, an executable edit-plan intermediate, evidence provenance from ChEMBL/PubChem, and integrated property evaluation are combined in a way I haven't seen before. The formative study with six medicinal chemists is well done, and the design requirements it produces are concrete and sensible. The pipeline logs (95% executable edit plans, 57% valid candidates, 91% constraint satisfaction) also suggest the back-end works as claimed. This is a real contribution to HCI for AI-assisted molecular design.\n\nThe soft spot is the evaluation, and it is load-bearing. Section 5.1.4 assigns the two systems to two different formal tasks (Task A: nalidixic acid to quinolone; Task B: cefazolin to cephalosporin), counterbalanced only for order. That means every paired comparison — all 13 questionnaire items, the SUS score, and the expert-rated final molecule quality — is a within-subject comparison across two different tasks, not across two systems on the same task. If one task is intrinsically easier or more satisfying, the observed superiority of MolecularCanvas could be entirely an artifact of task assignment. The paper reports no per-task means and includes no task term in any analysis, so the reader cannot separate system effects from task effects. This is more fundamental than the uncontrolled baseline-toolset issue the reader flagged; even a perfectly matched alternative interface would not fix the task confound.\n\nThe secondary concerns are real but minor by comparison: the baseline was “any tools participants were familiar with,” no correction was made for multiple comparisons across 13 questionnaire items, no component-level ablation was performed (the authors admit this in Section 6.2), and no code or data are released. None of those would sink the paper alone.\n\nIf the task confound were removed — say, with a within-task comparison or a proper factorial design — the core insight about structure-guided constraints would probably hold; the qualitative themes and system logs point in that direction. As reported, the headline effect sizes (0.73 vs. 0.58, all p<0.05) should be treated with real skepticism.\n\nThis paper deserves a serious referee. It is a well-written systems contribution with a plausible design rationale and a fixable but significant methodological flaw. I would not cite the quantitative results, but I would cite the interaction paradigm and the formative findings. It should go to peer review, with a strong request to scrutinize the study design and likely require substantial revision.","headline":"The system design is thoughtful and the qualitative findings are plausible, but the paired user study is confounded by task, so the headline quantitative claims don't stand as reported.","tokens_in":21026,"tokens_out":1575,"would_cite":true,"duration_ms":17725,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"MolecularCanvas claims that letting chemists annotate molecules directly on a canvas makes LLM-assisted drug optimization more effective, with expert-rated final molecules beating a familiar-tool baseline.","keywords":["large language models","molecular optimization","small-molecule drug discovery","structure-guided constraints","human-AI interaction","interactive systems","in-context prompting","candidate generation"],"falsifier":"A matched controlled experiment would settle it: same LLM, same property calculators, same evidence database, and only the input modality varied (structured canvas annotations vs. plain text prompts), with chemists randomly assigned and unaware of the hypothesis. If expert-rated final-molecule quality, structural-constraint satisfaction, and SUS scores are not higher in the structured condition, the central claim fails. A cheaper log-based check: track whether anchor/avoid constraints are ever violated by top-ranked candidates in the current system; if violations are rare even in text-only bas","tokens_in":20152,"feed_emoji":"🧪","tokens_out":9447,"duration_ms":90780,"temperature":0.7,"pith_summary":"MolecularCanvas is an interactive system for AI-assisted small-molecule optimization. The paper's claim is that letting chemists build a structured optimization context—high-level goals, on-molecule annotations (anchor, avoid, free-form edits), property target ranges, and reference-based preferences—makes LLM-generated candidates align better with expert intent than single free-form prompts. It also claims that attaching evidence from public drug databases to each candidate and integrating property calculators into one interface builds trust and removes the need to switch between separate tools. A 12-participant user study is offered as evidence: MolecularCanvas was rated significantly higher than the familiar-tool baseline on all 13 questionnaire items, scored higher on the System Usability Scale, and its final molecules were rated higher by two independent medicinal chemists (normalized 0.73 vs 0.58, p<0.001). The paper further argues this points to a broader design principle: in structure-intensive domains, the domain object itself should be the medium for instructing generative AI.","feed_headline":"Structure-guided annotations steer AI drug design better","feed_subtitle":"A canvas interface with anchor, avoid, and property targets scored higher with chemists than their usual tools.","key_machinery":"The load-bearing mechanism is the structured optimization context plus a two-stage generation pipeline. The context is a formal specification that captures the current molecule in SMILES, a text goal, property constraints with target ranges, and atom-resolved structural annotations (anchor/avoid/free-form) taken directly from the Ketcher molecular editor, without vision-language parsing. The pipeline first asks an LLM to propose an edit plan as discrete operations, then runs those operations through RDKit to generate concrete molecules, checks chemical validity, removes duplicates and near-duplicates by Morgan fingerprint similarity (Tanimoto >0.85), scores candidates on property satisfactio","core_discovery":"At the center of the paper is a proposed fix for a workflow mismatch: chemists think in molecular structures, but generative AI tools mostly accept language. MolecularCanvas lets users select atoms and bonds on a 2D canvas and tag them as anchor (preserve), avoid (do not touch), or free-form instructions, combine this with desired property ranges and reference molecules, and feed the whole structured package to an LLM. The LLM is asked to produce an explicit edit plan of executable operations—replace atom X with fluorine, add a polar group at position Y—rather than a finished molecule, and RDKit executes and validates the plan. The paper's claim is that this structured, evidence-backed, inte","pith_inferences":["The same structure-as-interface pattern should transfer to other structure-intensive domains—protein engineering, materials discovery, reaction planning—where users can point at the object to constrain generation; the authors gesture at this but do not test it.","The reported gains may partly reflect novelty and tutorial warm-up rather than the interface alone; a component-level ablation would show whether anchor/avoid annotations, property sliders, or evidence cards each contribute independently.","The History Panel plus logged annotations is effectively an auto-generated design notebook; if it records decisions and evidence, it could make molecular optimization audits and reproducibility checks far cheaper than manual note-taking.","The edit-plan architecture is a reusable safety layer for LLM chemistry tools: restricting the model to executable operations and validating them deterministically bounds hallucination in a way that free-form molecule generation does not."],"forward_implications":["Chemists can specify precise modifications—keep this ring, replace this chloride with fluorine—by pointing at the molecule, removing the need to translate structural intent into prompt prose.","Candidates arrive with provenance (known compounds, safety signals, clinical phase, patents), so users can judge and justify AI suggestions in collaborative settings.","Branching history lets users revert, compare, and pursue multiple design directions, turning optimization from a linear sequence into an explorable graph.","Because the LLM only proposes edit plans and RDKit executes them, invalid or duplicate molecules are filtered before display; logs suggest roughly one valid, non-duplicate, constraint-satisfying candidate per two generation attempts.","Integrating property calculators and evidence lookup into one interface removes the repeated switching between separate generation, property, and visualization tools that participants described."],"supporting_citations":[{"why":"Supplies Ketcher, the open-source molecular editor embedded in the canvas that users annotate with anchor, avoid, and free-form constraints.","marker":"[13]"},{"why":"Supplies RDKit, the deterministic toolkit that executes edit plans, validates molecules, computes properties, and calculates fingerprint similarity.","marker":"[26]"},{"why":"SwissADME is an external ADMET platform cited as part of the fragmented multi-tool workflow that MolecularCanvas consolidates.","marker":"[11]"},{"why":"ADMETlab 2.0 is another external property-prediction platform whose separate use motivates the integrated evaluation panel.","marker":"[55]"},{"why":"Provides the System Usability Scale used to quantify usability in the user study.","marker":"[6]"},{"why":"Provides the intraclass correlation coefficient used to measure agreement between the two expert raters of final molecules.","marker":"[45]"}],"fun_headline_variants":["Chemists annotate molecules, AI executes edits","Structure tags guide LLM to rewrite molecules","Canvas tool turns structure intents into AI actions","LLM plans edits, RDKit verifies with anchors","MolecularCanvas: anchor-avoid tags steer AI design"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The load-bearing premise is that the baseline—each participant using whatever familiar tools they chose—is a fair comparator, so the measured gains reflect MolecularCanvas's design rather than unequal familiarity, tutorial exposure, or a novelty effect.","fun_headline_variants_meta":{"raw":{"variants":["Chemists annotate molecules, AI executes edits","Structure tags guide LLM to rewrite molecules","Canvas tool turns structure intents into AI actions","LLM plans edits, RDKit verifies with anchors","MolecularCanvas: anchor-avoid tags steer AI design"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000149,"raw_usage":{"total_tokens":1017,"prompt_tokens":721,"completion_tokens":296,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":465,"completion_tokens_details":{"reasoning_tokens":223}},"tokens_in":465,"tokens_out":296,"duration_ms":3837,"temperature":1.0,"reasoning_tokens":223,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T04:12:13.184755+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A matched controlled experiment would settle it: same LLM, same property calculators, same evidence database, and only the input modality varied (structured canvas annotations vs. plain text prompts), with chemists randomly assigned and unaware of the hypothesis. If expert-rated final-molecule quality, structural-constraint satisfaction, and SUS scores are not higher in the structured condition, the central claim fails. A cheaper log-based check: track whether anchor/avoid constraints are ever violated by top-ranked candidates in the current system; if violations are rare even in text-only bas","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies Ketcher, the open-source molecular editor embedded in the canvas that users annotate with anchor, avoid, and free-form constraints."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies RDKit, the deterministic toolkit that executes edit plans, validates molecules, computes properties, and calculates fingerprint similarity."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"SwissADME is an external ADMET platform cited as part of the fragmented multi-tool workflow that MolecularCanvas consolidates."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"ADMETlab 2.0 is another external property-prediction platform whose separate use motivates the integrated evaluation panel."},{"cited_title":"Quick and Dirty","cited_arxiv_id":null,"evidence_quote":"Provides the System Usability Scale used to quantify usability in the user study."},{"cited_title":"Shrout and Joseph L","cited_arxiv_id":null,"evidence_quote":"Provides the intraclass correlation coefficient used to measure agreement between the two expert raters of final molecules."}],"review_version":1}