{"id":"47446e9e-fbd0-46a0-8de3-d9b75dfb8a35","arxiv_id":"2608.08430","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"CELLens is a human-in-the-loop visualization system that helps domain experts correct auto-mined causal graphs for virtual cells, shown on rice single-cell data.","lead":"The authors present CELLens, an interactive visual analytics tool that lets biologists explore, validate, and refine the causal graphs used to guide virtual cell models. It combines a gene-similarity-aware causal layout with counterfactual visualizations, and the paper evaluates it through a rice gene-expression case study and expert feedback.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The refinement loop validates against the same counterfactual model being refined, with no independent ground-truth regulatory check, so the effectiveness claim currently demonstrates workflow feasibility rather than causal correctness.","rationale":"The reader's weakest assumption and my concern are the same at the core: the counterfactual outputs of the virtual cell are used as evidence for refining the causal graph, and there is no independent ground-truth validation. My stress-test pass makes the circularity more explicit: the expert's decisions in Sec. 6.1.1 rely on expression changes and causal paths generated by CausCell after it was trained on the auto-mined graph, so an erroneous structure can reinforce itself. The paper does provide real evidence: a real rice dataset, a detailed case study, a reproducible open-source codebase, and a small but genuine literature match for one gene pair. I am not arguing that the tool is useless or that the authors are dishonest; the workflow may well be effective in practice. However, the headline claim as stated in the Abstract and Conclusion asserts that the method's effectiveness is demonstrated by case studies and expert feedback. Those demonstrations establish usability and plausibility, not causal correctness. Because the missing external validation is directly load-bearing for the central claim, the paper should be conditionally accepted rather than fully accepted. The conditional path is clear: add an independent evaluation against curated regulatory knowledge or held-out perturbation data, or soften the central claim to workflow feasibility. Since the reader already assigned CONDITIONAL, my recommendation is UNCHANGED: retain the conditional verdict and require the external validation as a condition for full acceptance.","tokens_in":20613,"tokens_out":3541,"duration_ms":45391,"concrete_test":"Re-run the rice case study edits against a curated rice regulatory network, such as PlantRegMap/PRR or RiceFREND: map member genes of the 15 concepts to curated TF-target interactions, build a ground-truth edge set over concepts, then compute precision and recall for the added edges (photosynthesis to biosynthesis-and-stimulus, chromatin to cellular metabolism) and for the moved clusters' new concept assignments relative to the initial auto-mined graph. If the refined graph does not show higher agreement with the curated network than the initial graph, or if the moved clusters are not significantly enriched for curated interactions with their new concepts, the refinement loop is not validated.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central effectiveness claim depends on expert edits being guided by counterfactual responses generated by the same CausCell virtual cell that is being refined (Sec. 5.2, Sec. 6.1.1). The expert decides that a link is missing or a gene grouping is wrong based on expression changes produced by a model that was trained to conform to the auto-mined causal graph. This is a closed loop: the model being refined is also the source of evidence for refining it. If the counterfactual outputs are biased, the expert edits can move the graph away from biological truth while still appearing internally consistent. The only statistical check in Sec. 6.1.1 compares counterfactual mean expression of \"biosynthesis and stimulus\" across control, initial, and refined models; a higher mean after intervention confirms the model learned the injected positive link, not that the link is biologically correct. The literature verification in Sec. 6.1.2 is a single gene-level precedent (Os07g0605200-Os10g0463800), not a network-level validation of the added concept links or the reassigned gene-concept memberships. The paper's own Limitations section (Sec. 7.2) concedes the case study used one expert under pair-analytics and a small number of participants, but the deeper gap is external: there is no curated regulatory network, held-out perturbation dataset, or independent expert ground truth against which the refined graph or its counterfactual predictions are scored. The result is that the paper supports the claim that experts can operate the tool and produce plausible narratives, but not the stronger claim that the refined causal graphs are more biologically reliable.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript presents CELLens, a visual analytics system for human-guided refinement of causal graphs used in causal-driven virtual cells. The system combines a gene-similarity-aware causal graph layout, which encodes concept-level causal links as regions and gene-level similarity as points inside those regions, with a causal path view and a counterfactual clustering view. A domain expert can intervene on a concept, inspect counterfactual expression changes, and then edit concept-concept links or gene-concept assignments before retraining the virtual cell. The evaluation is a rice single-cell case study conducted with one co-developing expert, plus qualitative feedback from five experts, and the paper claims that the tool supports exploration, validation, and refinement of causal graphs and yields scientifically meaningful insights.","tokens_in":20949,"tokens_out":9612,"duration_ms":103235,"significance":"If the effectiveness claims held, CELLens would address a genuine gap: existing causal graph visualization tools focus on concept-level structure, while causal-driven virtual cells also require inspection and correction of gene-concept correspondences and gene-gene similarities. The paper has real strengths: the layout optimization is described in enough detail to be reimplemented, the visual encodings are well motivated, the design was grounded in a 12-month collaboration with domain experts, and the source code is released. The proposed visualizations are also novel in integrating counterfactual analysis with interactive graph editing. However, the current evidence supports only workflow-level feasibility: the evaluation rests on a single case study with an expert who co-designed the tool, the statistical check is a model-consistency test rather than a biological validation, and there is no independent ground-truth comparison. The significance is therefore moderate and depends on the authors either adding external validation or explicitly tempering the effectiveness claims.","major_comments":[{"comment":"The refinement loop validates the causal graph using counterfactual responses generated by the same CausCell virtual cell that is being refined. In the rice case study, B1 decides that a link is missing or a concept is mislabeled after inspecting expression changes produced by the initial model, and the only statistical check then compares control, initial, and refined models on the same counterfactual pathway. The Friedman test (p<0.001) shows that the refined model gives a higher mean expression for \"biosynthesis and stimulus\" after intervention; this is a consistency check showing that the model has learned the injected positive link. It does not establish that the added link or the reassigned gene-concept memberships are biologically correct. The manuscript does not compare the refined graph with a curated regulatory network, held-out perturbation data, or independent expert ground truth, and Sec. 7.2 does not list this self-referential validation as a limitation. I recommend adding an external validation of the refined graph or, at minimum, explicitly reframing the effectiveness claim as demonstrating workflow-level feasibility rather than biological correctness.","section":"Sec. 5.2 and Sec. 6.1.1"},{"comment":"The evaluation is based on a single case study conducted with B1, who participated in the 12-month requirement-analysis and prototype-development process described in Sec. 4, using a pair-analytics protocol in which the authors operated the tool. The abstract's claim of \"two real-world case studies\" is inaccurate: the two subsections of Sec. 6.1 are two analytical objectives from the same rice session. The follow-up interviews in Sec. 7.1 include two independent experts (B4-B5), but they received only a 30-minute introduction and their feedback is qualitative. This evidence supports a design-study demonstration, not the stronger effectiveness claim made in the abstract. I ask the authors to add an independent case study or a controlled user study with experts not involved in the development, and to correct the \"two case studies\" wording.","section":"Sec. 6.1 and Sec. 7"},{"comment":"The counterfactual-aware clustering concatenates a multi-hot GO-term vector with a single scalar expression change and then L2-normalizes the combined vector. Because the semantic part is a high-dimensional binary vector while the response part is a single continuous coordinate, the relative influence of the counterfactual response on the K-means objective depends on the number of GO terms per gene and is not controlled by any weighting. As written, the method does not guarantee that the resulting clusters are actually \"counterfactual-aware.\" The paper should either justify the weighting, add an ablation showing that the response channel changes the clustering outcome, or revise the claim about this visualization.","section":"Sec. 5.3, Eq. (6)"}],"minor_comments":[{"comment":"The sentence \"experts can they can filter out implausible causal paths\" contains a duplicated phrase; please revise to \"experts can filter out implausible causal paths.\"","section":"Sec. 5.2"},{"comment":"The parameters λ1, λ2, λ3, δx, and ε are said to be determined by grid search, but the main text does not report the chosen values or the search ranges; please provide these values and a sensitivity analysis in the appendix.","section":"Sec. 5.1.1"},{"comment":"The Friedman test indicates a significant difference across the three conditions, but no post-hoc pairwise tests are reported; please identify which pairwise differences are significant and report effect sizes, especially because the paper singles out the refined model as the best.","section":"Sec. 6.1.1"},{"comment":"The aggregation of concept-level expression change using the top 5% of member genes ranked by absolute change is an arbitrary threshold; please justify it empirically or report sensitivity to this parameter.","section":"Sec. 5.1"},{"comment":"The related work claims that existing tools such as CausalVis and CausalMap do not preserve gene-level information, but the evaluation does not compare CELLens against these tools; adding even a qualitative comparison would strengthen the novelty argument.","section":"Sec. 2.2"}],"recommendation":"major_revision","confidential_remarks":"To the editor: the paper is a competent design study with a clearly described system, but the evaluation is not strong enough for the claims as stated. I would accept a revised version that either adds external validation (e.g., against a curated regulatory network or perturbation data) or substantially tempers the effectiveness claims and corrects the 'two case studies' wording. The circularity issue is the main risk and should be addressed head-on in the revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"CELLens is a genuinely new visual analytics tool for refining auto-mined causal graphs in virtual cells. The layout algorithm—t-SNE-style gene similarity, left-to-right causal ordering, and Voronoi margin injection—is not present in CausalVis, CausalMap, or D-BIAS, and the two-step global-to-local optimization is a sensible way to avoid nested convex-hull computations. The counterfactual clustering view, which groups genes by both GO semantics and counterfactual response and lets experts drag clusters between concepts, is a useful interaction. The integration with CausCell is concrete, and the source code is available.\n\nThe paper's real weakness is the evaluation. The case study is one expert (B1) who co-designed the tool, using pair analytics. The abstract says 'two real-world case studies,' but the second phase is just a continuation of the same rice analysis. The expert feedback from B4–B5, who received a 30-minute introduction, is a positive-opinion poll, not evidence of effectiveness.\n\nMore importantly, the refinement loop is self-referential. Experts decide whether a causal link is missing or wrong based on counterfactual responses from the same CausCell model they are refining. If those counterfactuals are biased by the auto-mined graph, edits can reinforce errors while looking internally consistent. The Friedman test in Sec. 6.1.1 only shows the retrained model learned the injected positive relationship; it does not validate the relationship biologically. The literature check is one gene-pair precedent, not a network-level validation.\n\nThese are addressable. The authors could compare the refined graph against a curated regulatory network or held-out perturbation data, or reframe the claim as demonstrating workflow feasibility. The limitations section mentions the small participant count but does not address the circularity, which should be flagged.\n\nWho is this for? Visualization researchers working on causal interaction, and computational biologists building virtual cells. It is a solid design study with a novel layout technique, but the effectiveness claim needs toning down or external validation. I'd accept it for peer review under major revision.","headline":"Novel visualization method for refining causal graphs in virtual cells; evaluation shows feasibility, not causal correctness, yet merits peer review with major revisions.","tokens_in":21495,"tokens_out":4551,"would_cite":true,"duration_ms":45511,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A human-in-the-loop visual workflow lets biologists correct the causal graphs inside virtual cells and retrain models that respond as biology expects.","keywords":["virtual cell","causal graph","visual analytics","human-in-the-loop","counterfactual analysis","gene expression","single-cell data","causal knowledge injection"],"falsifier":"Apply the CELLens workflow to an organism or cell type whose regulatory network is already established independently, compare the auto-mined and expert-refined graphs against that network, and measure agreement in edge direction and gene membership. If the refined graph is not closer to the known network than the auto-mined graph—or if the newly added links fail targeted perturbation tests—the central claim that human-guided injection improves biological plausibility is falsified.","tokens_in":20438,"feed_emoji":"🧬","tokens_out":8074,"duration_ms":80014,"temperature":0.7,"pith_summary":"Virtual cells—machine-learning models that predict gene expression changes under new cell conditions—become more interpretable when guided by a causal graph, but such graphs are rarely available. Auto-mined graphs group genes into concepts and infer causal links, yet the unsupervised process produces mixed concepts, wrong links, and missing links. This paper claims that a visual analysis tool, CELLens, lets domain experts explore, validate, and refine those graphs by showing gene similarities alongside causal edges, generating counterfactual responses to concept interventions, and supporting direct drag-and-drop edits. In the rice single-cell case study, the expert disentangled mixed concepts, added two missing causal links, and retrained the virtual cell so that a photosynthesis intervention produced the expected positive effect on a downstream concept; the refined model also pointed to candidate regulatory genes, one backed by published evidence. If this is right, expert biological judgment is a practical complement to automatic causal discovery for virtual cells, not a replacement for it.","feed_headline":"Biologists can hand-correct the causal graphs inside virtual cells","feed_subtitle":"Counterfactual checks and drag-and-drop editing turn expert knowledge into better cell simulations.","key_machinery":"The load-bearing mechanism is the gene-similarity-aware causal graph layout, computed by a two-stage hybrid optimization. In the global stage, a t-SNE-style KL-divergence term preserves gene-gene similarities and gene-concept membership while geometric penalties enforce causal direction left-to-right, keep linked concepts close, and separate concepts sharing a target. In the local stage, a center-attractive force pulls boundary genes back into their own concept when convex hulls overlap. The resulting positions are turned into polygonal regions with a GMap Voronoi construction modified by a margin-injection trick: invisible virtual points along shared borders of non-causally-linked regions create spatial gaps, and adding or removing those points gives real-time causal link/separate editing. On the validation side, the second main mechanism is the counterfactual analysis strategy: the virtual cell, implemented as a diffusion-based structural causal model, generates predicted expression changes under intervention, the causal path view extracts and ranks shortest paths by saliency (response magnitude divided by path length), and the counterfactual clustering view groups genes by combining semantic GO features with predicted response so the expert can see whether response and semantics agree.","core_discovery":"The paper's central claim is that the errors in auto-mined causal graphs for virtual cells are correctable by an expert-in-the-loop process, and that the right visual interface makes that process efficient. The system it builds, CELLens, treats a causal graph as a multi-level object: concept nodes carry causal links, each concept contains member genes, and those genes carry pairwise similarities. Its gene-similarity-aware layout solves a joint optimization that preserves gene-gene similarities, keeps genes inside their concept regions, and arranges causally linked concepts left-to-right with non-adjacent concepts separated by Voronoi margins. Validation is driven by counterfactual generation: intervening on a concept produces predicted expression changes, which are summarized along candidate causal paths ranked by a saliency score, and inside concepts genes are clustered by both semantic annotation and predicted response so that an expert can see which clusters respond anomalously. Refinement is direct: split a cluster into a new concept, merge it into another, add a link, remove a link, rename a concept, then retrain the virtual cell. The paper demonstrates the cycle on rice data, where the expert's edits produced a retrained model whose intervention response matched the hypothesized regulation with a statistically significant difference (p < 0.001), and where the same views surfaced a candidate upstream regulator with literature support.","pith_inferences":["[Editorial inference] The most decisive evaluation would apply the same workflow to a system whose true regulatory network is known, then measure whether expert edits increase agreement with that gold standard; the current case study relies on literature support and plausibility rather than exhaustive ground truth.","[Editorial inference] Because the counterfactual predictions used to guide edits come from the very model being refined, any systematic bias in those predictions could steer expert corrections in the wrong direction; a model-agnostic validation against held-out perturbation experiments would test this.","[Editorial inference] The saliency score's inverse-length penalty assumes causal influence decays with each step; if that assumption is wrong, path rankings could mislead, and one testable extension is to learn attenuation weights from data.","[Editorial inference] Annotation gaps are a hidden variable in the workflow: genes labeled N/A under GO cannot contribute semantic signal, so clusters may be split or merged based on missing annotations rather than true function; incorporating independent interaction databases could sharpen the clustering."],"forward_implications":["A refined causal graph can be retrained into a virtual cell whose intervention behavior matches the expert's hypothesized mechanism, as the rice experiment showed for photosynthesis positively regulating a downstream concept.","The same concept-level and gene-level views let experts nominate candidate regulatory genes, including one with independent literature support and one the expert flagged as potentially novel.","The low-level similarity view is claimed to make semantic purity of concepts visible, so experts can spot mixed concepts before running expensive counterfactual simulations.","The paper expects the workflow to transfer to other multi-level causal analysis tasks where high-level causal relationships, low-level similarities, and their correspondences all matter."],"supporting_citations":[{"why":"Supplies the diffusion-based virtual cell model whose structural causal model generates the counterfactual expression changes used for validation.","marker":"[19]"},{"why":"Groups genes into concepts via weighted correlation network analysis, creating the high-level nodes of the causal graph.","marker":"[27]"},{"why":"Provides the PC algorithm used to infer concept-level causal links from the derived concept expression matrix.","marker":"[47]"},{"why":"Contributes the t-SNE-style similarity formulation that the layout uses to preserve gene-gene and gene-concept relationships in 2D.","marker":"[52]"},{"why":"Supplies the GMap Voronoi-based map construction that turns gene positions into contiguous, editable concept regions.","marker":"[18]"},{"why":"Provides the rice single-cell dataset used in the case study to demonstrate the refinement workflow and its biological outcomes.","marker":"[55]"},{"why":"Reports the literature evidence connecting Os07g0605200 to Os10g0463800 that supports the expert-identified regulatory link.","marker":"[56]"}],"fun_headline_variants":["Human-guided editing corrects causal graphs in virtual cells","Experts hand-fix causal graphs to improve virtual cell models","Counterfactual validation refines causal graphs for virtual cells","Interactive causal graph correction boosts virtual cell fidelity","Human expertise injected into virtual cells via causal graph edits"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The refinement loop assumes that the virtual cell's simulated responses to concept interventions are reliable enough to tell a correct causal link from a wrong one; if those simulations are biased, expert edits could move the causal graph further from biological truth.","fun_headline_variants_meta":{"raw":{"variants":["Human-guided editing corrects causal graphs in virtual cells","Experts hand-fix causal graphs to improve virtual cell models","Counterfactual validation refines causal graphs for virtual cells","Interactive causal graph correction boosts virtual cell fidelity","Human expertise injected into virtual cells via causal graph edits"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000252,"raw_usage":{"total_tokens":1578,"prompt_tokens":983,"completion_tokens":595,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":599,"completion_tokens_details":{"reasoning_tokens":519}},"tokens_in":599,"tokens_out":595,"duration_ms":6073,"temperature":1.0,"reasoning_tokens":519,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T04:35:17.054955+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Apply the CELLens workflow to an organism or cell type whose regulatory network is already established independently, compare the auto-mined and expert-refined graphs against that network, and measure agreement in edge direction and gene membership. If the refined graph is not closer to the known network than the auto-mined graph—or if the newly added links fail targeted perturbation tests—the central claim that human-guided injection improves biological plausibility is falsified.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the diffusion-based virtual cell model whose structural causal model generates the counterfactual expression changes used for validation."},{"cited_title":"Langfelder and S","cited_arxiv_id":null,"evidence_quote":"Groups genes into concepts via weighted correlation network analysis, creating the high-level nodes of the causal graph."},{"cited_title":"Sakai, S","cited_arxiv_id":null,"evidence_quote":"Provides the PC algorithm used to infer concept-level causal links from the derived concept expression matrix."},{"cited_title":"Van der Maaten and G","cited_arxiv_id":null,"evidence_quote":"Contributes the t-SNE-style similarity formulation that the layout uses to preserve gene-gene and gene-concept relationships in 2D."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the GMap Voronoi-based map construction that turns gene positions into contiguous, editable concept regions."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the rice single-cell dataset used in the case study to demonstrate the refinement workflow and its biological outcomes."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Reports the literature evidence connecting Os07g0605200 to Os10g0463800 that supports the expert-identified regulatory link."}],"review_version":1}