{"id":"6ea9f226-42bc-4ffb-9f40-12c1b5d643a0","arxiv_id":"2608.05900","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"Small, targeted removal of companion cells can change a single cell's refined annotation without touching the cell itself, revealing query cohort composition as an attack surface in single-cell annotation.","lead":"This paper shows that removing a small number of non-target cells from a single-cell RNA sequencing query cohort can flip the final cell-type label of an unchanged target cell, while leaving its expression profile and base prediction intact. The result matters because it exposes a target-preserving attack surface and a reproducibility risk in cohort-dependent annotation tools.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Headline flip rates omit the audit-set qualifier; the practical frequency of companion-removal flips across a typical query cohort is not established by the reported percentages.","rationale":"The reader's weakest assumption correctly identifies the gap between the audit-set flip rates and the abstract's unqualified headline percentages. My read of the full text confirms the qualifier is present in Section III.A but absent from the abstract and contributions, so the reader's CONDITIONAL verdict is appropriate. I do not see a deeper soundness flaw: the controlled experiments preserve target features and base predictions, the lambda = 0 ablation isolates the refinement stage, and the CellTypist evaluation independently shows that cohort-dependent majority voting can change even when independent predictions are frozen. The main unresolved issue is population-level frequency, which the paper explicitly does not claim to estimate. Therefore the verdict should remain CONDITIONAL, requiring either a base-rate estimate or a consistently qualified abstract before acceptance.","tokens_in":8943,"tokens_out":4633,"duration_ms":48052,"concrete_test":"Run the Search V2 multi-start procedure on an independently drawn uniform random sample of 100 correctly annotated test cells per dataset, classifier, and seed, using the same 5% budget and stopping rules, and compare the achieved flip rate with the reported 24.33% / 19.67% and with the random-removal baseline, reporting per-seed rates and bootstrap confidence intervals. If the flip rate on the unselected sample is close to the random-removal rate, the headline percentages are artifacts of the lower-confidence target enrichment; if it remains at a comparable level, the concern does not land.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central quantitative evidence is the abstract's claim that multi-start search changed 24.33% of linear-SVM targets and 19.67% of logistic-regression targets while removing a small fraction of the cohort. However, Section III.A explicitly states that these rates were measured on 100 lower-confidence, correctly annotated targets per dataset, classifier, and seed, and that they 'should be interpreted as vulnerability rates within this targeted audit set, rather than as estimates over all cells in the dataset.' The abstract and introduction do not carry this qualifier, so a reader can reasonably infer that the rates describe typical query cells. The paper's own CellTypist results show that initially stable cells flip far less often than context-sensitive cells, so the selection policy strongly enriches for vulnerability. The existence of a target-preserving attack is supported by the controlled pipeline and by the CellTypist majority-voting result, and the lambda = 0 ablation confirms the mechanism; the load-bearing uncertainty is therefore not whether such flips can occur, but how often they occur in ordinary query cohorts. Without a base-rate estimate, the headline percentages cannot be generalized, and the practical significance of 'query cohort composition as a target-preserving attack surface' remains unquantified.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper proposes CohortHijack, a robustness audit for single-cell annotation pipelines that include a cohort-dependent refinement stage. The method removes a small set of non-target 'companion' cells from the query cohort while keeping the target cell's expression profile, base classifier, and trained parameters unchanged, and it measures whether the refined label changes. The authors evaluate random, same-class, and nearest-cell removal, plus greedy, multi-start greedy, and beam search, on PBMC3K and Paul15 with logistic regression and calibrated linear SVM. They supplement the controlled pipeline with a CellTypist majority-voting validation. They report that structured removal is stronger than random removal on Paul15, that multi-start search changes 24.33% of linear-SVM targets and 19.67% of logistic-regression targets with low collateral damage, that the effect disappears when the context weight is zero, and that CellTypist refined labels are similarly unstable while its independent labels remain unchanged.","tokens_in":9064,"tokens_out":6021,"duration_ms":53180,"significance":"If the result is taken at face value, the paper identifies a genuinely new target-preserving attack surface for cohort-dependent single-cell annotation and offers a reusable audit procedure. The controlled experimental design is sound, the lambda = 0 ablation is a clean mechanistic check, and the CellTypist validation provides independent evidence that the phenomenon is not an artifact of the authors' own pipeline. The main weakness is that the headline flip rates are measured on a deliberately constructed audit set of lower-confidence, correctly annotated targets, and the paper does not quantify the base rate of such cells; the practical frequency of flips in an ordinary query cohort is therefore not established. As stated in the paper, the rates are vulnerability rates within the targeted audit set, not estimates over all cells. The contribution is best characterized as demonstrating that cohort-removal flips exist and can be induced, with the practical significance depending on how common context-sensitive cells are in routine cohorts.","major_comments":[{"comment":"The headline claim in the Abstract—'Multi-start search changed 24.33% of linear-SVM targets and 19.67% of logistic-regression targets'—omits the qualifier stated in Section III.A: these rates were measured on 100 correctly annotated, lower-confidence targets per dataset, classifier, and seed, and 'should be interpreted as vulnerability rates within this targeted audit set, rather than as estimates over all cells in the dataset.' Because the target-selection policy deliberately enriches for context-sensitive cells, the abstract's percentages cannot be read as population-level flip rates and should carry the audit-set qualifier. The paper should also discuss how the base rate of lower-confidence, context-sensitive cells affects the practical frequency of flips in a typical query cohort.","section":"Abstract and Section III.A"},{"comment":"The search-based success rates are reported as single point estimates aggregated over seeds 13, 37, and 73, with no per-seed breakdown, confidence interval, or p-value. Because the 100 targets per dataset, classifier, and seed are sampled without replacement, the 24.33% and 19.67% figures could vary across target samples and seeds; the paper should report per-seed results or otherwise demonstrate that the aggregated rates are stable.","section":"Section III.C and Table I"},{"comment":"There is a numerical inconsistency in the CellTypist validation. Section III.A describes 10 context-sensitive and 10 initially stable targets per seed over 3 seeds, i.e., 60 targets total; with 3 removal methods and 2 budgets, this gives 360 evaluations for each outcome, not the '840 evaluations' claimed in Section III.E. In addition, the 'Context-sensitive 2%' row in Table II reports 27.33%, which would correspond to 8.2 out of 30 targets and is not an integer count. These inconsistencies should be corrected because they directly affect the credibility of the validation statistics.","section":"Section III.A and Section III.E"},{"comment":"The search methods optimize the refined-probability margin m_t(S) defined in Section II.F, and the attack-success indicator in Section II.G is exactly a negative margin. This is a legitimate audit procedure, but it means the multi-start success rates are an upper-bound-style worst-case measure under a matched objective, not an estimate of naturally occurring flips. The paper should state this interpretation explicitly and should not present the search-based rates on the same footing as the structured-removal rates, which do not optimize the success criterion.","section":"Section II.F, Section II.G, and Section III.C"}],"minor_comments":[{"comment":"The statement that 'paired tests confirmed' the Paul15 differences is not verifiable without naming the test (e.g., paired t-test, Wilcoxon signed-rank) and reporting the test statistics or p-values; please add these details.","section":"Section III.B"},{"comment":"The phrase 'The same target manifests are reused' appears to be a typo for 'target sets' or 'target lists'; the intended meaning is clear but the wording should be corrected.","section":"Section II.D"},{"comment":"Please state explicitly whether each structured-removal method was run once per seed or repeated; the ten repeats are specified for random removal only, and this affects how the reported flip rates should be interpreted.","section":"Section III.A"},{"comment":"The percentages in Table II are based on only 30 targets per group across three seeds, so each 3.33% increment corresponds to one target; this should be stated in the caption to prevent over-interpretation of small differences between conditions.","section":"Table II"}],"recommendation":"major_revision","confidential_remarks":"I recommend major revision. The central result appears defensible and the ablation and CellTypist validation provide good supporting evidence. The main corrections are to align the abstract and introduction with the audit-set qualifier, repair the CellTypist arithmetic, and add per-seed or uncertainty information for the search rates. I do not see a basis for rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper is a clean empirical audit of a threat that hasn't been studied before: removing non-target companion cells from a query cohort to flip a single cell's refined annotation without touching the cell itself, its base prediction, or the trained model. That's genuinely new, and the authors did the right experiments to support it. The mechanism is confirmed by the lambda = 0 ablation (no flips when neighborhood refinement is off), by the CellTypist validation where independent predictions never change, and by measuring collateral damage. The search methods are thorough, and the comparison to random removal is fair. The paper also honestly states the audit-set qualifier in Section III.A, which is more than many papers do.\n\nWhere it gets soft is in the packaging. The abstract quotes 24.33% and 19.67% without the qualifier that these are rates on 100 lower-confidence, correctly annotated targets per dataset/classifier/seed. The reader can reasonably infer these are typical cell rates, but the paper's own CellTypist results show context-sensitive cells flip 5-15x more often than initially stable cells. So the practical frequency of flips in a routine query cohort is unknown. That's a presentation and generalizability issue, not a fatal one, but it matters for how the paper will be cited.\n\nThe stats are also thin in places: no per-seed variability or p-values for the search results, and the search-based flip rates are partly by construction since the search optimizes the same refined-probability margin used to define success. The CellTypist result is independent of that and helps, but it's on only 60 targets (10 per group per seed).\n\nWho is this for? People building single-cell annotation tools and anyone thinking about reproducibility of cohort-dependent pipelines. The paper would benefit from an abstract that says 'targeted audit set' instead of implying all cells, per-seed numbers, and a code commit hash so others can reproduce the exact search. The authors already provide a GitHub link, which is good.\n\nOverall verdict: the existence result is solid, the mechanism is clear, and the paper earns a serious referee. Send it to peer review, but ask for the qualifier in the abstract and better statistical reporting.","headline":"The core finding is real and well supported, but the headline flip rates are measured on a deliberately vulnerable audit set, so the paper overstates the typical risk and under-reports variability.","tokens_in":9671,"tokens_out":1940,"would_cite":false,"duration_ms":20369,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Removing selected non-target companion cells from a query cohort can change a single cell's refined annotation even when its expression profile, base prediction, and trained model are unchanged.","keywords":["single-cell RNA sequencing","cell-type annotation","adversarial robustness","cohort dependence","label refinement","neighborhood voting","companion-cell removal","stability testing"],"falsifier":"Apply the same removal budget and search procedure to the full test cohort, including high-confidence and misclassified cells, and compare target flip rates to the preselected target set. If the full-cohort flip rate is indistinguishable from random subsampling noise, the claimed attack surface is an artifact of target selection rather than a general property of the refinement stage.","tokens_in":8632,"feed_emoji":"🧬","tokens_out":6835,"duration_ms":53167,"temperature":0.7,"pith_summary":"This paper tries to establish that cohort-dependent label refinement is a target-preserving attack surface in single-cell annotation. Removing a small set of non-target companion cells from the query cohort can change a target cell's refined annotation, even though the target cell's expression profile, the base classifier, and all trained parameters are left untouched. The authors test this with random, structured, and search-based removal strategies on two annotated single-cell datasets, finding that structured removal flips substantially more targets than random removal and that a multi-start search changes about 24% of targeted linear-SVM cells and about 20% of logistic-regression cells while mean collateral flip rates stay below 0.4%. Ablations show that the effect vanishes when the neighborhood-refinement stage is disabled, which isolates that stage as the mechanism. If the paper is right, clean annotation accuracy alone does not capture how reliable a refined label is, and cells with uncertain identities can have labels that depend on which other cells happen to be in the query cohort.","feed_headline":"Removing a few companion cells flips single-cell labels","feed_subtitle":"A new audit shows refined annotations change even when the target cell, its expression, and the trained model are untouched.","key_machinery":"The load-bearing mechanism is the neighborhood-refinement stage, in which the refined probability $r_{ic}$ for cell $i$ and class $c$ combines the frozen base probability $p_{ic}$ with a neighborhood support term $h_{ic}$ that is a confidence-weighted average of neighboring cells' base probabilities (equivalently, a uniform average under majority voting). When a removal set $S$ is deleted from the query cohort, the neighbor sets $\\mathcal{N}_k(i)$ are recomputed and $h_{ic}$ changes, which can shift the argmax of $r_{ic}$. The paper's removal strategies differ in how they choose $S$: random baseline, removing same-class cells, removing nearest cells, and search procedures (greedy, multi-start greedy, and beam search) that lexicographically optimize the target's margin while penalizing collateral flips.","core_discovery":"The central claim is that the final label produced by a cohort-dependent annotation pipeline is not determined by the target cell alone. In the paper's formulation, the refined probability for a target cell is a weighted combination of its base-classifier probability and a neighborhood-support term computed from the surrounding cells; removing any non-target cell changes the neighbor set and therefore the support term. The paper demonstrates, for two datasets and two frozen linear classifiers, that small structured removals flip the refined labels of a meaningful fraction of low-confidence but correctly annotated cells, and that the same pattern appears in an established majority-voting annotation tool where independent predictions never change but cohort-voted labels do. The paper concludes by identifying query cohort composition as a target-preserving attack surface and by recommending that pipelines report independent and refined predictions separately.","pith_inferences":["The target set is deliberately biased toward lower-confidence correctly annotated cells, so the headline flip rates are likely upper bounds for a typical query cohort; the paper does not measure the base rate of such vulnerable cells.","If vulnerable cells are concentrated along developmental transitions or overlapping populations, then cohort-removal flips could also distort downstream trajectory or composition analyses, not just the labels themselves.","A natural testable extension is to compute a stability score for every cell from an ensemble of controlled subsamples and check whether that score predicts disagreement between independent and refined labels or manual-review outcomes."],"forward_implications":["If the claim holds, any pipeline that refines labels by neighborhood voting or cluster majority is potentially sensitive to which companion cells survive quality control and downsampling.","The mechanism ablation implies that the fix should target the refinement stage: when the context weight is zero, no removal changes any target label.","In the majority-voting validation, all target flips occurred with unchanged independent predictions, so cohort-level voting is the point of failure, not the base classifier.","Reporting independent and refined labels separately, and flagging cells where they disagree, would give users a practical warning that a label may be cohort-dependent.","Repeated controlled subsampling could serve as a stability check, with unstable cells assigned broader lineage labels or routed to manual review."],"supporting_citations":[{"why":"provides the data structures and preprocessing software used to prepare the two evaluation datasets.","marker":"[1]"},{"why":"provides the majority-voting annotation pipeline used to show the threat appears in an established tool with unchanged independent predictions.","marker":"[2]"},{"why":"documents that random subsampling can distort single-cell dataset structure, motivating the question of whether small removals can alter labels.","marker":"[13]"},{"why":"establishes the adversarial-perturbation paradigm that motivates target-preserving robustness audits.","marker":"[15]"},{"why":"evaluates feature-level adversarial attacks on single-cell RNA-seq classifiers, the contrast that CohortHijack extends by leaving the target cell unchanged.","marker":"[20]"},{"why":"motivates focusing on low-confidence or uncertain annotations, which the target-selection policy relies on.","marker":"[24]"}],"fun_headline_variants":["CohortHijack: small companion-cell removals flip single-cell labels","Remove a few neighbors, flip single-cell annotations","CohortHijack shows refined labels hinge on companion cells","Target cell untouched, but a few neighbor cells change its label"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing assumption is that the low-confidence correctly annotated cells selected as targets represent the practically relevant vulnerable population; if such cells are rare in a real query cohort, the measured flip rates overstate how often companion-cell removal will change an arbitrary cell's label.","fun_headline_variants_meta":{"raw":{"variants":["CohortHijack: small companion-cell removals flip single-cell labels","Remove a few neighbors, flip single-cell annotations","CohortHijack shows refined labels hinge on companion cells","Target cell untouched, but a few neighbor cells change its label"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000446,"raw_usage":{"total_tokens":2228,"prompt_tokens":894,"completion_tokens":1334,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":510,"completion_tokens_details":{"reasoning_tokens":1261}},"tokens_in":510,"tokens_out":1334,"duration_ms":35227,"temperature":1.0,"reasoning_tokens":1261,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T21:26:08.165364+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Apply the same removal budget and search procedure to the full test cohort, including high-confidence and misclassified cells, and compare target flip rates to the preselected target set. If the full-cohort flip rate is indistinguishable from random subsampling noise, the claimed attack surface is an artifact of target selection rather than a general property of the refinement stage.","supporting_citations":[{"cited_title":"SCANPY: Large-scale single- cell gene expression data analysis,","cited_arxiv_id":null,"evidence_quote":"provides the data structures and preprocessing software used to prepare the two evaluation datasets."},{"cited_title":"Cross-tissue immune cell analysis reveals tissue-specific features in humans,","cited_arxiv_id":null,"evidence_quote":"provides the majority-voting annotation pipeline used to show the threat appears in an established tool with unchanged independent predictions."},{"cited_title":"scsampler: fast diversity- preserving subsampling of large-scale single-cell transcriptomic data,","cited_arxiv_id":null,"evidence_quote":"documents that random subsampling can distort single-cell dataset structure, motivating the question of whether small removals can alter labels."},{"cited_title":"adverscarial: assessing the vulnerability of single-cell rna-sequencing classifiers to adversarial attacks,","cited_arxiv_id":null,"evidence_quote":"evaluates feature-level adversarial attacks on single-cell RNA-seq classifiers, the contrast that CohortHijack extends by leaving the target cell unchanged."},{"cited_title":"Interpreting single-cell and spatial omics data using deep neural network training dynamics,","cited_arxiv_id":null,"evidence_quote":"motivates focusing on low-confidence or uncertain annotations, which the target-selection policy relies on."}],"review_version":1}