{"id":"8c0d191d-99a8-40c5-959a-dc22f5ce1e7e","arxiv_id":"2603.27465","paper_version":2,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":7.0,"correctness_risk":"high","formal_verification":"none","parameter_count":0,"one_line_summary":"Under 1% adversarially crafted DNA sequences can selectively degrade or backdoor genomic foundation models at pre-training and fine-tuning with little collateral damage.","lead":"DNA foundation models can be poisoned with under 1% crafted sequences so they fail on chosen genomic motifs while behaving normally elsewhere. This matters because public genome corpora are huge, hard to audit, and feed clinical and design tools.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.5","headline":"Full manuscript body is the wrong paper (Frobenius 1877 elliptic functions), so the <1% selective poisoning and LoRA backdoor claims remain uncheckable.","rationale":"The reader correctly diagnosed that the supplied full text is Frobenius & Stickelberger (1877), not the DNA-poisoning manuscript, and therefore correctly set UNVERDICTED with LOW confidence on abstract-only grounds. That diagnosis still holds: without methods and results for the stated Evo 2 / GENERator / ClinVar / BRCA1 experiments, neither the <1% selective pre-training attack nor the near-exclusive LoRA trigger can be assessed for soundness or for the transfer assumption the reader flagged. No new technical soft spot inside a real methods section can be identified because that section is not present. Verdict remains UNVERDICTED; no adjustment toward ACCEPT, CONDITIONAL, or REJECT is warranted until the correct manuscript body is available. The concrete test is simply to obtain and check that body against the abstract's quantitative claims.","tokens_in":5354,"tokens_out":604,"duration_ms":14252,"concrete_test":"Retrieve the actual arXiv:2603.27465 PDF/source (not 2603.27466). Confirm presence of quantitative results for (i) poison fraction <1% with selective degradation on TATA/CTCF/synthetic contexts vs. controls, (ii) LoRA backdoor activation nearly exclusive to the trigger, (iii) BRCA1 label-poison effect size. If those tables/figures are missing or contradict the abstract, the central claim fails; if present and consistent, re-evaluate transferability next.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim (abstract: <1% adversarially crafted sequences selectively degrade generative performance on TATA-box / CTCF / synthetic contexts while leaving unrelated sequences unaffected; LoRA backdoor activates almost exclusively on trigger; targeted label corruption compromises BRCA1 VEP) is an experimental claim. It requires inspectable methods, poison construction, training corpora, metrics, and ablations for Evo 2 / GENERator pre-training and the ClinVar LoRA / frozen-Evo2-7B BRCA1 fine-tuning setups. The CACHEABLE full text supplied for 2603.27465 is instead the English translation of Frobenius & Stickelberger 1877 on elliptic functions (arXiv 2603.27466), containing only determinant identities for σ and ℘. No attack rates, selectivity curves, trigger exclusivity numbers, or BRCA1 AUROC shifts can be verified. Until the correct body is present, the strongest claim is unsupported assertion, not demonstrated result. The reader's transferability concern is secondary; the primary load-bearing failure is absence of the evidence the claim depends on.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The submission’s abstract and title claim the first systematic study of training-data poisoning against DNA foundation models (Evo 2, GENERator), asserting that <1% adversarially crafted sequences can selectively degrade generative performance on TATA-box, CTCF, and synthetic contexts, that a LoRA fine-tune on a ClinVar-derived corpus installs a near-exclusive trigger backdoor, and that targeted label corruption of frozen Evo 2 7B embeddings compromises BRCA1 variant-effect prediction. The body supplied as the full manuscript is instead an English translation of Frobenius & Stickelberger (1877), “On the Theory of Elliptic Functions,” deriving determinant identities for the Weierstrass σ and ℘ functions (Eqs. 1–4). No methods, corpora, poison constructions, metrics, figures, or results on genomic models appear.","tokens_in":5589,"tokens_out":686,"duration_ms":11217,"significance":"If the abstract’s empirical claims were supported by a matching manuscript, the work would be significant: it would be the first systematic demonstration that genomic language models are vulnerable to low-footprint, context-selective data poisoning at both pre-training and fine-tuning, with direct implications for clinical variant interpretation and genome design. The supplied body, however, contains none of those results, so the claimed significance cannot be assessed or credited.","major_comments":[{"comment":"Title/abstract vs. full text: the entire experimental programme (Evo 2 / GENERator pre-training poisons at <1%, TATA-box / CTCF / synthetic-insert scenarios, ClinVar LoRA backdoor, frozen-Evo2-7B BRCA1 label corruption) is absent. The body is Frobenius–Stickelberger 1877 (arXiv:2603.27466), containing only elliptic-function determinant identities. No attack rates, selectivity curves, trigger-exclusivity numbers, or AUROC shifts can be verified. The central claims are therefore unsupported assertions, not demonstrated results.","section":null},{"comment":"Because the load-bearing evidence (methods, poison construction, training corpora, evaluation metrics, ablations, baselines) is missing, the abstract’s transferability claim—that genomic foundation models in general are susceptible with minimal footprint—cannot be evaluated. A correct manuscript body is required before any scientific assessment is possible.","section":null}],"minor_comments":[{"comment":"The supplied body is a clean, well-annotated historical translation of a classical paper; presentation quality of that text is not at issue. The mismatch with the claimed DNA-poisoning paper is the sole defect.","section":null}],"recommendation":"reject","confidential_remarks":"The CACHEABLE full-text block for paper_id 2603.27465 is the wrong document (it is the Frobenius–Stickelberger translation, arXiv 2603.27466). This appears to be a content-routing or arXiv-id collision error rather than an author submission of elliptic-function history under a genomics title. The editor should request the correct PDF/source for 2603.27465 before any further review; until then the submission cannot be refereed on its scientific merits."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The one thing you need to know: we do not have the manuscript. The body attached under 2603.27465 is Frobenius & Stickelberger 1877 on σ and ℘ determinants (arXiv 2603.27466). Everything below is abstract-only for the DNA-poisoning paper.\n\nWhat the abstract claims is new and, if true, worth the field’s attention. First systematic poisoning study on genomic foundation models, both pre-training (Evo 2, GENERator) and fine-tuning (LoRA CTCF backdoor; frozen Evo 2 7B embeddings + BRCA1 label flip). The headline empirical claim is selective degradation of targeted contexts (TATA-box, CTCF, synthetic inserts) with under 1% crafted sequences, unrelated sequences left alone, plus a near-exclusive trigger backdoor and a clinically framed VEP compromise. That framing is clear, the attack surface (opaque DNA tokens, public multi-trillion-token corpora) is real, and the call for provenance and integrity checks is the right policy ask.\n\nSoft spots are not subtle: there are no methods, poison recipes, corpora, metrics, ablations, error bars, or baselines to inspect. Transfer from controlled Evo 2 / GENERator / ClinVar setups to real public pipelines is asserted, not shown. Circularity is not the issue; missing evidence is. The reader’s low soundness score and the stress-test note are correct on that point; I am not going to invent a verdict from an abstract.\n\nWho it is for: people building or auditing genomic LMs, and anyone who already tracks data-poisoning in foundation models. A serious editor should send a real full paper with those experiments to referees—the topic and the claimed effect sizes justify referee time. I would not cite or run a reading group on the abstract alone. Get the correct PDF; until then this is an unverified security claim, not a demonstrated result.","headline":"Abstract-only security claim on DNA model poisoning; the supplied “full text” is the wrong paper (1877 elliptic functions), so the <1% selective-attack results are still uncheckable.","tokens_in":6211,"tokens_out":511,"would_cite":false,"duration_ms":10259,"reading_group":"no","serious_thinker":"unclear","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Less than 1% adversarially crafted DNA sequences can selectively degrade genomic foundation models on targeted contexts while leaving unrelated sequences untouched.","keywords":["DNA foundation models","data poisoning","backdoor attacks","genomic language models","adversarial robustness","variant effect prediction","training data integrity","CTCF"],"falsifier":"Retrain a genomic foundation model from scratch on a large public corpus after inserting the same <1% crafted sequences and measure whether generative metrics drop selectively on the targeted TATA-box or CTCF contexts while control contexts remain intact; absence of that selective degradation would falsify the central claim.","tokens_in":6255,"feed_emoji":"🧬","tokens_out":854,"duration_ms":22007,"temperature":0.7,"pith_summary":"Genomic foundation models learn from enormous public DNA datasets that lack the semantic cues that make poisoned text easy to spot. This paper shows that inserting a tiny fraction of carefully designed sequences into pre-training data is enough to selectively ruin the model's generative behavior on chosen biological motifs—TATA-box promoters, CTCF binding sites, or synthetic inserts—without harming performance elsewhere. The same idea extends to fine-tuning: a few poisoned CTCF sites install a conditional backdoor that fires almost only when a trigger is present, and targeted label flips on frozen embeddings can selectively break a clinically relevant BRCA1 variant classifier. The authors argue that these results establish a real vulnerability of DNA language models and that data provenance, integrity checks, and adversarial testing must become standard practice before such models are trusted for genome design or clinical prediction.","feed_headline":"Under 1% poisoned DNA selectively breaks genomic AI","feed_subtitle":"Crafted sequences sabotage TATA, CTCF and BRCA1 tasks while the rest of the model looks fine.","key_machinery":"Targeted insertion of crafted DNA sequences into the pre-training corpus and, at fine-tuning, subset poisoning of CTCF sites or label corruption of downstream data; these rare but consistent corruptions are absorbed into the model's representations of the targeted motifs or classes.","core_discovery":"Genomic foundation models are susceptible to targeted training-data poisoning with a footprint under one percent: adversarially crafted sequences inserted into pre-training (demonstrated on Evo 2 and GENERator) selectively degrade generative performance on chosen genomic contexts while leaving unrelated sequences unaffected, and fine-tuning poisons can install nearly exclusive conditional backdoors or selectively compromise clinically relevant variant classification such as BRCA1.","pith_inferences":["Similar low-footprint poisoning risks likely apply to protein and RNA foundation models trained on public sequence databases.","Federated or multi-lab genomic training may need cryptographic commitments to training sets to detect unauthorized insertions.","The opacity of nucleotide tokens may make influence-function or gradient-based defenses harder to apply than in natural-language models.","Clinical AI regulators may eventually require documented poison-resistance testing for models used in variant interpretation."],"forward_implications":["Public genomic training corpora require systematic provenance tracking and integrity verification before model training.","Adversarial robustness evaluation against data poisoning should become a standard step in genomic model development.","Conditional backdoors can be installed with minimal poison so a model behaves normally except when a trigger sequence is present.","Clinically used downstream tasks such as BRCA1 variant effect prediction can be selectively compromised by targeted label corruption.","Because DNA lacks semantic transparency, conventional text-poison detectors will not transfer; new DNA-specific curation methods are needed."],"fun_headline_variants":["Under 1% DNA poison selectively sabotages genomic foundation models","Sub-1% crafted sequences backdoor DNA models at pretrain and fine-tune","Tiny adversarial DNA installs nearly exclusive CTCF and BRCA1 backdoors","Less than 1% poisoned genomes break targeted tasks while rest stays fine","Adversarial DNA under 1% degrades TATA CTCF and BRCA1 in foundation models"],"cache_read_input_tokens":128,"weakest_assumption_plain":"The attacks that succeed with under 1% poison in the authors' controlled Evo 2, GENERator, ClinVar and BRCA1 experiments will still work against real multi-trillion-token public pipelines that use different curation, deduplication and filtering.","fun_headline_variants_meta":{"raw":{"variants":["Under 1% DNA poison selectively sabotages genomic foundation models","Sub-1% crafted sequences backdoor DNA models at pretrain and fine-tune","Tiny adversarial DNA installs nearly exclusive CTCF and BRCA1 backdoors","Less than 1% poisoned genomes break targeted tasks while rest stays fine","Adversarial DNA under 1% degrades TATA CTCF and BRCA1 in foundation models"]},"model":"grok-4.5","effort":"low","cost_usd":0.004188,"raw_usage":{"total_tokens":1314,"prompt_tokens":828,"num_sources_used":0,"completion_tokens":85,"cost_in_usd_ticks":41880000,"prompt_tokens_details":{"text_tokens":828,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":401,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":828,"tokens_out":85,"duration_ms":4718,"temperature":1.0,"reasoning_tokens":401,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-13T16:54:06.336584+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Retrain a genomic foundation model from scratch on a large public corpus after inserting the same <1% crafted sequences and measure whether generative metrics drop selectively on the targeted TATA-box or CTCF contexts while control contexts remain intact; absence of that selective degradation would falsify the central claim.","supporting_citations":[],"review_version":1}